Literary work geographic mark presentation system based on literature, history and geography fusion
By integrating literary works and historical data, using natural language processing and geographic information systems, building a literary density thermal value matrix and generating causal assumptions, the correlation problem of unspecified geographical areas in literary geography research is solved, and automated detection and in-depth analysis of literary blank areas are realized.
Patent Information
- Application Number
- CN202510463768.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to effectively and dynamically relate to geographical areas not clearly mentioned in literary works, resulting in inefficient literary geography research and the inability to automatically infer the causes of literary blank areas, which is difficult to support the exploration of deep causal relationships.
By integrating literary works, historical geographical data and population density data, natural language processing models are used to extract geographical entities and label spatiotemporal labels, a literary density thermal value matrix is constructed, negative space candidates are screened, and causal assumptions are generated with historical events, and a dynamic visual map is finally generated.
It significantly improves the degree of automation and empiricality of literary geographic analysis, can efficiently detect and reveal the distribution rules of the "negative space" of literary geography and its deep correlation with historical background, and supports cultural research, education and heritage protection.
Smart Images

Figure CN120371910A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a geographical marking presentation system for literary works based on the integration of literature, history, and geography. Background Art
[0002] In the field of literary geography research, existing technologies usually rely on manual sorting of the geographical locations explicitly mentioned in the text and static annotation through a Geographic Information System (GIS). Although such methods can intuitively display the spatial distribution of literary works, they have the following significant limitations:
[0003] On the one hand, a large number of geographical regions that are vague or not mentioned in literary works (such as the Yunnan region rarely described in Tang Dynasty poems) are often overlooked, and these "silent regions" may imply important historical, political, or cultural information.
[0004] On the other hand, existing tools cannot dynamically associate interdisciplinary data such as historical events and population migrations, making it difficult for researchers to automatically infer the causes of literary blank areas.
[0005] The above problems lead to low efficiency in literary geography research and are difficult to support the mining requirements for deep causal relationships in application scenarios such as education and cultural protection. Summary of the Invention
[0006] Based on this, it is necessary to provide a geographical marking presentation system for literary works based on the integration of literature, history, and geography in view of the above technical problems.
[0007] In a first aspect, the present application provides a geographical marking presentation system for literary works based on the integration of literature, history, and geography, including: a data collection module, a geographical entity processing module, a grid analysis module, a heat matrix construction module, a negative space determination module, a causal association module, and a visualization display module;
[0008] The data collection module is used to obtain literary works, historical geographical data, historical time period length, and historical population density data;
[0009] The geographical entity processing module is used to extract geographical entities in literary works through a natural language processing model, annotate spatio-temporal tags to the geographical entities, and generate a structured spatio-temporal tag data set;
[0010] The grid analysis module is used to generate multiple geographical grids based on historical geographical data; based on the structured spatio-temporal tag data set, count the number of mentions of geographical entities in each geographical grid in literary works;
[0011] The heat matrix construction module is used to generate a literary density heat value matrix based on the spatio-temporal dimension according to the number of mentions, historical time period length, and historical population density data;
[0012] A negative space determination module, configured to screen, based on a literary density heat value matrix, geographical grids with a literary density lower than a preset literary density threshold and a historical population greater than zero as negative space candidate areas;
[0013] A causal association module, configured to associate the negative space candidate areas with each historical event, generate causal hypotheses, and calculate the confidence levels of the causal hypotheses; screen historical events for which causal hypotheses with confidence levels greater than a preset threshold are generated as candidate historical events;
[0014] A visual display module, configured to dynamically superimpose the negative space candidate areas and the candidate historical events on a map interface to generate a dynamic visualization map.
[0015] In a second aspect, the present application further provides a method for presenting geographical markers of literary works based on the integration of literature, history, and geography, applied to the system in the first aspect, including:
[0016] S1: A data acquisition module acquires literary works, historical geographical data, the length of a historical time period, and historical population density data;
[0017] S2: A geographical entity processing module extracts geographical entities in the literary works through a natural language processing model, and labels spatio-temporal tags for the geographical entities to generate a structured spatio-temporal tag data set;
[0018] S3: A grid analysis module generates multiple geographical grids based on the historical geographical data; based on the structured spatio-temporal tag data set, counts the number of mentions of geographical entities in each geographical grid in the literary works;
[0019] S4: A heat matrix construction module generates a literary density heat value matrix based on the spatio-temporal dimension according to the number of mentions, the length of the historical time period, and the historical population density data;
[0020] S5: The negative space determination module screens, based on the literary density heat value matrix, geographical grids with a literary density lower than a preset literary density threshold and a historical population greater than zero as negative space candidate areas;
[0021] S6: The causal association module associates the negative space candidate areas with each historical event, generates causal hypotheses, and calculates the confidence levels of the causal hypotheses; screens historical events for which causal hypotheses with confidence levels greater than a preset threshold are generated as candidate historical events;
[0022] S7: The visual display module dynamically superimposes the negative space candidate areas and the candidate historical events on a map interface to generate a dynamic visualization map.
[0023] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a method for presenting geographical markers of literary works based on the integration of literature, history, and geography as described in the first aspect.
[0024] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for presenting geographical markers of literary works based on the integration of literature, history, and geography as described in the first aspect.
[0025] The above-mentioned system for presenting geographical markers of literary works based on the integration of literature, history, and geography integrates the original text of literary works, historical and geographical data, the length of historical time periods, and historical population density data; uses a natural language processing model to extract geographical entities and annotate spatio-temporal tags to generate a structured data set; constructs a matrix of literary density heat values in the spatio-temporal dimension to quantitatively analyze geographical distribution characteristics, combines a historical event database to dynamically screen negative space candidate areas with literary density lower than the threshold and active historical data, and generates causal hypotheses, and generates a dynamic visualization map by overlaying a visualization heat map and a timeline on a map interface. Thereby, it efficiently detects and reveals the distribution law of the "negative space" of literary geography and its deep connection with the historical background, significantly improves the automation degree and empirical nature of literary geography analysis, and provides intelligent decision-making support for cultural research, education, and heritage protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 It is a schematic structural diagram of a system for presenting geographical markers of literary works based on the integration of literature, history, and geography provided by the present invention;
[0028] Figure 2 It is a schematic flowchart of a method for presenting geographical markers of literary works based on the integration of literature, history, and geography provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0030] Refer to Figure 1, which shows a schematic structural diagram of a geographical marking presentation system 10 for literary works based on the integration of literature, history, and geography provided by this application. The system includes: a data collection module 11, a geographical entity processing module 12, a grid analysis module 13, a heat matrix construction module 14, a negative space determination module 15, a causal association module 16, and a visualization display module 17;
[0031] The data collection module 11 is used to obtain literary works, historical and geographical data, the length of historical time periods, and historical population density data.
[0032] Specifically, in the data collection process, data fusion technology is comprehensively used to unify and integrate structured (such as database tables) and unstructured (such as literary texts) data. Through the ETL (Extract, Transform, Load) process, data from different sources and with different formats can be extracted, transformed, and loaded. In the transformation link, according to predefined data models and semantic rules, text descriptions in literary works, coordinate information of historical and geographical data, numerical statistics of population density, etc. are transformed into a common data format within the system.
[0033] 1) For the acquisition of literary works: API connections can be established with major libraries and literature databases, such as the digital resource interface of the National Library of China and academic databases such as JSTOR. By setting retrieval conditions such as keywords, author names, and work names, relevant literary texts are regularly crawled. The crawling process follows corresponding data transfer protocols, such as OAuth 2.0 for identity authentication and authorization, to ensure legal data acquisition. For paper ancient books, optical character recognition (OCR) technology is used for digital conversion. A professional ancient book OCR software with an accuracy higher than 98% is adopted. For classical Chinese with rich traditional Chinese characters and variant characters, as well as aging documents with different printing fonts and layouts, the text recognition algorithm is optimized, and an artificial proofreading link is paired to ensure conversion accuracy.
[0034] 2) For the collection of historical and geographical data: Modern geographical data can be imported from professional geographical information platforms, such as the Amap Open Platform, to obtain vector data including administrative division boundaries, topographic features, and water system distributions, with the accuracy controlled at the sub-meter level to ensure clear geographical outlines and accurate positions. Ancient geographical information in historical documents is mined and connected to databases such as the "China Historical Geographic Information System" (CHGIS), which covers data such as territorial changes, the rise and fall of cities and towns, and transportation routes from the pre-Qin period to modern times, with the time granularity accurate to the period of dynasty changes or major historical event nodes, and the spatial granularity can be located at the county-level administrative division.
[0035] 3) For the integration of historical time - period lengths and population density data: Information on historical time - period lengths can be extracted from historical statistical yearbooks and archaeological reports. For different dynasties and periods of complex regime changes, a time - series database can be established, with centuries as the major units, subdivided into decades or even finer year intervals, and key time tags such as the start of dynasties and turning points between prosperous and chaotic eras are marked. For historical population density data, multiple sources of statistical materials can be integrated, such as ancient household registers, tax records, and modern census reports. Data cleaning algorithms are used to remove duplicates, errors, and missing values, and then combined with the spatial analysis function of Geographic Information System (GIS) to match population data with geographical regions in corresponding historical periods, generating a rasterized population density distribution map. The raster accuracy can be dynamically adjusted between 1km×1km and 10km×10km according to the detail level of the data.
[0036] The geographical entity processing module 12 is used to extract geographical entities from literary works through a natural language processing model, annotate spatio - temporal tags for the geographical entities, and generate a structured spatio - temporal tag data set.
[0037] Specifically, this module is based on the cross - technology of natural language processing and knowledge engineering. The natural language processing model mines potential semantic and entity information from the text, while the knowledge base provides prior knowledge constraints and supplements. The two work together to achieve the accurate identification of geographical entities and spatio - temporal semantic annotation. During the annotation process, geographical coding technology is used to convert the name text of geographical entities into standardized geographical coordinates to ensure the precise positioning of entities in the geographical space.
[0038] 1) For the extraction of geographical entities:
[0039] A natural language processing (NLP) model based on deep learning can be used, especially pre - trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and its variants. According to the characteristics of literary texts, the model is fine - tuned. Inputting text fragments of literary works, the model captures the semantic features of the text through the encoder, outputs the vector representation of each word, and then through the Named Entity Recognition (NER) layer, using the Conditional Random Field (CRF) algorithm and combining context semantic associations, accurately identifies geographical entities, including the names of mountains, rivers, cities, villages, etc.
[0040] Build a geographical entity knowledge base, collecting various geographical nouns at home and abroad, covering common names (such as "mountain", "river") and proper names (such as "Mount Tai", "Yellow River"), and recording their aliases and abbreviations. When the model identifies potential geographical entities, it is compared and verified with the knowledge base to supplement and improve entity information, such as correcting misidentifications caused by interchangeable characters and different meanings between ancient and modern times, and enhancing the robustness of entity extraction.
[0041] 2) Labeling of spatiotemporal labels:
[0042] For the extracted geographical entities, based on the time clues in the literary works, such as the author's creation year, the narrative time background of the work, and the clear time words in the text ("the Kaiyuan Period", "the fifteenth year of Jiaqing"), combined with the timeline of historical events, precise time labels are added to the entities. The time accuracy is refined from the span of centuries to specific years, seasons, and even shorter periods, such as the several months of a battle.
[0043] In terms of spatial positioning, the names of geographic entities are jointly searched with historical geographic databases and modern geographic information systems to obtain their latitude and longitude coordinates and geographic area boundary information. In the event of changes in ancient place names, the relationship between their historical locations and modern corresponding locations is recorded to construct an entity mapping in the space-time coordinate system.
[0044] The grid analysis module 13 is used to generate multiple geographic grids based on historical geographic data; based on the structured spatiotemporal label data set, count the number of times the geographic entities in each geographic grid are mentioned in literary works.
[0045] Specifically, grid analysis relies on spatial segmentation and statistical analysis techniques. By discretizing continuous geographic space into grids, complex geographic distribution problems are transformed into statistical problems within grid units, which facilitates quantitative analysis and calculation. Spatial overlay analysis technology realizes the spatial association between geographic entities and grids, and statistical algorithms quantitatively summarize the citations of geographic entities in the text. The combination of the two comprehensively presents the focus of literary works in geographic space and the distribution of blank areas.
[0046] 1) For the generation of geographic grid:
[0047] Based on the boundary range of historical geographic data, the benchmark area for grid division can be determined, which can be based on the national territory or the administrative area of a specific historical period. The Fishnet tool of the Geographic Information System (GIS) is used to set the latitude and longitude intervals of the grid. For example, a coarser interval (0.5°×0.5°) is used in areas with vast areas and large differences in geographical features, and it is refined to 0.1°×0.1° in areas with rich geographical details and intensive literary activities. A rectangular geographic grid covering the target area is generated, and each grid cell has a unique identifier and spatial location information.
[0048] Consider the actual significance of topography to the grid. For example, in mountainous areas, in order to avoid analysis distortion caused by a grid spanning multiple mountains, the grid is fine-tuned based on terrain data to make the grid boundary fit the natural geographic unit boundary as closely as possible, thereby enhancing the spatial fit between the grid and the geographic entity.
[0049] 2) For the statistics of geographic entity mentions:
[0050] Perform a spatial overlay analysis on the geospatial entity dataset with spatio-temporal tags generated by the geospatial entity processing module and the geospatial grid. For each geospatial entity, determine the geospatial grid cell it belongs to based on its coordinates, and use a spatial query algorithm to count the number of times the geospatial entity is mentioned in literary works within each grid. The counting process differentiates between different literary genres (poetry, prose, novels, etc.) and different author groups (region, era, genre), forming a multi-dimensional mention frequency matrix, and recording the frequency of the entity being mentioned under a specific grid and specific literary classification.
[0051] The heat matrix construction module 14 is used to generate a literary density heat value matrix based on spatio-temporal dimensions according to the mention frequency, historical time period length, and historical population density data.
[0052] Specifically, using multi-index comprehensive evaluation technology, decompose complex social and cultural phenomena (literary creation density) into quantifiable and interpretable sub-indicators, reflect the importance of each factor through weight allocation, standardize to eliminate dimension and magnitude differences, and finally fuse and calculate to obtain a comprehensive heat value. Organize the data in matrix form for subsequent modules to screen and analyze according to multiple dimensions, revealing the density pattern of literary creation in the geographical space and the long river of time.
[0053] 1) For weight setting and data standardization:
[0054] Allocate weights to the three key indicators of mention frequency, historical time period length, and population density. Based on the expert experience in literary geography research and statistical correlation analysis, for example, determine the weight of mention frequency as 0.5, the weight of historical time period length as 0.3, and the weight of population density as 0.2, reflecting the comprehensive influence of literary creation by population aggregation, time accumulation, and text attention.
[0055] Perform standardization processing on the data of each indicator. The mention frequency uses Min-Max standardization to map the mention frequencies under different grids and different literary classifications to the [0, 1] interval; calculate the proportion of the historical time period length in a specific dynasty or historical stage, such as the proportion of the existence time of a certain grid in the Tang Dynasty to the total length of the Tang Dynasty; standardize the population density through Z-score to eliminate dimension differences in different historical periods and regions, ensuring the comparability of the numerical values of each indicator.
[0056] 2) For the calculation of literary density heat value:
[0057] A weighted calculation model can be constructed. After multiplying the standardized mention count, historical time period length, and population density value by their corresponding weights respectively and then summing them up, the literary density heat value of each geographical grid can be obtained. The formula can be: Literary density heat value = Standardized mention count × 0.5 + Standardized historical time period length × 0.3 + Standardized population density × 0.2. The calculation process is executed in parallel on a distributed computing framework such as Hadoop or Spark to efficiently process a large amount of grid data and generate a complete matrix of literary density heat values based on the spatio-temporal dimension. The matrix dimensions cover multiple labels such as grid ID, time period, literary genre, etc., and finely depict the distribution of literary popularity.
[0058] The negative space determination module 15 is used to screen, based on the matrix of literary density heat values, geographical grids with a literary density lower than a preset literary density threshold and a historical population greater than zero as candidate areas for the negative space.
[0059] Specifically, it is based on statistical threshold analysis and spatial data filtering techniques. The statistical method starts from the data distribution law to scientifically define the abnormal low-value interval, and spatial data filtering combines population elements to eliminate areas that do not conform to the real logic. The double screening ensures that the determined negative space not only conforms to the statistical anomalies of literary creation but also has practical significance for historical and cultural research, accurately delimiting the geographical scope that needs to be explored in depth.
[0060] 1) For the setting of the literary density threshold and preliminary screening:
[0061] Analyze the global distribution characteristics of the matrix of literary density heat values, calculate its mean and standard deviation, and based on the principle of normal distribution, set the literary density threshold as the mean minus 1 times the standard deviation to ensure that areas significantly lower than the average level are screened out. At the same time, referring to the "edge effect" and "core-periphery structure" in literary research, for areas such as the edge of the cultural center and around transportation arteries that should theoretically have a certain degree of literary attention, appropriately lower the threshold to enhance the pertinence of the screening.
[0062] For each geographical grid in the matrix, compare its literary density heat value with the threshold, and initially select the grids lower than the threshold as potential negative space areas, generating a candidate list containing grid ID, heat value, and geographical location information. The list is sorted in ascending order of heat value to highlight areas with extremely low density.
[0063] 2) For historical population filtering and negative space confirmation:
[0064] Spatially associate the candidate negative space grid with historical population density data, and use spatial query statements to filter out the grids with historical population greater than zero. Because even if literary creation is scarce, as long as there is a certain scale of population, it means that there may be a cultural activity foundation. Exclude the desert areas where there is no literature due to uninhabitedness. The remaining grids are the real negative space candidate areas, forming the final negative space determination list, which details the spatio-temporal scope, population scale, surrounding geographical environment, etc. of the negative space.
[0065] The causal association module 16 is used to associate the negative space candidate areas with various historical events, generate causal hypotheses, and calculate the confidence levels of the causal hypotheses; filter out the historical events that generate causal hypotheses with confidence levels greater than the preset threshold as candidate historical events.
[0066] Specifically, this module integrates interdisciplinary knowledge and statistical causal inference techniques. Interdisciplinary knowledge provides theoretical support for constructing reasonable causal hypotheses, and statistical models quantify the credibility of the hypotheses. By combining data-driven and expert experience, it deeply analyzes the internal causal relationship between the negative space and historical events, accurately extracts key causes from a large amount of historical information, and helps in-depth interpretation of cultural phenomena and decision-making for inheritance and protection.
[0067] 1) For historical event collection and spatio-temporal matching:
[0068] Historical events in multiple fields such as politics, military, economy, and culture can be sorted out, event information can be extracted from historical literature databases, archaeological excavation reports, and academic research results, and a historical event library can be constructed to record key attributes such as event names, time ranges of occurrence, geographical areas involved, and event types (such as wars, capital relocations, natural disasters, cultural policy reforms).
[0069] For the negative space candidate areas, according to their spatio-temporal scope, retrieve the events that overlap or are closely adjacent to them in time and space in the historical event library. Use spatio-temporal buffer analysis technology to generate a buffer range that extends before and after in time and expands around in space for the negative space, filter out all relevant historical events within the buffer area, and establish an association set between the negative space and the events. The set elements include event details and spatio-temporal distance indicators from the negative space.
[0070] 2) For causal hypothesis generation and confidence level calculation:
[0071] For each associated event, generate causal hypotheses by combining interdisciplinary knowledge such as literary creation laws, cultural dissemination theories, and population migration patterns. For example, if literary creation is scarce in a certain negative space during a specific period, and if the area experiences large-scale wars during that period, assume that the wars led to a sharp reduction in the population and damage to cultural facilities, thus inhibiting literary development; if there is an economic recession, assume that the weakening of the economic foundation caused a decline in cultural consumption and creative motivation.
[0072] Using statistical causal inference methods such as Bayesian networks and structural equation models to calculate the confidence of causal hypotheses. Taking the Bayesian network as an example, a directed acyclic graph containing negative space literary density and historical event variables is constructed. The network parameters are learned through historical data, and according to Bayes' theorem, the probability that the negative space presents a low literary density under the given historical events is calculated, that is, the confidence value. At the same time, an expert scoring mechanism can also be introduced to invite historians and literary research experts to score the rationality of the hypotheses, and integrate the scoring results into the confidence calculation to comprehensively determine the reliability of the causal hypotheses.
[0073] 3) For the screening of candidate historical events:
[0074] According to the preset confidence threshold (such as 0.6), historical events with a confidence higher than the threshold are screened out from all causal hypotheses as candidate historical events with a strong causal association with the negative space literary blank phenomenon, generating a final list of candidate historical events. The list content covers event details, causal hypothesis descriptions, confidence values, etc., providing in-depth interpretation materials for the visualization display module.
[0075] The visualization display module 17 is used to dynamically superimpose the negative space candidate area and candidate historical events on the map interface to generate a dynamic visualization map.
[0076] Specifically, this module is based on geographic information system visualization technology and interaction design principles. GIS software provides powerful map rendering and spatial analysis functions to achieve accurate visualization of geographical data; the interaction design follows the user-centered principle, and through diversified interaction controls and response mechanisms, it meets the personalized exploration needs of users, transforms complex data into an intuitive, easy-to-understand and convenient-to-operate visualization product, and strongly supports multiple application scenarios such as education communication and cultural decision-making.
[0077] 1) For the construction of the map interface and data loading:
[0078] A professional geographic information system (GIS) software platform such as ArcGIS or QGIS can be selected to build a basic map interface. Import modern geographic base maps and historical geographic data layers, set multi-level map zoom scales, from the macroscopic national perspective to the microscopic county details, to meet the observation needs of users at different scales. During the map loading process, the tile map technology is used to divide the map into multiple tiles and load them dynamically as needed, improving the loading speed and interaction fluency.
[0079] Convert the negative space candidate area data output by the negative space determination module and the candidate historical event data screened by the causal association module into the data formats supported by GIS software (such as Shapefile, GeoJSON) to ensure that the data can be accurately located and rendered in the map interface. The negative space is identified by special filling patterns (such as diagonal lines, dot matrices) or colors (low-saturation gray tones), and historical events are displayed on the corresponding geographical coordinates with icons (such as an explosion icon representing war, a gold coin icon representing economic events) or marking symbols.
[0080] 2) For dynamic overlay and interaction design:
[0081] Implement the dynamic overlay display function of the negative space and historical events. Users can adjust the historical period through the time slider to observe the distribution changes of the negative space and the evolution of related historical events in different periods. During the overlay process, use layer control technology to allow users to customize the opening or closing of the negative space and different types of historical event layers, and freely combine the content to view.
[0082] Design rich interaction operations. Click on the negative space area to pop up an information box to display details such as the spatio-temporal range, literary density heat value, and list of related historical events in this area; click on the historical event icon to present in-depth information such as event brief, causal hypothesis, and confidence level. At the same time, support users for spatial queries, such as selecting a specific area, counting the number of negative spaces and the distribution of historical event types within the area. The query results are presented in the form of charts (bar charts, pie charts) to enhance data readability.
[0083] The above-mentioned literary work geographical marking presentation system based on the integration of literature, history, and geography integrates the original text of literary works, historical and geographical data, historical time period length, and historical population density data; uses a natural language processing model to extract geographical entities and annotate spatio-temporal tags to generate a structured data set; constructs a literary density heat value matrix in the spatio-temporal dimension to quantitatively analyze geographical distribution characteristics, combines with a historical event database to dynamically screen negative space candidate areas with literary density lower than the threshold and active historical data and generate causal hypotheses, and generates a dynamic visualization map by overlaying a heat map and a timeline on the map interface. Thus, it can efficiently detect and reveal the distribution law of the literary geography "negative space" and its deep association with the historical background, significantly improve the automation degree and empirical nature of literary geography analysis, and provide intelligent decision-making support for cultural research, education, and heritage protection.
[0084] In an optional embodiment, the heat matrix construction module 14 includes:
[0085] A normalization processing unit 141, which is used to perform normalization calculations on the mention times based on the historical time period length and historical population density to generate a normalized literary density value.
[0086] Specifically, the normalization process is a key technical step in solving multi-source data fusion and comparability. Due to the significant differences in dimension and magnitude among the mention count, historical time period length, and population density data, direct comprehensive calculation will lead to data deviation and information distortion. Through normalization technology, these heterogeneous data are mapped to a unified numerical interval, eliminating the differences in dimension and magnitude, enabling the fusion calculation of different indicator data under the same benchmark, and laying a foundation for accurately constructing the literary density heat value matrix.
[0087] 1) Regarding the selection and application of data standardization methods:
[0088] For the mention count data, the Min-Max normalization method is adopted. Let the original value of the mention count be x ij (where i represents the geographical grid and j represents the literary work category), then the normalized mention count value x' ij is calculated as follows:
[0089]
[0090] where, min(x j ) and max(x j ) respectively represent the minimum and maximum mention counts of the j-th category of literary works in all geographical grids. This method maps the mention count data to the [0,1] interval, highlighting the relative differences.
[0091] For the historical time period length, ratio normalization is adopted. The calculation formula is: Let the existence duration of a certain geographical grid in the k-th historical period be t ik , and the total duration of this period be T k , then the normalized time length value t' ik is:
[0092]
[0093] This method reflects the time proportion of the geographical grid in a specific historical period and its degree of influence by time.
[0094] Regarding the population density data, Z-score normalization is used. The calculation formula is: Let the population density of geographical grid i in the k-th historical period be p ik , then the normalized population density value p' ik is calculated as:
[0095]
[0096] where, μ k and σ kThey respectively represent the mean and standard deviation of the population density of all geographical grids in the k-th historical period. This method eliminates the distribution differences of population density data in different historical periods and makes them comparable.
[0097] 2) For the fusion calculation to generate the normalized literary density value:
[0098] Set the weights of each index. Based on the expert experience in literary geography research and statistical correlation analysis, determine that the weight of the mention frequency is w1, the weight of the historical time period length is w2, and the weight of the population density is w3.
[0099] For each geographical grid i, fuse the three normalized data to calculate the normalized literary density value D i , and the formula is:
[0100] D i = w1 × x′ ij + w2 × t′ ik + w3 × p′ ik ;
[0101] This calculation is executed in parallel on a distributed computing framework such as Spark, efficiently processing a large amount of grid data, generating a dataset containing the normalized literary density values of all geographical grids, and providing accurate data input for the subsequent three-dimensional matrix construction.
[0102] The three-dimensional matrix construction unit 142 is used to map the normalized literary density value into a three-dimensional matrix of the time dimension, longitude dimension, and latitude dimension, and obtain the literary density heat value matrix.
[0103] Specifically, the three-dimensional matrix construction is a key step in organizing the normalized literary density data into a clear spatio-temporal semantic structure. By mapping the literary density value to the three dimensions of time, longitude, and latitude, a comprehensive and intuitive literary creation density model is constructed, realizing an all-round characterization of literary creation under the intersection of geographical space and historical time, and facilitating in-depth analysis and visual display by subsequent modules.
[0104] 1) For the matrix dimension definition and initialization:
[0105] The time dimension can divide the entire research time span into N time periods according to the historical period division standard, and assign a unique time identifier t = 1, 2,..., N to each time period.
[0106] The longitude dimension can divide the longitude grid points according to the longitude range of the research area with a set longitude interval (such as 0.1°), and generate M longitude identifiers lon = 1, 2,..., M.
[0107] Similarly for the latitude dimension, divide the latitude grid points and generate L latitude identifiers lat = 1, 2,..., L.
[0108] Initialize a three-dimensional matrix of N×M×L, with the initial values of elements set to 0, for storing the literary density heat values at each spatio-temporal position.
[0109] 2) For data mapping and matrix filling:
[0110] For each geographical grid i, obtain its longitude i , latitude i and the normalized literary density value D i .
[0111] Determine the corresponding position of the grid in the three-dimensional matrix. One can find the longitude grid point identifier lon i nearest to longitude j in the longitude dimension, and the latitude grid point identifier lat i nearest to latitude k in the latitude dimension. Combine the time period identifier t m it belongs to, and fill D i into the position (t m , lon j , lat k ) of the three-dimensional matrix.
[0112] If multiple geographical grids are mapped to the same matrix position, then perform a weighted average on the corresponding D i values. The weights can be determined based on factors such as grid area and population density. The formula is:
[0113]
[0114] where K is the number of geographical grids mapped to the same position, ω i is the weight of the i-th grid, ensuring that the matrix value can comprehensively reflect the literary creation density characteristics of the region.
[0115] 3) For matrix optimization and storage:
[0116] Perform sparsification on the filled three-dimensional matrix. For the large number of elements with value 0 in the matrix (representing areas without literary creation records), adopt a sparse matrix storage format such as the compressed sparse row (CSR) format to reduce storage space and improve data processing efficiency.
[0117] Store the optimized three-dimensional matrix in a distributed file system (such as HDFS) or a relational database (such as the array type of PostgreSQL), facilitating subsequent modules to read and analyze on demand, and at the same time supporting data updates and expansions to adapt to the dynamic changes of newly added literary work data or historical geographical data.
[0118] In an alternative embodiment, the causal association module 16 includes:
[0119] A data extraction module 161 for extracting war information, policy information, and economic indicator data that overlap with the spatio-temporal range of the negative space candidate region from each historical event.
[0120] Specifically, the data extraction module filters out key information such as wars, policies, and economies that are spatio-temporally correlated with the negative space candidate region from a large amount of historical event data based on spatio-temporal matching and semantic analysis techniques. Through precise spatio-temporal range comparison and keyword retrieval, it ensures that the extracted data is highly relevant to the negative space candidate region in terms of time and space, providing accurate and comprehensive data support for subsequent causal hypothesis generation.
[0121] 1) For historical event database integration and preprocessing:
[0122] Integrate multi-source historical event data, including historical literature databases (such as the Biographical Database of Chinese Historical Figures), archaeological excavation reports, academic research results, and economic statistical yearbooks, etc., to construct a unified historical event database. Clean, deduplicate, and format the data to ensure the accuracy and consistency of the data.
[0123] Annotate detailed spatio-temporal information for each historical event, including the time range (start year and end year) when the event occurred, the geographical area involved (latitude and longitude coordinates, administrative division names, etc.), and the event type (war, policy change, economic boom / recession, etc.). At the same time, extract the key descriptive information of the event, such as the event name, main participants, event results, etc., for subsequent analysis and matching.
[0124] 2) For the spatio-temporal range matching algorithm:
[0125] For each negative space candidate region, obtain its spatio-temporal range (time interval and geographical boundary). In the historical event database, use spatio-temporal query statements to filter out historical events that overlap or are closely adjacent to the spatio-temporal range of the negative space candidate region. The specific algorithm is as follows:
[0126] Time matching: Calculate the overlap degree between the time range of the historical event and the time interval of the negative space candidate region, measured by the time overlap coefficient. The formula is:
[0127]
[0128] Set a time overlap coefficient threshold (such as 0.3) to filter out temporally relevant historical events.
[0129] Spatial matching: Conduct a spatial intersection analysis between the geographical region of historical events and the geographical boundaries of negative space candidate areas, calculate the proportion of the overlapping area to the area of negative space, and set an area proportion threshold (such as 0.1) to ensure that the extracted events are significantly associated with negative space spatially.
[0130] 3) For information extraction and structured storage:
[0131] Extract war information, policy information, and economic indicator data from the selected historical events. War information includes the name of the war, the warring parties, the nature of the war, the intensity of the war (such as the scale of troops involved, the number of battles), etc.; policy information covers the name of the policy, the policy maker, the policy goal, the scope of policy implementation, etc.; economic indicator data involves the scale of population migration, tax changes, agricultural production indicators, handicraft industry development index, etc.
[0132] Store the extracted information in a structured format. For example, construct a data table with fields including event ID, negative space ID, details of war information, details of policy information, economic indicator values, etc., to ensure the standardization and usability of the data, facilitating the calling and analysis by the subsequent causal hypothesis generation module.
[0133] The causal hypothesis generation module 162 is used to generate causal hypotheses based on war information, policy information, and economic indicator data.
[0134] Specifically, the causal hypothesis generation module integrates interdisciplinary knowledge and statistical causal inference techniques, and constructs a reasonable causal logic chain based on the extracted historical event information. By combining multidisciplinary knowledge such as the laws of literary creation, cultural communication theory, and population migration models, analyze the potential impact mechanism of historical events on literary creation, and generate explanatory causal hypotheses. At the same time, use statistical models to quantify the credibility of the hypotheses to ensure the scientificity and reliability of the hypotheses.
[0135] 1) For interdisciplinary knowledge application and causal logic construction:
[0136] Apply the laws of literary creation, considering that literary creation is affected by factors such as social environment, cultural atmosphere, and author experience. For example, war may lead to cultural rupture and the flight of literati, thus inhibiting literary creation; policy support may promote cultural prosperity, attract the gathering of literati, and stimulate the creative enthusiasm.
[0137] Combine cultural communication theory to analyze the promoting role of economic activities such as population migration and commercial trade in literary dissemination and innovation. For example, large-scale population migrations may integrate literary styles and themes from different regions, giving birth to new literary genres; the rise and fall of transportation arteries may affect the dissemination scope and audience groups of literary works.
[0138] Based on historical research results, understand the profound transformation of historical events on social structure, and then affect the evolution of literary themes and styles. For example, the change of dynasties may trigger literary expressions of the legitimacy of the new regime.
[0139] 2) For the causal hypothesis generation algorithm and model:
[0140] Based on the extracted historical event information, construct a causal graph model, where the nodes represent historical events, the negative space literary blank phenomenon, and related influencing factors, and the edges represent potential causal relationships. Through a combination of expert knowledge and data-driven methods, determine the connection relationships and weights between nodes.
[0141] Use Bayesian networks for causal reasoning, calculate the probability of low literary density in the negative space under the condition of given historical events, that is, the confidence of the causal hypothesis. At the same time, introduce a structural equation model to analyze the comprehensive influence path and degree of multiple historical event factors on literary creation, and generate a quantitative description of causal relationships.
[0142] 3) For hypothesis verification and optimization:
[0143] Compare and verify the generated causal hypothesis with historical literature records and literary research conclusions to check the rationality and accuracy of the hypothesis. For hypotheses that conflict with existing research results, re-examine the historical event data and the causal logic construction process to find possible data errors or improper knowledge applications.
[0144] According to the verification results, optimize the causal hypothesis, adjust the weights of historical event factors, supplement missing influencing factors, or correct causal logic relationships until the hypothesis has high credibility and explanatory power. The finally generated causal hypothesis is described in clear language, including the specific mechanism of how historical events affect literary creation, the expected trend of literary density changes, etc., providing users with in-depth and accurate causal interpretations.
[0145] In an alternative embodiment, the causal association module 16 further includes:
[0146] A time overlap calculation unit 163 for calculating the overlap ratio between the time period of a historical event and the time period of a negative space candidate area.
[0147] Specifically, the time overlap calculation unit, based on time series analysis technology, accurately quantifies the overlap degree between a historical event and a negative space candidate area in the time dimension. By calculating the ratio relationship between the intersection of their time intervals and the time interval of the negative space, determine the influence coverage range of the historical event on the negative space in time, providing a quantitative basis in the time dimension for subsequent confidence evaluation.
[0148] 1) For time data structured processing:
[0149] For the negative space candidate region, extract its clear time period, expressed as [negative space start time, negative space end time].
[0150] For each historical event, extract the time period when it occurred, expressed as [event start time, event end time].
[0151] 2) For the calculation of the time overlap ratio:
[0152] Calculate the intersection time period between the time period of the historical event and the time period of the negative space candidate region. If the intersection time period is empty, the overlap ratio is 0.
[0153] Calculate the time overlap ratio. The formula is:
[0154]
[0155] This ratio reflects the degree of fit between the historical event and the negative space in terms of time. The higher the ratio, the more likely it is that the historical event has an impact on the literary creation of the negative space in terms of time.
[0156] The spatial coverage evaluation unit 164 is used to calculate the grid ratio of the negative space candidate regions covered by the historical event based on the influence range of the historical event.
[0157] Specifically, the spatial coverage evaluation unit uses the spatial analysis technology of Geographic Information System (GIS) to accurately calculate the overlapping degree between the influence range of the historical event and the negative space candidate region in the geographical space. By quantifying the grid ratio of the negative space candidate regions covered by the historical event, it evaluates the influence of the historical event on the negative space in terms of space, providing quantitative support in the spatial dimension for the comprehensive confidence evaluation.
[0158] 1) For the vectorization processing of geographical data:
[0159] Convert the geographical boundaries of the negative space candidate region and the influence range of the historical event into vector data formats to ensure the accuracy and compatibility of the data.
[0160] 2) For the spatial intersection analysis and grid ratio calculation:
[0161] Use the spatial intersection tool of GIS software to calculate the intersection area between the influence range of the historical event and the negative space candidate region.
[0162] Perform an overlay analysis of the intersection area with the negative space candidate region, and count the proportion of the intersection area in the total area of the negative space candidate region, that is, the spatial coverage ratio. The formula is:
[0163]
[0164] This ratio intuitively reflects the degree of coverage of historical events in space on negative space. The higher the ratio, the more likely it is that historical events have an impact on negative space literary creation in space.
[0165] The confidence comprehensive evaluation unit 165 is used to perform weighted summation on the overlap ratio and the grid ratio to generate a confidence score.
[0166] Specifically, the confidence comprehensive evaluation unit integrates two key indicators, the time overlap ratio and the space coverage ratio, and comprehensively quantifies the association strength between historical events and the negative space candidate area through weighted summation to generate the final confidence score. This score reflects the credibility of the causal hypothesis. The higher the score, the more likely it is that historical events are potential causes leading to the negative space literary blank.
[0167] 1) For weight setting:
[0168] The weight of the time overlap ratio can be set as δ1, and the weight of the space coverage ratio can be set as δ2 according to the relative importance of the impact of time factors and space factors on literary creation. Time factors usually have a more direct and significant impact on literary creation, so a higher weight is given.
[0169] 2) For calculating the confidence score by weighted summation:
[0170] For each historical event, multiply its time overlap ratio and space coverage ratio by the corresponding weights respectively and then add them to obtain the confidence score. The formula is:
[0171] Confidence score = δ1 × time overlap ratio + δ2 × space coverage ratio;
[0172] This score ranges between [0, 1]. The higher the score, the higher the degree of association between historical events and the negative space candidate area, and the more credible the causal hypothesis.
[0173] 3) For setting the score threshold and result screening:
[0174] Set the confidence score threshold, such as 0.6, and screen out historical events with scores higher than the threshold as strongly associated events.
[0175] Sort the screened strongly associated events in descending order of confidence score to form the final list of candidate historical events. The list content includes information such as historical event details, time overlap ratio, space coverage ratio, and confidence score, providing users with clear and accurate causal association analysis results to assist in deeply exploring the causes of negative space literary blanks.
[0176] In an optional embodiment, the visualization display module 17 includes:
[0177] The heat map generation module 171 is used to generate a literary density heat map on the map interface based on the literary density heat value matrix.
[0178] Specifically, the heat map generation module converts the data in the literary density heat value matrix into a visually patterned color gradient based on geographic information system (GIS) visualization technology and heat map rendering algorithms. By intuitively reflecting the level of literary creation density through the depth of color, it accurately presents the spatial distribution of literary creation and provides a visual basis for users to grasp the overall picture of literary geography.
[0179] 1) For map interface construction and data loading:
[0180] You can choose a professional GIS software platform (such as ArcGIS API for JavaScript or Leaflet) to build an interactive map interface, import basic geographical base maps (such as modern administrative divisions, topography) and historical geographical data layers to ensure the accuracy and richness of the map.
[0181] Convert the data of the literary density heat value matrix output by the heat matrix construction module into the heat map data format supported by the GIS software (such as GeoJSON) to ensure that the data can be correctly rendered on the map.
[0182] 2) For heat map rendering parameter configuration:
[0183] Set the color gradient scheme for the heat map, using a color gradient system from light blue (low density) to dark red (high density) to ensure smooth color transitions and good visual differentiation.
[0184] Adjust the radius parameter of the heat map to dynamically control the size of the heat map spots according to the map zoom level to ensure the display effect of the heat map at different scales.
[0185] Set the transparency of the heat map so that the underlying geographical information is still visible, enhancing the readability of the map.
[0186] 3) For heat map generation and rendering:
[0187] Enable the heat map layer on the map interface and load the data of the literary density heat value matrix into the heat map layer. The literary density heat value of each geographical grid corresponds to a color point on the heat map, and the depth of color is determined according to the size of the heat value.
[0188] Use WebGL technology to accelerate heat map rendering to ensure the rapid visualization of large-scale data sets and enhance the user interaction experience.
[0189] The negative space overlay module 172 is used to overlay and label the negative space candidate areas on the literary density heat map.
[0190] Specifically, the negative space overlay module accurately overlays the negative space candidate areas on the literary density heat map based on the layer overlay technology and spatial annotation method of GIS. Through special visual symbols and color identifications, the negative space stands out in the heat map, guiding users to focus on these scarce areas of literary creation and providing visual clues for subsequent correlation analysis of historical events.
[0191] 1) For negative space data processing and symbolization:
[0192] Convert the negative space candidate area data output by the negative space determination module into a vector data format supported by GIS software (such as Shapefile).
[0193] Perform symbolization settings on the negative space data, using filling patterns (such as diagonal filling, dot matrix filling) or colors (such as low-saturation gray systems) that contrast sharply with the colors of the heat map, and set appropriate border colors and widths to ensure that the negative space is clearly distinguishable on the heat map.
[0194] 2) For layer overlay and interaction control:
[0195] In the map interface, overlay the negative space layer on top of the literary density heat map layer. Through the layer control tool, allow users to freely switch the display and hiding of the negative space layer, facilitating users' comparative analysis.
[0196] Add interaction events to the negative space, such as popping up an information box when the mouse hovers, displaying detailed information about the negative space, including the time-space range, literary density heat value, list of associated historical events, etc.
[0197] The dynamic interaction module 173 is used to generate an interactive sliding timeline. Based on the sliding operation of the timeline, dynamically display the positions and description information of candidate historical events on the literary density heat map to obtain a dynamic visualization map.
[0198] Specifically, the dynamic interaction module comprehensively uses timeline interaction technology and a dynamic data update mechanism based on event listening to enable users to dynamically display the literary density heat map and historical event information in different periods through the sliding operation of the timeline. Through front-end and back-end data interaction and real-time rendering, this module ensures that users can smoothly explore the spatio-temporal evolution and historical background of literary creation.
[0199] 1) For timeline component design and implementation:
[0200] Use a front-end framework (such as Vue.js or React) to design an interactive sliding timeline component. The range of the timeline covers the entire research period, marking key historical periods and nodes.
[0201] The timeline slider supports drag operations and triggers an event listener function when sliding, sending a time change request to the backend.
[0202] 2) For the coordination of backend data processing and frontend rendering:
[0203] The backend receives the position information of the timeline slider and filters the data for the corresponding period from the literary density heat value matrix and the historical event database based on the selected time, performing necessary data processing and format conversion.
[0204] The frontend receives the data returned by the backend and dynamically updates the heat map and historical event annotations. For the heat map, re-render the literary density data for the corresponding time; for historical events, clear the historical event annotations for the previous time, filter out the historical events associated with the negative space candidate area according to the new time, regenerate the annotations on the map, and display brief description information near the event location.
[0205] 3) For the display and interaction of historical event details:
[0206] When the user clicks on a historical event annotation, a detailed information box pops up, showing the complete description of the historical event, including the event name, occurrence time, scope involved, event type, potential impact on literary creation, etc.
[0207] Support the user to quickly locate specific historical events or negative space areas through auxiliary tools such as a search box, enhancing the flexibility and convenience of interaction.
[0208] The above-mentioned literary work geographical marking presentation system based on the integration of literature, history, and geography extracts and structures geographical entities and spatio-temporal tags in literary works through multi-source heterogeneous data fusion and natural language processing technologies, constructs a literary density heat value matrix in the spatio-temporal dimension to quantify geographical distribution characteristics, combines dynamic thresholds to screen negative space candidate areas with literary density lower than the benchmark and active historical data, generates high-confidence causal hypotheses through historical event database association and spatio-temporal coverage weighted calculation, and finally generates an interactive dynamic visualization map that combines a heat map and a timeline comparison, realizing the automated detection, cause reasoning, and visualization verification of literary geography "negative space", significantly improving the depth and efficiency of literary geography research, breaking through the dependence on explicit information in texts, and providing intelligent tool support for interdisciplinary exploration of the historical and cultural values of literary blank areas.
[0209] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0210] Based on the same inventive concept, an embodiment of the present application also provides a method applied to the above-mentioned literary work geographical marking presentation system based on the integration of literature, history and geography. The implementation solution provided by this method to solve the problem is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the literary work geographical marking presentation method based on the integration of literature, history and geography provided below can refer to the limitations on a literary work geographical marking presentation system based on the integration of literature, history and geography in the above text, and will not be repeated here.
[0211] In an exemplary embodiment, as Figure 2 shown, a method for presenting geographical marks of literary works based on the integration of literature, history and geography is provided, including:
[0212] S1: The data acquisition module acquires literary works, historical geographical data, historical time period length, and historical population density data.
[0213] S2: The geographical entity processing module extracts geographical entities in the literary works through a natural language processing model, annotates spatio-temporal labels for the geographical entities, and generates a structured spatio-temporal label data set.
[0214] S3: The grid analysis module generates multiple geographical grids based on the historical geographical data; based on the structured spatio-temporal label data set, it counts the number of mentions of geographical entities in each geographical grid in the literary works.
[0215] S4: The heat matrix construction module generates a literary density heat value matrix based on the spatio-temporal dimension according to the number of mentions, the historical time period length, and the historical population density data.
[0216] S5: The negative space determination module, based on the literary density heat value matrix, screens geographical grids with a literary density lower than a preset literary density threshold and a historical population greater than zero as negative space candidate areas.
[0217] S6: The causal association module associates the negative space candidate regions with each historical event, generates causal hypotheses, and calculates the confidence levels of the causal hypotheses; filters out the historical events that generate causal hypotheses with confidence levels greater than a preset threshold as candidate historical events.
[0218] S7: The visualization display module dynamically superimposes the negative space candidate regions and the candidate historical events on the map interface to generate a dynamic visualization map.
[0219] An embodiment of the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.
[0220] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0221] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0222] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the application embodiments. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. A literary work geographical marking presentation system based on the integration of literature, history and geography, characterized in that, The system includes: a data acquisition module, a geographical entity processing module, a grid analysis module, a heat matrix construction module, a negative space determination module, a causal association module, and a visualization display module; The data acquisition module is used to obtain literary works, historical geographical data, historical time period length, and historical population density data; The geographical entity processing module is used to extract geographical entities in the literary works through a natural language processing model, annotate spatio-temporal tags to the geographical entities, and generate a structured spatio-temporal tag data set; The grid analysis module is used to generate a plurality of geographical grids based on the historical geographical data; and based on the structured spatio-temporal tag data set, count the number of mentions of the geographical entities in each of the geographical grids in the literary works; The heat matrix construction module is used to generate a literary density heat value matrix based on spatio-temporal dimensions according to the number of mentions, the historical time period length, and the historical population density data; The negative space determination module is used to, based on the literary density heat value matrix, screen the geographical grids with a literary density lower than a preset literary density threshold and a historical population greater than zero as negative space candidate areas; The causal association module is used to associate the negative space candidate areas with each historical event to generate causal hypotheses, and calculate the confidence levels of the causal hypotheses; screen the historical events that generate causal hypotheses with confidence levels greater than a preset threshold as candidate historical events; The visualization display module is used to dynamically superimpose the negative space candidate areas and the candidate historical events on a map interface to generate a dynamic visualization map.
2. The system according to claim 1, characterized in that, The heat matrix construction module includes: A normalization processing unit, which is used to perform normalization calculation on the number of mentions based on the historical time period length and the historical population density to generate a normalized literary density value; A three-dimensional matrix construction unit, which is used to map the normalized literary density value into a three-dimensional matrix of time dimension, longitude dimension, and latitude dimension to obtain the literary density heat value matrix.
3. The system according to claim 1, characterized in that, The causal association module includes: A data extraction module, which is used to extract war information, policy information, and economic indicator data that overlap with the spatio-temporal range of the negative space candidate areas from each of the historical events; A causal hypothesis generation module, which is used to generate the causal hypotheses based on the war information, the policy information, and the economic indicator data.
4. The system according to claim 1, wherein The causal association module further includes: A time overlap calculation unit, which is used to calculate the overlap ratio of the time period of the historical event and the time period of the negative space candidate area; A space coverage evaluation unit, which is used to calculate the grid ratio of the negative space candidate area covered by the historical event based on the influence range of the historical event; A confidence level comprehensive evaluation unit, which is used to perform weighted summation on the overlap ratio and the grid ratio to generate the confidence level score.
5. The system according to any one of claims 1 to 4, characterized in that, The visualization display module includes: A heat map generation module, which is used to generate a literary density heat map on the map interface based on the literary density heat value matrix; A negative space superimposition module, which is used to superimpose and label the negative space candidate areas on the literary density heat map; A dynamic interaction module for generating an interactive and slidable timeline, and dynamically displaying the positions and description information of the candidate historical events on the literary density heat map based on the sliding operation of the timeline, thereby obtaining the dynamic visualization map.
6. A method for presenting geographical markers of literary works based on the integration of literature, history, and geography, applied to the system described in any one of claims 1 to 5, characterized in that, The method includes: S1: A data acquisition module acquires literary works, historical geographical data, historical time period length, and historical population density data; S2: A geographical entity processing module extracts geographical entities from the literary works through a natural language processing model, annotates spatio-temporal tags to the geographical entities, and generates a structured spatio-temporal tag data set; S3: A grid analysis module generates a plurality of geographical grids based on the historical geographical data; and based on the structured spatio-temporal tag data set, counts the number of mentions of the geographical entities within each geographical grid in the literary works; S4: A heat matrix construction module generates a literary density heat value matrix based on spatio-temporal dimensions according to the number of mentions, the historical time period length, and the historical population density data; S5: A negative space determination module filters out the geographical grids with literary density lower than a preset literary density threshold and historical population greater than zero as negative space candidate areas based on the literary density heat value matrix; S6: A causal association module associates the negative space candidate areas with each historical event, generates causal hypotheses, and calculates the confidence levels of the causal hypotheses; and filters out the historical events for which the generated causal hypotheses have confidence levels greater than a preset threshold as candidate historical events; S7: A visualization display module dynamically superimposes the negative space candidate areas and the candidate historical events on a map interface to generate a dynamic visualization map.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method described in claim 6 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in claim 6 is implemented.