Chinese language and literature computer networking query reading system
Through mixed feature analysis, distributed retrieval and knowledge graph visualization technologies, the problem of insufficient semantic correlation analysis of different characters and traditional Chinese characters in the existing technology is solved, and efficient query and visual display of knowledge of Chinese literary works is realized, improving the accuracy and user experience of query results.
Patent Information
- Application Number
- CN202510472432.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Chinese language and literature computer network query and reading technology lacks semantic correlation analysis of different characters and traditional Chinese characters. The results show that there is a lack of structure of knowledge networks and cannot meet users' needs for knowledge systematization and visual reading.
The query request processing module performs mixed feature analysis to generate mixed query vectors that combine radicals and semantics; the distributed search module searches in the space-time shard index structure, and the preliminary search optimization module performs adaptive conversion of heterogeneous characters and traditional and simple characters; the visual query output module generates a knowledge graph based on the space-time relationship of literary entities.
It improves the accuracy and comprehensiveness of query results, enhances users' experience of systematic and visual reading of Chinese language and literature knowledge, and realizes efficient query and visual display of Chinese language and literature works.
Smart Images

Figure CN120386860A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a computer network query and reading system for Chinese language and literature. Background Art
[0002] With the development of computer technology and natural language processing technology, there has emerged a computer network query and reading technology for Chinese language and literature. This technology integrates a vast amount of literary resources through digital means, realizes a text retrieval function based on keyword matching, and can provide users with basic literary work query services. However, the processing of users' query requests mainly relies on single-modal input such as pure text keywords, and realizes information matching by constructing an inverted index or a full-text retrieval engine. The processing of Chinese characters only stays at the level of character encoding conversion, lacking semantic association analysis of variant Chinese characters and simplified and traditional Chinese characters. Moreover, in terms of result display, it usually presents fragmented text in a list form, lacking a structured display of the spatio-temporal relationships and knowledge networks among literary entities, and unable to meet users' needs for systematic and visual reading of knowledge systems. Summary of the Invention
[0003] Based on this, it is necessary to provide a computer network query and reading system for Chinese language and literature to address the above technical problems, so as to improve the accuracy, comprehensiveness, and efficiency of Chinese language and literature queries, and enhance users' experience of systematic and visual reading of Chinese language and literature knowledge systems.
[0004] In a first aspect, the present application provides a computer network query and reading system for Chinese language and literature, including:
[0005] A query request processing module, configured to perform hybrid feature parsing processing on a user's query request to generate a hybrid query vector integrating radicals and semantics, where the query request includes text and / or pictures;
[0006] A distributed retrieval module, configured to perform distributed retrieval in a preset spatio-temporal sharding index structure based on the hybrid query vector to obtain a candidate literary dataset, where the preset spatio-temporal sharding index structure includes a time dimension, a genre dimension, and an author influence dimension;
[0007] A preliminary retrieval optimization module, configured to perform adaptive conversion processing on variant Chinese characters and simplified and traditional Chinese characters in the candidate literary dataset according to Chinese character evolution rules and context to obtain standardized text data;
[0008] A visual query output module, configured to perform knowledge graph enhancement processing on the standardized text data according to a spatio-temporal relationship model of literary entities to generate a visual query result.
[0009] In one embodiment, the query request processing module includes:
[0010] An image processing sub-unit, configured to:
[0011] When the query request includes an image, use an adversarial generative network to correct the distortion of the handwritten font of the image to obtain an intermediate image with normalized glyphs;
[0012] When the intermediate image is a scanned ancient book, use a bidirectional LSTM-CRF model to reconstruct the layout of the vertical text in the intermediate image, output the horizontal text, and combine the preset prior knowledge base of Chinese character structures to perform stroke topology verification on the horizontal text to obtain the standardized OCR text;
[0013] A text processing sub-unit, configured to:
[0014] When the query request includes text, perform semantic parsing on the text to obtain an analysis structure including a character-level radical decomposition tree and a word-level dependency relationship graph;
[0015] Based on the context semantics in the analysis structure, use a dynamic weight allocation algorithm to adjust the radical contribution degree of each character in the text, generate a radical feature vector, and input the analysis structure into the BERT-wwm pre-training model for semantic encoding to output a semantic vector;
[0016] A hybrid query vector generation sub-unit, configured to:
[0017] Input the standardized OCR text and the radical feature vector into an adversarial feature alignment network, output a cross-modal intermediate vector, and perform topological structure encoding on the cross-modal intermediate vector based on a preset radical relationship graph to generate a feature matrix including the glyph spatial relationship;
[0018] Fuse the feature matrix with the semantic vector to obtain a hybrid query vector.
[0019] In one embodiment, the distributed retrieval module includes:
[0020] A time dimension sub-unit, configured to perform dynamic sharding processing on a preset literary dataset according to the time dimension, and use a sliding time window algorithm to segment the cross-dynasty text in the preset literary dataset to generate a time shard set. Among them, for the text with time label conflicts in the preset literary dataset, generate overlapping shard identifiers through a Gaussian mixture model;
[0021] A genre dimension sub-unit, configured to build a genre weight calculation model based on the genre dimension, dynamically adjust the retrieval weights of poems, prose, and novels through the genre feature intensity in the hybrid query vector, and generate a weighted genre shard index;
[0022] An author influence dimension sub-unit, configured to:
[0023] Establish a dynamic authority evaluation system according to the author influence dimension, construct a citation network through the literature citation relationship between authors, and calculate the authority value of author nodes in the citation network based on the PageRank algorithm;
[0024] Conduct sentiment analysis on historical review data to obtain the sentiment analysis results, and calculate the spatio-temporal decay factor in combination with the text publication time;
[0025] Multiply the authority value by the spatio-temporal decay factor to generate the author influence index;
[0026] The sharding retrieval subunit is used to input the time sharding set, genre sharding index, and author influence index into the cross-dimensional joint retrieval engine, construct a three-dimensional sharding space composed of the time dimension, genre dimension, and author influence dimension based on the preset spatio-temporal sharding index structure, calculate the similarity of the mixed query vector through the three-dimensional sharding space, and generate a candidate literature dataset.
[0027] In one embodiment, the authority value is calculated by the following formula:
[0028]
[0029] where a i is the authority value of author node i, μ is the decay factor, τ is the time decay coefficient, Cite(i) is the set of literatures citing author i, t i and t j are the eras in which authors i and j are located respectively, and |Ref(j)| is the number of times the literature of author j is cited.
[0030] In one embodiment, the preliminary retrieval optimization module includes:
[0031] The variant character mapping library construction subunit is used to construct a cross-dynasty variant character mapping library and establish a variant character relationship graph using the evolved glyph similarity algorithm;
[0032] The mode selection and semantic analysis subunit is used to select the conversion mode according to the creation era of the candidate literature dataset and perform semantic analysis through a pre-trained context semantic model to generate the current semantic context vector. Among them, the traditional Chinese benchmark conversion mode is enabled for literature before the Tang Dynasty, and the simplified Chinese enhancement mode is enabled for modern texts;
[0033] The genetic optimization path search subunit is used to perform optimal replacement path search using the genetic algorithm based on the variant character relationship graph and the current semantic context vector to generate a preliminary standardized text;
[0034] The phono-semantic substitution conversion suppression and font structure verification subunit is used for:
[0035] Analyze the semantic context of the preliminary standardized text through the Bi-GRU model, suppress the conversion of interchangeable characters and ancient and modern characters, and generate an intermediate standardized text;
[0036] Use a generative adversarial network to construct a Chinese character topological structure validator, perform stroke-level structure matching on the intermediate standardized text, and when the detected stroke topology deviation exceeds the 5% threshold, repair it based on a preset character source database to generate standardized text data.
[0037] In one embodiment, the visual query output module includes:
[0038] A literary entity extraction and analysis subunit, which is used for:
[0039] Through the literary entity spatio-temporal relationship model, use the multi-head spatio-temporal attention mechanism to extract literary entities from the standardized text data, and generate three-dimensional spatio-temporal coordinates including dynasty numbers, geographical codes, and timestamps;
[0040] Based on the citation frequency of literary entities in historical documents and the modern social media dissemination heat, calculate the spatio-temporal dual-domain influence value through the LSTM time series network, where the spatio-temporal dual-domain influence value includes historical authority weight and modern dissemination weight;
[0041] A knowledge graph enhancer subunit, which is used for:
[0042] Construct a multi-scale knowledge graph. Among them, the original text relationship between literary entities is retained in the micro layer of the multi-scale knowledge graph, and in the macro layer of the multi-scale knowledge graph, a hierarchical graph convolutional network is used to perform spatio-temporal clustering on the three-dimensional spatio-temporal coordinates and spatio-temporal dual-domain influence values to generate literary genre clustering labels and regional cultural diffusion paths;
[0043] Map the multi-scale knowledge graph to the coordinate system corresponding to the three-dimensional spatio-temporal coordinates, and generate a three-dimensional heat cloud map through the WebGL engine. The cultural dissemination heat cloud map is used to interact with users through a visual graphical interface;
[0044] A target data generation subunit, which is used for:
[0045] Obtain the user's selection operations on the time range, geographical area, and literary genre clustering labels through the visual graphical interface, and convert the selection operations into spatio-temporal slice parameters;
[0046] Extract target data from the three-dimensional heat cloud map based on the spatio-temporal slice parameters to generate a visual query result.
[0047] In one embodiment, the target data generation subunit includes:
[0048] A semantic verification subunit, configured to perform semantic verification on spatio-temporal slice parameters. If the time range in the spatio-temporal slice parameters exceeds the dynasty interval in which the literary entity exists, it generates a warning signal containing an out-of-bounds timestamp, and the warning signal is used to prompt the user to make corrections;
[0049] A hierarchical sampling output subunit, configured to:
[0050] Determine the accuracy level according to the data density distribution of the three-dimensional thermal cloud map, where the accuracy level includes a high-density level and a low-density level;
[0051] When it is detected that the target data volume belongs to the high-density level, a simplified view is generated using an aggregation algorithm based on K-means;
[0052] When it is detected that the target data volume belongs to the low-density level, the original data view is output;
[0053] Align the simplified view or the original data view with the coordinate system of the three-dimensional thermal cloud map to generate a visual query result.
[0054] In a second aspect, the present application also provides a computer networking query and reading method for Chinese language and literature, including:
[0055] Perform hybrid feature parsing processing according to the user's query request to generate a hybrid query vector that combines radicals and semantics, where the query request includes text and / or pictures;
[0056] Based on the hybrid query vector, perform distributed retrieval in a preset spatio-temporal shard index structure to obtain a candidate literary dataset, where the preset spatio-temporal shard index structure includes a time dimension, a genre dimension, and an author influence dimension;
[0057] According to the Chinese character evolution rules and context, perform adaptive conversion processing on the variant characters and complex and simplified characters in the candidate literary dataset to obtain standardized text data;
[0058] According to the spatio-temporal relationship model of literary entities, perform knowledge graph enhancement processing on the standardized text data to generate a visual query result.
[0059] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and the steps in the first aspect are implemented when the processor executes;
[0060] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and the steps in the first aspect are implemented when the computer program is executed by the processor.
[0061] The above-mentioned Chinese language and literature computer networking query and reading system realizes the efficient query and knowledge visualization display of Chinese language and literature works through the coordinated operation of multiple modules. Among them, the query request processing module performs hybrid feature parsing and processing on the user's query request containing text and / or pictures, can effectively integrate different modal input information, generate a hybrid query vector integrating radicals and semantics, provides a richer query basis for subsequent retrieval, and improves the relevance and comprehensiveness of query results. The distributed retrieval module can perform distributed retrieval in a preset spatio-temporal shard index structure including time dimension, genre dimension, and author influence dimension based on the hybrid query vector, and obtain a candidate literature dataset. Compared with the traditional single-dimensional retrieval method, this module can more flexibly and comprehensively meet the user's query requirements for literature works in different dimensions, improves the retrieval efficiency and the accuracy of results, and reduces the missed detection situation. The preliminary retrieval optimization module can perform adaptive conversion processing on the variant characters and simplified and traditional Chinese characters in the candidate literature dataset according to the Chinese character evolution rules and context, can effectively eliminate the semantic ambiguity problem caused by the evolution of Chinese character glyphs, obtain standardized text data, and ensure the consistency and accuracy of the query results in terms of text expression. The visualization query output module can perform knowledge graph enhancement processing on the standardized text data according to the spatio-temporal relationship model of literary entities, generate visualization query results, further improve the operability and information depth of the query results, enable users to quickly master the relevance of literary works, the relationships between authors, and the development trajectories of works in the time and space dimensions through the graph, and enhance the user's experience of systematic and visual reading of Chinese language and literature knowledge.
[0062] Compared with the traditional Chinese language and literature query and reading methods, this system realizes the efficient query and in-depth reading support for Chinese language and literature works based on multi-modal input processing, multi-dimensional index retrieval, intelligent text processing, and knowledge graph visualization technology. Among them, the comprehensive multi-modal input processing provides richer information for the query, improves the adaptability to different input forms, and the multi-dimensional index retrieval and intelligent text processing further enhance the accuracy and comprehensiveness of the query results and improve the retrieval efficiency of complex literary resources. The knowledge graph visualization display provides a more intuitive and personalized reading experience for users, effectively improves the efficiency and quality of Chinese language and literature query and reading, and provides strong support for users to deeply understand the Chinese language and literature knowledge system. Brief Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 Schematic diagram of the structure of a Chinese language and literature computer network query and reading system provided for an exemplary embodiment of the present invention;
[0065] Figure 2 Flowchart of a Chinese language and literature computer network query and reading method provided for an exemplary embodiment of the present invention. Detailed implementation manners
[0066] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0067] In one embodiment, as Figure 1 shown, a Chinese language and literature computer network query and reading system 100 is provided. In this embodiment, it is exemplified that the system is applied to a terminal. It can be understood that the system can also be applied to a server, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the system includes:
[0068] A query request processing module 101, configured to perform hybrid feature parsing processing according to a user's query request to generate a hybrid query vector that combines radicals and semantics, where the query request includes text and / or pictures.
[0069] Specifically, for text input, this module can extract the radical information and semantic features of the text through a combination of radical parsing and semantic analysis to generate a hybrid query vector. Among them, radical parsing can identify the structural features of Chinese characters, while semantic analysis can extract the semantic information of the text through natural language processing technology. For picture input, this module can extract the text content in the picture through image recognition technology and further perform radical and semantic parsing to generate a hybrid query vector.
[0070] A distributed retrieval module 102, configured to perform distributed retrieval in a preset spatio-temporal sharding index structure based on the hybrid query vector to obtain a candidate literature data set, where the preset spatio-temporal sharding index structure includes a time dimension, a genre dimension, and an author influence dimension.
[0071] Specifically, the time dimension can be divided by the dynasty or period in which literary works were created, the genre dimension can be divided by poetry, prose, novels, etc., and the author influence dimension can be evaluated by the historical status of the author or the citation frequency of the works. Schematically, this module can adopt a distributed computing framework such as MapReduce or Spark, and can thus quickly retrieve candidate literary data sets that match the hybrid query vector in a large data set.
[0072] The preliminary retrieval optimization module 103 is used to adaptively convert variant Chinese characters and simplified and traditional Chinese characters in the candidate literary data set according to Chinese character evolution rules and context, to obtain standardized text data.
[0073] Specifically, since there are a large number of variant Chinese characters and simplified and traditional Chinese characters in Chinese language and literature, this module can standardize the characters in the candidate literary data set according to Chinese character evolution rules such as simplified and traditional Chinese character conversion and variant Chinese character recognition. For example, convert traditional Chinese characters to simplified Chinese characters, or unify variant Chinese characters into standard glyphs. And through context analysis, the candidate literary data set can be further optimized to generate standardized text data. For example, the specific meaning of a polysemous character can be judged according to the context to ensure the accuracy of the retrieval results.
[0074] The visual query output module 104 is used to perform knowledge graph enhancement processing on the standardized text data according to the spatio-temporal relationship model of literary entities, to generate visual query results.
[0075] Specifically, a knowledge graph is a graphical knowledge structure that can clearly present the relationships between different literary entities such as authors, works, events, genres, etc. Through this knowledge graph, this module can intuitively display the associations between different literary data in the standardized text data according to the spatio-temporal relationships of literary entities, which helps users better understand the interactions between literary works and characters. Schematically, when a user queries the works of a certain writer, the system not only displays all the works of this writer, but can also display information such as the historical background, literary genre, and creation period related to these works, further enhancing the user's reading experience effect.
[0076] The above system mainly includes four modules. In the query request processing module 101, by performing hybrid feature parsing and processing on the user's query request, it can effectively integrate multimodal information of text and images, generate a hybrid query vector that combines radicals and semantics, providing a high-quality query basis for subsequent retrieval, and can effectively improve the accuracy and pertinence of the query. In the distributed retrieval module 102, based on the hybrid query vector, it performs distributed retrieval in the preset spatio-temporal shard index structure. Its multi-dimensional index in the time dimension, genre dimension, and author influence dimension can effectively improve the retrieval efficiency and diversity of literary data compared with traditional single-dimensional retrieval. In the preliminary retrieval optimization module 103, by applying Chinese character evolution rules and context, it adaptively converts variant characters and simplified and traditional Chinese characters in the candidate literary dataset, not only optimizing the accuracy of the query results, but also solving the retrieval obstacles caused by text changes, dialect differences, or the use of simplified and traditional Chinese characters. In the visualization query output module 104, according to the spatio-temporal relationship model of literary entities, it performs knowledge graph enhancement processing on the standardized text data to generate a visualization query result, which can more intuitively display the spatio-temporal relationship and knowledge association of literary works compared with the traditional text list-style result.
[0077] Through the collaborative work of the above four modules, the system can effectively improve the accuracy, efficiency, and visualization degree of Chinese language and literature data retrieval, providing users with a more intelligent query experience. In addition, the system greatly enhances the processing ability of complex literary data through a multi-dimensional retrieval mechanism and text standardization processing, enabling users to obtain the required knowledge more efficiently when conducting literary research or learning.
[0078] In one embodiment, the query request processing module 101 includes:
[0079] An image processing sub-unit, used for:
[0080] When the query request includes an image, use an adversarial generative network to correct the handwritten font distortion of the image to obtain an intermediate image with normalized glyphs;
[0081] When the intermediate image is a scanned ancient book, use a bidirectional LSTM-CRF model to reconstruct the layout of the vertical text in the intermediate image, output the horizontal text, and combine a preset Chinese character structure prior knowledge base to perform stroke topology verification on the horizontal text to obtain standardized OCR text;
[0082] A text processing sub-unit, used for:
[0083] When the query request includes text, perform semantic parsing on the text to obtain an analysis structure including a character-level radical decomposition tree and a word-level dependency relationship graph;
[0084] Based on the context semantics in the parsing structure, the dynamic weight allocation algorithm is used to adjust the radical contribution degrees of each character in the text, generate a radical feature vector, and input the parsing structure into the BERT-wwm pre-trained model for semantic encoding to output a semantic vector;
[0085] The hybrid query vector generation subunit is used for:
[0086] Input the normalized OCR text and the radical feature vector into the adversarial feature alignment network, output a cross-modal intermediate vector, and construct a graph based on the preset radical relationship to perform topological structure encoding on the cross-modal intermediate vector to generate a feature matrix containing the glyph spatial relationship;
[0087] Fuse the feature matrix with the semantic vector to obtain a hybrid query vector.
[0088] Specifically, the adversarial generation network consists of a generator and a discriminator. In this embodiment, the generator is responsible for attempting to generate a canonical glyph image with the distortion of the handwritten font removed, and the discriminator differentiates the generated image from the real canonical image. The two are in an adversarial relationship until the image generated by the generator can pass the detection of the discriminator, thereby obtaining an intermediate image with normalized glyphs. And when this intermediate image is recognized as an ancient book scanned copy, the LSTM part in the bidirectional LSTM-CRF model can understand the order and association relationship between characters in the vertical text. The CRF part can, based on the output of the LSTM part, utilize the context features in the text to convert the vertical text into a horizontal text that is convenient for subsequent processing. In addition, in combination with the preset prior knowledge base of Chinese character structures, stroke topology verification is performed on the horizontal text. This preset prior knowledge base of Chinese character structures stores a large amount of information such as the stroke structures and orders of standard Chinese characters. By comparing the reconstructed horizontal text with the information in the knowledge base, possible stroke errors can be corrected, and finally, normalized OCR (Optical Character Recognition) text is obtained to ensure that the text information in the image can enter the subsequent processing process in a standard and accurate form.
[0089] Specifically, in the word processing subunit, the character-level radical decomposition tree disassembles each Chinese character according to its radical structure, showing the internal composition relationship of the Chinese character, which helps to understand the text from the perspective of the glyph structure. The word-level dependency graph depicts the syntactic and semantic dependency relationships between words, such as the subject-predicate relationship and the verb-object relationship, revealing the logic of text combination from the semantic level. Based on the context semantics in the parsing structure, a dynamic weight assignment algorithm can be used to adjust the weights of the radicals of each character, and then generate a radical feature vector that better fits the context. And the parsing structure can be input into a pre-trained BERT-wwm (Bidirectional Encoder Representations from Transformers-Whole Word Masking) model for semantic encoding. This BERT-wwm model has been pre-trained on a large-scale text and can understand the semantic representation of the text. By processing the parsing structure, it outputs a semantic vector reflecting the semantics of the text.
[0090] Specifically, in the hybrid query vector generation subunit, the adversarial feature alignment network can find and establish an alignment relationship between different modalities, such as text derived from images and the feature vectors related to the semantics of the text, and output a cross-modal intermediate vector. Subsequently, based on a preset radical relationship graph, topological structure encoding can be performed on the cross-modal intermediate vector. This radical relationship graph describes information such as the spatial position and combination relationship between different radicals. Through this encoding method, a feature matrix containing the spatial relationship of the glyphs can be generated, which integrates different modality information and glyph spatial features. Finally, this feature matrix can be fused with the semantic vector output by the word processing subunit to obtain a hybrid query vector that can comprehensively reflect the intention of the user's query request. This vector provides a high-quality input basis for subsequent distributed retrieval, improving the accuracy and comprehensiveness of the system's understanding of the user's query request.
[0091] In one embodiment, the distributed retrieval module 102 includes:
[0092] The time dimension subunit is used to perform dynamic sharding processing on a preset literary dataset according to the time dimension, and use a sliding time window algorithm to segment the cross-dynasty texts in the preset literary dataset to generate a time shard set. Among them, for the texts with time label conflicts in the preset literary dataset, overlapping shard identifiers are generated through a Gaussian mixture model;
[0093] The genre dimension subunit is used to construct a genre weight calculation model based on the genre dimension, and dynamically adjust the retrieval weights of poems, prose, and novels according to the genre feature intensity in the hybrid query vector to generate a weighted genre shard index;
[0094] The author influence dimension subunit is used for:
[0095] Establish a dynamic authority evaluation system based on the dimension of author influence, construct a citation network through the literature citation relationship between authors, and calculate the authority value of author nodes in the citation network based on the PageRank algorithm;
[0096] Conduct sentiment analysis on historical review data to obtain sentiment analysis results, and calculate the spatio-temporal decay factor in combination with the text publication time;
[0097] Multiply the authority value by the spatio-temporal decay factor to generate the author influence index;
[0098] The sharding retrieval sub-unit is used to input the time shard set, genre shard index, and author influence index into the cross-dimensional joint retrieval engine, construct a three-dimensional shard space composed of the time dimension, genre dimension, and author influence dimension based on the preset spatio-temporal shard index structure, calculate the similarity of the mixed query vector through the three-dimensional shard space, and generate a candidate literary data set.
[0099] Specifically, the time dimension sub-unit can adopt a dynamic sharding processing strategy, fully considering the characteristics of literary development in different periods and the data distribution. For example, in the period when ancient literature developed relatively slowly and the number of works was relatively small, the time slices may be divided wider; while in the modern and contemporary literature prosperous stage with a sharp increase in the number of works, the time slices are divided more finely. And for literary works spanning multiple dynasties, this sub-unit can adopt a sliding time window algorithm, and reasonably slice them into parts belonging to different dynasty time slices according to the text content and time clues. In addition, during the processing, there may be cases where there are time label conflicts in some literary works, that is, there are multiple statements or uncertainties in the definition of the creation time of the works. For this situation, this sub-unit analyzes the relevant characteristics of the time label conflict text such as content style, time reference of allusions, etc., calculates the probability distribution of different time possibilities using the Gaussian mixture model, thereby determining the overlapping degree of the text on different time slices, and generating corresponding overlapping shard identifiers. Finally, this sub-unit can generate a detailed and accurate time shard set, providing the basic data division in the time dimension for subsequent cross-dimensional retrieval.
[0100] Specifically, in the genre dimension sub-unit, the genre weight calculation model takes the genre feature intensity in the mixed query vector as the key input. That is, when the user initiates a query, the mixed query vector contains clues about the genres that the user may be interested in. For example, if the user frequently mentions keywords such as rhyme and imagery in the query, the poetry genre feature intensity in the mixed query vector may be higher; if the query content focuses on plot narration and character portrayal, the novel genre feature intensity will increase accordingly.
[0101] The genre weight calculation model can dynamically adjust the weights of different genres such as poetry, prose, and novels during the retrieval process based on these genre feature intensities, and generate a weighted genre shard index. This index not only clarifies the distribution of literary works of different genres in the dataset, but also assigns different degrees of importance to various genres according to the user's query intention, providing a basis for weight allocation in the genre dimension for subsequent cross-dimensional joint retrieval, so that the retrieval results can better meet the user's preference for a specific genre.
[0102] Specifically, in the field of literature, the influence of an author is one of the important factors in measuring the value of their works. The citation network in the author influence dimension subunit reflects the academic and creative associations and inheritance relationships among authors. Based on this citation network, the PageRank algorithm can be used to calculate the authority value of each author node, that is, if an author's works are cited by a large number of other influential authors, their authority value will increase accordingly. Moreover, as new literary works are published and citation relationships change, the citation network can be updated in real time, and the author's authority value will also be dynamically adjusted accordingly to accurately reflect the changing influence of the author in the literary field.
[0103] Schematically, the authority value is calculated by the following formula:
[0104]
[0105] where a i is the authority value of author node i, μ is the attenuation factor, τ is the time decay coefficient, Cite(i) is the set of documents citing author i, t i and t j are the eras in which authors i and j are located respectively, and |Ref(j)| is the number of citations of author j's documents. In addition, this subunit can also perform sentiment analysis on historical review data, judge the sentiment tendency towards the works and authors in the reviews through natural language processing techniques and sentiment analysis algorithms, and quantify the degree of sentiment. Combining with the text publication time, considering that the direct influence of works that are older may decay in the current era, a specific mathematical model is used to calculate the spatio-temporal decay factor. This factor comprehensively reflects the change in the influence of the work over time and under different historical period review sentiments. Finally, the calculated author authority value can be multiplied by the spatio-temporal decay factor to obtain a value that comprehensively reflects the author's influence in the current retrieval scenario, and generate an author influence index. This index can sort and associate authors according to their influence sizes, providing key information in the author influence dimension for subsequent cross-dimensional retrieval, so that the works of authors with higher influence can be given priority consideration during the retrieval process, meeting the user's query needs for high-quality and highly influential literary works.
[0106] Specifically, the sharded retrieval sub-unit can input the time shard set, genre shard index, and author influence index into the cross-dimensional joint retrieval engine, and construct a three-dimensional shard space composed of the time dimension, genre dimension, and author influence dimension based on the preset spatio-temporal shard index structure. Subsequently, specific similarity measurement algorithms such as cosine similarity and Euclidean distance can be used to calculate the similarity between the hybrid query vector and the positions of various literary works in the three-dimensional space. The higher the similarity, the more the work matches the user's query intention. According to the calculation results, literary works with higher similarity are screened out to generate a candidate literary data set. This data set provides a high-quality data basis for subsequent preliminary retrieval optimization and visual query output, greatly improving the accuracy and relevance of the retrieval results and meeting the complex and diverse query needs of users in the field of Chinese language and literature.
[0107] In one embodiment, the preliminary retrieval optimization module 103 includes:
[0108] The variant character mapping library construction sub-unit is used to construct a cross-dynasty variant character mapping library and establish a variant character relationship graph using the evolved glyph similarity algorithm;
[0109] The mode selection and semantic analysis sub-unit is used to select a conversion mode according to the creation time of the candidate literary data set and perform semantic analysis through a pre-trained context semantic model to generate the current semantic context vector. Among them, the traditional character benchmark conversion mode is enabled for literature before the Tang Dynasty, and the simplified character enhancement mode is enabled for modern texts;
[0110] The genetic optimization path search sub-unit is used to perform an optimal replacement path search using a genetic algorithm based on the variant character relationship graph and the current semantic context vector to generate a preliminary standardized text;
[0111] The borrowed-character conversion suppression and font structure verification sub-unit is used for:
[0112] Analyze the semantic context of the preliminary standardized text through a Bi-GRU model, suppress the conversion of borrowed characters and ancient and modern characters, and generate an intermediate standardized text;
[0113] Use an adversarial generative network to construct a Chinese character topology structure validator to perform stroke-level structure matching on the intermediate standardized text. When it is detected that the stroke topology deviation exceeds the 5% threshold, repair it based on the preset character source database to generate standardized text data.
[0114] Specifically, examples of variant Chinese characters in the literature of each dynasty can be used to construct a cross-dynasty variant-character mapping library. And to further clarify the internal relationships between variant characters, the variant-character mapping library construction subunit can adopt an evolutionary glyph similarity algorithm to deeply analyze the stroke structure, component composition, and evolution trajectory of Chinese characters, calculate the similarity values between different variant characters, and construct a variant-character relationship graph. In this graph, each variant character can be used as a node, and the similarity connections between variant characters can be used as edges, so as to visually display the close and distant relationships of variant characters in the process of glyph evolution, providing a solid data foundation for subsequent text conversion work. Due to the huge span of the creation time of Chinese language and literature works, there are significant differences in the writing norms and usage habits in different periods. Therefore, in the mode selection and semantic analysis subunit, the works in the candidate literary dataset can be judged according to their creation time. For the literature before the Tang Dynasty, since traditional Chinese characters were widely used and had high standardization during this period, the traditional-character benchmark conversion mode can be enabled, with traditional Chinese characters as the core, and the variant characters and non-standard glyphs in the dataset can be uniformly converted into the corresponding traditional-character standard forms. For modern texts, given that simplified Chinese characters have become the mainstream writing norm, the simplified-character enhancement mode can be enabled to convert the traditional characters and variant characters in the dataset into the modern popular simplified-character forms, and ensure the accuracy and coherence of semantics. And when selecting the conversion mode, this subunit can perform semantic analysis on the text through a pre-trained context semantic model to generate a context vector that can accurately reflect the current text semantic environment. This vector not only contains the semantic information of the words themselves, but also incorporates their context association information in the whole text, providing a key semantic basis for the subsequent search for the optimal replacement path.
[0115] Specifically, in the genetic optimization path search subunit, based on the variant-character relationship graph and the current semantic context vector, the genetic algorithm first randomly generates an initial population of replacement schemes, and then evaluates the rationality of variant-character replacements in each scheme according to the variant-character relationship graph, and combines the semantic context vector to judge whether the text after replacement is semantically coherent and natural. And the quality of each replacement scheme is quantified by calculating a fitness function value. The higher the fitness value, the more compliant the scheme is. Secondly, the algorithm selects the better individuals from the current population with a certain probability for "crossover" and "mutation" operations to generate a new generation of replacement scheme population. After multiple rounds of iterative optimization, the replacement path that can make the text reach the optimal effect in terms of glyph standardization and semantic coherence is finally obtained, and then a preliminary standardized text is generated.
[0116] In addition, due to the existence of homophones and ancient and modern characters in Chinese language and literature, semantic confusion is easily caused. Therefore, the homophone conversion suppression and font structure verification subunit can perform semantic context analysis on the preliminary standardized text through the Bi-GRU model, that is, the bidirectional gated recurrent unit model. The Bi-GRU model can accurately identify which words are homophones or ancient and modern characters by analyzing the text word by word, and suppress possible incorrect conversions based on their semantic functions in a specific context. For example, in some contexts, "说" is the same as "悦". If the conventional conversion rules are simply followed, "说" may be mistakenly converted into other glyphs, but the Bi-GRU model can determine that this is a homophone usage based on the context, thereby keeping its original glyph unchanged and generating semantically accurate intermediate standardized text. In addition, a generative adversarial network can be used to construct a Chinese character topology structure verifier to ensure the accuracy of the intermediate standardized text in font structure. When it is detected that the stroke topology deviation of the generated image exceeds the pre-set threshold of 5%, it is determined that there is a problem with the font structure of the Chinese character. The correct glyph structure is retrieved based on the preset character source database, and the deviated Chinese character is repaired. Finally, standardized text data that conforms to the font structure specifications is generated, providing a high-quality text foundation for subsequent knowledge graph construction and visualization.
[0117] In one embodiment, the visual query output module 104 includes:
[0118] The literary entity extraction and analysis subunit is used to:
[0119] Through the spatiotemporal relationship model of literary entities, a multi-head spatiotemporal attention mechanism is used to extract literary entities from standardized text data and generate three-dimensional spatiotemporal coordinates containing dynasty numbers, geocodes, and timestamps;
[0120] Based on the frequency of citations of literary entities in historical documents and the popularity of dissemination on modern social media, the spatiotemporal dual-domain influence value is calculated through the LSTM time series network, where the spatiotemporal dual-domain influence value includes the historical authority weight and the modern dissemination weight.
[0121] Knowledge graph enhancement subunit, used to:
[0122] Construct a multi-scale knowledge graph. In the micro-layer of the multi-scale knowledge graph, the original textual relationships between literary entities are preserved. In the macro-layer of the multi-scale knowledge graph, a hierarchical graph convolutional network is used to perform spatiotemporal clustering of three-dimensional spatiotemporal coordinates and spatiotemporal dual-domain influence values to generate literary genre cluster labels and regional cultural diffusion paths.
[0123] Map the multi-scale knowledge graph to the coordinate system corresponding to the three-dimensional space-time coordinates, and generate a three-dimensional heat cloud map through the WebGL engine. The cultural communication heat cloud map is used to interact with users through a visual graphical interface;
[0124] A target data generation subunit, configured to:
[0125] Obtain the user's selection operations on the time range, geographical area, and literary genre clustering labels through a visual graphical interface, and convert the selection operations into spatio-temporal slice parameters;
[0126] Extract target data from the three-dimensional heat cloud map based on the spatio-temporal slice parameters to generate a visual query result.
[0127] Specifically, literary entities include key elements such as authors, works, literary events, and specific literary images. The multi-head spatio-temporal attention mechanism can simultaneously focus on information at different positions and different dimensions in the text, and extract literary entities from the standardized text data. And during the extraction process, in order to accurately identify the position of literary entities in the spatio-temporal dimension, three-dimensional spatio-temporal coordinates including dynasty numbers, geographical codes, and timestamps can be generated for each literary entity. The dynasty number clarifies the historical dynasty in which the literary entity is located, facilitating positioning and classification in the time dimension. The geographical code is accurate to geographical location information such as the place where the event related to the literary entity occurred or the native place of the author. The timestamp can further refine the time dimension, accurate to the specific creation time or event occurrence time, thereby constructing the accurate position of the literary entity in the three-dimensional spatio-temporal coordinate system, providing key spatio-temporal positioning information for subsequent analysis and knowledge graph construction. For the citation frequency of historical documents, by sorting and counting a large number of historical literary materials, the number of times each literary entity is cited can be determined to reflect its importance in academic research and literary inheritance. In terms of the popularity of modern social media dissemination, web crawler technology and data analysis tools can be used to collect data such as the topic popularity, number of likes, and number of comments of literary entities on mainstream social media platforms, and quantify their dissemination influence in contemporary society. Subsequently, by inputting the citation frequency of historical documents and the data of modern social media dissemination popularity into the LSTM network, a spatio-temporal dual-domain influence value is output. Among them, the historical authority weight reflects the academic status and influence of literary entities in the long history, and the modern dissemination weight reflects the dissemination breadth and popularity of literary entities in the current social media environment.
[0128] Specifically, in this embodiment, the multi-scale knowledge graph is a semantic network that reveals the relationships between entities. At the micro level, the original text relationships between literary entities are retained. At the macro level, in order to more comprehensively display the distribution and correlation characteristics of literary entities in the time-space dimension, a hierarchical graph convolutional network can be used to perform time-space clustering on the three-dimensional time-space coordinates and the time-space dual-domain influence values. The graph convolutional network can perform convolutional operations on graph-structured data and effectively extract the spatial features between nodes, that is, literary entities. And through hierarchical processing, the information of literary entities is gradually aggregated and abstracted, and literary entities with similar time-space coordinates and influence characteristics are clustered together. For example, the authors and works of the same literary genre can be grouped into one category during the time-space clustering process because of their similarities in creation time, region, and literary style inheritance, thus generating a literary genre clustering label. By analyzing the connections and evolutions between different clusters, the regional cultural diffusion path can be depicted, and the dissemination and development context of literary culture in different regions and times can be shown. Finally, this subunit can map the knowledge graph to the coordinate system corresponding to the three-dimensional time-space coordinates and use the WebGL engine for rendering, that is, vividly presenting the information in the knowledge graph in the form of a three-dimensional heat cloud map. In the three-dimensional heat cloud map, different literary entities can be displayed in different shades of color and shapes according to their influence levels and distributions in the time-space dimension, and the literary genre clustering and regional cultural diffusion path can be intuitively shown through color gradients and spatial position relationships.
[0129] The user selects the time range, geographical area, and literary genre clustering label according to their own interests in the visualization graphical interface. Subsequently, the target data generation subunit can parse and convert this selection operation to generate corresponding time-space slice parameters. These time-space slice parameters precisely define the data range of interest to the user and provide clear screening conditions for extracting target data from the three-dimensional heat cloud map. That is, based on the generated time-space slice parameters, this subunit can traverse all literary entity data in the three-dimensional heat cloud map and filter out the literary entities and their related information that meet the user's selection conditions. By organizing and formatting the extracted target data and presenting it to the user in an intuitive and clear visualization form, the final visualization query result is generated.
[0130] In one embodiment, the target data generation subunit includes:
[0131] A semantic verification subunit, which is used to perform semantic verification on the time-space slice parameters. If the time range in the time-space slice parameters exceeds the dynasty interval in which the literary entities exist, an early warning signal containing an out-of-bounds timestamp is generated, and the early warning signal is used to prompt the user to make corrections;
[0132] A hierarchical sampling output subunit, which is used for:
[0133] Determine the accuracy level according to the data density distribution of the three-dimensional thermal cloud map. The accuracy level includes a high-density level and a low-density level;
[0134] When it is detected that the target data volume belongs to the high-density level, use the aggregation algorithm based on K-means to generate a simplified view;
[0135] When it is detected that the target data volume belongs to the low-density level, output the original data view;
[0136] Align the simplified view or the original data view with the coordinate system of the three-dimensional thermal cloud map to generate a visual query result.
[0137] Specifically, literary entities have their specific existence dynasty intervals in history. The semantic verification subunit can compare the time range in the spatio-temporal slice parameter with the dynasty interval in which the literary entity actually exists. If it is found that the time range exceeds the dynasty interval to which the literary entity belongs, for example, the time range set by the user is a certain period before the Tang Dynasty, but the literary entity queried clearly belongs to a dynasty after the Tang Dynasty, an early warning signal containing the out-of-bounds timestamp will be generated and presented to the user in the form of popping up a prompt box, emitting a warning sound, etc., to guide the user to make corrections. Through this verification mechanism, it can effectively avoid the situation of obtaining incorrect or irrelevant data due to user misoperation or unfamiliarity with the time background of literary entities, and ensure the semantic accuracy and rationality of the query results. Due to the large amount of data and uneven distribution in the three-dimensional thermal cloud map, there are significant differences in the data density of different regions, and the accuracy levels of high-density level and low-density level can be divided according to its data density distribution.
[0138] When it is detected that the target data volume belongs to the high-density level, if the original data is directly output, it may lead to overly complex information and it is difficult for users to quickly obtain key information. Therefore, the aggregation algorithm based on K-means can be used to process the data, that is, cluster the literary entity data in the high-density area according to features such as its spatio-temporal coordinates, influence value, and mutual association relationship to generate a simplified view. While retaining the key information, this simplified view significantly reduces the amount of data, making the information presentation more concise and clear, and facilitating users to quickly grasp the core literary knowledge of this area. When it is detected that the target data volume belongs to the low-density level, the original data view can be directly output. Finally, align the simplified view or the original data view with the coordinate system of the three-dimensional thermal cloud map to generate a visual query result, which can further improve the user's query experience and knowledge acquisition efficiency.
[0139] Based on the same inventive concept, such as Figure 2As shown in the figure, the embodiments of the present application also provide a method for computer network query and reading of Chinese language and literature. The implementation solution provided by this method to solve problems is similar to the implementation solution described in the above system. Therefore, the specific limitations in one or more embodiments of the method for computer network query and reading of Chinese language and literature provided below can refer to the limitations on the system for computer network query and reading of Chinese language and literature in the above text, and will not be repeated here. The method includes:
[0140] S201: Perform hybrid feature parsing processing according to the user's query request to generate a hybrid query vector that combines radicals and semantics. The query request includes text and / or pictures;
[0141] S202: Based on the hybrid query vector, perform distributed retrieval in a preset spatio-temporal shard index structure to obtain a candidate literature dataset. The preset spatio-temporal shard index structure includes a time dimension, a genre dimension, and an author influence dimension;
[0142] S203: According to the Chinese character evolution rules and context, perform adaptive conversion processing on the variant characters and simplified and traditional Chinese characters in the candidate literature dataset to obtain standardized text data;
[0143] S204: According to the spatio-temporal relationship model of literary entities, perform knowledge graph enhancement processing on the standardized text data to generate a visual query result.
[0144] In an exemplary embodiment, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the system for computer network query and reading of Chinese language and literature of the present application are implemented. A multi-core processor is preferably used to improve the parallel processing ability of the system. Memory: Provide sufficient temporary storage space to support the operation of the program and the processing of data. The memory capacity should be large enough to accommodate a large amount of supply information and computing tasks.
[0145] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the system for computer network query and reading of Chinese language and literature of the present application are implemented. The computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid-state drive (SSD, Solid State Drives), or optical disc, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance RandomAccess Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory).
[0146] The above-described embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. A computer network query and reading system for Chinese language and literature, characterized in that, The system includes: A query request processing module, configured to perform hybrid feature parsing processing according to a user's query request to generate a hybrid query vector integrating radicals and semantics, where the query request includes text and / or pictures; A distributed retrieval module, configured to perform distributed retrieval in a preset spatio-temporal shard index structure based on the hybrid query vector to obtain a candidate literature dataset, where the preset spatio-temporal shard index structure includes a time dimension, a genre dimension, and an author influence dimension; A preliminary retrieval optimization module, configured to perform adaptive conversion processing on variant Chinese characters and simplified and traditional Chinese characters in the candidate literature dataset according to Chinese character evolution rules and context to obtain standardized text data; A visual query output module, configured to perform knowledge graph enhancement processing on the standardized text data according to a literature entity spatio-temporal relationship model to generate a visual query result.
2. The system according to claim 1, wherein The query request processing module includes: An image processing sub-unit, configured to: When the query request includes the image, use an adversarial generation network to correct the handwritten font distortion of the image to obtain an intermediate image with normalized glyphs; When the intermediate image is a scanned ancient book, use a bidirectional LSTM-CRF model to reconstruct the layout of the vertical text in the intermediate image, output horizontal text, and perform stroke topology verification on the horizontal text in combination with a preset Chinese character structure prior knowledge base to obtain standardized OCR text; A text processing sub-unit, configured to: When the query request includes the text, perform semantic parsing on the text to obtain a parsing structure including a character-level radical decomposition tree and a word-level dependency relationship graph; Based on the context semantics in the parsing structure, use a dynamic weight assignment algorithm to adjust the radical contribution degrees of each character in the text to generate a radical feature vector, and input the parsing structure into a BERT-wwm pre-trained model for semantic encoding to output a semantic vector; A hybrid query vector generation sub-unit, configured to: Input the standardized OCR text and the radical feature vector into an adversarial feature alignment network, output a cross-modal intermediate vector, and perform topological structure encoding on the cross-modal intermediate vector based on a preset radical relationship composition diagram to generate a feature matrix including glyph spatial relationships; Fuse the feature matrix with the semantic vector to obtain the hybrid query vector.
3. The system according to claim 1, wherein The distributed retrieval module includes: A time dimension sub-unit, configured to perform dynamic sharding processing on a preset literature dataset according to the time dimension, and use a sliding time window algorithm to segment cross-dynasty texts in the preset literature dataset to generate a time shard set, where overlapping shard identifiers are generated for texts with time label conflicts in the preset literature dataset through a Gaussian mixture model; A genre dimension sub-unit, configured to construct a genre weight calculation model based on the genre dimension, dynamically adjust the retrieval weights of poetry, prose, and novels according to the genre feature intensity in the hybrid query vector, and generate a weighted genre shard index; An author influence dimension sub-unit, configured to: A dynamic authority evaluation system is established according to the author influence dimension. A citation network is constructed through the literature citation relationship between authors, and the authority value of the author nodes in the citation network is calculated based on the PageRank algorithm; Sentiment analysis is performed on historical review data to obtain sentiment analysis results, and a spatio-temporal decay factor is calculated in combination with the text publication time; The authority value is multiplied by the spatio-temporal decay factor to generate an author influence index; A sharding retrieval sub-unit is used to input the time sharding set, the genre sharding index, and the author influence index into a cross-dimensional joint retrieval engine. Based on the preset spatio-temporal sharding index structure, a three-dimensional sharding space composed of the time dimension, the genre dimension, and the author influence dimension is constructed, and the hybrid query vector is used to calculate similarity through the three-dimensional sharding space to generate a candidate literary dataset.
4. The system according to claim 3, characterized in that, The authority value is calculated by the following formula: Among them, a i is the authority value of the author node i, μ is the attenuation factor, τ is the time decay coefficient, Cite(i) is the set of documents citing author i, t i and t j are the eras in which authors i and j are located respectively, and |Ref(j)| is the number of citations of the documents of author j.
5. The system according to claim 1, wherein The preliminary retrieval optimization module includes: A variant character mapping library construction sub-unit is used to construct a cross-dynasty variant character mapping library and establish a variant character relationship graph using an evolutionary glyph similarity algorithm; A mode selection and semantic analysis sub-unit is used to select a conversion mode according to the creation era of the candidate literary dataset and perform semantic analysis through a pre-trained context semantic model to generate a current semantic context vector. Among them, the traditional Chinese benchmark conversion mode is enabled for literature before the Tang Dynasty, and the simplified Chinese enhancement mode is enabled for modern texts; A genetic optimization path search sub-unit is used to search for the optimal replacement path using a genetic algorithm based on the variant character relationship graph and the current semantic context vector to generate a preliminary standardized text; A borrowed-character conversion suppression and font structure verification sub-unit is used for: Analyze the semantic context of the preliminary standardized text through a Bi-GRU model, suppress the conversion of borrowed characters and ancient and modern characters, and generate an intermediate standardized text; Construct a Chinese character topology structure validator using an adversarial generation network to perform stroke-level structure matching on the intermediate standardized text. When it is detected that the stroke topology deviation exceeds the 5% threshold, repair is performed based on a preset character source database to generate the standardized text data.
6. The system according to claim 1, characterized in that, The visualization query output module includes: A literary entity extraction and analysis sub-unit is used for: Extract literary entities from the standardized text data using a multi-head spatio-temporal attention mechanism through the literary entity spatio-temporal relationship model, and generate a three-dimensional spatio-temporal coordinate including a dynasty number, a geographical code, and a timestamp; Based on the citation frequency of the literary entity in historical literature and the modern social media dissemination popularity, calculate the spatio-temporal dual-domain influence value through an LSTM time series network, where the spatio-temporal dual-domain influence value includes a historical authority weight and a modern dissemination weight; A knowledge graph enhancement sub-unit is used for: Construct a multi-scale knowledge graph. Among them, the original text relationship between literary entities is retained in the micro layer of the multi-scale knowledge graph, and in the macro layer of the multi-scale knowledge graph, a hierarchical graph convolutional network is used to perform spatio-temporal clustering on the three-dimensional spatio-temporal coordinate and the spatio-temporal dual-domain influence value to generate literary genre clustering labels and regional cultural diffusion paths; Map the multi-scale knowledge graph to the coordinate system corresponding to the three-dimensional spatio-temporal coordinates, and generate a three-dimensional heat cloud map through the WebGL engine. The cultural dissemination heat cloud map is used to interact with users through a visual graphical interface; The target data generation subunit is used for: Obtain the user's selection operations on the time range, geographical area, and the literary genre clustering label through the visual graphical interface, and convert the selection operations into spatio-temporal slice parameters; Extract target data from the three-dimensional heat cloud map based on the spatio-temporal slice parameters, and generate the visual query result.
7. The system according to claim 6, wherein The target data generation subunit includes: The semantic verification subunit is used to perform semantic verification on the spatio-temporal slice parameters. If the time range in the spatio-temporal slice parameters exceeds the dynasty interval in which the literary entity exists, an early warning signal containing an out-of-bounds timestamp is generated. The early warning signal is used to prompt the user to make corrections; The hierarchical sampling output subunit is used for: Determine the accuracy level according to the data density distribution of the three-dimensional heat cloud map. The accuracy level includes a high-density level and a low-density level; When it is detected that the amount of target data belongs to the high-density level, a simplified view is generated using an aggregation algorithm based on K-means; When it is detected that the amount of target data belongs to the low-density level, the original data view is output; Align the simplified view or the original data view with the coordinate system of the three-dimensional heat cloud map to generate the visual query result.
8. A computer network query and reading method for Chinese language and literature, characterized in that, The method includes: Perform hybrid feature parsing processing according to the user's query request to generate a hybrid query vector that combines radicals and semantics. The query request includes text and / or pictures; Based on the hybrid query vector, perform distributed retrieval in a preset spatio-temporal sharding index structure to obtain a candidate literary dataset. The preset spatio-temporal sharding index structure includes a time dimension, a genre dimension, and an author influence dimension; According to the Chinese character evolution rules and context, perform adaptive conversion processing on the variant characters and simplified and traditional characters in the candidate literary dataset to obtain standardized text data; According to the spatio-temporal relationship model of literary entities, perform knowledge graph enhancement processing on the standardized text data to generate a visual query result.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the system according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the system according to any one of claims 1 to 7.
Citation Information
Cited By
Vector retrieval method based on language model
CN121327102A
Vector retrieval method based on language model
CN121327102B
Ancient book digitalization and research expert collaborative labeling method
CN121582754A
Patent retrieval system and method based on dynamic weight adjustment and multi-modal semantics
CN122262183A