Chinese language and literature database online query reading method and system
By generating semantic candidate sets, contextual semantic mapping paths, and version paragraph retrieval scheduling tables, the accuracy and consistency issues of Chinese language and literature document retrieval in the existing technology are solved, and efficient information positioning and reading experience are achieved.
Patent Information
- Application Number
- CN202511333101.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-18
AI Technical Summary
When processing complex Chinese language and literature documents, existing technologies are unable to effectively cope with the changing language structures and complex semantic relationships, resulting in a decrease in the accuracy and relevance of retrieval results, and a lack of dynamic adjustment mechanisms, which affects the comprehensiveness and accuracy of information.
By obtaining the meanings of the search terms and integrating them with the temporal weights of the documents, a semantic candidate set is generated, the subject pronoun and verb collocation structure is extracted, the contextual semantic mapping path is analyzed, the paragraph position is located, the structural offset between versions is adjusted, and word meaning comparison prompt blocks are generated to optimize the positioning, display and structural alignment of the documents.
It improves the accuracy and consistency of target documents, reduces reliance on manual keyword settings, improves the accuracy and efficiency of information acquisition, and optimizes the search experience.
Smart Images

Figure CN120821831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information retrieval technology, and in particular to an online query and reading method and system for a Chinese language and literature database. Background Art
[0002] The field of information retrieval technology involves core issues such as the organization, storage, indexing, and querying of text data. It primarily encompasses methods such as natural language-based text analysis, structured semantic processing, keyword matching, and related information location. These methods aim to enable users to quickly access specific information targets by building retrieval logic and query mechanisms. The methodological characteristics of this field are reflected in the efficient location and identification of massive information resources through steps such as lexical analysis, text comparison, and language model construction. Traditional online query and reading methods for Chinese language and literature databases refer to data-based management solutions for ancient and modern Chinese language and literature text resources. These technical issues address how to process the complex linguistic structures and unstructured text data in Chinese language and literature during information retrieval. Traditional methods utilize a term indexing mechanism based on specific document classification numbers, such as the TP3913 information retrieval system. Text matching and distribution are performed by manually setting keywords or employing general keyword extraction algorithms to enable online browsing and preliminary screening of target materials. These methods generally employ fixed classification matching, word frequency statistics, or pattern recognition to retrieve and display information.
[0003] The shortcomings of existing technologies are mainly reflected in the limitations of processing complex text structures and semantic associations. Traditional information retrieval methods often rely on fixed vocabulary classification and pattern matching, and screen texts through manually set keyword entries, but this cannot fully cope with the changing language structures and complex semantic relationships in documents. Especially when dealing with documents with rich context and context dependence, such as ancient Chinese language and literature, keyword matching is often misleading, resulting in a decrease in the accuracy and relevance of retrieval results. Existing systems also fail to effectively deal with differences between document versions and lack a dynamic adjustment mechanism, resulting in an inability to automatically handle structural offsets and content inconsistencies when comparing the same information in different versions. This processing method limits the comprehensiveness and accuracy of information, making it impossible for users to quickly and accurately locate information during use, affecting retrieval efficiency and reading experience. Summary of the Invention
[0004] In order to solve the technical problems existing in the prior art, an embodiment of the present invention provides an online query and reading method for a Chinese language and literature database.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: The online query and reading method of the Chinese Language and Literature Database includes the following steps: S1: Obtain the word meanings of the search terms, extract the citation frequency of core ancient books, the length of the interpretation text and the number of meaning boundary annotations, integrate them with the document time sequence weight, screen the word meanings that meet the conditions, and generate a semantic candidate set; S2: Based on the word meanings in the semantic candidate set, extract the subject pronoun and verb collocation structure, analyze and sort the paths associated with the semantic subject and action verb in the interpretation, and generate a contextual semantic mapping path sequence; S3: Based on the word meaning and chapter information in the priority path in the contextual semantic mapping path sequence, locate the corresponding paragraph in the version index, mark the target word meaning paragraph number and document location, build a retrieval structure, and generate a version paragraph retrieval scheduling table; S4: calling the version paragraph search scheduling table content, extracting the target word sense associated phrase structure and order, comparing the structure order offset between versions, adjusting the position of the paragraph that exceeds the set standard, and generating a word sense mapping paragraph alignment result; S5: Based on the paragraph alignment result of the word sense mapping, the word sense number and position of the keywords in the paragraph are marked, the reading area display structure is organized, the word sense interpretation, document source and semantic path information are output, and the word sense comparison prompt block is generated.
[0006] As a further solution of the present invention, the semantic candidate set includes a list of candidate word senses, a word sense feature parameter table, a selection threshold identifier, and a score summary; the context semantic mapping path sequence includes a path sequence number set, a relationship strength level set, a subject-verb collocation snapshot set, and an interpretation subject mapping mark set; the version paragraph retrieval scheduling table includes a version identification code, a paragraph number list, a document location coordinate, and a retrieval routing item; the word sense mapping paragraph alignment result includes an associated phrase structure table, a position sequence sequence, a structure order offset result, and a position adjustment record; the word sense comparison prompt block includes a word sense number marking area, a keyword position index, a document source entry, a path prompt information card, and a reading area display layout.
[0007] As a further solution of the present invention, the specific steps of S1 are: S101: Obtain the meaning of the search term in the Chinese language and literature database, collect the citation frequency of the meaning in the core ancient books and documents, the length of the interpretation text and the number of meaning boundary marks, organize the data into a unified format, and generate the document meaning parameters; S102: Based on the document word meaning parameters, the document time sequence weight is used to process the citation frequency of word meaning items in ancient documents of different ages, and the length of the interpretation text and the number of boundary marks of the meaning item are compared to generate a word meaning index; S103: Screening the word senses according to the word sense index, retaining a set of word senses that meet the conditions, and arranging the word senses in the set according to a predetermined sorting standard to generate a semantic candidate set.
[0008] As a further solution of the present invention, the specific steps of S2 are: S201: Obtain word senses in the semantic candidate set, extract the subject pronoun and verb collocation structure of the paragraph where the word senses are located, organize the subject pronouns and verbs into combination records, and generate paragraph collocation combination quantity; S202: Based on the paragraph collocation combination quantity, calling the semantic subject and action verb in the word meaning interpretation, comparing the subject pronoun and verb collocation structure in the paragraph, determining the corresponding relationship and recording the strength result to obtain the semantic association strength value; S203: Sort the paths corresponding to the word meanings according to the semantic association strength values, merge the path results according to the paragraph positions, and generate a context semantic mapping path sequence.
[0009] As a further solution of the present invention, the specific steps of S3 are: S301: Obtaining a priority path in the context semantic mapping path sequence, extracting word senses and chapter information under the path, detecting the paragraph position corresponding to the word sense in the version index, and recording the paragraph number and index mark to generate a paragraph index number; S302: Based on the paragraph index number, the paragraph number and document position containing the target word meaning are retrieved, the corresponding relationship between the paragraph number and the position data is determined, and a combination structure that can uniquely identify the paragraph is selected to obtain a paragraph positioning matching value; S303: According to the paragraph positioning matching value, a retrieval structure of paragraph number and document position is established in the version index, all combination relationships are sequentially merged into a search path, and a version paragraph search scheduling table is generated.
[0010] As a further solution of the present invention, the specific steps of S4 are: S401: Calling the paragraph number and document position in the version paragraph search scheduling table, extracting the associated phrase structure and appearance order of the target word meaning in the paragraph, comparing the number and position relationship between adjacent phrases, and generating phrase sequence annotations; S402: Based on the phrase sequence annotation amount, calculate the deviation degree of the structural sequence in the difference version, compare the paragraph number difference with the set sequence deviation standard value one by one, obtain the paragraph set that exceeds the standard, and generate the sequence deviation difference value; S403: According to the sequence offset difference value, adjust the position of the paragraphs that exceed the standard, call the number and document position to rearrange the paragraph order, map all adjusted paragraph positions to corresponding word sense combinations, and generate a word sense mapping paragraph alignment result.
[0011] As a further solution of the present invention, the specific steps of S5 are: S501: Based on the paragraph alignment result of the word sense mapping, annotate the word sense number and the position index within the paragraph corresponding to the keyword paragraph by paragraph, call the number of occurrences and position difference of the keyword in the paragraph, calculate the combination of the number and the position index, and generate the word sense position annotation amount; S502: Sort the annotated word meaning numbers in the paragraph according to the word meaning position annotation quantity, match the position index of the keyword corresponding to the number in the reading area with the paragraph structure, calculate the sorting difference and compare it with the set display reference value to obtain the reading area sequence difference; S503: Calling the reading zone sequence difference, reorganizing the paragraphs corresponding to the difference exceeding the display benchmark value, outputting the word meaning interpretation, document source and path information in paragraph order, combining the content into prompt blocks, and generating a word meaning comparison prompt block.
[0012] As a further embodiment of the present invention, the core ancient books and documents refer to representative historical documents in the field of Chinese language and literature, including the Book of Documents, Zuo Zhuan, and Records of the Grand Historian, classic documents of the Han Dynasty; The document time sequence weight refers to the weighted value assigned to the meaning of words in different documents according to the age, genre and influence of the document in language evolution; The semantic candidate set refers to a set of all candidate word meanings screened by multi-dimensional parameters; The subject pronoun refers to the pronoun that represents the subject in the sentence; The verb collocation structure refers to the grammatical structure consisting of a verb and an associated noun, adjective or other modifying element; The context semantic mapping path sequence refers to a path set generated by sorting the semantic relationships in the context.
[0013] As a further solution of the present invention, the version index is an index including different document versions, paragraph positions and chapter structures; The section number and document position are identifiers in the document; The version paragraph search scheduling table refers to the search order and retrieval structure generated based on the different version documents and target paragraphs during the query process; The word meaning item related phrase structure refers to a phrase, word group or sentence structure that is closely related to the target word meaning item in the document paragraph, and the structure is consistent with the semantics of the word meaning item; The structural order deviation refers to the order difference of phrases or word groups in sentences in different versions of documents. The order is sorted and adjusted by calculating the structural order difference between different versions. The paragraph alignment result of the word sense mapping refers to the paragraph alignment result generated by adjusting the structural order between versions; The word meaning comparison prompt block refers to the word meaning comparison framework displayed during the user's reading process, including word meaning interpretation, document source and path information.
[0014] Based on the same inventive concept, a Chinese language and literature database online query and reading system is also proposed, which is used to implement the above-mentioned Chinese language and literature database online query and reading method, including: The word meaning candidate module obtains the citation frequency, meaning length, meaning boundary marking number and document time sequence weight of the search terms in the core ancient books and documents, and sets the screening threshold to screen the shortlisted word meanings and generate a semantic candidate set; The semantic path module extracts subject pronouns and verb collocations in the corresponding paragraphs based on the word meanings in the semantic candidate set, calculates the matching strength of the semantic subject and the action verb and sorts them, and generates a contextual semantic mapping path sequence; The version positioning module locates the section number and document position in the version index and records the coordinates based on the priority path in the context semantic mapping path sequence, combined with the word meaning and the chapter information, and generates a version paragraph retrieval scheduling table; The paragraph alignment module calls the version paragraph search scheduling table, extracts the phrase structure and position order of the target word sense, calculates the order offset between the different versions, and adjusts the paragraphs that exceed the threshold, thereby generating a word sense mapping paragraph alignment result; The reading prompt module marks the keyword meaning number and position in the reading area based on the paragraph alignment result of the word meaning mapping, organizes the display of document source and path information, and generates a word meaning comparison prompt block.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, semantic mapping is used to optimize the positioning, display and structural alignment of documents, thereby improving the accuracy and consistency of target documents. Contextual information is combined with version comparison to accurately screen relevant content, and structural offsets between versions are automatically adjusted, reducing dependence on manual keyword settings. At the same time, the generated word meaning comparison prompt block optimizes the document retrieval experience, enabling users to understand the meaning of words and contextual relationships more quickly and accurately, thereby improving the accuracy and efficiency of information acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 Schematic diagram of the steps of the present invention; Figure 2 This is a schematic diagram of the refinement of S1 of the present invention; Figure 3 This is a schematic diagram of the refinement of S2 of the present invention; Figure 4 This is a schematic diagram of the refinement of S3 of the present invention; Figure 5 This is a schematic diagram of the refinement of S4 of the present invention; Figure 6 This is a schematic diagram of the refinement of S5 of the present invention; Figure 7 It is a system module diagram of the present invention. DETAILED DESCRIPTION
[0018] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0019] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0020] See also Figure 1 The embodiment of the present invention provides an online query and reading method for a Chinese language and literature database, comprising the following steps: S1: Obtain the word meanings corresponding to the search terms in the Chinese language and literature database, extract the citation frequency, length of the interpretation text, and the number of meaning boundary annotations in the core ancient books and documents, integrate the parameters based on the document time sequence weight, screen the word meanings that meet the requirements, and generate a semantic candidate set; The Chinese Language and Literature Database is a digital document database that includes ancient books and modern Chinese academic data. The database includes multiple versions of ancient Chinese vocabulary, grammar, and changes. Core ancient books and documents refer to representative historical documents in the field of Chinese language and literature, including classic Han Dynasty texts such as "Shangshu", "Zuo Zhuan", and "Shiji"; Document temporal weight refers to the weighted value assigned to word meanings in different documents based on the document's age, genre, and influence in language evolution; The semantic candidate set refers to the set of all candidate word meanings that are screened by multi-dimensional parameters (including frequency and meaning length); S2: Based on the word meanings in the semantic candidate word meaning set, extract the subject pronoun and verb collocation structure in the corresponding paragraph, determine the association with the semantic subject and action verb in the word meaning interpretation, sort the paths according to the association strength, and generate a contextual semantic mapping path sequence; Subject pronouns refer to the pronouns that represent the subject in a sentence and are used to reflect the theme of the sentence in the context; Verb collocation structure refers to the grammatical structure consisting of verbs and associated nouns, adjectives or other modifying elements, which is used to capture the usage pattern of verbs in context; Association strength refers to the degree of match between the subject pronoun and the semantic subject in the word meaning, and between the verb and the action verb in the word meaning, which is used to judge the closeness of the association between the word meaning and the context; The context semantic mapping path sequence refers to a set of paths generated by sorting the semantic relationships in the context, which is used to accurately match the query word meanings; S3: Based on the word meanings and text information of the priority path in the contextual semantic mapping path sequence, the corresponding paragraph is located in the version index, the paragraph number and document location containing the target word meaning are marked, the retrieval structure is constructed, and a version paragraph retrieval scheduling table is generated; The version index is an index that includes different document versions, paragraph positions, and chapter structures, and is used to locate the content in the document; The paragraph number and document location are identifiers in the document, which are used to mark the target paragraph and its specific location in the document, and help to retrieve the content of the target paragraph; The version paragraph retrieval schedule refers to the retrieval order and retrieval structure generated based on the different version documents and target paragraphs during the query process, which is used to organize the order and index of paragraph content in the retrieval system; S4: Call the content in the version paragraph search schedule to extract the target word sense associated phrase structure and position sequence, compare the structure sequence offset between the different versions, adjust the position of the paragraphs whose sequence offset exceeds the set standard, and generate the word sense mapping paragraph alignment result; The word meaning related phrase structure refers to the phrase, word group or sentence structure that is closely related to the target word meaning in the document paragraph, and the structure is consistent with the semantics of the word meaning; Structural order shift refers to the order difference of phrases or word groups in sentences in different versions of documents. The order difference between different versions is calculated to perform sorting and adjustment. The paragraph alignment result of word sense mapping refers to the paragraph alignment result generated by adjusting the structural order between versions to ensure semantic consistency and structural matching; S5: Based on the paragraph alignment results of word sense mapping, mark the word sense numbers and positions corresponding to the keywords in the paragraph, organize the reading area display structure, output word sense interpretation, document source and path information, and generate word sense comparison prompt blocks; The word meaning comparison prompt block refers to the word meaning comparison framework displayed during the user's reading process, including word meaning interpretation, document source and path information, which is used to improve the convenience of document comparison and understanding.
[0021] The semantic candidate set includes a list of candidate word senses, a table of word sense feature parameters, a selection threshold identifier, and a score summary. The context semantic mapping path sequence includes a path sequence number set, a relationship strength level set, a subject-verb collocation snapshot set, and a interpretation subject mapping tag set. The version paragraph retrieval scheduling table includes a version identification code, a section number list, a document location coordinate, and a retrieval routing item. The word sense mapping paragraph alignment result includes a table of associated phrase structures, a position sequence sequence, a structure order offset result, and a position adjustment record. The word sense comparison prompt block includes a word sense number annotation area, a keyword position index, a document source entry, a path prompt information card, and a reading area display layout.
[0022] See also Figure 2 , the specific steps of S1 are: S101: Obtain the meaning of the search term in the Chinese language and literature database, collect the citation frequency of the meaning in the core ancient books and documents, the length of the interpretation text and the number of meaning boundary marks, organize the data into a unified format, and generate the document meaning parameters; When obtaining the word meanings corresponding to the search terms in the Chinese language and literature database, the terms are first mapped to the numbers in the database one by one. For example, the meaning items corresponding to "ren" are numbered A1, A2, and A3. Then, the citation frequencies of these meanings are collected in core ancient books and documents. For example, A1 appears 15 times in "The Analects of Confucius" and 8 times in "Mencius". A2 appears a total of 6 times in different documents. Then, the length of the interpretation text is counted. The interpretation of A1 is 120 words, the interpretation of A2 is 85 words, and the interpretation of A3 is 60 words. At the same time, the number of meaning boundary annotations is counted. A1 has 4 boundaries, A2 has 2 boundaries, and A3 has 3 boundaries. The above-collected frequency, length, and number of boundary annotations are unified and organized into corresponding records, and arranged into a data set in a fixed order. Each record consists of three parameters to obtain the document word meaning parameters.
[0023] S102: Based on the document word meaning parameters, the document time sequence weight is used to process the citation frequency of word meaning items in ancient documents of different ages, and the length of the interpretation text and the number of meaning boundary marks are compared to generate a word meaning index; Based on the word meaning parameters of the document, it is necessary to weight the ancient books and documents of different ages and set the time sequence weights of different eras, such as 1.2 for the pre-Qin period, 0.9 for the Han Dynasty, and 0.7 for the Tang Dynasty. When a certain meaning appears 10 times in the pre-Qin documents and 5 times in the Han Dynasty, the citation frequency obtained after weighting is 16.5. Then, the length of the interpretation text and the number of meaning boundary annotations are compared. The text length threshold is set to 100 words. If the number of interpretation words exceeds 100, it is recorded as 1, otherwise it is 0. The baseline value of the number of boundary annotations is 3. If it is greater than or equal to 3, it is recorded as 1, otherwise it is 0. For example, the length of A1 interpretation is 120 words, which is greater than the threshold, and it is recorded as 1. The number of boundary annotations is 4, which is greater than the baseline value, is recorded as 1. The weighted frequency is added to these two recorded values to obtain a comprehensive value. The larger the value, the more prominent its weight in the parameter set. The same operation is performed on all meanings in turn to obtain the word meaning index.
[0024] S103: Screening word senses according to word sense indicators, retaining a set of word senses that meet the conditions, and arranging the word senses in the set according to a predetermined sorting standard to generate a semantic candidate set; According to the word meaning index, all word meanings are screened, and the screening threshold is set to 10. If the value is greater than or equal to 10, it is retained, otherwise it is discarded. For example, A1 with a value of 17.5 meets the conditions and is retained, while A2 with a value of 7.3 is discarded. The retained meanings after screening are combined into a set, such as C={A1, A3}, and then arranged in descending order according to the index values. If the values are the same, the order is adjusted from short to long according to the length of the interpretation text. For example, if A1 has a value of 17.5 and A3 has a value of 11.2, then A1 is ranked before A3. If A3 and A4 have the same value, the number of words in the interpretation text is compared, and the one with fewer words is ranked first. After sorting, a new set sequence is obtained to generate a semantic candidate set.
[0025] See also Figure 3 , the specific steps of S2 are: S201: Obtain word senses in the semantic candidate set, extract the subject pronoun and verb collocation structure of the paragraph where the word sense is located, organize the subject pronouns and verbs into combination records, and generate paragraph collocation combination quantity; When obtaining the semantic items in the semantic candidate set, it is necessary to check the sentence structure of the paragraph where each semantic item is located item by item, extract the subject pronouns and verb collocations in it, first identify the pronouns in the sentence through syntactic decomposition methods, and then locate the verb actions that collocate with them. For example, in the sentence "What he said is good", "he" forms a collocation with the verb "said" as the subject pronoun, and in the sentence "He attacks the city", "he" forms a collocation with "attacks". Record these collocations one by one according to the paragraph number to form a unified combined record. Each record contains three elements: paragraph number, subject pronoun, and verb. To maintain consistency, a fixed text format can be used for recording. For example, "D1 - he - said", "D2 - he - attacks". If the same collocation structure appears in the same paragraph, the count is incremented. For example, if "he - said" appears 2 times in D1, it is recorded as "D1 - he - said × 2". As the processed paragraphs increase, these collocation combinations are aggregated into a whole data set, which can be regarded as the paragraph collocation combination quantity and is used for subsequent semantic comparison operations.
[0026] S202: Based on the paragraph collocation combination quantity, call the semantic subject and action verb in the semantic interpretation, compare the subject pronoun and verb collocation structure in the paragraph, judge the corresponding relationship and record the intensity result to obtain the semantic association intensity value; Based on the paragraph collocation combination quantity, it is necessary to call the semantic subject and action verb in the semantic interpretation for comparison processing. First, organize the interpretation text, identify the semantic subjects in it, such as "person", "scholar", "lord", etc., and then extract the verbs related to the subject from the interpretation sentence. For example, "love", "guard", "act", and then form a semantic combination table, such as "person - love", "scholar - guard", "lord - act". Then, map the subject pronoun and verb in the paragraph collocation combination to the semantic combination table in turn. If the object referred to by the pronoun matches the semantic subject and the verb meaning is consistent, it is determined as a strong correspondence. If only partial matching occurs, it is a medium correspondence. If there is no match at all, it is a weak correspondence. To ensure quantification, a strong correspondence can be set as 100%, a medium correspondence as 50%, and a weak correspondence as 0%. For example, when "he - act" in D1 corresponds to the interpretation "person - act", it can be considered 100%. When "he - attacks" in D2 corresponds to "scholar - guard" in the interpretation, only the verb part does not match, so it is recorded as 0%. Another example is that when "someone - follow" in D3 corresponds to "scholar - act" in the interpretation, the subject does not match and the verb is close, so it can be marked as 50%. Compare all the paragraph combinations one by one, record the results and organize them into a unified semantic association intensity value to provide a basis for subsequent sorting.
[0027] S203: According to the semantic association intensity value, sort the paths corresponding to the semantic items, merge the path results according to the paragraph position, and generate a context semantic mapping path sequence; According to the semantic association strength value, the paths corresponding to each word sense need to be sorted, and each path is arranged in order according to the strength value. If a sense appears in multiple strength results in different paragraphs, for example, A1 corresponds to D1 with a strength of 100% and D2 with a strength of 50%, D1 is ranked first and then D2. If A2 corresponds to D3 with a strength of 80% and D4 with a strength of 30%, D3 is ranked first and then D4. When encountering the same strength value, they are sorted according to the order of the paragraphs. For example, if the strength of two paragraphs is 70%, the paragraph with the smaller number is ranked first. After the sorting is completed, the results are merged according to the paragraph position, and the paths of different senses in the same paragraph are merged. For example, D1 contains A1, D2 contains A1 and A3, and D3 contains A2. After integration, a continuous contextual semantic mapping path sequence is formed. This sequence can cover all paragraphs and reflect the arrangement results of each sense, thus obtaining a complete contextual semantic mapping path sequence.
[0028] See also Figure 4 , the specific steps of S3 are: S301: Obtaining a priority path in a context semantic mapping path sequence, extracting word senses and chapter information under the path, detecting the paragraph position corresponding to the word sense in the version index, and recording the paragraph number and index mark to generate a paragraph index number; To obtain the priority path in the context semantic mapping path sequence, it is necessary to search the paths in the sequence one by one and compare the number of occurrences of the path with the semantic association strength to determine the priority. For example, when a path appears more than 5 times in the sequence and the strength value is greater than 70%, it can be marked as a priority path. Then, the relevant word senses and chapter information are extracted under the priority path, and these word senses are matched with the corresponding paragraphs one by one. The paragraph position of the word sense is detected in the version index. The detection process requires scanning the paragraph numbers in sequence and recording the corresponding paragraph entries, and then adding index tags to the entries. For example, when the paragraph numbers are D1 and D3, the corresponding word sense A can be recorded as D1-A and D3-A. If the same word sense appears multiple times in a paragraph, the number of occurrences is accumulated and recorded as D1-A×2. Such records are uniformly classified into the paragraph index number. After accumulating data from multiple paragraphs, a set containing paragraph numbers and index tags is formed for subsequent retrieval calls.
[0029] S302: Based on the paragraph index number, the paragraph number and document position containing the target word meaning are retrieved, the corresponding relationship between the paragraph number and the position data is determined, and a combination structure that can uniquely identify the paragraph is selected to obtain a paragraph positioning matching value; Based on the paragraph index number quantity, it is necessary to call the paragraph number and document position data containing the target word meaning, establish a corresponding relationship between the paragraph number and the document position and make a judgment. First, expand the index number quantity, for example, including D1-A, D3-A and D2-B, and then retrieve the corresponding document position data from the document index table. For example, paragraph D1 corresponds to P12, paragraph D3 corresponds to P27, and paragraph D2 corresponds to P15. Compare the paragraph number and the document position one by one. If a paragraph number corresponds to multiple positions, it is necessary to filter the combination structure that can uniquely identify the paragraph. The screening criteria can be based on the uniqueness of the position number and the matching degree of the paragraph number. When there is an offset between the position number and the paragraph number, the unique value can be selected by comparing the size of the offset. For example, when paragraph D1 corresponds to positions P12.1 and P12.3, P12.1 can be determined as the unique matching value based on the minimum offset. Each paragraph number only retains one corresponding document position to form a paragraph positioning matching value. After processing, unique combination results such as D1-P12.1, D3-P27.2, and D2-P15.1 can be obtained.
[0030] S303: Based on the paragraph positioning matching value, a retrieval structure of paragraph number and document location is established in the version index, all combination relationships are sequentially merged into a search path, and a version paragraph search scheduling table is generated; According to the paragraph positioning matching value, it is necessary to establish a retrieval structure of paragraph number and document position in the version index. First, the paragraph number and matching value are combined into a mapping relationship, for example, D1 corresponds to P12.1, D2 corresponds to P15.1, and D3 corresponds to P27.2. Then all mapping relationships are merged in sequence to form a retrieval path. During the merging process, the paragraph number sequence is kept consistent with the original chapter sequence. If there are identical matching values, they are arranged in sequence according to the paragraph number. After integration, a complete retrieval path link is obtained, and then these paths are recorded one by one in a scheduling table. The scheduling table contains three types of information: paragraph number, document position, and positioning matching value. For example, it can be recorded as D1-P12.1, D2-P15.1, and D3-P27.2. This tabular text structure can be regarded as a version paragraph retrieval scheduling table, which is used for subsequent rapid positioning of each paragraph.
[0031] See also Figure 5 , the specific steps of S4 are: S401: Calling the paragraph number and document position in the version paragraph search scheduling table, extracting the associated phrase structure and appearance order of the target word meaning in the paragraph, comparing the number and position relationship between adjacent phrases, and generating phrase sequence annotations; To call the version paragraph retrieval schedule for the paragraph number and document position, it is necessary to read the schedule entries one by one, use the paragraph number and the corresponding document position as the basic reference items, extract the associated phrase structure of the target word meaning in the paragraph in sequence, and record it in the order of actual appearance. When extracting the phrase structure, it can be separated sentence by sentence in combination with the natural semantic flow within the paragraph. For example, in paragraph D1, the word meaning X appears in the document position P12, forming the phrases "operation unit X1" and "operation unit X2", which are marked as Z1 and Z2 respectively. When there are multiple phrases in the paragraph, the numbering needs to be extended to form continuous marks such as Z1, Z2, and Z3, and then the numbers and The positional relationship is compared based on the line spacing difference in the document position. The distance between each pair of phrases is recorded as a numerical range. For example, when the line spacing between adjacent phrases is less than 10 lines, it is marked as a close relationship, when the line spacing is between 10 and 20 lines, it is marked as an intermediate relationship, and when the line spacing is greater than 20 lines, it is marked as a distant relationship. All phrase relationships are summarized in sequence to form a phrase sequence annotation quantity, which includes not only the phrase number and sequence, but also a qualitative description of the positional difference, such as close, intermediate or distant. Each record can be regarded as a sequence relationship of a group of adjacent phrases, and a complete phrase sequence annotation set is obtained for subsequent structural comparison and sequence difference judgment.
[0032] S402: Based on the phrase sequence annotation amount, the deviation degree of the structural order in the difference version is calculated, and the paragraph number difference is compared with the set sequence deviation standard value one by one to obtain the paragraph set that exceeds the standard and generate the sequence deviation difference value; Based on the phrase order annotation quantity, it is necessary to calculate the offset of the structural order between different versions. First, the phrase order annotations of the two versions are expanded one by one, and the distance differences between the same phrase pairs are compared one by one. For example, the distance between Z1 and Z2 in version V1 is marked as a close relationship, and the distance between Z1 and Z2 in version V2 is marked as an intermediate relationship. It is determined that there is an order offset in the phrase pair, and then this difference is converted into an offset value. The offset value can be set according to the difference level. For example, the difference between close and intermediate is recorded as 2, the difference between intermediate and far is recorded as 3, and the difference between close and far is recorded as 5. The obtained values are compared one by one with the pre-set sequence offset standard value. The standard value can be set to 5. When the difference value is less than or equal to 5, it is determined to be within the standard range. When the difference value is greater than 5, it is determined to be out of range. The paragraphs to which the phrase pairs that are out of range belong are included in the offset paragraph set. For example, if the difference value between Z2 and Z3 is 7 in paragraph D2, the paragraph is recorded as an offset paragraph, and a sequence offset difference value of 7 is generated and archived. After comparison, all paragraphs form a set, which contains all paragraphs that exceed the standard and the corresponding difference values, providing a basis for subsequent sequence adjustment.
[0033] S403: Adjust the positions of the paragraphs that exceed the standard according to the sequence offset difference value, rearrange the paragraph order by using the number and document position, map all adjusted paragraph positions to corresponding word sense combinations, and generate a word sense mapping paragraph alignment result; According to the sequence offset difference value, the position of the paragraphs in the collection needs to be adjusted. First, read the collection items one by one, and analyze the paragraph number and document position in combination with the difference value. When the difference value is between 1 and 5, the original position can be kept unchanged. When the difference value is between 6 and 10, the paragraph needs to be moved forward or backward by one position. When the difference value is between 11 and 20, it needs to be moved two positions. If the value exceeds 20, the movement amplitude is further increased. In this way, the order is adjusted paragraph by paragraph, and the adjusted paragraphs are rearranged into a new sequence, and the paragraph numbers are kept in the arrangement process. The logical connection between them is not disrupted. Then the new paragraph position is mapped with the original word sense combination. For example, paragraph D3 is located at P27 in the original document. Because the difference value is 8, it is moved forward to P26 and a corresponding relationship is re-established with the combination of word senses X and Y, recorded as D3-P26-(X, Y). Then the other paragraphs are processed in turn to generate a new word sense mapping paragraph alignment result. The result includes three aspects: paragraph number, updated position number and word sense combination, which completely covers all adjusted paragraphs and realizes a unified structure after sequential alignment.
[0034] See also Figure 6 , the specific steps of S5 are: S501: Based on the paragraph alignment results of the word sense mapping, the word sense number and the position index corresponding to the keyword are annotated paragraph by paragraph, the number of occurrences of the keyword in the paragraph and the position difference are called, and the combination of the number and the position index is calculated to generate the word sense position annotation amount; Based on the paragraph alignment results of word sense mapping, it is necessary to mark the word sense number and the position index within the paragraph corresponding to the keyword. First, the paragraph numbers contained in the alignment results are read one by one, and the keywords under each paragraph number are assigned numbers according to the order of appearance. For example, paragraph D1 contains the keywords "search", "compare", and "sort", which are marked as W1, W2, and W3 respectively. Then, the position index of each keyword in the text is determined according to the order of words within the paragraph. Assuming that "search" is at the 8th word position, "compare" is at the 15th word position, and "sort" is at the 21st word position, the indexes are P8, P15, and P21 respectively. Then, the number of times the keyword appears is counted. When a keyword appears multiple times, the specific position of each appearance needs to be recorded. For example, "search" is at the 8th word position, "compare" is at the 15th word position, and "sort" is at the 21st word position. If a word appears twice in a paragraph, and the indexes are P8 and P24 respectively, the index record of the number is {P8, P24}. At the same time, the difference between the indexes needs to be calculated. If the difference of P24 minus P8 is 16, the value is recorded and bound to the number to form the difference 16 corresponding to W1. If W2 and W3 only appear once, the difference is 0. When a paragraph contains multiple word meaning numbers, it is necessary to form a set structure of numbers, indices and differences, such as {W1: P8, P24-difference 16}{W2: P15-difference 0}{W3: P21-difference 0}. This set is the word meaning position annotation, which can fully record the keyword number, occurrence position and relative difference. The same processing needs to be performed in different paragraphs so that each paragraph has a corresponding word meaning position annotation.
[0035] S502: Sort the annotated word meaning numbers in the paragraph according to the word meaning position annotation quantity, match the position index of the keyword corresponding to the number in the reading area with the paragraph structure, calculate the sorting difference and compare it with the set display benchmark value to obtain the reading area sequence difference; According to the amount of word meaning position annotation, the annotated word meaning numbers in the paragraph need to be sorted. First, all numbers are arranged from small to large according to the position index. For example, in paragraph D2, the position index of W1 is P6, the position index of W2 is P14, and the position index of W3 is P19. The sorting order is W1, W2, and W3. Then the numbering order is compared with the natural word order of the paragraph. If the numbering order is exactly the same as the actual order, the difference is judged to be 0. If the order is different, the difference needs to be calculated. The difference is measured by the number of displacements between the position indexes. For example, W1 is ranked first in the sorting, but appears in the natural word order of the paragraph. For the second digit, the difference is recorded as 1. After completing the comparison of all numbers in sequence, a difference set within a paragraph is formed, for example, {W1: 1, W2: 0, W3: 0}. The set is then compared with the display benchmark value, which can be set to 3. When the difference is less than or equal to 3, it is recorded as being within the standard range. When the difference is greater than 3, it is recorded as being out of range. For example, in paragraph D2, the W1 difference is 1, the W2 difference is 0, and the W3 difference is 4, then W3 is judged to be out of range, forming a reading area sequence difference set. This set clearly marks the difference corresponding to each number and compares it with the benchmark value, thereby obtaining the deviation items in each paragraph.
[0036] S503: calling the reading zone sequence difference, reorganizing the paragraphs corresponding to the difference exceeding the display benchmark value, outputting the word meaning interpretation, document source and path information in paragraph order, combining the content into prompt blocks, and generating a word meaning comparison prompt block; To call the reading area sequence difference, the paragraphs corresponding to the difference that exceeds the display benchmark value need to be reorganized. First, the abnormal numbers and differences are extracted. For example, the difference set of paragraph D3 is {W1: 0, W2: 5, W3: 2}, where the W2 difference is 5, which is greater than the benchmark value 3. It is determined that D3 needs to be adjusted. The adjustment method is to rearrange the position of the number in the paragraph according to the difference value. The numbers with differences between 1 and 3 remain in place, the numbers with differences between 4 and 6 move forward or backward by 1 index interval, and the numbers with differences between 7 and 10 move by 2 index intervals. If it exceeds 10, the movement amplitude continues to increase according to the difference size. For example, if the W2 difference is 5, it needs to move by 1 index interval, adjusting from the original position P20 to P1 9. After completing the position correction, the number is combined with the word meaning and the document source. For example, if the meaning of W2 is "search condition", the document source is the second paragraph of page 15 of a certain document, and the path information is "path / chapter / 15 / 2", it is recorded as W2-meaning-source-path. Then all numbers are combined in sequence to form a prompt block. The structure of the prompt block is the combination of all word meaning numbers corresponding to the paragraph number and their meaning, source, and path. For example, the prompt block of paragraph D3 is {W1-meaning-source-path, W2-meaning-source-path, W3-meaning-source-path}. The prompt block collection of all paragraphs constitutes the online reading word meaning comparison prompt block, which is arranged continuously in paragraph order and retains complete interpretation information.
[0037] See also Figure 7 Based on the same inventive concept, a Chinese language and literature database online query and reading system is also proposed, which is used to implement the above-mentioned Chinese language and literature database online query and reading method, including: The word meaning candidate module obtains the citation frequency, meaning length, meaning boundary marking number and document time sequence weight of the search terms in the core ancient books and documents, and sets the screening threshold to screen the shortlisted word meanings and generate a semantic candidate set; The semantic path module extracts subject pronouns and verb collocations in the corresponding paragraphs based on the word meanings in the semantic candidate set, calculates the matching strength between the semantic subject and the action verb, and sorts them to generate a contextual semantic mapping path sequence; The version positioning module maps the priority path in the path sequence based on the contextual semantics, combines the word meaning and the text information, locates the section number and document position in the version index, records the coordinates, and generates a version paragraph search schedule; The paragraph alignment module calls the version paragraph retrieval schedule to extract the phrase structure and position order of the target word sense, calculates the order offset between the different versions, and adjusts the paragraphs that exceed the threshold to generate the word sense mapping paragraph alignment result; The reading prompt module maps the paragraph alignment results based on the word meanings, marks the keyword meaning numbers and positions in the reading area, organizes the display of document sources and path information, and generates word meaning comparison prompt blocks.
[0038] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for online query and reading of Chinese language and literature database, characterized in that: The following steps are involved: S1: Obtain the word meanings of the search terms, extract the citation frequency of core ancient books, the length of the interpretation text and the number of meaning boundary annotations, integrate them with the document time sequence weight, screen the word meanings that meet the conditions, and generate a semantic candidate set; S2: Based on the word meanings in the semantic candidate set, extract the subject pronoun and verb collocation structure, analyze and sort the paths associated with the semantic subject and action verb in the interpretation, and generate a contextual semantic mapping path sequence; S3: Based on the word meaning and chapter information in the priority path in the contextual semantic mapping path sequence, locate the corresponding paragraph in the version index, mark the target word meaning paragraph number and document location, build a retrieval structure, and generate a version paragraph retrieval scheduling table; S4: calling the version paragraph search scheduling table content, extracting the target word sense associated phrase structure and order, comparing the structure order offset between versions, adjusting the position of the paragraph that exceeds the set standard, and generating a word sense mapping paragraph alignment result; S5: Based on the paragraph alignment result of the word sense mapping, the word sense number and position of the keywords in the paragraph are marked, the reading area display structure is organized, the word sense interpretation, document source and semantic path information are output, and the word sense comparison prompt block is generated.
2. The online query and reading method for the Chinese language and literature database according to claim 1 is characterized in that: The semantic candidate set includes a list of candidate word senses, a table of word sense feature parameters, a selection threshold identifier, and a score summary; the context semantic mapping path sequence includes a path sequence number set, a relationship strength level set, a subject-verb collocation snapshot set, and a meaning subject mapping tag set; the version paragraph retrieval scheduling table includes a version identification code, a section number list, a document location coordinate, and a retrieval routing item; the word sense mapping paragraph alignment result includes an associated phrase structure table, a position sequence sequence, a structure sequence offset result, and a position adjustment record; the word sense comparison prompt block includes a word sense number marking area, a keyword position index, a document source entry, a path prompt information card, and a reading area display layout.
3. The online query and reading method for the Chinese language and literature database according to claim 1 is characterized in that: The specific steps of S1 are: S101: Obtain the meaning of the search term in the Chinese language and literature database, collect the citation frequency of the meaning in the core ancient books and documents, the length of the interpretation text and the number of meaning boundary marks, organize the data into a unified format, and generate the document meaning parameters; S102: Based on the document word meaning parameters, the document time sequence weight is used to process the citation frequency of word meaning items in ancient documents of different ages, and the length of the interpretation text and the number of boundary marks of the meaning item are compared to generate a word meaning index; S103: Screening the word senses according to the word sense index, retaining a set of word senses that meet the conditions, and arranging the word senses in the set according to a predetermined sorting standard to generate a semantic candidate set.
4. The online query and reading method for the Chinese language and literature database according to claim 3 is characterized in that: The specific steps of S2 are: S201: Obtain word senses in the semantic candidate set, extract the subject pronoun and verb collocation structure of the paragraph where the word senses are located, organize the subject pronouns and verbs into combination records, and generate paragraph collocation combination quantity; S202: Based on the paragraph collocation combination quantity, calling the semantic subject and action verb in the word meaning interpretation, comparing the subject pronoun and verb collocation structure in the paragraph, determining the corresponding relationship and recording the strength result to obtain the semantic association strength value; S203: Sort the paths corresponding to the word meanings according to the semantic association strength values, merge the path results according to the paragraph positions, and generate a context semantic mapping path sequence.
5. The online query and reading method of the Chinese language and literature database according to claim 4 is characterized in that: The specific steps of S3 are: S301: Obtaining a priority path in the context semantic mapping path sequence, extracting word senses and chapter information under the path, detecting the paragraph position corresponding to the word sense in the version index, and recording the paragraph number and index mark to generate a paragraph index number; S302: Based on the paragraph index number, the paragraph number and document position containing the target word meaning are retrieved, the corresponding relationship between the paragraph number and the position data is determined, and a combination structure that can uniquely identify the paragraph is selected to obtain a paragraph positioning matching value; S303: According to the paragraph positioning matching value, a retrieval structure of paragraph number and document position is established in the version index, all combination relationships are sequentially merged into a search path, and a version paragraph search scheduling table is generated.
6. The online query and reading method for the Chinese language and literature database according to claim 5 is characterized in that: The specific steps of S4 are: S401: Calling the paragraph number and document position in the version paragraph search scheduling table, extracting the associated phrase structure and appearance order of the target word meaning in the paragraph, comparing the number and position relationship between adjacent phrases, and generating phrase sequence annotations; S402: Based on the phrase sequence annotation amount, calculate the deviation degree of the structural sequence in the difference version, compare the paragraph number difference with the set sequence deviation standard value one by one, obtain the paragraph set that exceeds the standard, and generate the sequence deviation difference value; S403: According to the sequence offset difference value, adjust the position of the paragraphs that exceed the standard, call the number and document position to rearrange the paragraph order, map all adjusted paragraph positions to corresponding word sense combinations, and generate a word sense mapping paragraph alignment result.
7. The online query and reading method for the Chinese language and literature database according to claim 6 is characterized in that: The specific steps of S5 are: S501: Based on the paragraph alignment result of the word sense mapping, annotate the word sense number and the position index within the paragraph corresponding to the keyword paragraph by paragraph, call the number of occurrences and position difference of the keyword in the paragraph, calculate the combination of the number and the position index, and generate the word sense position annotation amount; S502: Sort the annotated word meaning numbers in the paragraph according to the word meaning position annotation quantity, match the position index of the keyword corresponding to the number in the reading area with the paragraph structure, calculate the sorting difference and compare it with the set display reference value to obtain the reading area sequence difference; S503: Calling the reading zone sequence difference, reorganizing the paragraphs corresponding to the difference exceeding the display benchmark value, outputting the word meaning interpretation, document source and path information in paragraph order, combining the content into prompt blocks, and generating a word meaning comparison prompt block.
8. The online query and reading method for the Chinese language and literature database according to claim 1 is characterized in that: The core ancient books and documents refer to the representative historical documents in the field of Chinese language and literature, including the classic Han Dynasty documents such as Shangshu, Zuo Zhuan, and Shiji; The document time sequence weight refers to the weighted value assigned to the meaning of words in different documents according to the age, genre and influence of the document in language evolution; The semantic candidate set refers to a set of all candidate word meanings screened by multi-dimensional parameters; The subject pronoun refers to the pronoun that represents the subject in the sentence; The verb collocation structure refers to the grammatical structure consisting of a verb and an associated noun, adjective or other modifying element; The context semantic mapping path sequence refers to a path set generated by sorting the semantic relationships in the context.
9. The online query and reading method for the Chinese language and literature database according to claim 1, characterized in that: The version index is an index including different document versions, paragraph positions and chapter structures; The section number and document position are identifiers in the document; The version paragraph search scheduling table refers to the search order and retrieval structure generated based on the different version documents and target paragraphs during the query process; The word meaning item related phrase structure refers to a phrase, word group or sentence structure that is closely related to the target word meaning item in the document paragraph, and the structure is consistent with the semantics of the word meaning item; The structural order deviation refers to the order difference of phrases or word groups in sentences in different versions of documents. The order is sorted and adjusted by calculating the structural order difference between different versions. The paragraph alignment result of the word sense mapping refers to the paragraph alignment result generated by adjusting the structural order between versions; The word meaning comparison prompt block refers to the word meaning comparison framework displayed during the user's reading process, including word meaning interpretation, document source and path information.
10. A Chinese language and literature database online query and reading system, used to implement the Chinese language and literature database online query and reading method according to any one of claims 1 to 9, characterized in that: include: The word meaning candidate module obtains the citation frequency, meaning length, meaning boundary marking number and document time sequence weight of the search terms in the core ancient books and documents, and sets the screening threshold to screen the shortlisted word meanings and generate a semantic candidate set; The semantic path module extracts subject pronouns and verb collocations in the corresponding paragraphs based on the word meanings in the semantic candidate set, calculates the matching strength of the semantic subject and the action verb and sorts them, and generates a contextual semantic mapping path sequence; The version positioning module locates the section number and document position in the version index and records the coordinates based on the priority path in the context semantic mapping path sequence, combined with the word meaning and the chapter information, and generates a version paragraph retrieval scheduling table; The paragraph alignment module calls the version paragraph search scheduling table, extracts the phrase structure and position order of the target word sense, calculates the order offset between the different versions, and adjusts the paragraphs that exceed the threshold, thereby generating a word sense mapping paragraph alignment result; The reading prompt module marks the keyword meaning number and position in the reading area based on the paragraph alignment result of the word meaning mapping, organizes the display of document source and path information, and generates a word meaning comparison prompt block.
Citation Information
Patent Citations
Method for automatically creating keyword index table
CN103064969A
Big data analysis-based legal history document classified storage method
CN120596589A
Digital literature visual generation and interaction device based on Chinese language resources
CN120610640A
Method and system for semantic search and retrieval of electronic documents
US20060235843A1