Digital archive arrangement system based on semantic analysis

The semantic analysis-based digital archive organization system solves the problem of identifying field hierarchy and semantic associations in digital archives, enabling efficient document structure understanding and archiving review, and generating accurate document summaries and tags.

CN121301567AInactive Publication Date: 2026-01-09GUANGDONG XINGZHI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511476389.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to identify hierarchical relationships between fields when dealing with digital archives that are complex in format and varied in structure. This leads to inaccurate information classification, fragmentation of semantically related information, and affects the accuracy and efficiency of document content during archiving, review, and reuse.

Method used

A digital archive organization system based on semantic analysis is adopted. Through a hierarchical extraction module, a semantic evaluation module, a co-occurrence phrase identification module, and a word combination merging module, the system accurately determines the nesting level and structural position of fields in the document, filters out keywords with high semantic credibility, merges redundant or semantically repetitive keyword combinations, calculates semantic weights, and generates semantic organization results.

Benefits of technology

It achieves accurate identification of semantically related content, ensures the semantic independence and representativeness of keyword groups, improves the accuracy of document structure understanding and the efficiency of archiving and review, and generates efficient document structure summaries and tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301567A_ABST
    Figure CN121301567A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data arrangement, in particular to a digital archive arrangement system based on semantic analysis, which comprises a hierarchical extraction module, a semantic evaluation module, a co-occurrence phrase recognition module, a phrase merging module and an arrangement output module. According to the method, numerical combination and hierarchical mapping are carried out on the structural features of the first row of the field, the nesting hierarchy and the structural position of the field in the document can be accurately determined, confusion interference to different semantic hierarchies is avoided, keywords with high semantic credibility are screened out in a word segmentation statistics and field structure linkage mode, and the semantic reliability of the keywords is improved. Dynamic recognition and accurate extraction of information density and semantic weight are achieved, co-occurrence structure extraction of keywords in cross-field context is combined, the recognition capability of semantic association content is enhanced, vectorization and similarity measurement are carried out on the word order, context and appearing field features of high-frequency keywords, and the semantic association content recognition efficiency is improved. Redundant or semantically repeated keyword combinations are effectively combined, and semantic representativeness of extracted phrases is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data organization technology, and in particular to a digital archive organization system based on semantic analysis. Background Technology

[0002] The field of data processing technology involves classifying, cataloging, integrating, and archiving multi-source heterogeneous information. Its core aspects include data collection, content parsing, format standardization, information extraction, classification and database construction, and indexing.

[0003] Among them, the digital archive organization system refers to a systematic tool for collecting, classifying, cataloging and indexing digital archive data, and uniformly classifying and identifying scattered electronic documents so that they can quickly complete the organization process according to preset rules.

[0004] Because existing technologies rely solely on collecting, classifying, cataloging, and indexing the basic fields of electronic documents, they struggle to effectively identify hierarchical relationships between fields when faced with large amounts of digital archives with complex formats and varied structures. In documents with chaotic or ambiguous field structures, inaccurate information classification can easily occur. For example, fields with multiple semantic levels that are misclassified as belonging to the same category level will render subsequent search tags ineffective. Cataloging based solely on surface field keywords without considering their semantic context leads to the fragmentation of information related to fields, making it impossible to identify semantic clusters of content distributed across different fields. This is especially true when dealing with repetitive or nested records across multiple fields or cross-paragraph themes, as the lack of understanding of semantic connections within the context can result in isolated keywords or redundant classifications, thus affecting the accuracy and efficiency of document content during archiving, review, and reuse. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a digital archive organization system based on semantic analysis.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a digital archive organization system based on semantic analysis, the system comprising: The hierarchical extraction module detects the structural features of the first line of each field in the digital archive and converts the structural features into the structural hierarchy value of the field in the document, which serves as the field's structural location information. The semantic evaluation module segments the text content within all fields of the digital archive, extracts words whose frequency exceeds a set threshold as a keyword set, extracts the set of field numbers where the keyword set is located, and calculates the semantic confidence of each keyword by combining the field structure position information of each field, and filters the effective keyword semantic set. The co-occurrence phrase identification module analyzes the sentence structure of the effective keywords in the semantic set of effective keywords within each field, extracts the keyword combinations that exist simultaneously in multiple fields, and integrates them into a candidate co-occurrence phrase set; The word combination merging module performs vectorization processing on the candidate co-occurring word group set, calculates the semantic vector similarity between each pair of co-occurring word groups, and merges similar co-occurring word groups to construct a merged word group set. The output module calculates the frequency of occurrence of the merged phrase set in the entire field, assigns a semantic weight to each merged phrase in the field based on the frequency of occurrence, and obtains the semantic processing result of the digital archive.

[0007] The present invention improves upon the following: the field structure location information specifically includes a structure hierarchy identifier value, a field number index, and a relative nesting depth; the keyword semantic set includes semantic tags, keyword field mapping relationships, and keyword representativeness scores; the candidate co-occurring phrase set specifically includes a multi-field keyword combination structure, keyword co-occurrence paths, and structural order encoding; the merged phrase set includes a unified phrase expression form, a semantic vector association index, and a phrase combination mapping table; and the digital archive semantic organization result includes field semantic tags, phrase weight distribution, and a structured keyword index.

[0008] The present invention is improved in that the structural features of the first line of each field include the number of spaces, the number of tabs, or the number of punctuation marks in the paragraph numbering format of the first line of each field.

[0009] The present invention is improved in that the hierarchical extraction module includes: The structural feature extraction submodule detects the structural features on the left side of the first row of each field in the digital archive and performs numerical combination to establish the structural feature combination of the field; The hierarchical interval division submodule counts the numerical distribution characteristics of the structural feature combination of all the fields, sets the dividing point according to the density of the values, divides the continuous numerical distribution interval into multiple non-overlapping hierarchical segments, configures a unique hierarchical identifier for each hierarchical segment, and establishes a structural hierarchical mapping benchmark. The structure hierarchy conversion submodule calls the structural feature combination of a single field, and determines the hierarchy segment to which the structural feature combination value belongs based on the structure hierarchy mapping benchmark, assigns the corresponding hierarchy identifier to the hierarchy segment, and generates field structure position information.

[0010] The present invention is improved in that the semantic evaluation module includes: The keyword preliminary extraction submodule obtains the internal text content of all fields in the digital archive and performs word segmentation, counts the frequency of each word and extracts the set of field numbers, compares the frequency of each word with the set frequency threshold, filters out words that exceed the frequency threshold, and establishes a multi-frequency keyword set. The semantic confidence calculation submodule obtains the set of multi-frequency keywords and the field structure position information of each field, combines the number of times the keyword appears in the field with the field structure position information of the corresponding field, and inputs it as a feature into the support vector regression algorithm to calculate the semantic confidence of each keyword. The keyword semantic filtering submodule compares the semantic confidence of each keyword with a set keyword confidence threshold, and filters keywords whose semantic confidence is not lower than the confidence threshold to obtain an effective keyword semantic set.

[0011] The present invention is improved in that the co-occurrence phrase identification module includes: The context structure analysis submodule analyzes the sentence structure of the effective keywords in the semantic set within each field, extracts the order relationship, connection method, and modification relationship position between the keywords and adjacent words, and establishes the context structure features of the keywords. The cross-field combination extraction submodule calls the contextual structure features of each keyword, starting with a single valid keyword, identifies the continuous occurrence order of valid keywords in multiple fields, extracts keyword combinations that exist simultaneously in multiple fields, and obtains multi-field keyword combinations. The candidate word group generation submodule summarizes and merges all multi-field keyword combinations that exist simultaneously in multiple fields, eliminates duplicate combinations, and establishes a candidate co-occurring word group set.

[0012] The present invention is improved in that the word combination and merging module includes: The phrase vectorization submodule extracts the word order position, association frequency and field number of the keywords in each phrase in the candidate co-occurring phrase set as features, and constructs a corresponding vector expression for each co-occurring phrase, thus establishing a co-occurring phrase feature vector. The vector similarity calculation submodule combines the vector representations of all the co-occurring word group feature vectors in pairs and calculates the similarity between the two vectors in each pair as a measure of semantic association between word groups, thereby obtaining semantic vector similarity information. The similar word combination and merging submodule calls the semantic vector similarity information of each pair of co-occurring word groups, compares the similarity with the set path normalization threshold, filters out co-occurring word group pairs whose similarity reaches the path normalization threshold condition and merges them to obtain a merged word group set.

[0013] The present invention is improved in that the sorting and output module includes: The phrase frequency statistics submodule counts the frequency of each merged phrase in the merged phrase set in the field, and obtains the frequency information of the phrase field. The semantic weight calculation submodule obtains the frequency information of the word group field and the field structure position information of the corresponding field, uses the field structure position information as a semantic influence factor, and combines it with the frequency of the word group field to calculate the semantic weight value of each merged word group in the field using the TF-IDF algorithm. The semantic result integration submodule calls each merged phrase and its corresponding semantic weight value, assigns semantic weights to each merged phrase within the fields, and generates digital archive semantic organization results, which are used to establish document structure summaries, archive tags, or assist human reviewers in archiving and review.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by numerically combining and hierarchically mapping the structural features of the first row of a field, the nesting level and structural position of the field in the document can be accurately determined, avoiding confusion and interference between different semantic levels. By linking word segmentation statistics with field structure, keywords with high semantic credibility are selected, achieving dynamic identification and accurate extraction of information density and semantic weight. Combined with the extraction of co-occurrence structure of keywords in cross-field context, the ability to identify semantically related content is enhanced. By vectorizing the word order, context, and features of the fields in which high-frequency keywords appear and performing similarity measurement, redundant or semantically repetitive keyword combinations are effectively merged, ensuring the semantic independence and representativeness of the finally extracted phrases. Based on frequency, a hierarchical influence factor is introduced for weight calculation, so that keywords appearing in high semantic levels receive higher semantic weights. This achieves the unity between semantic distribution, field structure, and actual document intent in the semantic extraction results, making the semantic processing results usable for efficiently assisting in document structure understanding, summary generation, and archiving review judgment. Attached Figure Description

[0015] Figure 1 This is a system module diagram of the present invention; Figure 2 This is a system framework diagram of the present invention; Figure 3 This is a schematic diagram of the hierarchical extraction module of the present invention; Figure 4 This is a schematic diagram of the semantic evaluation module of the present invention; Figure 5 This is a schematic diagram of the co-occurrence phrase recognition module of the present invention; Figure 6 This is a schematic diagram of the word combination and merging module of the present invention; Figure 7 This is a schematic diagram of the output module of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] Please see Figure 1 This invention provides a technical solution: a digital archive organization system based on semantic analysis, the system comprising: The hierarchical extraction module detects the structural features of the first line of each field in the digital archive and converts the structural features into the structural hierarchy value of the field in the document, which serves as the field's structural location information. The semantic evaluation module segments the text content within all fields of the digital archive, extracts words whose frequency exceeds a set threshold as a keyword set, extracts the set of field numbers where the keyword set is located, and calculates the semantic confidence of each keyword by combining the field structure position information of each field, and filters the effective keyword semantic set. The co-occurrence phrase identification module analyzes the sentence structure of effective keywords within the context of each field in the semantic set of effective keywords, extracts keyword combinations that exist simultaneously in multiple fields, and integrates them into a candidate co-occurrence phrase set; The word combination merging module vectorizes the candidate co-occurring word group set, calculates the semantic vector similarity between each pair of co-occurring word groups, and merges similar co-occurring word groups to construct a merged word group set. The output module calculates the frequency of occurrence of the merged phrase set in the entire field, assigns a semantic weight to each merged phrase in the field based on the frequency of occurrence, and obtains the semantic organization result of the digital archive. The field structure location information specifically includes the structure level identifier value, field number index, and relative nesting depth. The keyword semantic set includes semantic tags, keyword field mapping relationships, and keyword representativeness scores. The candidate co-occurring phrase set specifically includes the multi-field keyword combination structure, keyword co-occurrence path, and structural order encoding. The merged phrase set includes the unified expression form of phrases, semantic vector association index, and phrase combination mapping table. The semantic organization results of digital archives include field semantic tags, phrase weight distribution, and structured keyword index.

[0018] The structural characteristics of the first line of each field include the number of spaces, the number of tabs, or the number of punctuation marks in the paragraph numbering format.

[0019] Please see Figure 2 and Figure 3 The hierarchical extraction module includes: The structural feature extraction submodule detects the structural features on the left side of the first row of each field in the digital archive and performs numerical combination to establish the structural feature combination of the field; First, a digitized "Project Establishment Approval Report" file is read. This file contains one hundred fields, such as field 2: "1.1 Current Status of Domestic Research". This submodule detects the blank area on the left side of the first line and the paragraph number. Specifically, it detects that there are two half-width spaces in the first line of field 2, no tabs, and the paragraph number "1.1" contains one period. Then, these structural features are numerically combined by multiplying the number of spaces and the number of tabs by a fixed conversion coefficient of 4 and the number of periods, and then summing them. The conversion coefficient for tabs is set to 4 because, in a regular text editing environment, the display width of one tab is equivalent to 4 spaces. This coefficient is used to ensure that its contribution to the structural depth calculation is consistent with the equivalent number of spaces. Therefore, the structural feature combination value of field 2 is 2 plus 0 multiplied by 4 plus 1, resulting in 3. By performing the same detection and numerical combination operation on all one hundred fields in the file, a structural feature combination value is generated for each field.

[0020] The hierarchical interval division submodule counts the numerical distribution characteristics of the structural feature combination of all fields, sets the boundary points according to the density of the values, divides the continuous numerical distribution interval into multiple non-overlapping hierarchical segments, and configures a unique hierarchical identifier for each hierarchical segment to establish a structural hierarchical mapping benchmark. Next, the one hundred structural feature combination values ​​generated in the previous steps were obtained, and the distribution characteristics of these values ​​were statistically analyzed. Specifically, by drawing a histogram of the value distribution, it was found that the values ​​mainly clustered around the three value points 1, 3, and 5, forming three independent value clusters that were densely packed internally but sparsely spaced externally. This clearly indicates that the document structure has three significant levels. The boundary points were set according to the gaps between the value clusters. The method of setting the boundary was to take the arithmetic mean between the center points of two adjacent value clusters. This was based on the fact that the arithmetic mean is the boundary between two data points without any gaps. The off-center position maximizes the distance from the dividing point to the two cluster centers, thus maximizing the fault tolerance of the hierarchical division. Specifically, the first dividing point is set at point 2, the midpoint between values ​​1 and 3, and the second dividing point is set at point 4, the midpoint between values ​​3 and 5. This divides the continuous numerical interval into three non-overlapping hierarchical segments: [1,2], [3,4], and [5, positive infinity]. Each of these three hierarchical segments is assigned a unique hierarchical identifier, namely "Level 1", "Level 2", and "Level 3", respectively. This method establishes the structural hierarchical mapping benchmark.

[0021] The structure hierarchy conversion submodule calls the structure feature combination of a single field, and determines the level segment to which the structure feature combination value belongs based on the structure hierarchy mapping benchmark, assigns the level identifier corresponding to the level segment, and generates the field structure position information. The system retrieves the structural feature combination value of a single field and, based on the previously established structural hierarchy mapping benchmark, determines the hierarchical segment to which the combination value belongs, and then assigns it the corresponding hierarchical identifier. For example, if the structural feature combination value 3 of field 2 is retrieved, it is determined that the value falls within the hierarchical segment interval [3,4], and therefore the field is assigned the hierarchical identifier "Level 2". For another field, if its structural feature combination value is 5, it is determined that it falls within the hierarchical segment interval [5, positive infinity), and it is assigned the hierarchical identifier "Level 3". By performing this judgment and assignment operation on all fields, a unique field structure position information is generated for each field in the file. This information accurately reflects its depth and position in the overall document structure.

[0022] Please see Figure 2 and Figure 4 The semantic evaluation module includes: The keyword preliminary extraction submodule obtains the internal text content of all fields in the digital archive and performs word segmentation, counts the frequency of each word and extracts the set of field numbers, compares the frequency of each word with the set frequency threshold, filters out words that exceed the frequency threshold, and establishes a multi-frequency keyword set. The process involves retrieving the internal text content of all 100 fields from a digital archive, such as a "Project Approval Report," and performing word segmentation using a pre-built financial dictionary. The cumulative frequency of each word across all fields is then calculated and compared to a preset frequency threshold. This threshold is based on a statistical analysis of the word frequency distribution of 1,000 similar "Project Approval Report" archives. The analysis results show a long-tail distribution, meaning a small number of words occupy the vast majority of occurrences. These high-frequency words are typically the core concepts of the document. Selecting the top 5% of words as candidate keywords ensures that the majority of words are retrieved. The core concept and the best practice of balancing the size of the candidate set to reduce subsequent computational complexity are to set the frequency threshold to the minimum frequency value that can just filter out the top 5% of high-frequency words. In this embodiment, assuming that the total number of words in the file is 5,000 and there are 800 unique words, the 5% filter needs to retain the 40 most frequent words. If the 40th most frequent word "budget" appears 15 times, then the frequency threshold is set to 15. All words with a frequency of not less than 15 times are filtered out, and the set of numbers of all fields in which they appear is recorded to establish a multi-frequency keyword set.

[0023] The semantic confidence calculation submodule obtains a set of multi-frequency keywords and the field structure and position information of each field. It combines the frequency of occurrence of keywords in the field with the field structure and position information of the corresponding field, and inputs it as a feature into the support vector regression algorithm to calculate the semantic confidence of each keyword. Obtain the set of high-frequency keywords and the field structure and position information of each field. Perform a normalized combination operation on the frequency of a keyword in a specific field and the value corresponding to the hierarchical identifier of that field. This operation follows the semantic confidence formula: ,in, It is the calculation target, representing keywords. The final semantic confidence score; a higher score indicates greater importance of the keyword within the entire document; summation symbol. Indicates keywords It iterates through all fields that have appeared and sums up the confidence scores of each field to determine its overall importance throughout the document; Frequency weighting is based on the assumption that the repetition of keywords in a local text is a direct reflection of their importance; the more times they appear, the higher their importance, hence the weighting. Keywords In the field The original number of occurrences in; It is the maximum number of times all keywords appear in any single field. Divide by This involves normalizing the frequency so that its value falls within the [0,1] range, thereby eliminating the impact of differences in field length. The hierarchical weighting is based on the fact that the position of a keyword in the document structure determines its macro-level importance. Words appearing in higher-level positions such as headings are more important than words appearing deep within body paragraphs, hence they are assigned a weight. and The values ​​of 0.6 and 0.4 were determined through regression analysis of fifty files whose importance had been assigned by experts. These values ​​resulted in the frequency contributing slightly more than the structural location, which is consistent with conventional understanding. It is a field The numerical value corresponding to the level identifier; It is the highest level value in the file, expressed as: Used to reverse the level number, so that level 1 has the highest score, level The lowest score, then divide by Normalization is performed so that the values ​​also fall within the range of [0,1], ensuring that the frequency and level components have consistent dimensions when added together. Taking the keyword "budget" as an example, assuming it appears in field 4 (3 times, level 1) and field 25 (5 times, level 2), and the highest level of the file is... The maximum number of times all keywords appear in a single field is 3. If the value is 5, then the calculation process for its semantic confidence is as follows: The result indicates that the semantic confidence of the keyword "budget" is 1.627.

[0024] The keyword semantic filtering submodule compares the semantic confidence of each keyword with the set keyword confidence threshold, and filters keywords with a semantic confidence of not lower than the confidence threshold to obtain an effective keyword semantic set. The semantic confidence score calculated for each keyword is compared with a set keyword confidence threshold. This threshold is set based on the following: from fifty training files pre-calibrated by experts, all keywords deemed "important" by the experts are used to calculate their confidence scores using the aforementioned confidence formula. The minimum of these scores is then taken as the final confidence threshold. The logic of this method is to ensure that any important keyword previously recognized by experts can pass this threshold screening, thereby maximizing recall while maintaining precision. If this value is calculated to be 1.50, the confidence score of the keyword "budget" (1.627) is compared with 1.50. Since 1.627 is not lower than 1.50, "budget" is selected as a valid keyword. This screening process is then performed on all words in the multi-frequency keyword set to ultimately form a set of valid keyword semantics.

[0025] Please see Figure 2 and Figure 5 The co-occurrence phrase recognition module includes: The context structure analysis submodule analyzes the sentence structure of effective keywords within the context of each field, extracts the order relationship, connection method, and modification relationship position between keywords and adjacent words, and establishes the context structure features of keywords. After obtaining the set of effective keyword semantics, for each effective keyword, we analyze its contextual sentence structure within each field. For example, in the field "4.1 Project Budget Approval Process", for the keyword "approval", we extract the word immediately to its left as "budget" and the word immediately to its right as "process", and record the sequential relationship of "budget-approval-process" and the way they are directly connected to each other, thus establishing a contextual structure feature of the keyword "approval" in field 4.1. This process traverses all effective keywords in all fields, establishing a set of detailed contextual structure features for each keyword.

[0026] The cross-field combination extraction submodule calls the contextual structure features of each keyword, takes a single valid keyword as the starting point, identifies the continuous occurrence order of valid keywords in multiple fields, extracts the keyword combinations that exist in multiple fields at the same time, and obtains multi-field keyword combinations. The contextual structure features of all effective keywords established in the preceding steps are invoked, and a single effective keyword is used as the starting point for retrieval. The continuous occurrence order of effective keywords is identified and tracked between different fields. Specifically, the word "project" is identified to appear in the field "1. Project Background" at level 1, while the word "budget" appears in the field "2.1 Budget Preparation" at the subsequent level 2, and the word "approval" appears in the field "2.1.1 Budget Approval Principles" at the deeper level 3. Since these three keywords present a top-down progressive relationship in the macro structure of the document, a keyword combination spanning multiple fields is extracted, namely (project, budget, approval), resulting in a multi-field keyword combination.

[0027] The candidate word group generation submodule summarizes and merges all multi-field keyword combinations that exist simultaneously in multiple fields, eliminates duplicate combinations, and establishes a candidate co-occurrence word group set. All multi-field keyword combinations identified by the cross-field combination extraction submodule, such as (project, budget), (project, approval), (budget, approval), (project, budget, approval), are summarized and merged. The exact matching of strings is used to determine whether two combinations are completely identical. If they are identical, only one is kept. In this way, combinations with completely duplicate content are eliminated, and finally a set of candidate co-occurring word groups without duplicates is established.

[0028] Please see Figure 2 and Figure 6 The word combination module includes: The phrase vectorization submodule extracts the word order position, association frequency, and field number of keywords in each phrase in the candidate co-occurring phrase set as features, and constructs a corresponding vector representation for each co-occurring phrase, thus establishing a co-occurring phrase feature vector. First, assign a unique numerical ID to each valid keyword in the dictionary, for example, "project" is 101, "budget" is 205, and "funds" is 208. Then, extract the internal features of each phrase in the candidate co-occurrence phrase set. Taking the phrase (project, budget) as an example, extract the sequence [101, 205] composed of the IDs of its internal keywords. The association frequency of this phrase in the entire file, that is, the number of times the two words appear together in the same field is 5, and the set of field numbers where the phrase appears together {4, 28, 71}. Based on these features, construct a fixed-length numerical feature vector for this phrase, for example, by expanding and padding the set of field numbers to a fixed length, and finally forming a comprehensive feature vector.

[0029] The vector similarity calculation submodule combines the vector representations of all co-occurring word phrase feature vectors in pairs and calculates the similarity between the two vectors in each pair as a measure of semantic association between word phrases, thus obtaining semantic vector similarity information. The feature vectors of all co-occurring word groups are paired, and the cosine similarity is calculated for each pair of vectors, specifically for the vectors of word groups (project, budget). Vectors of phrases (projects, funds) The calculation process involves summing the element-wise products of the two vectors and then dividing by the product of the Euclidean norms of the two vectors to obtain a similarity value between 0 and 1. For example, the result is 0.88. This value is used as a quantitative measure of the semantic association between the two word pairs. By performing this calculation on all word pairs, a complete semantic vector similarity information table is obtained.

[0030] The similar word combination and merging submodule calls the semantic vector similarity information of each pair of co-occurring word groups, compares the similarity with the set path normalization threshold, filters out co-occurring word group pairs whose similarity reaches the path normalization threshold condition and merges them to obtain a merged word group set; The semantic vector similarity information of each co-occurring word pair is calculated and compared with a set path normalization threshold. This threshold is set based on hierarchical clustering experiments on a large number of test datasets, and the silhouette coefficient is calculated under different thresholds. The silhouette coefficient is an indicator of clustering effect. The larger the value, the better the cohesion and separation of the clustering results. The similarity value that maximizes the average silhouette coefficient is selected as the final threshold to ensure that the merged word pairs are semantically the most cohesive and reasonable. If the experiment finally determines that the threshold is 0.85, the similarity of the word pair (project, budget) and (project, funds) is compared between 0.88 and 0.85. Since 0.88 is not lower than the threshold condition, the two co-occurring word pairs are merged. The specific rule for merging is to create a new word pair containing the two original word pairs, denoted as "project-budget / funds". All word pairs with similarity reaching the threshold condition are processed in this way to obtain the final merged word pair set.

[0031] Please see Figure 2 and Figure 7 The output processing module includes: The phrase frequency statistics submodule counts the frequency of each merged phrase in the merged phrase set in the field, and obtains the frequency information of the phrase field. First, iterate through the set of merged phrases obtained in the previous steps and count the frequency of each merged phrase, such as "project-budget / funds", in each field of the file. Specifically, check the text content of field 4 and find that "project" and "budget" appear together twice. Therefore, the frequency of "project-budget / funds" in field 4 is recorded as 2 times. By iterating through all fields in this way, the frequency information of each merged phrase in each field can be obtained.

[0032] The semantic weight calculation submodule obtains the frequency information of the phrase field and the field structure position information of the corresponding field. It uses the field structure position information as a semantic influence factor and combines it with the frequency of the phrase field. It then calculates the semantic weight value of each merged phrase in the field using the TF-IDF algorithm. The frequency of occurrence of phrases within a field and the corresponding field structure position information are obtained. An improved TF-IDF algorithm is then used to calculate the semantic weight of each merged phrase within a specific field. The specific calculation follows the formula: ,in, It is the calculation target, representing the merged phrases. In a specific field The final semantic weight within; the part within parentheses It's the term frequency (TF) part. It is a phrase In the field The number of times it appears in the denominator It is a field The total number of all phrases appearing in the text is used to normalize the word frequency by dividing the total number of occurrences by the total number of occurrences. This normalization is based on the assumption that the higher the proportion of a phrase in a specific field, the greater its semantic contribution to that field. The logarithmic part... It is the inverse document frequency (IDF) part. It is the total number of fields in the file. The entire file contains phrases The total number of fields is set based on the principle that the fewer fields a phrase appears in, the stronger its uniqueness and representativeness, and the higher its weight should be. The logarithmic function is used to smooth out this effect. It is based on field hierarchy The calculated semantic influence factor is specifically: The basis for this setting is that the higher the structural hierarchy of the field ( The smaller the value, the stronger its macroscopic importance; therefore, it should be given a higher weight. Using the reciprocal form directly reflects this inverse relationship. The innovation of this formula lies in quantifying structural information into factors. The weights are then multiplied and calculated, resulting in a final weight that integrates the local importance, global rarity, and structural position importance of the phrase. Taking the calculation of the semantic weight of the merged phrase "project-budget / funds" in field 4 (level 1) as an example, it is known that... Total frequency of all phrases in field 4 The total number of fields including "Project-Budget / Funding" Total number of fields The hierarchy of field 4 First, calculate the semantic influence factor. Then substitute the values ​​into the complete formula to perform the calculation: Calculate using the natural logarithm The value is approximately 3.219, so the final weight is... The result indicates that the semantic weight of the combined phrase “project-budget / funding” in field 4 is 1.2876.

[0033] The semantic results integration submodule calls each merged phrase and its corresponding semantic weight value, assigns semantic weights to each merged phrase within the fields, and generates digital archive semantic organization results, which are used to create document structure summaries, archive tags, or assist human reviewers in archiving and reviewing. The system retrieves all calculated merged phrases and their corresponding semantic weight values ​​in each field. Based on these weight values, each merged phrase is assigned a weight within the corresponding field to generate the final semantic organization result of the digital archive. Specifically, each calculated result with a weight value greater than a preset output threshold is combined into a triple containing "merged phrase", "field number", and "semantic weight value". The output threshold is set based on retaining the top 20% of the weighted results, aiming to filter out a large amount of low-weight background information and output only the organization result that best represents the core semantics of the archive. For example, if 1.2876 ranks in the top 20% of all calculated weight values, it forms a triple ("Project-Budget / Funds", 4, 1.2876). All the triples that pass this threshold are merged into a list as the final output.

[0034] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A digital archive organization system based on semantic analysis, characterized in that: The system includes: The hierarchical extraction module detects the structural features of the first line of each field in the digital archive and converts the structural features into the structural hierarchy value of the field in the document, which serves as the field's structural location information. The semantic evaluation module segments the text content within all fields of the digital archive, extracts words whose frequency exceeds a set threshold as a keyword set, extracts the set of field numbers where the keyword set is located, and calculates the semantic confidence of each keyword by combining the field structure position information of each field, and filters the effective keyword semantic set. The co-occurrence phrase identification module analyzes the sentence structure of the effective keywords in the semantic set of effective keywords within each field, extracts the keyword combinations that exist simultaneously in multiple fields, and integrates them into a candidate co-occurrence phrase set; The word combination merging module performs vectorization processing on the candidate co-occurring word group set, calculates the semantic vector similarity between each pair of co-occurring word groups, and merges similar co-occurring word groups to construct a merged word group set. The output module calculates the frequency of occurrence of the merged phrase set in the entire field, assigns a semantic weight to each merged phrase in the field based on the frequency of occurrence, and obtains the semantic processing result of the digital archive.

2. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The field structure location information specifically includes structure hierarchy identifier value, field number index, and relative nesting depth. The keyword semantic set includes semantic tags, keyword field mapping relationship, and keyword representativeness score. The candidate co-occurring phrase set specifically includes multi-field keyword combination structure, keyword co-occurrence path, and structural order encoding. The merged phrase set includes unified phrase expression form, semantic vector association index, and phrase combination mapping table. The digital archive semantic organization result includes field semantic tags, phrase weight distribution, and structured keyword index.

3. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The structural features of the first line of each field include the number of spaces, the number of tabs, or the number of periods in the paragraph numbering format.

4. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The hierarchical extraction module includes: The structural feature extraction submodule detects the structural features on the left side of the first row of each field in the digital archive and performs numerical combination to establish the structural feature combination of the field; The hierarchical interval division submodule counts the numerical distribution characteristics of the structural feature combination of all the fields, sets the dividing point according to the density of the values, divides the continuous numerical distribution interval into multiple non-overlapping hierarchical segments, configures a unique hierarchical identifier for each hierarchical segment, and establishes a structural hierarchical mapping benchmark. The structure hierarchy conversion submodule calls the structural feature combination of a single field, and determines the hierarchy segment to which the structural feature combination value belongs based on the structure hierarchy mapping benchmark, assigns the corresponding hierarchy identifier to the hierarchy segment, and generates field structure position information.

5. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The semantic evaluation module includes: The keyword preliminary extraction submodule obtains the internal text content of all fields in the digital archive and performs word segmentation, counts the frequency of each word and extracts the set of field numbers, compares the frequency of each word with the set frequency threshold, filters out words that exceed the frequency threshold, and establishes a multi-frequency keyword set. The semantic confidence calculation submodule obtains the set of multi-frequency keywords and the field structure position information of each field, combines the number of times the keyword appears in the field with the field structure position information of the corresponding field, and inputs it as a feature into the support vector regression algorithm to calculate the semantic confidence of each keyword. The keyword semantic filtering submodule compares the semantic confidence of each keyword with a set keyword confidence threshold, and filters keywords whose semantic confidence is not lower than the confidence threshold to obtain an effective keyword semantic set.

6. The digital archive organization system based on semantic analysis according to claim 5, characterized in that: The formula for calculating the semantic confidence score of each keyword is as follows: ; in, Keywords The final semantic confidence score is the highest of which indicates the greater importance of the keyword within the entire document. Indicates keywords Iterate through all fields that have appeared. For frequency weights, Keywords In the field The original number of occurrences in It is the maximum number of times all keywords appear in any single field. For hierarchical weights, It is a field The numerical value corresponding to the level identifier. It is the highest level value in the archive. Used to reverse the level number, so that level 1 has the highest score, level The lowest score.

7. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The co-occurrence phrase identification module includes: The context structure analysis submodule analyzes the sentence structure of the effective keywords in the semantic set within each field, extracts the order relationship, connection method, and modification relationship position between the keywords and adjacent words, and establishes the context structure features of the keywords. The cross-field combination extraction submodule calls the contextual structure features of each keyword, starting with a single valid keyword, identifies the continuous occurrence order of valid keywords in multiple fields, extracts keyword combinations that exist simultaneously in multiple fields, and obtains multi-field keyword combinations. The candidate word group generation submodule summarizes and merges all multi-field keyword combinations that exist simultaneously in multiple fields, eliminates duplicate combinations, and establishes a candidate co-occurring word group set.

8. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The word combination module includes: The phrase vectorization submodule extracts the word order position, association frequency and field number of the keywords in each phrase in the candidate co-occurring phrase set as features, and constructs a corresponding vector expression for each co-occurring phrase, thus establishing a co-occurring phrase feature vector. The vector similarity calculation submodule combines the vector representations of all the co-occurring word group feature vectors in pairs and calculates the similarity between the two vectors in each pair as a measure of semantic association between word groups, thereby obtaining semantic vector similarity information. The similar word combination and merging submodule calls the semantic vector similarity information of each pair of co-occurring word groups, compares the similarity with the set path normalization threshold, filters out co-occurring word group pairs whose similarity reaches the path normalization threshold condition and merges them to obtain a merged word group set.

9. The digital archive organization system based on semantic analysis according to claim 1, characterized in that: The sorting and output module includes: The phrase frequency statistics submodule counts the frequency of each merged phrase in the merged phrase set in the field, and obtains the frequency information of the phrase field. The semantic weight calculation submodule obtains the frequency information of the word group field and the field structure position information of the corresponding field, uses the field structure position information as a semantic influence factor, and combines it with the frequency of the word group field to calculate the semantic weight value of each merged word group in the field using the TF-IDF algorithm. The semantic result integration submodule calls each merged phrase and its corresponding semantic weight value, assigns semantic weights to each merged phrase within the fields, and generates digital archive semantic organization results, which are used to establish document structure summaries, archive tags, or assist human reviewers in archiving and review.

10. The digital archive organization system based on semantic analysis according to claim 9, characterized in that: To calculate the semantic weight value of each merged phrase within the field, the formula is as follows: ; in, It is a combined phrase In a specific field semantic weights within, It refers to word frequency, specifically the TF portion. It is a phrase In the field The number of times it appears in It is a field The total number of times all phrases appear in the text. It is the inverse document frequency, i.e., the IDF part. It is the total number of fields in the file. The entire file contains phrases The total number of fields, It is based on field hierarchy The calculated semantic influence factor is specifically: .