Digital archive multi-modal data semantic enhancement fusion retrieval method and system

By constructing a policy cycle timeline and a terminology evolution map, identifying seal combination patterns and authority levels, and generating an evolution record of authority validity, the problem of retrieval misjudgment caused by the historical evolution of seal authority was solved, and the accuracy and recall rate of archive retrieval were improved.

CN122019797APending Publication Date: 2026-05-12MID-RANGE INFORMATION (GUANGDONG) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MID-RANGE INFORMATION (GUANGDONG) CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot accurately determine the legal validity of seal authority in different historical periods when processing document retrieval, and cannot match historical documents with obsolete terminology, leading to missed detections and misjudgments.

Method used

By constructing a policy cycle timeline and terminology evolution map, we can identify seal combination patterns and authority levels, generate records of authority validity evolution, and fuse temporal authority feature vectors with content semantic vectors to achieve multimodal representation vectors for query expansion and filtering.

Benefits of technology

Accurately determine the legal effect of the same seal combination in different historical policy cycles, cover historical archives using obsolete terminology, and improve retrieval accuracy and recall rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019797A_ABST
    Figure CN122019797A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital archive management and information retrieval, and discloses a digital archive multi-modal data semantic enhancement fusion retrieval method and system.The method comprises the steps that a policy cycle time axis and a policy term evolution graph are constructed, tense logical reasoning is conducted on archive seals, and permission effectiveness evolution is derived; the temporal permission feature vector and the content semantic vector are fused to generate a multi-modal representation vector, and cross-policy-cycle semantic enhancement retrieval is realized by combining query expansion and temporal permission filtering, so that the problems of missing detection and misjudgment of policy and regulation archives in seal permission historical evolution and term cross-cycle retrieval are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of digital archives management and information retrieval technology, and more specifically, to a method and system for semantic enhancement fusion retrieval of multimodal digital archives data. Background Technology

[0002] In the field of digital archives management for administrative approvals, seals, as the carriers of approval authority, have their approval validity changing with the evolution of policies and regulations. Simultaneously, the official terminology used in policy and regulatory archives also evolves with policy changes.

[0003] Existing technologies for processing document retrieval typically employ keyword matching or semantic vector retrieval methods. However, these methods handle seal authority identification and policy terminology evolution separately, without establishing a correlation between seal authority semantics and policy cycles, or a mapping relationship between the historical evolution of policy terms.

[0004] The existing technology has the following drawbacks: when users search for "historical archives with independent approval effect" or use current policy terminology for retrieval, the system cannot determine the actual legal effect of the same seal combination in different historical periods. At the same time, the use of current policy terminology cannot match historical archives expressed using obsolete terminology, which results in the inability to accurately respond to users' search needs based on historical status of permissions and across policy cycles, causing technical problems of missed detections and misjudgments in retrieval. Summary of the Invention

[0005] This invention provides a method and system for semantic enhancement and fusion retrieval of multimodal digital archives, which solves the technical problems of missed detections and misjudgments caused by the historical evolution of seal authority and cross-cycle changes in policy terminology in related technologies.

[0006] This invention provides a semantically enhanced fusion retrieval method for multimodal data of digital archives, comprising the following steps: Obtain scanned images of archives and their creation time metadata; extract the release time, scope of application, and repeal time of policy documents associated with the archives; generate a policy cycle timeline based on the time attributes of each policy document; match the archive creation time with the policy cycle timeline to determine the policy cycle to which the archives belong. The document extracts terms from the archival texts of each policy cycle, identifies policy proper nouns and authority-related terms, calculates the semantic similarity of the term sets of each policy cycle, identifies term evolution pairs with semantic continuity, and generates a policy term evolution graph by using the terms of each policy cycle as nodes and term evolution pairs as edges. Seal detection and recognition are performed on scanned images of archives. The position coordinates and features of each seal are extracted. The seal features are matched with a seal knowledge base to generate identity tags and basic permission levels for each seal. Analyze the spatial relationship of multiple seals on the same file, identify seal combination patterns, associate and match the identity tags of each seal with the set of permission rules for the current policy period, and obtain the permission level definition of each seal within the policy period. Based on the seal combination pattern and the current authority level definition, the composite authority validity label of the seal combination in the current policy cycle is derived, and an authority validity evolution record is generated. The evolution record of authority effectiveness is encoded into a temporal authority feature vector. The archive text content is semantically encoded to generate a content semantic vector. The temporal authority feature vector and the content semantic vector are fused to generate a multimodal representation vector with enhanced temporal authority. The system receives user queries, identifies policy terms and permission status constraints in the queries, performs graph traversal starting from the identified policy terms in the policy term evolution graph, retrieves equivalent terms for each historical policy cycle, replaces the policy terms in the original queries with historical equivalent terms, generates an extended query set, performs semantic encoding on the extended query set to generate an extended query semantic vector, and generates temporal permission filtering conditions based on permission status constraints. The database is filtered using temporal permission filtering conditions to obtain a set of candidate files that meet the permission status constraints. The similarity between the extended query semantic vector and the temporal permission enhanced multimodal representation vector of each file in the candidate file set is calculated. The candidate files are sorted according to the comprehensive similarity and the search results are output.

[0007] Furthermore, the process of generating the policy cycle timeline includes: The policy documents are sorted by their release date, with the release date of each policy document serving as the starting point of a new cycle and the expiration date or the release date of the next policy document serving as the end point of that cycle, thus forming a continuous policy cycle sequence.

[0008] Furthermore, the identification of term evolution pairs with semantic continuity relationships includes: Calculate the semantic similarity between the terms in the current policy cycle and the terms in adjacent policy cycles. When the similarity exceeds a preset threshold, determine that the two terms constitute a term evolution pair and add an evolution type label to the term evolution pair. The evolution type tags include name replacement, concept merging, and concept splitting.

[0009] Furthermore, the identification stamp combination pattern includes: Based on the relative positions and overlaps of multiple seals, seal combination patterns are classified into parallel stamping, overlapping stamping, and continuous stamping with interlocking seams. The parallel stamping pattern refers to multiple stamps arranged sequentially in a horizontal or vertical direction without overlapping. The overlapping stamping pattern refers to the presence of partially overlapping areas among multiple stamps. The continuous stamping pattern refers to the stamp continuously stamping across the boundaries of the document page.

[0010] Furthermore, the derivation of the composite authority validity label of the seal combination in the current policy cycle based on the seal combination pattern and the current authority level definition includes: Retrieve the permission level definition and seal combination mode of each seal in the seal combination; Based on the seal combination pattern, retrieve the corresponding permission stacking rules from the permission combination rules of the current policy cycle; Substitute the permission level definitions of each seal into the permission stacking rules to calculate the combined permission effectiveness value; Based on the comparison between the composite permission effectiveness value and the permission threshold, a composite permission effectiveness label is generated.

[0011] Furthermore, the composite authority validity label includes full validity, partial validity, and invalid validity; The record of the evolution of authority validity includes seal combination identifier, source policy cycle identifier, target policy cycle identifier, source cycle authority validity label, target cycle authority validity label, and validity change type.

[0012] Furthermore, the fusion of the temporal permission feature vector and the content semantic vector includes: performing one-hot encoding on each field in the permission validity evolution record, mapping the seal combination identifier, permission validity label, and validity change type to vectors of corresponding dimensions, and concatenating them to form a temporal permission feature vector; multiplying the temporal permission feature vector and the content semantic vector by their respective weight coefficients and then concatenating them to generate a multimodal representation vector for enhanced temporal permissions.

[0013] Furthermore, the permission status constraints include permission validity type constraints and policy cycle range constraints; The permission validity type constraint specifies the permission validity tags that the target file should possess, and the policy cycle range constraint specifies the policy cycle range to which the target file was created.

[0014] Furthermore, the step of calculating the similarity between the extended query semantic vector and the temporal permission-enhanced multimodal representation vector of each file in the candidate file set includes: For each extended query semantic vector in the extended query set, calculate the cosine similarity with the temporal permission-enhanced multimodal representation vector of the candidate file. Multiply each cosine similarity by the corresponding extended query weight and sum them to obtain the comprehensive similarity score.

[0015] This invention provides a semantically enhanced fusion retrieval system for multimodal digital archive data, used to execute the aforementioned semantically enhanced fusion retrieval method for multimodal digital archive data, comprising: The policy cycle determination module is used to acquire scanned images of archives and their creation time metadata, extract the release time, scope of application and repeal time of policy documents associated with the archive, generate a policy cycle timeline based on the time attributes of each policy document, match the archive creation time with the policy cycle timeline, and determine the policy cycle to which the archive belongs. The policy terminology evolution graph construction module is used to extract terms from the archival texts in each policy cycle, identify policy proper nouns and authority-related terms, calculate the semantic similarity of the term sets in each policy cycle, identify term evolution pairs with semantic continuity, and generate a policy terminology evolution graph by using the terms in each policy cycle as nodes and term evolution pairs as edges. The seal recognition and permission matching module is used to detect and recognize seals in scanned images of documents, extract the position coordinates and features of each seal, match the seal features with the seal knowledge base, and generate the identity label and basic permission level of each seal. The temporal permission validity derivation module is used to analyze the spatial relationship of multiple seals on the same file, identify seal combination patterns, associate and match the identity tags of each seal with the permission rule set of the current policy period, and obtain the permission level definition of each seal within the policy period. Based on the seal combination pattern and the current authority level definition, the composite authority validity label of the seal combination in the current policy cycle is derived, and an authority validity evolution record is generated. The temporal permission feature fusion module is used to encode the permission validity evolution record into a temporal permission feature vector, perform semantic encoding on the archive text content to generate a content semantic vector, and fuse the temporal permission feature vector with the content semantic vector to generate a temporal permission enhanced multimodal representation vector. The query extension module is used to receive user query statements, identify policy terms and permission status constraints in the query statements, perform graph traversal in the policy term evolution graph starting from the identified policy terms, retrieve equivalent terms of the term in each historical policy cycle, replace the policy terms in the original query statement with historical equivalent terms, generate an extended query set, perform semantic encoding on the extended query set to generate an extended query semantic vector, and generate temporal permission filtering conditions based on permission status constraints. The retrieval output module is used to filter the archive database using temporal permission filtering conditions, obtain a set of candidate archives that meet the permission status constraints, calculate the similarity between the extended query semantic vector and the temporal permission enhanced multimodal representation vector of each archive in the candidate archive set, sort the candidate archives according to the comprehensive similarity, and output the retrieval results.

[0016] The beneficial effects of this invention are as follows: By constructing a policy cycle timeline and associating archives with corresponding policy cycles, this invention uses a temporal logic reasoning algorithm to deduce the authority validity of seal combinations and generate an authority validity evolution record, enabling the accurate determination of the actual legal validity of the same seal combination in different historical policy cycles. By constructing a policy terminology evolution graph to establish semantic continuity relationships between terms in each policy cycle, and by retrieving historical equivalent term sets from the graph for query expansion, queries initiated by users using current policy terms can automatically cover historical archives using obsolete terms. By fusing temporal authority feature vectors with content semantic vectors to generate multimodal representation vectors, deep integration of authority information and semantic information is achieved. This invention solves the technical problems in the prior art of insufficient accuracy in retrieval based on historical authority status due to the historical evolution of seal authority and cross-cycle changes in policy terms, as well as the technical problem of missed detection in cross-cycle retrieval of policy and regulation archives, achieving the technical effect of improving the accuracy and recall rate of archive retrieval. Attached Figure Description

[0017] Figure 1 This is a flowchart of the digital archive multimodal data semantic enhancement fusion retrieval method of the present invention; Figure 2 This is a bar chart of the number of archives and average authority effectiveness in different policy cycles of the present invention, showing the distribution of the number of archives and the average authority effectiveness value in different policy cycles, corresponding to the policy cycle division in step 1 and the authority effectiveness derivation in step 5. Figure 3 This is a scatter plot of the overall similarity distribution of the search results of the present invention, showing the distribution of the overall similarity scores of the file search results, corresponding to the similarity calculation and sorting process in step 8; Figure 4 This is a grouped bar chart comparing the retrieval performance before and after query expansion according to the present invention. It shows that by constructing a policy term evolution map and performing query expansion, the retrieval system has achieved significant improvements in recall and F1 score. Detailed Implementation

[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0019] Example 1

[0020] This embodiment provides a semantically enhanced fusion retrieval method for multimodal data of digital archives, such as... Figure 1 As shown, it includes the following steps: Step 1: Obtain the scanned image of the archive and its creation time metadata, extract the release time, scope of application and repeal time of the policy documents associated with the scanned image of the archive, generate a policy cycle timeline based on the time attributes of each policy document, match the creation time of the archive with the policy cycle timeline, and determine the policy cycle to which the scanned image of the archive belongs.

[0021] It should be noted that the process of generating the policy cycle timeline is as follows: sort the policy documents by their release date, take the release date of each policy document as the starting point of the new cycle, and take the repeal date or the release date of the next policy document as the ending point of the policy cycle, thus forming a continuous policy cycle sequence.

[0022] Step 2: Extract terms from the archival text within the current policy cycle, identify policy proper nouns and authority-related terms, calculate the semantic similarity between the term set of the current policy cycle and the term set of adjacent policy cycles, identify term evolution pairs with semantic continuity, use the terms of each policy cycle as nodes and term evolution pairs as directed edges, add evolution type labels and timestamps, and generate a policy term evolution graph.

[0023] The aforementioned terminology extraction refers to extracting policy-specific terms and authority-related terms from archival texts within a policy cycle, and outputting a terminology set for the current policy cycle.

[0024] Furthermore, the specific implementation method for terminology extraction is as follows: First, the archival text is segmented into words. Then, a combination of Named Entity Recognition (NER) technology and Term Frequency-Inverse Document Frequency (TF-IDF) statistical methods is used to combine high-frequency words with entity recognition results. For policy-specific terms, string matching can be performed by constructing a dedicated policy terminology database. For permission-related terms, permission attributes are labeled in the terminology database to identify terms with permission characteristics. Finally, the matched terms are compiled to form a terminology set for the current policy cycle.

[0025] The aforementioned semantic similarity calculation refers to calculating the similarity between two terms, with the input being terms from the current policy cycle. Terminology for adjacent policy cycles The output is the semantic similarity value between the two. When the similarity exceeds a preset threshold At that time, the judgment and This constitutes a term evolution pair. Evolution type labels are assigned by analyzing the semantic relationship between two terms, marking them as name replacement, concept merging, or concept splitting.

[0026] Furthermore, the criteria for judging the evolution type are: (1) Name replacement: When a term is replaced by another semantically similar but differently expressed term in a new policy cycle, and the semantic scope of the two terms is basically the same, it is judged as name replacement; (2) Concept merging: When multiple different terms from the old policy cycles are merged into one term expression in the new policy cycle, that is, multiple Corresponding to one In the case of (3) Concept merging: when a term from an old policy cycle is split into multiple more specific terms in a new policy cycle, that is, a Corresponding to multiple In cases like these, it is determined to be a concept split. This classification standard allows for a clear distinction and labeling of term evolution relationships.

[0027] Furthermore, the temporal attribute of directed edges in the policy terminology evolution graph is defined as follows: each directed edge points from a source term to a target term, indicating that the source term appeared earlier than the target term in time; the timestamp of the edge is recorded as follows. ,in The starting time of the policy cycle in which the source term is located. The starting time of the policy cycle in which the target term is located satisfies the time sequence constraint. Such time constraints ensure that all evolution paths in the policy terminology evolution graph follow the unidirectional nature of time, so that when traversing along directed edges from any term, the process proceeds in the forward direction of time, thereby ensuring that the historical evolution trajectory of terms can be accurately traced when retrieving historical equivalent terms.

[0028] Furthermore, preset threshold The range of values ​​is Due to similarity value The similarity is calculated by weighting word vector similarity and context similarity. Since both input similarity values ​​are within the range [0, 1], the weighted result is also within the range [0, 1]. The value should be taken within this range.

[0029] semantic similarity value The calculation method is as follows:

[0030] in The word vector similarity between two terms is represented by the cosine similarity of their word vectors. The similarity between two terms in their contextual context is determined by comparing the degree of overlap of the surrounding vocabulary in the policy text. is a weighting coefficient used to balance word vector similarity and contextual similarity.

[0031] Furthermore, the word vectors for terms are obtained as follows: Pre-trained word vector models (such as Word2Vec, GloVe, or BERT language models) are used to vectorize the terms; specifically, the terms are input into the pre-trained model, and their vector representations are extracted as word vectors; for multi-character terms, the average of the word vectors of each character can be taken, or contextual embeddings can be used. The overlap of contextual vocabulary is calculated as follows: extract the term... The vocabulary within a fixed window size (e.g., 5 words before and after) in the text in which it appears forms a context vocabulary set. , for terminology Similarly, extract the context vocabulary set. Then, calculate the ratio of the intersection to the union of the two sets (Jaccard similarity) or use word overlap counts to quantify contextual similarity.

[0032] Furthermore, weighting coefficients The range of values ​​is This is because and All similarity values ​​are within the range of [0, 1], and the weighting coefficients are... and The sum of is 1, thus ensuring that the weighted sum is also in the range [0, 1].

[0033] In this embodiment of the application, in order to accurately identify semantic continuity relationships, when calculating semantic similarity, in addition to considering the word vector similarity of terms, the context of terms in policy texts is also considered for comprehensive judgment, thereby improving the accuracy of term evolution in identification.

[0034] Step 3: Process the scanned image of the document using a seal detection algorithm to locate all seal areas in the scanned image of the document and extract the position coordinates and size information of each seal; normalize the extracted coordinates and size information to make it independent of the resolution and size of the original scanned image of the document; use a seal recognition algorithm to extract text and identify the type of each seal area, match the extracted seal features with the seal knowledge base, and generate the identity label and basic permission level of each seal.

[0035] The aforementioned seal detection algorithm processes scanned images of documents, taking the scanned images as input and outputting the position coordinates and size information of each seal. The aforementioned seal recognition algorithm identifies the located seal areas, taking the image areas of each seal as input and outputting the text content, seal type, and other feature information of the seal. The extracted seal features are then matched against a seal knowledge base to retrieve the identity tag and basic permission level of each seal.

[0036] It should be noted that the seal knowledge base pre-stores the standard features of various seals, along with their corresponding departmental affiliation, seal type, and basic permission level information. Seal types include official seals, contract seals, financial seals, and legal representative seals. When normalizing the position coordinates and dimensions, the coordinates are normalized to the relative position (0-1 range) of the archival image, and the dimensions are normalized to a ratio to the image size, ensuring that the extracted spatial relationship features are independent of the original dimensions of the archive.

[0037] Step 4: Analyze the spatial relationship of multiple seals on the same scanned image of the document, identify the seal combination pattern based on the relative position and overlap between the seals, and associate and match the identified seal identity tags with the set of permission rules for the current policy period to obtain the permission level definition of each seal within the current policy period.

[0038] It should be noted that the seal combination patterns include three types: parallel stamping, overlapping stamping, and continuous stamping across the seam. Parallel stamping refers to multiple seals arranged sequentially in the horizontal or vertical direction without overlap; overlapping stamping refers to multiple seals having some overlapping areas; and continuous stamping across the seam refers to seals being stamped continuously across the boundaries of the document page.

[0039] Furthermore, the specific method for analyzing the spatial relationship of the seals is as follows: Calculate the Euclidean distance and relative direction angle between the center points of each seal based on the normalized seal position coordinates obtained in step 3; determine the overlap by using the Intersection over Union (IoU) of the seal area, and determine that there is overlap when the IoU is greater than the preset overlap threshold; for the identification of parallel stamping patterns, determine whether the center points of multiple seals are arranged in the horizontal or vertical direction and the IoU is close to zero; for the identification of overlapping stamping patterns, determine whether the IoU is greater than the overlap threshold and less than the complete coverage threshold; for the identification of continuous patterns across the seam, determine whether the distance between the seal area and the page boundary position (normalized coordinates close to 0 or 1) is less than the preset threshold.

[0040] Step 5: Based on the seal combination pattern and the current authority level definition, use temporal logic reasoning algorithm to derive the composite authority validity label of the seal combination in the current policy cycle, compare and analyze the authority validity labels of the same seal combination in adjacent policy cycles, and generate an authority validity evolution record.

[0041] The aforementioned temporal logical reasoning algorithm refers to an algorithm that derives the composite authority validity label of a seal combination within the current policy cycle based on the definition of its authority level and spatial relationship. Its processing includes: Step 501: Obtain the permission level definition of each seal in the seal combination. Combination patterns with seals ,in The number of seals in the seal set. They are respectively the 1st to the 1st The permission level definition for each seal is defined; the permission level definition is preprocessed, the categorized permission levels are encoded and converted into numerical values, and each permission level definition is standardized to a uniform numerical range (e.g., 0-1) for subsequent mathematical calculations. Step 502: Based on the seal combination pattern Retrieve the corresponding permission stacking rules from the permission combination rules of the current policy cycle. The permission stacking rules define how the permissions of multiple seals are stacked under different combination modes; Step 503: Substitute the preprocessed seal permission level definitions into the permission stacking rules. Calculate the validity value of composite permissions The aforementioned rules for stacking permissions. This refers to the combination pattern of seals For the collection of seals The permission level definitions of each seal are aggregated. The calculation process involves performing mathematical operations on the permission effectiveness of multiple seals under different combination modes, using weighted summation or other aggregation strategies (such as finding the maximum value of the permission effectiveness of multiple seals when they are stamped side by side, finding the average value when they are stamped in layers, and applying special permission effectiveness transfer rules when they are stamped consecutively across the seam), and finally obtaining a single composite permission effectiveness value. This value is a scalar, representing the overall authority and effectiveness of the seal combination; Furthermore, the validity value of composite permissions The range of values ​​is This is because in step 501, the permission level definition of each seal has been standardized to the range of [0, 1]. Regardless of the aggregation strategy used by the permission superposition rule, such as weighted summation, finding the maximum value, or finding the average value, the result should remain within the range of [0, 1], thus ensuring that the combined permission effectiveness value is also within the range of [0, 1].

[0042] Step 504: Based on the composite permission validity value The comparison results with the preset permission threshold generate a composite permission validity label. When the permission validity value reaches the full approval permission threshold, the label is "full validity"; when the permission validity value is within the range of partial permission threshold, the label is "partial validity"; and when the permission validity value does not reach the partial permission threshold, the label is "invalid".

[0043] Furthermore, the preset permission threshold is defined as follows: Let the partial permission threshold be... The threshold for full approval authority is ,in .when When, generate a "full efficacy" label; when When, generate a "partial effectiveness" label; when When this happens, an "invalid" label is generated. This design ensures a monotonic correspondence between the permission validity value and the permission validity label, making the permission determination results clearly distinguishable.

[0044] In this embodiment of the application, in order to track the historical changes in the effectiveness of permissions, the record content when generating the permission effectiveness evolution record includes: a seal combination identifier, a source policy cycle identifier, a target policy cycle identifier, a source cycle permission effectiveness label, a target cycle permission effectiveness label, and a type of effectiveness change. The types of effectiveness change include effectiveness enhancement, effectiveness weakening, and effectiveness remaining unchanged.

[0045] Step 6: Encode the evolution record of authority effectiveness into a temporal authority feature vector, perform semantic encoding on the archive text content to generate a content semantic vector, perform data preprocessing on the two vectors to eliminate the difference in units, and then fuse them to generate a multimodal representation vector for enhanced temporal authority.

[0046] The aforementioned temporal permission feature vector encoding process is as follows: Each field in the permission validity evolution record is one-hot encoded. The input to one-hot encoding consists of discrete attribute fields such as seal combination identifier, permission validity label, and validity change type, with the output being a corresponding binary vector. Simultaneously, the policy cycle identifier is time-encoded, outputting a continuous value vector representing time information. The encoded vectors of each field and the time-encoded vector are concatenated to form the temporal permission feature vector. The aforementioned semantic encoding refers to encoding the archival text content. The input is the archival text, and the output is a content semantic vector. The content semantic vector uses a point in the vector space to represent the semantic information of the archival text content.

[0047] Furthermore, the time coding method for policy cycle identifiers is as follows: for the source policy cycle and target policy cycle in the record of the evolution of authority effectiveness, their start times are extracted respectively. and end time Calculate the time span of the policy cycle and the relative position of the policy cycle on the entire timeline. ,in This represents the starting time of the policy cycle timeline. The total span of the time axis is represented by this information. This time information is combined into a time encoding vector, which explicitly expresses the time position and time span information of the policy cycle in the temporal authority feature vector, so that the time differences of different policy cycles can be reflected in the vector space.

[0048] Furthermore, the specific method for obtaining content semantic vectors is as follows: use a pre-trained text encoding model (such as BERT, RoBERTa, and other Transformer models) to encode the content of the archive text; specifically, input the archive text into the encoder of the pre-trained model, extract the vector representation corresponding to the [CLS] token in the last layer of the model, or perform average pooling on the vector representations of all tokens to obtain a fixed-dimensional content semantic vector; the content semantic vector can capture the overall semantic information of the archive text, providing a text semantic representation for subsequent vector fusion and similarity calculation.

[0049] Data preprocessing includes the following: Since the temporal permission feature vector is obtained through one-hot encoding, its element values ​​are either 0 or 1, while the value range of the content semantic vector may be different. Therefore, the two vectors need to be normalized before fusion. Specifically, the temporal permission feature vector and the content semantic vector are both standardized using the L2 norm, mapping both vectors to a unit hypersphere. This eliminates the difference in the vector value ranges and ensures that the two vectors have comparable dimensions during fusion.

[0050] In this embodiment, to preserve the distinguishability of temporal permission information in the fused representation, the preprocessed temporal permission feature vector and the content semantic vector are fused using a weighted concatenation method. The fusion formula is as follows:

[0051] in, This is the fused multimodal representation vector. This is the temporal permission feature vector after L2 norm standardization. This is the content semantic vector after L2 norm standardization. and These are the weighting coefficients. This indicates a vector concatenation operation.

[0052] Furthermore, weighting coefficients and The value constraints are as follows , Both coefficients should be positive. This is because the standardized vector retains its direction and relative relationships after being scaled with positive coefficients. Furthermore, to preserve the discriminative power of temporal authorization information, information from both modalities should be retained in the fused representation. and The relative size relationships can be adjusted according to the specific application's emphasis on temporal permission information and content semantic information, but generally... and They can be equal or set according to specific task requirements.

[0053] Step 7: Receive the user query statement, identify the policy terms and permission status constraints in the query statement, perform graph traversal in the policy term evolution graph starting from the identified policy terms, retrieve the equivalent terms of the term in each historical policy cycle, summarize them to form a historical equivalent term set, replace the policy terms in the original query statement with the terms in the historical equivalent term set, generate an extended query set, perform semantic encoding on the extended query set to generate an extended query semantic vector, and generate temporal permission filtering conditions based on the permission status constraints.

[0054] The aforementioned policy terminology recognition refers to identifying terms belonging to the policy terminology evolution graph from user queries. The input is the user query, and the output is the identified policy terms. The aforementioned graph traversal refers to traversing the policy terminology evolution graph starting from the identified policy terms and following the evolution edges. The output is a set of terms associated with the starting term in each historical policy cycle. The aforementioned extended query semantic encoding refers to semantically encoding each query statement in the extended query set. The input is the query statement, and the output is the corresponding semantic vector.

[0055] Furthermore, the specific strategy for policy terminology identification is as follows: First, the user query is segmented into a word sequence; then, the segmentation results are matched against all nodes (terms) in the policy terminology evolution graph, prioritizing the longest match; for successfully matched words, they are marked as policy terms and their node identifiers in the graph are recorded; for words that do not match completely, their edit distance or semantic similarity with terms in the graph is calculated to identify approximately matched terms; finally, all policy terms identified in the query and their corresponding node identifiers are output. The extended query semantic encoding uses the same method as the archival text content encoding in step 6, employing a pre-trained text encoding model to extract the semantic vector representation of the extended query.

[0056] It should be noted that the permission status constraints include permission effectiveness type constraints and policy cycle range constraints. The permission effectiveness type constraint specifies the permission effectiveness label that the target file should have, such as "full effectiveness" or "partial effectiveness"; the policy cycle range constraint specifies the policy cycle range to which the target file was created.

[0057] In this embodiment of the application, in order to improve the coverage of the query expansion, when performing graph traversal, not only are the evolved terms directly connected to the starting term retrieved, but also multi-hop related terms are recursively retrieved along the evolution path to form a complete historical equivalent term chain.

[0058] Furthermore, the definition of a multi-hop related term is as follows: In a policy terminology evolution graph, starting from the initial term, a multi-step graph traversal is performed along directed evolution edges. Each directed edge traversed is called a hop, and a term reached by skipping multiple directed edges is called a multi-hop related term. For example, if the term... In policy cycle 1, it evolved into (1 jump) In policy cycle 2, it evolved into (2 jumps), then and There is a two-hop association relationship between them. The recursive retrieval process is as follows: first, find the starting term. All directly connected terms (1 jump) are traversed, and then the 2-jump terms are found from these terms. This process continues until all reachable terms in the entire term evolution graph have been traversed, thus forming a complete historical equivalent term chain that includes the starting term and its related terms at each level.

[0059] Furthermore, the temporal directionality constraint of graph traversal is as follows: because the directed edges of the policy terminology evolution graph follow a temporal order constraint. Graph traversal can be performed along two time directions: (1) Forward traversal: Traversing from the starting term along the directed edges in the forward direction to retrieve new terms evolved from the starting term, i.e., searching for equivalent terms that are later than the starting term in time; (2) Reverse traversal: Traversing from the starting term along the directed edges in the reverse direction to retrieve old terms that evolved into the starting term, i.e., searching for historical terms that are earlier than the starting term in time. When performing query expansion, in order to cover all historical policy cycles, it is necessary to obtain historical terms through reverse traversal or subsequent terms through forward traversal, thereby forming a complete set of equivalent terms covering the starting term in all policy cycles.

[0060] Step 8: Filter the archive database using temporal permission filtering conditions to obtain a set of candidate archives that meet the permission status constraints. Calculate the similarity between the extended query semantic vector and the temporal permission enhanced multimodal representation vector of each archive in the candidate archive set. Summarize the matching scores for each policy cycle, sort the candidate archives according to the comprehensive similarity, and output the search results.

[0061] The aforementioned similarity calculation employs the cosine similarity algorithm. This algorithm takes the extended query semantic vector and the temporal permission-enhanced multimodal representation vector of the candidate files as input, and outputs the similarity value between the two. The formula for calculating the overall similarity is:

[0062] in, To calculate the overall similarity score, To expand the number of query statements in the query set, For the first The weight of each extended query, For the first The semantic vector of each extended query. Enhance the multimodal representation vector of temporal permissions for candidate files. Let be the cosine similarity function, which calculates the cosine value of two vectors.

[0063] Furthermore, in the similarity calculation process, the first... Semantic vectors of extended queries This is achieved through step 7. The temporal permission-enhanced multimodal representation vector of candidate files is obtained by semantically encoding an extended query statement. This vector, generated through step 6, integrates the temporal permission features and content semantic information of the archive. For each extended query, the cosine similarity between its semantic vector and the candidate archive vector is calculated, and then the vectors are weighted according to the weights of each extended query. The weighted sum is then used to obtain the overall similarity score of the candidate file. .

[0064] Furthermore, expand query weights The value constraints are as follows , This is because of the overall similarity. The similarity score, representing a weighted average, should be a weighted combination of the similarities between each expanded query and the candidate profile. (Setting...) The constraints ensure that the overall similarity value remains within the range of cosine similarity (i.e., [-1, 1] or the normalized [0, 1]), thus guaranteeing the comparability and interpretability of the overall similarity value. The specific value can be set based on factors such as the semantic relevance of each extended query and its position in the terminology evolution chain.

[0065] Here is an example of an application of the present invention: Organization X established a digital records management system to manage administrative approval documents from 2010 to 2023. During this period, the city underwent policy and regulatory adjustments and institutional reforms: Policy A was implemented from 2010 to 2015, Policy B from 2016 to 2020, and Policy C was implemented after 2021. Simultaneously, in 2018, Organization Y and Organization Z merged to form Organization X, resulting in changes to seal authority. In January 2024, the records administrator needs to retrieve all historical documents related to Policy C that have full approval validity.

[0066] In step 1, the system first extracts the time information of policy documents from the archive and generates a policy cycle timeline. The metadata for the creation time of a certain scanned image is June 15, 2014. The relevant policy information extracted by the system from the policy document archive is shown in the table below:

[0067] Based on the matching of the archive's creation date of June 15, 2014, with the policy cycle timeline, it was determined that the scanned image of the archive belongs to Policy Cycle 1 (2010-2015), and the corresponding policy document is the "A Policy Management Regulations".

[0068] In step 2, the system extracts terms from the archival texts across the three policy cycles and calculates semantic similarity. For the term "Policy A" from policy cycle 1 and the term "Policy B" from policy cycle 2, the system calculates word vector similarity. Context similarity Set weighting coefficients. Then the semantic similarity value is:

[0069] because The system determines whether two terms constitute a term evolution pair. A partial policy term evolution map constructed by the system is shown in the table below:

[0070] In step 3, the system performs seal detection on the scanned image of the document (resolution 1200×1800 pixels), locating two seal areas. The original position coordinates of seal 1 are (300, 1500), and its size is 120×120 pixels; the original position coordinates of seal 2 are (450, 1500), and its size is 115×115 pixels. After normalization, the center coordinates of seal 1 are (0.25, 0.83), and the normalized size is 0.10; the center coordinates of seal 2 are (0.38, 0.83), and the normalized size is 0.096. The seal recognition results are shown in the table below:

[0071] In step 4, the system analyzes the spatial relationship between the two seals. The Euclidean distance between the center points of the seals is calculated as follows: The relative direction angle is (Horizontal arrangement). The Intersection over Union (IoU) of the seal areas is 0.02, which is less than the preset overlap threshold of 0.1. Since the center points of the two seals are arranged horizontally, it is determined to be a "parallel stamping" mode. The system matches the seal identity tags with the permission rules of policy cycle 1 to obtain the permission level definition of each seal within the current policy cycle: the permission level of institution Y's seal is 0.6, and the permission level of institution Z's seal is 0.7.

[0072] In step 5, the system uses a temporal logic reasoning algorithm to derive the validity of composite permissions. In step 501, the permission level definition is obtained. The seal combination mode $M = "Parallel Seal". In step 502, the permission stacking rule for the retrieved parallel seal mode is the maximum value strategy. In step 503, the combined permission effectiveness value is calculated. In step 504, set , ,because The generated compound permission validity label is "partial validity".

[0073] The system further analyzes the authority effectiveness of the same seal combination during policy cycle 2 (2016-2020). Due to the institutional reform in 2018, institution Y and institution Z merged, and the seal combination was identified as a "historical seal combination" in the authority rules of policy cycle 2, with its authority level reduced to 0.3 and 0.3 respectively, resulting in a composite authority effectiveness value. The permission validity label is "invalid". The generated permission validity evolution record is shown in the table below:

[0074] In step 6, the system encodes the evolution of authority validity into a temporal authority feature vector. The seal combination identifier, the authority validity labels "partial validity" and "invalidity," and the validity change type "weakened validity" are generated into a binary vector after one-hot encoding; the policy cycle identifier is time-encoded, indicating the start time of the source policy cycle. End time Time span Year, setting the policy cycle timeline starting time Total span Year, relative position is The time-coded vector is The one-hot encoded vector and the temporal encoded vector are concatenated to form a temporal permission feature vector (32 dimensions).

[0075] Simultaneously, the system semantically encodes the archival text content "Decision on Approving Zhang's Application for Related Matters," using the BERT model to extract the vector representation of the [CLS] marker, generating a content semantic vector (768 dimensions). After L2 norm standardization of the two vectors respectively, they are merged using a weighted concatenation method, and the following settings are applied. , Generate temporally-permission-enhanced multimodal representation vectors (800 dimensions).

[0076] In step 7, the archivist enters the query "Retrieve archives related to Policy C and with full validity". The system performs word segmentation on the query, identifying the policy term "Policy C" and the permission status constraint "full validity". Starting from "Policy C", the system performs a reverse traversal in the policy terminology evolution graph, retrieving historical equivalent terms. Reversing by 1 hop yields "Policy B", and continuing by 2 hops yields "Policy A". The set of historical equivalent terms is... "Policy C", "Policy B", "Policy A" The generated extended query set is shown in the table below:

[0077] The system performs semantic encoding on the three extended query statements respectively, generating extended query semantic vectors. , , (All dimensions are 768). Based on the permission status constraint "Full Validity", generate temporal permission filtering conditions: Permission Validity Tag. "Full effectiveness".

[0078] In step 8, the system uses temporal permission filtering conditions to filter the archive database, filtering out a set of candidate archives (containing 523 archives) with the permission validity label "full validity". The aforementioned archive (archive number DOC_2014_0615) has the permission validity label "partial validity", does not meet the filtering conditions, and is not included in the candidate archive set.

[0079] For files that meet the filtering criteria (file number DOC_2014_1028, created on October 28, 2014, with the permission validity label "full validity"), the system will expand the query semantic vector and the temporal permission-enhanced multimodal representation vector of the file. (With a dimension of 800) Similarity calculations were performed, and the cosine similarities between the three extended queries and this file were as follows: , , The overall similarity score is:

[0080] The system sorts all files in the candidate file set in descending order of their overall similarity score, and outputs the Top-10 search results as shown in the table below:

[0081] Through the above processing flow, the system successfully expanded the user's query using the current policy term "Policy C" into a query set that includes historical equivalent terms. Combined with temporal permission filtering conditions, it accurately retrieved historical archives with full approval effect that span multiple policy cycles, thus solving the problem of missed detection in cross-cycle retrieval of policy and regulation archives.

[0082] It is understood that data preprocessing methods known to those skilled in the art include data cleaning, data transformation, and data reduction. Data transformation includes type conversion and normalization and standardization. Although the dimensions and types of data were omitted in the description of the preceding embodiments, data preprocessing is a technical knowledge known to those skilled in the art and a prerequisite step in data processing. Therefore, the previously described well-known data preprocessing steps were not described independently.

[0083] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A semantically enhanced fusion retrieval method for multimodal digital archive data, characterized in that, Includes the following steps: Obtain scanned images of archives and their creation time metadata; extract the release time, scope of application, and repeal time of policy documents associated with the archives; generate a policy cycle timeline based on the time attributes of each policy document; match the archive creation time with the policy cycle timeline to determine the policy cycle to which the archives belong. The document extracts terms from the archival texts of each policy cycle, identifies policy proper nouns and authority-related terms, calculates the semantic similarity of the term sets of each policy cycle, identifies term evolution pairs with semantic continuity, and generates a policy term evolution graph by using the terms of each policy cycle as nodes and term evolution pairs as edges. Seal detection and recognition are performed on scanned images of archives. The position coordinates and features of each seal are extracted. The seal features are matched with a seal knowledge base to generate identity tags and basic permission levels for each seal. Analyze the spatial relationship of multiple seals on the same file, identify seal combination patterns, associate and match the identity tags of each seal with the set of permission rules for the current policy period, and obtain the permission level definition of each seal within the policy period. Based on the seal combination pattern and the current authority level definition, the composite authority validity label of the seal combination in the current policy cycle is derived, and an authority validity evolution record is generated. The evolution record of authority effectiveness is encoded into a temporal authority feature vector. The archive text content is semantically encoded to generate a content semantic vector. The temporal authority feature vector and the content semantic vector are fused to generate a multimodal representation vector with enhanced temporal authority. The system receives user queries, identifies policy terms and permission status constraints in the queries, performs graph traversal starting from the identified policy terms in the policy term evolution graph, retrieves equivalent terms for each historical policy cycle, replaces the policy terms in the original queries with historical equivalent terms, generates an extended query set, performs semantic encoding on the extended query set to generate an extended query semantic vector, and generates temporal permission filtering conditions based on permission status constraints. The database is filtered using temporal permission filtering conditions to obtain a set of candidate files that meet the permission status constraints. The similarity between the extended query semantic vector and the temporal permission enhanced multimodal representation vector of each file in the candidate file set is calculated. The candidate files are sorted according to the comprehensive similarity and the search results are output.

2. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 1, characterized in that, The process of generating the policy cycle timeline includes: The policy documents are sorted by their release date, with the release date of each policy document serving as the starting point of a new cycle and the expiration date or the release date of the next policy document serving as the end point of that cycle, thus forming a continuous policy cycle sequence.

3. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 1, characterized in that, The identification of term evolution pairs with semantic continuity relationships includes: Calculate the semantic similarity between the terms in the current policy cycle and the terms in adjacent policy cycles. When the similarity exceeds a preset threshold, determine that the two terms constitute a term evolution pair and add an evolution type label to the term evolution pair. The evolution type tags include name replacement, concept merging, and concept splitting.

4. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 1, characterized in that, The identification stamp combination patterns include: Based on the relative positions and overlaps of multiple seals, seal combination patterns are classified into parallel stamping, overlapping stamping, and continuous stamping with interlocking seams. The parallel stamping pattern refers to multiple stamps arranged sequentially in a horizontal or vertical direction without overlapping. The overlapping stamping pattern refers to the presence of partially overlapping areas among multiple stamps. The continuous stamping pattern refers to the stamp continuously stamping across the boundaries of the document page.

5. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 1, characterized in that, Based on the seal combination pattern and the current authority level definition, the composite authority effectiveness label of the seal combination in the current policy cycle is derived as follows: Retrieve the permission level definition and seal combination mode of each seal in the seal combination; Based on the seal combination pattern, retrieve the corresponding permission stacking rules from the permission combination rules of the current policy cycle; Substitute the permission level definitions of each seal into the permission stacking rules to calculate the combined permission effectiveness value; Based on the comparison between the composite permission effectiveness value and the permission threshold, a composite permission effectiveness label is generated.

6. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 5, characterized in that, The composite authority validity label includes full validity, partial validity, and invalidity; The record of the evolution of authority validity includes seal combination identifier, source policy cycle identifier, target policy cycle identifier, source cycle authority validity label, target cycle authority validity label, and validity change type.

7. The digital archive multimodal data semantic enhancement fusion retrieval method according to claim 1, characterized in that, The process of fusing the temporal permission feature vector with the content semantic vector includes: performing one-hot encoding on each field in the permission validity evolution record, mapping the seal combination identifier, permission validity label, and validity change type to vectors of corresponding dimensions, and concatenating them to form a temporal permission feature vector; multiplying the temporal permission feature vector and the content semantic vector by their respective weight coefficients and then concatenating them to generate a multimodal representation vector for enhanced temporal permissions.

8. The method for semantic enhancement and fusion retrieval of multimodal digital archive data according to claim 1, characterized in that, The permission status constraints include permission validity type constraints and policy cycle range constraints; The permission validity type constraint specifies the permission validity tags that the target file should possess, and the policy cycle range constraint specifies the policy cycle range to which the target file was created.

9. The method for semantic enhancement and fusion retrieval of multimodal digital archive data according to claim 1, characterized in that, The step of calculating the similarity between the extended query semantic vector and the temporal permission-enhanced multimodal representation vector of each file in the candidate file set includes: For each extended query semantic vector in the extended query set, calculate the cosine similarity with the temporal permission-enhanced multimodal representation vector of the candidate file. Multiply each cosine similarity by the corresponding extended query weight and sum them to obtain the comprehensive similarity score.

10. A digital archive multimodal data semantic enhancement fusion retrieval system, used to execute the digital archive multimodal data semantic enhancement fusion retrieval method according to any one of claims 1 to 9, characterized in that, include: The policy cycle determination module is used to acquire scanned images of archives and their creation time metadata, extract the release time, scope of application and repeal time of policy documents associated with the archive, generate a policy cycle timeline based on the time attributes of each policy document, match the archive creation time with the policy cycle timeline, and determine the policy cycle to which the archive belongs. The policy terminology evolution graph construction module is used to extract terms from the archival texts in each policy cycle, identify policy proper nouns and authority-related terms, calculate the semantic similarity of the term sets in each policy cycle, identify term evolution pairs with semantic continuity, and generate a policy terminology evolution graph by using the terms in each policy cycle as nodes and term evolution pairs as edges. The seal recognition and permission matching module is used to detect and recognize seals in scanned images of documents, extract the position coordinates and features of each seal, match the seal features with the seal knowledge base, and generate the identity label and basic permission level of each seal. The temporal permission validity derivation module is used to analyze the spatial relationship of multiple seals on the same file, identify seal combination patterns, associate and match the identity tags of each seal with the permission rule set of the current policy period, and obtain the permission level definition of each seal within the policy period. Based on the seal combination pattern and the current authority level definition, the composite authority validity label of the seal combination in the current policy cycle is derived, and an authority validity evolution record is generated. The temporal permission feature fusion module is used to encode the permission validity evolution record into a temporal permission feature vector, perform semantic encoding on the archive text content to generate a content semantic vector, and fuse the temporal permission feature vector with the content semantic vector to generate a temporal permission enhanced multimodal representation vector. The query extension module is used to receive user query statements, identify policy terms and permission status constraints in the query statements, perform graph traversal in the policy term evolution graph starting from the identified policy terms, retrieve equivalent terms of the term in each historical policy cycle, replace the policy terms in the original query statement with historical equivalent terms, generate an extended query set, perform semantic encoding on the extended query set to generate an extended query semantic vector, and generate temporal permission filtering conditions based on permission status constraints. The retrieval output module is used to filter the archive database using temporal permission filtering conditions, obtain a set of candidate archives that meet the permission status constraints, calculate the similarity between the extended query semantic vector and the temporal permission enhanced multimodal representation vector of each archive in the candidate archive set, sort the candidate archives according to the comprehensive similarity, and output the retrieval results.