Dynamic corpus elimination system and method based on large language model
By using a dynamic corpus elimination system based on a large language model, redundant and invalid content in international compliance corpora can be automatically identified and cleaned up. This solves the problems of lack of corpus elimination mechanism and cross-language recognition in existing technologies, and realizes efficient and automated updating of the knowledge base and cross-language consistency, reducing the reliance on manual intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies lack a corpus elimination mechanism in the field of international contract compliance, and cannot identify and clean up redundant, invalid or conflicting corpora. This leads to a decrease in knowledge base storage space and semantic retrieval accuracy. Furthermore, they lack the ability to identify cross-language and cross-textual structure substitution relationships, cannot ensure multi-source version comparison and trust level classification, rely on manual intervention, resulting in low efficiency, do not consider cross-language consistency, and cannot be automatically updated.
A dynamic corpus elimination system based on a large language model is adopted, including modules for new corpus acquisition, semantic analysis and substitution identification, multi-version comparison and confidence assessment, cross-language consistency verification, and corpus update and elimination execution. Through semantic elimination judgment, cross-version tracing and self-learning optimization, it automatically identifies and cleans up expired or replaced corpus, ensuring efficient updates of the knowledge base.
An automated corpus elimination mechanism has been implemented, which improves the storage efficiency and semantic retrieval accuracy of the corpus, reduces legal risks, enhances the ability to identify cross-language substitution relationships, ensures the consistency of comparison of multiple source versions, improves the automation of the update process, ensures the consistency of cross-language content, and continuously improves system performance through adaptive optimization.
Smart Images

Figure CN121860034A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer information processing and natural language processing, specifically to a dynamic corpus elimination system and method based on a large language model. Background Technology
[0002] Currently, artificial intelligence technology and natural language processing have developed rapidly, and text understanding, knowledge mining, and intelligent question answering systems based on large language models have been widely used in many fields. Particularly in the field of international compliance, various international organizations, government departments, and enterprises are increasingly relying on automated systems to parse, compare, and update legal provisions, international conventions, and policy documents. In this scenario, if the knowledge base cannot automatically determine which older corpora have been replaced or are invalid, it may retain redundant, conflicting, or invalid content for a long time, thereby reducing system efficiency and increasing the time cost of manual queries.
[0003] Existing dynamic knowledge update methods and systems based on industry-wide big data models automatically identify, integrate, and verify new and old knowledge through these models, thereby maintaining the timeliness and consistency of the industry knowledge base. However, existing technologies still have the following limitations and problems:
[0004] 1. Lack of a corpus elimination mechanism
[0005] The system only performs replacement or fusion operations when old and new knowledge coexist, without establishing a proactive mechanism for cleaning up redundant, invalid, and conflicting corpora. This leads to the accumulation of a large amount of outdated or overwritten corpora in the knowledge base over time, occupying storage space and reducing semantic retrieval accuracy. For international contract compliance corpora, the validity period and legal status of the clauses are highly sensitive, and the lack of an elimination mechanism will directly lead to semantic misrepresentation and legal risks.
[0006] 2. Insufficient recognition of semantic substitution relationships
[0007] Existing knowledge fusion algorithms mainly rely on semantic similarity calculation and manual verification, lacking the ability to handle complex cross-language and cross-textual structural substitution relationships. In international contract performance scenarios, clause substitution is often achieved through different languages, wording rewriting, or clause integration, making it difficult to identify the substitution logic through simple similarity comparison.
[0008] 3. No multi-source version comparison and trust level classification mechanism has been established.
[0009] In the field of international compliance, data often comes from documents or translations from different international organizations and member states. The lack of version comparison and trust rating mechanisms can lead to data conflicts and accidental deletions.
[0010] 4. The dynamic update process relies on manual intervention and lacks automation.
[0011] Human intervention is still required in key stages such as knowledge review, replacement confirmation, and version release. For a system like the International Convention Compliance Corpus, which is frequently updated and has a huge volume, manual intervention will significantly reduce efficiency.
[0012] 5. Failed to consider cross-language consistency and semantic drift issues.
[0013] Existing semantic understanding models are mainly based on monolingual corpora and cannot identify the consistency or substitution relationship of clauses in cross-language versions, while international performance documents usually exist in multiple languages such as Chinese, English, French, and Spanish. Summary of the Invention
[0014] To address the shortcomings of existing technologies, the present invention aims to provide a dynamic corpus elimination system and method based on a large language model. The present invention introduces semantic elimination judgment, cross-version tracing, confidence weighting, and self-learning optimization modules to automatically detect and eliminate corpus expired or replaced due to updates to international treaties or policy revisions, thereby achieving efficient iteration and accurate updating of corpus content.
[0015] To solve the above problems, the technical solution of the present invention is as follows:
[0016] A dynamic corpus elimination system based on a large language model includes:
[0017] The new corpus acquisition module is used to monitor and receive new or updated information on various performance documents;
[0018] The corpus storage module is used to store existing treaty corpora and their version information;
[0019] The semantic analysis and substitution identification module uses a large language model to perform semantic comparison between new and old corpora to discover substitution relationships.
[0020] The multi-version comparison and confidence assessment module is used to perform multi-angle comparison analysis and confidence calculation on candidate substitution relationships from the semantic analysis and substitution identification module in order to determine whether to substitute and perform elimination.
[0021] A cross-language consistency verification module is used to ensure the synchronous updating and elimination of multilingual versions of the corpus;
[0022] The corpus update and elimination module is used to perform actual data update and elimination operations on the knowledge base.
[0023] The monitoring and self-learning optimization module is used to continuously monitor and adaptively optimize the operation of the entire elimination mechanism.
[0024] Prior to this, when a new file or corpus is uploaded to the system, the new corpus acquisition module performs format parsing and preliminary metadata extraction on it, and generates an input corpus stream to be processed.
[0025] Prior to this, once the substitution relationship is confirmed by the multi-version comparison and confidence assessment module and the cross-language consistency verification module, the corpus update and elimination execution module performs the following actions on the knowledge base: First, it removes the old corpus and its related multilingual versions from the current valid corpus or marks them as obsolete; then, it adds the new corpus to the corpus storage module, stores it as the current valid version, and records its version information, source, and relationship link with the replaced clause; then, it updates the index and retrieval mechanism so that queries on this topic will prioritize retrieving the new corpus and no longer return the obsolete old corpus; finally, it generates an update log to record this substitution event. The corpus update and elimination execution module realizes dynamic updates to the corpus content, ensuring that new clauses take effect in a timely manner and old clauses are removed in a timely manner.
[0026] Prior to this, the monitoring and self-learning optimization module is used to continuously monitor and adaptively optimize the operation of the entire elimination mechanism. On the one hand, the monitoring and self-learning optimization module monitors the corpus update frequency, elimination accuracy, and the operating status of each module in the system in real time. When an abnormal situation occurs, the monitoring and self-learning optimization module will issue an alarm and execute remedial measures to ensure that the corpus content is credible and reliable. On the other hand, the monitoring and self-learning optimization module collects feedback from manual review and system operation data, and gradually optimizes the elimination rules and model parameters through self-learning algorithms.
[0027] Furthermore, the present invention also provides a dynamic corpus elimination method based on a large language model, comprising the following steps:
[0028] The system detects updates to the performance corpus and inputs and triggers new corpus data.
[0029] Semantic analysis and candidate matching are performed on the acquired new corpus;
[0030] Perform multi-version comparison and confidence assessment;
[0031] Perform cross-language consistency processing;
[0032] Perform corpus updates and deletion operations;
[0033] Perform results verification, continuous monitoring, and optimization.
[0034] Prior to this, the step of performing semantic analysis and candidate matching on the acquired new corpus specifically includes: performing in-depth analysis and semantic understanding on the acquired new corpus, and outputting a candidate replacement list, which contains the new corpus and one or more old corpus entries that it may replace.
[0035] Prior to this, the step of performing multi-version comparison and confidence assessment specifically includes: receiving the candidate substitution list, performing in-depth version history and confidence analysis on each pair of new and old corpora, calculating the substitution confidence through preset rules, and generating a confidence score for each candidate substitution relationship.
[0036] Prior to this, the step of performing cross-language consistency processing specifically includes: for multilingual old texts that are confirmed to be eliminated, they will all be included in the elimination list and processed together for cross-language updates. This can prevent the problem of inconsistency of information within the knowledge base caused by updating only one language, and ensure that users get the latest and valid terms information no matter which language they use to query.
[0037] Prior to this, the steps of updating and eliminating the corpus specifically include: processing old corpora one by one according to the elimination list, while adding new corpora to the corpus; generating an update log to record this replacement event, thereby realizing dynamic updates to the corpus content and ensuring that new clauses take effect in a timely manner and old clauses are removed in a timely manner.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) Establish an automated corpus elimination mechanism
[0040] To address the issue of existing knowledge bases lacking obsolete or outdated corpora, this invention introduces an intelligent identification method to automatically identify and remove expired or replaced corpora from newer treaties, preventing outdated content from remaining in the knowledge base for extended periods. This mechanism ensures that the corpus retains only valid and up-to-date treaties, reducing storage redundancy, improving the accuracy of corpus retrieval, and avoiding errors and potential legal risks caused by outdated information.
[0041] (2) Enhance the ability to identify semantic substitution relationships
[0042] This invention leverages the semantic understanding capabilities of large-scale models to establish multi-level semantic alignment and contextual logical reasoning methods, accurately identifying clause substitution relationships across different expressions and even languages. It can automatically determine whether a new document or clause substantially replaces or repeals an existing clause, thereby performing corresponding update or elimination operations and preventing conflicts arising from the coexistence of old and new clauses in the corpus.
[0043] (3) Introduce a multi-source version comparison and trust level classification mechanism
[0044] In response to the diverse sources and numerous versions of international compliance corpora, this invention establishes a version chain tracking system and a corpus credibility rating mechanism to compare and analyze relevant corpora from different sources or versions. Based on the authority of the source, publication time, and content relevance, different confidence levels are assigned to the corpora, automatically deciding which version to discard or retain in case of conflict.
[0045] (4) Improve the automation level of the update process
[0046] This invention employs a large-model-driven rule engine and an automatic verification module to automate the entire process of adding, replacing, and eliminating corpora. Through predefined rules and machine learning models, the system can autonomously complete replacement determination, conflict resolution, and version release, prompting human review only when necessary. This significantly reduces human intervention, not only improving processing efficiency but also mitigating delays and errors that may result from manual operations, enabling it to adapt to the high-frequency, large-scale needs of international treaty updates.
[0047] (5) Ensure consistent updates of cross-language content
[0048] This invention incorporates multilingual semantic alignment and translation consistency verification functions into its mechanism, ensuring that when treaty clauses in one language version are updated or repealed, the corresponding content in other languages can be identified and updated or phased out simultaneously. Through multilingual large-scale models or translation comparison technology, the system can detect equivalence relationships between texts in different languages, thereby maintaining the consistency of the corpus content across all languages and preventing inconsistencies and semantic drift caused by language differences.
[0049] (6) The system adapts and improves, resulting in excellent long-term performance.
[0050] By leveraging monitoring and self-learning optimization modules, the mechanism of this invention possesses the ability to gradually improve with use. Each updated experience is fed back to optimize the model and thresholds, enabling the system's capabilities to progressively increase. This also ensures that the invention maintains high performance and a low error rate during long-term operation. Simultaneously, the system can promptly detect and correct its own errors. Attached Figure Description
[0051] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0052] Figure 1 This is a block diagram of the dynamic corpus elimination system based on a large language model according to the present invention;
[0053] Figure 2 This is a flowchart of the dynamic corpus elimination method based on a large language model according to the present invention. Detailed Implementation
[0054] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0055] Specifically, this invention provides a dynamic corpus elimination system based on a large language model, such as... Figure 1 As shown, the system includes a new corpus acquisition module 101, a corpus storage module 102, a semantic analysis and substitution recognition module 103, a multi-version comparison and confidence assessment module 104, a cross-language consistency verification module 105, a corpus update and elimination execution module 106, and a monitoring and self-learning optimization module 107.
[0056] The new corpus acquisition module 101 is used to monitor and receive new or updated information on various implementation documents. This module can acquire new corpus from multiple sources, including convention update documents published by international organizations, implementation reports published by national governments, and related policy revision texts. When new documents or corpus are uploaded to the system, the new corpus acquisition module 101 performs format parsing and preliminary metadata extraction (such as treaty name, publication date, entry into force date, etc.) and generates an input corpus stream to be processed. The new corpus acquisition module 101 is also responsible for triggering the start signal of subsequent processing flows.
[0057] The corpus storage module 102 is used to store and manage existing treaty-related corpora, including historical versions of international treaty texts, legal provisions, and policy documents. The corpus storage module 102 can organize data using a relational database, document database, or graph database, and maintains metadata such as version information, language identifier, source, and effective / repealed status for each piece of corpus data. This module supports retrieval based on treaty number or topic and provides an interface for the alternative identification module to query potentially related older corpora.
[0058] The semantic analysis and substitution recognition module 103 is the core module of this invention, utilizing a large language model to perform in-depth semantic understanding and comparative analysis of the new corpus. The semantic analysis and substitution recognition module 103 first performs text preprocessing (such as word segmentation, syntactic parsing, named entity recognition, etc.) on the new corpus from the new corpus acquisition module 101, and then calls a pre-trained large model to generate the semantic representation vector or embedding of the text. Next, the semantic analysis and substitution recognition module 103 performs semantic comparison between the new corpus and existing clauses in the corpus storage module 102: on the one hand, it searches for potential old corpus related to the new corpus through keywords, treaty numbers, or citation relationships; on the other hand, it uses a large model to perform deep semantic matching of the new and old corpus texts, calculating similarity or semantic relevance scores. If certain existing clauses are highly similar in content to the new corpus, or if the new corpus explicitly mentions the repeal / replacement of an old clause (e.g., phrases such as "repealed..." or "replaced..."), the semantic analysis and substitution identification module 103 will determine the possibility of a substitution relationship and submit the candidate substitution pair to the multi-version comparison and confidence assessment module 104 for further verification.
[0059] The multi-version comparison and confidence assessment module 104 is used to perform multi-angle comparison analysis and confidence calculation on candidate substitution relationships from the semantic analysis and substitution identification module 103 to determine whether substitution is necessary and to eliminate the candidate. Specifically, the multi-version comparison and confidence assessment module 104 checks the historical version information and source reliability of the candidate old corpus through a "version chain tracking" mechanism. If the new corpus and the candidate old corpus belong to the same series (e.g., different years' revisions of the same treaty or revised versions of the same policy), the multi-version comparison and confidence assessment module 104 will follow the version chain to confirm whether the version number or publication date of the new corpus is later than the old version, to ensure that the identified substitution relationship is caused by an update iteration. At the same time, the multi-version comparison and confidence assessment module 104 calls a preset confidence level rule to compare the sources of the new and old corpus: if the new corpus comes from an authoritative institution and is published later, and the old corpus comes from an earlier version of the same source, then the substitution relationship is given a high confidence level; otherwise, if the source is unknown or the publication time is contradictory, the confidence level is reduced and it may be marked as requiring manual review. Furthermore, the multi-version comparison and confidence assessment module 104 comprehensively evaluates factors such as semantic matching score, content coverage (whether the new document content basically includes the main content of the old document), and domain knowledge (such as the continuity of legal clause numbers) to calculate a replacement confidence score. When this score is higher than a predetermined threshold, the multi-version comparison and confidence assessment module 104 confirms that the new corpus has indeed replaced the corresponding old corpus. For scores near the threshold, the system can decide whether to automatically eliminate the corpus or submit it for manual review based on its strategy.
[0060] The cross-language consistency verification module 105 is used to ensure the synchronous update of multilingual versions of the corpus. When the multi-version comparison and confidence assessment module 104 confirms that an old corpus needs to be eliminated, the cross-language consistency verification module 105 queries the corpus storage module 102 to retrieve other language versions (such as Chinese translations of English text) corresponding to the content of the old corpus. The cross-language consistency verification module 105 uses a multilingual semantic alignment model or machine translation technology to perform content matching on relevant clauses in different languages to determine whether they are actually different language expressions of the same clause. If it is confirmed that an old clause has equivalent versions in other languages, the cross-language consistency verification module 105 marks these equivalent versions as pending elimination, or selects to retain one language as the base version according to a strategy. Furthermore, if only some languages of the new corpus arrive in the system, while eliminating the old corpus, other language versions will be marked as "pending update" in the knowledge base, awaiting the input of new versions of the corpus in the corresponding languages. During the multilingual verification process, if it is found that the content of the new and old clauses is not completely consistent in different languages or there are differences in translation, the cross-language consistency verification module 105 can also call the large model for in-depth semantic comparison, and request manual review if necessary to ensure the accuracy of cross-language substitution.
[0061] The corpus update and elimination execution module 106 is responsible for executing the corpus update operation. After the replacement relationship is confirmed by the multi-version comparison and confidence assessment module 104 and the cross-language consistency verification module 105, the corpus update and elimination execution module 106 performs the following actions on the knowledge base: First, it removes or marks the old corpus and its related multi-language versions from the current valid corpus as "obsolete". This can be done through logical deletion or physical deletion, depending on system requirements. Then, it adds the new corpus to the corpus storage module 102 as the current valid version and records its version information, source, and relationship link with the replaced clause. Next, it updates the index and retrieval mechanism so that queries on this topic will prioritize retrieving the new corpus and no longer return obsolete old corpus. Finally, it generates an update log to record this replacement event, including the obsolete clause number, reason for obsolescence, and execution time. Through these steps, the corpus update and elimination execution module 106 achieves dynamic updates to the corpus content, ensuring that new clauses take effect promptly and old clauses are removed in a timely manner.
[0062] The monitoring and self-learning optimization module 107 is used to continuously monitor and improve the operation of the entire elimination mechanism. On the one hand, the monitoring and self-learning optimization module 107 monitors the corpus update frequency, elimination accuracy, and the operating status of each module in real time. When an anomaly occurs (such as mistakenly eliminating content that should not be eliminated, or omitting content that should be eliminated), the monitoring and self-learning optimization module 107 will issue an alarm and execute remedial measures to ensure the credibility and reliability of the corpus content. On the other hand, the monitoring and self-learning optimization module 107 collects feedback from manual reviewers and system operation data, and gradually optimizes the elimination rules and model parameters through self-learning algorithms. For example, when a manual reviewer identifies an incorrect substitution judgment, the system can feed that instance back to the large model for training and fine-tuning, enhancing the model's ability to judge similar contexts; or adjust the semantic matching threshold and confidence calculation weights to improve future judgment accuracy. Through continuous monitoring and adaptive optimization, the mechanism of this invention can continuously learn and improve during long-term operation, further enhancing the robustness and accuracy of the system.
[0063] Furthermore, this invention also provides a dynamic corpus elimination method based on a large language model, such as... Figure 2 As shown, the method includes the following steps:
[0064] S1: The system detects the update event of the performance corpus and inputs and triggers new corpus;
[0065] Specifically, the system monitors the update events of the treaty implementation corpus through the new corpus acquisition module 101. If a new treaty text or revised document is uploaded, the new corpus acquisition module 101 parses the new corpus, extracts basic information, and sends it into the subsequent process, while issuing an update trigger signal.
[0066] S2: Perform semantic analysis and candidate matching on the acquired new corpus;
[0067] Specifically, under the trigger signal, the semantic analysis and substitution identification module 103 is activated to perform in-depth analysis and semantic understanding of the new corpus obtained in step S1. Using a pre-trained large language model, the semantic analysis and substitution identification module 103 extracts the semantic features of the new corpus and retrieves existing clauses that may be affected or related in content from the corpus storage module 102. This retrieval includes: direct lookup based on metadata (such as matching treaty names and numbers), and fuzzy query based on vector semantics (by calculating the semantic similarity between the new corpus and each corpus in the library). For each retrieved candidate old clause, the semantic analysis and substitution identification module 103 performs a comparative analysis. If the new corpus is highly similar to the content of an old clause, or if the new corpus text contains statements referencing the old clause's number or name, the semantic analysis and substitution identification module 103 marks it as a possible substitute. Furthermore, if the new corpus belongs to a specific domain and is significantly later than important documents in the same domain in the library, the semantic analysis and substitution identification module 103 will also intelligently infer that the new corpus may be an update to an old document. In this process, the large model provides a high-level understanding of the meaning and context of the text, enabling the system to go beyond simple keyword matching and identify clause pairs that appear to differ in wording but have a substantive relationship of inheritance. Finally, the semantic analysis and substitution identification module 103 outputs a list of "candidate substitution pairs," which contains new corpus entries and one or more old corpus entries that they may substitute, and passes this list to the next step of processing.
[0068] S3: Perform multi-version comparison and confidence assessment;
[0069] Specifically, after receiving the candidate replacement list generated in step S2, the multi-version comparison and confidence assessment module 104 performs in-depth version history and confidence analysis on each pair of new and old corpora. The multi-version comparison and confidence assessment module 104 first uses version chain information to confirm the chronological relationship between the new and old corpora in the document sequence: if the new corpus and the candidate old corpus belong to the same treaty series, the multi-version comparison and confidence assessment module 104 checks whether the publication or signing date of the new corpus is later than that of the old corpus, and verifies details such as version number and amendment number to ensure that the new corpus is indeed the later version. In addition, the multi-version comparison and confidence assessment module 104 examines the status of the old corpus (such as whether the old corpus has been marked as expired but has not yet been cleaned up, or whether it has been updated), avoiding duplicate elimination or ignoring known obsolescence information. After verifying the version relationship, the multi-version comparison and confidence assessment module 104 calculates the substitution confidence level according to preset rules, including semantic matching score, the length coverage of the new and old texts (the degree to which the new text covers the key points of the old text), the differences in content between the new and old documents, and domain-specific characteristics (such as whether the legal clause numbering continues and whether the wording changes conform to revision conventions). Simultaneously, a trust leveling mechanism is introduced to consider the reliability of the data source: for example, if the new corpus comes from an official gazette and has legal effect, while the old corpus is only a draft or informal translation, the new substitution has a higher confidence level; conversely, if the source of the new corpus is questionable or has not been officially confirmed, the system tends to retain the old corpus for verification. After considering the above factors, the multi-version comparison and confidence assessment module 104 generates a confidence score for each candidate substitution relationship. If the score is higher than the threshold T, the old and new corpora are confirmed to have a substitution relationship, and the process proceeds to the next step. If the score is lower than the threshold, the automatic elimination process is abandoned (the old corpora can be retained and the new corpora temporarily stored, or the relationship can be marked for manual review). If the score is near the threshold, the system configuration determines whether to process automatically or review manually. This step ensures that the elimination operation is only performed under highly certain circumstances, thereby minimizing the risk of misjudgment.
[0070] S4: Perform cross-language consistency processing;
[0071] Specifically, for the substitution relationships confirmed in step S3, the cross-language consistency verification module 105 will intervene to ensure that all language versions involved are updated synchronously. The cross-language consistency verification module 105 queries the corpus storage module 102 to obtain a complete list of all stored language versions of the old corpus to be replaced. For example, if it is confirmed that Article 5 of the English version of an international convention needs to be phased out, the cross-language consistency verification module 105 will retrieve the corresponding Article 5 text of the convention in other languages such as Chinese and French. Subsequently, the cross-language consistency verification module 105 uses a multilingual alignment algorithm to check the new and old content: one case is that the new corpus itself contains multiple language versions (such as uploading Chinese, English, and French texts at the same time), then the cross-language consistency verification module 105 will perform substitution confirmation for each language separately; another case is that the current new corpus only has one language (for example, only the English revised version), then the cross-language consistency verification module 105 will attempt to translate the new text into other languages through machine translation or a multilingual large model for semantic comparison to determine whether the old texts that have not yet been updated in these languages are equivalent to the new text content. If they are equivalent, it is inferred that these old texts should also be phased out and marked as pending obsolescence. If the translation comparison reveals that the content of the old and new texts does not match in other languages, it indicates that there may be cross-language inconsistencies and marks them as abnormal for manual processing. For multilingual old texts confirmed to be phased out, the cross-language consistency verification module 105 will include them all in the phase-out list and process the cross-language update together. This can prevent the problem of inconsistencies in the internal information of the knowledge base caused by updating only one language, and ensure that users get the latest and valid terms information regardless of which language they use to query.
[0072] S5: Perform corpus update and elimination operations;
[0073] In this step, the corpus update and elimination execution module 106 modifies the corpus storage module 102 to implement elimination and update operations. First, the corpus update and elimination execution module 106 processes old corpora item by item according to the elimination list: for each eliminated clause and its multilingual version, its status is set to "eliminated" or "repealed." Specifically, the validity flag of these records can be invalidated in the database, or they can be moved to the historical archive area, but their content is retained for future reference. Simultaneously, new corpora are added to the corpus: if the new corpus is a revision or replacement of an old clause, the associated old clause number is referenced in the new record to establish a version association; if the new corpus is a completely new treaty clause, it is added as a new record normally. Subsequently, the corpus update and elimination execution module 106 generates a version update log for this update: the log records which clauses were added, which clauses were eliminated, the operation time, and the execution result status. During this process, if any elimination or addition operation fails, the corpus update and elimination execution module 106 will promptly report an error and attempt to revert to the previous operation to ensure that the knowledge base is not in an inconsistent state. With this step, the corpus content has been updated.
[0074] S6: Perform result verification, continuous monitoring, and optimization.
[0075] After the above update is executed, the monitoring and self-learning optimization module 107 verifies the update results and performs subsequent monitoring. The monitoring and self-learning optimization module 107 reads the update log and the current state of the corpus generated in step S5, checking whether the expected elimination and addition clauses are correctly reflected in the corpus. Then, the monitoring and self-learning optimization module 107 may perform an automatic verification. If any discrepancies are found (e.g., a certain elimination clause can still be retrieved by a normal query), the monitoring and self-learning optimization module 107 will mark the anomaly and notify maintenance personnel for handling. In addition, this step is also part of continuous monitoring. The monitoring and self-learning optimization module 107 will record this substitution judgment and execution status in the training dataset for subsequent optimization of the model and rules. For example, if a substitution decision is overturned during manual review, the system will add the negative sample to the large model training set, enabling the model to identify situations where substitution should not be performed in similar circumstances. Conversely, for successfully automatically replaced cases, the system accumulates positive samples to enhance the model's judgment ability. The monitoring and self-learning optimization module 107 also dynamically adjusts the threshold T or the weights of each evaluation factor mentioned in step S3 based on long-term accumulated data to adapt to changing environmental characteristics. Through this iterative monitoring and learning process, the system performance will become increasingly optimized, and the error rate will gradually decrease, achieving improvements in the dynamic updating of the dynamic corpus.
[0076] Furthermore, the above steps can be performed in parallel or iteratively as needed. For example, when multiple new corpora are entered simultaneously, the system can perform semantic analysis and matching on them in parallel, and consider the interrelationships between multiple documents during the version comparison phase. Moreover, for some complex large-scale treaty updates, which may require multiple rounds of semantic alignment and manual confirmation, the system can also break down the process into multiple iterative processes.
[0077] The following example illustrates this using the corpus elimination related to the updating of international tuna conservation measures:
[0078] Suppose that a regional organization's latest revision of its tuna protection measures is adopted in 2025, and a Chinese version of the revised text is published. The previous version of the measures (2017 version) exists in both Chinese and English versions in the corpus. The following will explain the entire process of how the old version of the measures is automatically identified and phased out after the new revised text is imported into this mechanism.
[0079] First, new corpus is imported. The Chinese text of "Tuna Conservation Measures X (Revised Edition)" released in 2025 is uploaded to the mechanism through the new corpus acquisition module 101. Module 101 identifies the basic information of the file, such as the file name "Tuna Conservation Measures", the revision number "2025 Revision", the language "Chinese", and the issuing organization "a certain regional organization", and stores this metadata along with the text content in the temporary storage area, while triggering the system to enter the processing flow.
[0080] Next, semantic analysis and matching, and the semantic analysis and substitution identification module 103 reads the Chinese text of the revised version of the new protection measure X. Through NLP preprocessing of the text, module 103 extracts a series of clauses and their numbers contained in the protection measure. Then, the large language model generates a semantic vector representation for each new clause. Module 103 accesses the corpus storage module 102 to retrieve existing records related to "Tuna Protection Measure X" in the region organization, finding clauses for "Tuna Protection Measure X (2017 Edition)" in the corpus, including both Chinese and English versions. Based on this, module 103 infers a version inheritance relationship between the new corpus and the 2017 version of the protection measure; therefore, it focuses on analyzing the correspondence between the clauses of the two. Through semantic comparison using the large model, module 103 finds that the new revision basically corresponds to the old version in terms of clause structure, but the content has been updated. For example, Article 10 of the new version is highly similar to Article 10 of the old version, except for changes in the numerical values of the fishing standards; Article 15 is a new clause added in the new version, while there is no corresponding content in the old version; Article 30 of the old version has been deleted or merged into other clauses in the new version. Based on these findings, Module 103 constructed a list of candidate replacement pairs, including: Article 10 of the old version <-> Article 10 of the new version (content updated replacement), Article 30 of the old version <-> (none, this clause is deleted in the new version), etc. At the same time, Module 103 noted that the introduction of the revised text of the new protection measure X explicitly states that "this revision repeals the corresponding clauses of the 2017 version from the date of its entry into force." This key content will be understood by the large model to strengthen the replacement judgment. Finally, Module 103 submits these candidate replacement relationships and their contextual information to Module 104 for confirmation.
[0081] Then, the version and confidence assessment module 104 reviews the aforementioned candidate relationships. Module 104 examines the version chain information of the 2017 and 2025 versions of protection measure X, confirming that the 2025 version is indeed the latest version of the protection measure, the 2017 version is the previous version, and the issuing regional organizations are the same, with the same level of legal force (both are official texts). Therefore, all candidate substitution relationships involving the 2017 version are reasonable in terms of version sequence. Next, module 104 calculates the confidence level of each candidate substitution based on multiple factors. For example, when the old Article 10 is replaced by the new Article 10, the semantic similarity is very high. The old and new articles have the same title and number, and the content changes are limited. The new article clearly inherits the position of the old article, and the confidence score is close to 1.0 (very certain). For the old Article 30 (deleted): the new version does not have an article with the same number, but through content comparison, it is found that the new Articles 29 and 31 cover most of the topics of the old Article 30. The new preamble also mentions the repeal of several old articles, which implicitly includes the old Article 30. In this case of indirect substitution, the semantic matching is slightly lower, but with the repeal statement in the preamble, Module 104 still gives a high confidence score (e.g., 0.85). At the same time, Module 104 considers the source factor: because the text of the new protection measures comes from the official website of the regional organization, it is a highly credible source, and the old article is also an official document version, so there is no source conflict issue, and the confidence score does not need to be deducted. For each substitution pair, Module 104 has reached the preset threshold, and therefore decides to confirm these substitution relationships.
[0082] Secondly, cross-language synchronization verification is performed. Since the old version of the protection measure X corpus exists in both Chinese and English, the cross-language consistency verification module 105 begins operation. Module 105 first locates the corresponding clauses (Articles 10 and 30) in the English version of the old 2017 protection measure X and verifies their correspondence with the Chinese version. It confirms that the old Chinese and English versions of Article 10 are the same clause in different languages. Since the new Chinese Article 10 replaces the old Chinese Article 10, it should also replace the old English Article 10. Therefore, module 105 adds "English Article 10 (2017 version)" to the obsolescence list. Similarly, the old English version of Article 30 also needs to be obsolescence. Since the currently input 2025 version is in Chinese, and the new English version has not yet been entered into the system, module 105, based on experience, anticipates that the new English version will be released soon, but it is not yet available. Therefore, Module 105 adopts a synchronous marking strategy: marking Article 30 of the old English version as obsolete, while simultaneously recording in the mechanism that "it needs to wait for the import of the English version of Protection Measure X 2025 to complete the transition between the old and new versions." For Article 15 (in Chinese) added in the new version, which is not present in the old version and has no corresponding content in the old English version, no obsolescence operation is required. However, Module 105 indicates that a corresponding English Article 15 should be added after the new English version is published. Through this cross-language processing, the system can ensure that both the Chinese and English clauses of the old 2017 Protection Measure X are properly handled, preventing a situation where the Chinese database is updated while the English database still retains outdated content.
[0083] Next, the update execution module 106 performs a batch update of the corpus based on the above confirmation results. In the corpus storage 102, module 106 marks all clause records of the 2017 version of Protection Measure X as "repealed (replaced by the 2025 version)". These records include: Article 10 of the old Chinese version, Article 10 of the old English version, Article 30 of the old Chinese version, Article 30 of the old English version, etc. The annotation process is implemented in the database by setting fields, while retaining their association with the 2025 new version for traceability (e.g., adding a reference "Article X of the 2025 version replaces this article" to the old clause metadata). Subsequently, module 106 adds the new clauses of the 2025 version of Protection Measure X to the corpus one by one: for each new record, it fills in its clause number, content, language, version information (2025 revised version), etc., and establishes version links with the old clauses when necessary (such as Article 10). Since the new English version was not included, only the new Chinese version clauses were added this time. Clause 15 is a completely new clause, and its "previous version" is marked as "none (newly added)". Module 106 updated the search index so that searching for "Protection Measure X Clause 10" will return the 2025 version content instead of the 2017 version. After all database update operations are completed, Module 106 generates an update log: recording information such as "Protection X (2025 revised version) Chinese clauses have been added, a total of N clauses; Protection Measure X (2017) Chinese Clause 10, English Clause 10, Chinese Clause 30, and English Clause 30 are marked as invalid; other clauses of the old version are still valid", and attaching a timestamp. The log shows that all operations were successful and without errors.
[0084] Subsequently, during the results verification and feedback phase, the monitoring and self-learning optimization module 107 checked the update. Module 107 automatically retrieved key content such as "Protection Measure X Article 10 (2017)," confirming that the search results no longer presented the old version or showed its status as repealed. Simultaneously, the search for "Protection Measure X Article 10" yielded the 2025 new version, verifying the replacement's effectiveness. The query results for the old version of Article 30 were also invalid, proving successful elimination. Module 107 noted that the new English version had not yet been imported; therefore, although the old English clause was marked as repealed, there were no corresponding new clauses, and this was highlighted in the health report. The entire update process was conducted without human intervention, and all judgments were correct. Module 107 added this case as a positive sample to the training set to enhance the large model's ability to recognize explicit replacement expressions such as "the revised version repeals the old version." The wording in the introduction of the new convention will also be learned by the model so that it can more quickly make elimination judgments when encountering similar phrases in the future (such as "the previous version is repealed from the date this revision comes into effect"). The update process is considered complete only after the system operators review the update logs and health report, approve the results, and raise no objections.
[0085] In practical applications, this invention's mechanism can efficiently and effectively update and eliminate treaty compliance corpora. For the introduction of Protective Measure X in the 2025 version, the system automatically identifies and eliminates the corresponding clauses of Protective Measure X in the 2017 version, including multilingual versions, achieving a synchronized update across the entire database. The entire process takes only a few minutes, far faster than the several days required for manual comparison and revision. Simultaneously, relying on a large model's understanding of the deep meaning of the text, the system can also properly handle implicit substitution relationships, ensuring no content that needs to be eliminated is overlooked. This not only reduces the burden of manual maintenance but also improves the reliability and timeliness of the knowledge base. Especially in the context of frequent updates to conventions and regulations, this invention's dynamic corpus elimination mechanism ensures that all participating parties obtain the latest and consistent treaty information at any time, possessing profound practical significance.
[0086] Furthermore, the mechanism of this invention is also applicable to the dynamic maintenance of other types of corpora. For example, in the field of international trade agreements, when a member state promulgates new regulations to replace old ones, the system can automatically identify and update the replacement; and duplicate clauses from different sources can also be selected to retain the authoritative version and replace the secondary version based on confidence level. Therefore, this invention has broad applicability. The technical solution of this invention can be applied to scenarios that utilize large-scale model semantic understanding and automatic decision-making capabilities to update and clean up corpora.
[0087] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A dynamic corpus elimination system based on a large language model, characterized in that, The system includes: The new corpus acquisition module is used to monitor and receive new or updated information on various performance documents; The corpus storage module is used to store existing treaty corpora and their version information; The semantic analysis and substitution identification module uses a large language model to perform semantic comparison between new and old corpora to discover substitution relationships. The multi-version comparison and confidence assessment module is used to perform multi-angle comparison analysis and confidence calculation on candidate substitution relationships from the semantic analysis and substitution identification module in order to determine whether to substitute and perform elimination. A cross-language consistency verification module is used to ensure the synchronous updating and elimination of multilingual versions of the corpus; The corpus update and elimination module is used to perform actual data update and elimination operations on the knowledge base. The monitoring and self-learning optimization module is used to continuously monitor and adaptively optimize the operation of the entire elimination mechanism.
2. The dynamic corpus elimination system based on a large language model according to claim 1, characterized in that, When a new file or corpus is uploaded to the system, the new corpus acquisition module performs format parsing and preliminary metadata extraction on it, and generates an input corpus stream to be processed.
3. The dynamic corpus elimination system based on a large language model according to claim 1, characterized in that, After the substitution relationship is confirmed by the multi-version comparison and confidence assessment module and the cross-language consistency verification module, the corpus update and elimination execution module performs the following actions on the knowledge base: First, it removes the old corpus and its related multilingual versions from the current valid corpus or marks them as obsolete; then, it adds the new corpus to the corpus storage module, stores it as the current valid version, and records its version information, source, and relationship links with the replaced clause; then, it updates the index and retrieval mechanism so that queries on this topic will prioritize retrieving the new corpus and no longer return the obsolete old corpus; finally, it generates an update log to record this substitution event. The corpus update and elimination execution module realizes dynamic updates to the corpus content, ensuring that new clauses take effect in a timely manner and old clauses are removed in a timely manner.
4. The dynamic corpus elimination system based on a large language model according to claim 1, characterized in that, The monitoring and self-learning optimization module is used to continuously monitor and adaptively optimize the operation of the entire elimination mechanism. On the one hand, the monitoring and self-learning optimization module monitors the corpus update frequency, elimination accuracy, and the operating status of each module in the system in real time. When an abnormal situation occurs, the monitoring and self-learning optimization module will issue an alarm and execute remedial measures to ensure that the corpus content is credible and reliable. On the other hand, the monitoring and self-learning optimization module collects feedback from manual review and system operation data, and gradually optimizes the elimination rules and model parameters through self-learning algorithms.
5. A dynamic corpus elimination method based on a large language model, characterized in that, The method includes the following steps: The system detects updates to the performance corpus and inputs and triggers new corpus data. Semantic analysis and candidate matching are performed on the acquired new corpus; Perform multi-version comparison and confidence assessment; Perform cross-language consistency processing; Perform corpus updates and deletion operations; Perform results verification, continuous monitoring, and optimization.
6. The dynamic corpus elimination method based on a large language model according to claim 5, characterized in that, The steps of semantic parsing and candidate matching of the acquired new corpus specifically include: performing in-depth parsing and semantic understanding of the acquired new corpus, and outputting a candidate replacement list, which contains the new corpus and one or more old corpus entries that it may replace.
7. The dynamic corpus elimination method based on a large language model according to claim 5, characterized in that, The steps of performing multi-version comparison and confidence assessment specifically include: receiving the candidate substitution list, performing in-depth version history and confidence analysis on each pair of new and old corpora, calculating the substitution confidence through preset rules, and generating a confidence score for each candidate substitution relationship.
8. The dynamic corpus elimination method based on a large language model according to claim 5, characterized in that, The steps for cross-language consistency processing specifically include: for multilingual old texts that are confirmed to be phased out, they will all be included in the phase-out list and processed together for cross-language updates. This can prevent the problem of inconsistency in the internal information of the knowledge base caused by updating only one language, and ensure that users get the latest and valid terms information no matter which language they use to query.
9. The dynamic corpus elimination method based on a large language model according to claim 5, characterized in that, The steps for updating and eliminating corpora specifically include: processing old corpora one by one according to the elimination list, while adding new corpora to the corpus; generating an update log to record the replacement event, thereby realizing dynamic updates to the corpus content and ensuring that new clauses take effect in a timely manner and old clauses are removed in a timely manner.