Updating driven file content identification optimization method and system

By constructing a hierarchical noun explanation rule base and distributed monitoring nodes, generating multi-scenario simulated training corpora, and optimizing the file content recognition model, the problem of low accuracy caused by dynamic updates of noun explanations is solved, and efficient adaptive and self-optimization capabilities are achieved.

CN121562606BActive Publication Date: 2026-04-10BEIJING MIANBI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, dynamic updates of terminology definitions are ineffective, resulting in low document recognition accuracy. Traditional rule bases struggle to adapt to knowledge changes, while static rule bases lead to poor recognition performance.

Method used

A hierarchical rule base for term explanation is constructed, including a basic rule layer and a dynamic update layer. Version consistency is checked through distributed monitoring nodes, multi-scenario simulated training corpus is generated, and quality filtering is performed through a semantic consistency validator. A hierarchical parameter update mechanism is adopted to optimize the model, and a self-correction process is introduced to adjust priority and association weights.

Benefits of technology

It significantly improves the accuracy and robustness of document content recognition systems in dynamically updated terminology environments, enabling them to proactively adapt to changes in external knowledge, coordinate internal conflicts, and continuously self-optimize, thus solving the problems of static terminology explanations and delayed updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562606B_ABST
    Figure CN121562606B_ABST
Patent Text Reader

Abstract

The application discloses a file content recognition optimization method and system driven by name explanation update, relates to the technical field of file content recognition, and comprises the following steps: constructing a hierarchical name explanation rule library; periodically checking the version consistency of name explanation data sources through a lightweight verification protocol; when it is detected that the version difference exceeds a preset fault tolerance threshold, triggering an incremental update process; based on a conflict resolution mechanism in the hierarchical name explanation rule library, generating multi-scenario simulation training corpus; using the multi-scenario simulation training corpus, optimizing a file content recognition model through a hierarchical parameter update mechanism, and performing name explanation recognition on input file text; based on the training effect feedback of the multi-scenario simulation training corpus, starting a self-correction process of the hierarchical name explanation rule library, and readjusting the priority and correlation weight of name explanation. The application solves the technical problem of low file recognition accuracy caused by poor dynamic update effect of name explanation in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of file content recognition, in particular to a file content recognition optimization method and system driven by term explanation update. BACKGROUND

[0002] In the field of natural language processing and file content recognition, accurately recognizing and understanding the term explanations in text is crucial for information extraction, knowledge management, and intelligent question answering. Traditional file content recognition systems usually rely on predefined term explanation rule libraries, which are built based on authoritative terminology libraries and standard documents. These rule libraries form a standardized storage structure by extracting term explanations. This method can provide high recognition accuracy in the initial stage. However, as knowledge continues to update and evolve, term explanations themselves also change, such as the emergence of new terms, the revision of old terms, or the expansion of semantics. This makes it difficult for static rule libraries to adapt to a dynamically changing environment, resulting in poor file content recognition results. SUMMARY

[0003] The present application provides a file content recognition optimization method and system driven by term explanation update, which is used to solve the technical problem of low file recognition accuracy caused by poor dynamic update of term explanations in the prior art.

[0004] In view of the above problems, the present application provides a file content recognition optimization method and system driven by term explanation update.

[0005] In a first aspect, the present application provides a file content recognition optimization method driven by term explanation update, comprising:

[0006] constructing a hierarchical term explanation rule library, wherein the hierarchical term explanation rule library includes a basic rule layer and a dynamic update layer;

[0007] deploying distributed monitoring nodes to periodically check the version consistency of term explanation data sources through a lightweight verification protocol, and triggering an incremental update process when detecting that the version difference exceeds a preset fault tolerance threshold;

[0008] generating multi-scenario simulation training corpus based on the conflict resolution mechanism in the hierarchical term explanation rule library, wherein the multi-scenario simulation training corpus contains mixed text instances of new and old term explanations, and is quality filtered by a semantic consistency verifier;

[0009] using the multi-scenario simulation training corpus to optimize the file content recognition model through a hierarchical parameter update mechanism, and applying the optimized file content recognition model to perform term explanation recognition on the input file text;

[0010] When any one of the accuracy, recall rate, or F1 value is lower than a preset threshold, a self-correction process of the hierarchical noun explanation rule library is started, and the priority and correlation weight of the noun explanation are readjusted.

[0011] In a second aspect, the application provides a file content recognition optimization system driven by noun explanation updating, comprising:

[0012] A rule library construction module is configured to construct a hierarchical noun explanation rule library, wherein the hierarchical noun explanation rule library comprises a basic rule layer and a dynamic updating layer.

[0013] A monitoring updating module is configured to deploy distributed monitoring nodes, periodically check the version consistency of the noun explanation data source through a lightweight verification protocol, and trigger an incremental updating process when detecting that the version difference exceeds a preset fault tolerance threshold.

[0014] A corpus generation module is configured to generate a multi-scenario simulation training corpus based on a conflict resolution mechanism in the hierarchical noun explanation rule library, wherein the multi-scenario simulation training corpus contains mixed text instances of new and old noun explanations, and is subjected to quality filtering by a semantic consistency verifier.

[0015] A noun explanation recognition module is configured to use the multi-scenario simulation training corpus to optimize a file content recognition model through a hierarchical parameter updating mechanism, and apply the optimized file content recognition model to perform noun explanation recognition on an input file text.

[0016] A self-correction module is configured to start a self-correction process of the hierarchical noun explanation rule library based on the training effect feedback of the multi-scenario simulation training corpus, and readjust the priority and correlation weight of the noun explanation when any one of the accuracy, recall rate, or F1 value is lower than a preset threshold.

[0017] One or more technical solutions provided in the application have at least the following technical effects or advantages:

[0018] The application provides a file content recognition optimization method and system driven by term explanation update, which significantly improves the accuracy, robustness and self-adaptive ability of a file content recognition system in a term explanation dynamic update environment by constructing a hierarchical term explanation rule library including a basic rule layer and a dynamic update layer, deploying a distributed monitoring node based on a lightweight authentication protocol, generating high-quality multi-scenario simulation training corpus by using a conflict resolution mechanism, optimizing the model by using a hierarchical parameter update mechanism based on conflict types, and introducing a self-correction process based on performance feedback. Compared with the traditional method, the technical scheme provided by the application significantly overcomes the problem of recognition performance decline caused by term explanation static, update lag and version conflict, and builds an intelligent file content recognition method which can actively adapt to external knowledge changes, has internal conflict coordination ability and can be continuously self-optimized, thereby providing a complete, efficient and reliable solution to the technical problem of term explanation dynamic evolution. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0020] Figure 1 The flowchart of the term explanation update driven file content recognition optimization method provided by the embodiment of the present application.

[0021] Figure 2 The structure diagram of the term explanation update driven file content recognition optimization system provided by the embodiment of the present application.

[0022] In the drawings, the components represented by the numbers are described as follows:

[0023] The rule library construction module 100, the monitoring update module 200, the corpus generation module 300, the term explanation recognition module 400 and the self-correction module 500. DETAILED DESCRIPTION

[0024] The present application provides a term explanation update driven file content recognition optimization method and system, which is used to solve the technical problem of low file recognition accuracy caused by poor term explanation dynamic update effect in the prior art.

[0025] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0026] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units need not be limited to only those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to the process, method, product or device.

[0027] Embodiment one, as shown in the present application provides a file content identification optimization method driven by term explanation update, wherein the method comprises: Figure 1

[0028] S10: Construct a hierarchical term explanation rule library, wherein the hierarchical term explanation rule library comprises a basic rule layer and a dynamic update layer.

[0029] In the initial construction stage of the file content identification system, how to construct a rule library that can guarantee the authority and stability of the basic term explanation and flexibly adapt to the dynamic changes of the term explanation in the future is a technical problem to be solved. The traditional single-structure rule library stores all term explanations in a mixed manner, lacking hierarchical management. When introducing externally added or revised term explanations, due to the lack of independent and safe storage areas and effective verification mechanisms, the core and stable basic rules are easily directly contaminated, resulting in a decrease in data reliability.

[0030] The step S10 in the method provided by the embodiment of the present application comprises:

[0031] Based on the authoritative term library and the standard document, the core term explanation and the corresponding associated semantic features are extracted, a normalized storage structure containing synonym sets, antonym sets, context usage patterns and domain-specific constraint rules is constructed as the basic rule layer;

[0032] Through a digital signature verification mechanism based on an asymmetric encryption algorithm, the externally added term explanations and the revised rules are received and verified, the identity of the update source is authenticated and the data integrity is checked, and the verified content is stored as the dynamic update layer;

[0033] A query priority rule, a conflict resolution mechanism and a dynamic update strategy are established between the basic rule layer and the dynamic update layer;

[0034] ​The query priority rule, the conflict resolution mechanism and the dynamic updating strategy between the basic rule layer and the dynamic updating layer are established, including:

[0035] The query priority rule is set, and the core term explanation of the basic rule layer is preferentially queried in the file content recognition process. The dynamic updating layer is automatically triggered for query only when the corresponding explanation is not matched in the basic rule layer;

[0036] The conflict resolution mechanism is constructed, and the term explanations in the basic rule layer and the dynamic updating layer are compared for consistency based on the time stamp and the version number. When the explanation content conflict is detected, the explanation with the latest version number and verified by the digital signature is used as the reference;

[0037] The dynamic updating strategy is configured. For the term explanation processed by the conflict resolution mechanism, the basic rule layer is updated in synchronization according to the preset rule, and the version iteration record is reserved.

[0038] In the embodiment of the application, a hierarchical term explanation rule library is constructed, wherein the hierarchical term explanation rule library includes a basic rule layer and a dynamic updating layer.

[0039] Specifically, first, the core term explanation and the corresponding associated semantic features are extracted based on the authoritative term library and the standard document, a normalized storage structure containing the synonym set, the antonym set, the context usage mode and the domain-specific constraint rule is constructed as the basic rule layer. For example, the authoritative term library can be a legal regulation database, a judicial interpretation compilation, etc. The core term such as “personal” is extracted, and the synonym set such as “natural person” is collected; the antonym set such as “collective” is collected; the context usage mode such as the typical description sentence in the legal clause is collected; and the domain-specific constraint rule such as the limited use scene is collected. The information is arranged into a normalized storage structure, and a relational database MySQL is used for storage. The table structure includes the term ID, the term text, the explanation text, the synonym list, the antonym list, the context mode, the domain-specific constraint rule and the like fields, so as to obtain the basic rule layer with stable content and clear structure, and provide reliable and fast query data basis for the core term explanation.

[0040] Further, the newly added external term explanation and revised rules are received and verified by a digital signature verification mechanism based on an asymmetric encryption algorithm, the identity of the update source is authenticated and the data integrity is checked, and the verified content is stored as a dynamic update layer. For example, a dynamic update layer is constructed, and the source of the newly added external term explanation, such as newly released judicial interpretation, and the asymmetric encryption algorithm RSA are used for digital signature verification. When receiving the newly added or revised term explanation, the digital signature is verified by the OpenSSL tool to ensure that the identity of the update source is trusted and the data is complete and intact; for example, the publisher signs the explanation content with a private key, and the system verifies the signature match with a public key, and then stores the content such as the explanation of “natural person”, version number v2.0, and timestamp 2025-10-01 in the MySQL database of the dynamic update layer, to ensure that the newly added term explanation can be stored in the database only after the source verification, and to prevent unauthorized modification.

[0041] Further, a query priority rule, a conflict resolution mechanism, and a dynamic update strategy are established between the basic rule layer and the dynamic update layer. Specifically, the query priority rule is set, and in the file content identification process, the core term explanation of the basic rule layer is queried first, and only when the basic rule layer does not match the corresponding explanation, the dynamic update layer is automatically triggered to query. For example, the query “personal data” does not have a record in the basic layer, and the dynamic update layer is triggered to query.

[0042] Further, a conflict resolution mechanism is constructed, and the term explanations in the basic rule layer and the dynamic update layer are compared for consistency based on the timestamp and the version number. When the explanation content conflict is detected, the explanation with the latest version number and passing the digital signature verification is used as the reference. For example, the version v1.0 of “personal data” in the basic rule layer is explained as “information that can identify personal identity”, and the version v2.0 in the dynamic update layer is explained as “information that can identify personal identity, including behavior characteristic data”. By comparing the version number and the timestamp, the explanation with the latest version number and passing the digital signature verification is used as the reference, and the explanation with the latest version number and passing the digital signature verification is used as the reference.

[0043] Further, a dynamic update strategy is configured, and for the term explanation processed by the conflict resolution mechanism, such as “personal data” v2.0, the preset rule such as periodic batch update is used to update the basic rule layer, and the version iteration record is retained, which includes the v1.0 explanation and the update time of v2.0, to support query traceability and version rollback.

[0044] By constructing a hierarchical noun explanation rule library, the rule data is structurally isolated and cooperatively managed in logic and physics, and significant technical effects are achieved. The establishment of the basic rule layer ensures the normativity and stability of the core noun explanation and its semantic network from the authoritative term library and standard documents, providing a reliable knowledge cornerstone for the entire system. The dynamic update layer ensures that all external additions or revisions can be entered into the library under the premise of reliable source and complete content, effectively isolating the potential risk of unreliable updates to the core knowledge.

[0045] S20: Deploy distributed monitoring nodes to periodically check the version consistency of the noun explanation data source through a lightweight authentication protocol. When a version difference exceeding a preset fault tolerance threshold is detected, an incremental update process is triggered.

[0046] Traditional methods often rely on periodic manual checks or full updates. The former is slow to respond and cannot meet the timeliness requirements in a rapidly iterating knowledge environment, while the latter consumes excessive network and computing resources when dealing with large amounts of data, resulting in low efficiency. In addition, the update process lacks effective triggering and control mechanisms, and minor version changes may trigger unnecessary frequent updates, while major updates may be delayed due to a lack of timely awareness.

[0047] The step S20 in the method provided by the embodiment of the present application comprises:

[0048] In a distributed architecture, multiple monitoring nodes are deployed. Each monitoring node establishes a secure connection with the noun explanation data source through a lightweight authentication protocol and performs version consistency verification at a preset time interval;

[0049] When the monitoring node detects that the version number of the noun explanation data source and the local rule library version difference exceeds a preset fault tolerance threshold, an incremental update process is automatically triggered, and only the noun explanation data with version difference is obtained;

[0050] During the incremental update process, the obtained noun explanation data is subjected to integrity check and digital signature verification. After verification, the corresponding noun explanation data is temporarily stored in the pending review area, and after passing the semantic consistency check, it is synchronized to the dynamic update layer.

[0051] In the embodiment of the present application, distributed monitoring nodes are deployed to periodically check the version consistency of the noun explanation data source through a lightweight authentication protocol. When a version difference exceeding a preset fault tolerance threshold is detected, an incremental update process is triggered.

[0052] Specifically, multiple monitoring nodes are deployed in a distributed architecture, each of which establishes a secure connection with the nomenclature data source through a lightweight authentication protocol and performs version consistency verification at a preset time interval. For example, multiple monitoring nodes are deployed in a distributed architecture, and a secure connection is established with the central nomenclature data source through a pre-shared key-based digest authentication protocol, and version consistency verification is automatically performed at a preset 24-hour time interval to ensure security.

[0053] Further, when the monitoring node detects that the version number of the nomenclature data source and the local rule library version difference exceeds the preset fault tolerance threshold, an incremental update process is automatically triggered, and only the nomenclature data of the version difference part is obtained. For example, the monitoring node periodically compares the data source version number, the data source version v2.1, and the local rule library version v1.0, and sets the preset fault tolerance threshold to 0.5. When the absolute value of the version difference is greater than or equal to the fault tolerance threshold, an incremental update process is automatically triggered, i.e. only the nomenclature data of the version difference part is requested from the nomenclature data source. By constructing an intelligent update triggering mechanism, only when a substantial version change occurs, the update is started, reducing unnecessary resource consumption.

[0054] Further, in the incremental update process, the obtained nomenclature data is subjected to integrity check and digital signature verification, and after verification, the corresponding nomenclature data is temporarily stored in the pending review area and synchronized to the dynamic update layer after passing the semantic consistency check. Illustratively, the digital signature of the data source is verified using the RSA asymmetric encryption algorithm. After verification, the corresponding nomenclature data such as the newly added "natural person" explanation is temporarily stored in the pending review area of the MySQL database. Further, the nomenclature data temporarily stored in the pending review area is subjected to semantic consistency check. The specific method is: through a pre-constructed semantic consistency verifier, the logical relationship between the newly added nomenclature explanation and the related concepts in the existing rule library is checked, and the contents with semantic contradiction such as the conflicting "collective" explanation of the "individual" definition are excluded. After passing the check, the data is synchronized to the dynamic update layer.

[0055] By deploying distributed monitoring nodes and combining lightweight authentication protocols and incremental update processes, an efficient, energy-saving and reliable rule library synchronization mechanism is effectively constructed. The deployment of distributed nodes ensures that the system continuously monitors the data source version from multiple locations, and the lightweight authentication protocol meets the basic security check requirements while minimizing the computational overhead caused by periodic verification. The introduction of the preset fault tolerance threshold and the automatic triggering mechanism realizes intelligent management of the update process, only starts the update when there is a substantial version difference, avoids unnecessary resource waste, and improves synchronization efficiency.

[0056] S30: generating multi-scenario simulation training corpus based on the conflict resolution mechanism in the hierarchical term explanation rule library, wherein the multi-scenario simulation training corpus contains mixed text instances of new and old term explanations, and quality filtering is performed by a semantic consistency verifier.

[0057] The prior art often directly uses original update data to train a model. Due to lack of systematic simulation of conflict scenarios, the model is difficult to handle complex text environments in which new and old term explanations coexist and alternate in reality. If the training corpus only contains a single version of explanation or simply mixes new and old explanations without identification and resolution, the model will be confused when encountering conflict instances, and the recognition accuracy will decrease.

[0058] The step S30 in the method provided in the embodiments of the present application includes:

[0059] extracting term explanation pairs with version differences or semantic conflicts from the basic rule layer and the dynamic update layer, and constructing mixed text instances containing new and old term explanations based on conflict types, wherein the conflict types include complete replacement type conflict, partial update type conflict and semantic expansion type conflict;

[0060] combining the mixed text instances into a training corpus set according to a preset proportion according to the feature distribution of the actual application scenario, wherein each training corpus unit contains a basic version explanation, an updated version explanation and a corresponding context environment description;

[0061] performing quality filtering on the training corpus set by a pre-constructed semantic consistency verifier, removing text instances with semantic contradictions or logical inconsistencies, and retaining training corpus that meets the semantic consistency requirements;

[0062] In the training corpus filtered by quality, the version source and conflict resolution result of each term explanation are labeled to form a multi-scenario simulation training corpus with version tracing capability.

[0063] In the embodiments of the present application, based on the conflict resolution mechanism in the hierarchical term explanation rule library, a multi-scenario simulation training corpus is generated, wherein the multi-scenario simulation training corpus contains mixed text instances of new and old term explanations, and quality filtering is performed by a semantic consistency verifier.

[0064] Specifically, the pairs of explanations of terms with version differences or semantic conflicts are extracted from the basic rule layer and the dynamic update layer, and mixed text instances containing new and old explanations of terms are constructed based on the conflict types, wherein the conflict types include complete replacement type conflict, partial update type conflict and semantic extension type conflict. For example, the v1.0 version of “natural person” in the basic rule layer is explained as “a person in the biological sense”, while the v2.0 version of “natural person” in the dynamic update layer is explained as “a legal subject of rights”, which constitutes a complete replacement type conflict pair; the v1.0 version of “personal data” in the basic rule layer is explained as “any information related to an identified or identifiable natural person”, while the v2.0 version of “personal data” in the dynamic update layer is explained as “any information related to an identified or identifiable natural person, but excluding anonymized information”, which constitutes a partial update type conflict pair; the v1.0 version of “personal information” in the basic rule layer is explained as “data that can identify a specific individual”, while the v2.0 version of “personal information” in the dynamic update layer is explained as “data that can identify a specific individual alone or in combination with other information”, which constitutes a semantic extension type conflict pair.

[0065] Further, according to the feature distribution of the actual application scene, the mixed text instances are combined into a training corpus set according to a preset ratio, wherein each training corpus unit contains a basic version explanation, an updated version explanation and a corresponding context environment description. Specifically, the scene distribution in the actual application is analyzed, and the mixed text instances are combined according to the ratio of complete replacement type:partial update type:semantic extension type = preset ratio A: preset ratio B: preset ratio C, wherein the preset ratio A, the preset ratio B and the preset ratio C are obtained based on the feature distribution of the actual application scene, for example, if there are more partial updates in the actual application scene, a higher preset ratio B is set, for example, the ratio is set to 3:4:3, and finally each training corpus unit contains a basic version explanation, an updated version explanation and a corresponding context environment description.

[0066] Further, the training corpus set is filtered for quality by a pre-constructed semantic consistency verifier, and text instances with semantic contradictions or logical inconsistencies are removed, and training corpus meeting the semantic consistency requirements is retained. Specifically, a semantic consistency verifier based on a rule base is used, a built-in semantic conflict rule base is loaded, and each training corpus unit input is subjected to lexical analysis and structure analysis, and conflict rule patterns are matched. For example, when it is detected that the new and old explanations have logical contradictions, such as defining “personal” to include minors and defining “personal” to exclude minors, the sample is automatically marked and filtered to remove text instances with semantic contradictions or logical inconsistencies, and training corpus meeting the semantic consistency requirements is retained.

[0067] Further, in the training corpus filtered by quality, the version source and conflict resolution result of each explanation of noun are labeled to form a multi-scenario simulation training corpus with version tracing capability. For example, the version source of the explanation of noun is the basic rule layer v1.0, and the conflict resolution result label, such as the merged explanation, has complete version tracing capability through labeling.

[0068] The method provided in the application can capture version differences and semantic conflicts from the rule library, and construct highly simulated mixed text instances accordingly, ensuring the diversity and representativeness of the training corpus in conflict types and scene distribution. Through the combination of preset proportions and strict semantic consistency verification, noise data with logical confusion and semantic contradiction is effectively eliminated, ensuring the overall quality and logical consistency of the training corpus set. Finally, the corpus is labeled with version source and conflict resolution result, providing the model with crucial learning signals, enabling it not only to recognize the explanation of noun itself, but also to understand its version evolution context and conflict resolution logic.

[0069] S40: using the multi-scenario simulation training corpus, optimizing the file content recognition model through a hierarchical parameter updating mechanism, and applying the optimized file content recognition model to perform recognition of the explanation of noun for the input file text.

[0070] After obtaining high-quality multi-scenario training corpus, how to utilize these corpora to optimize the existing file content recognition model without destroying its existing stable knowledge structure is a technical problem to be solved. The traditional global parameter updating strategy adjusts all weights of the model indiscriminately, which can learn new knowledge, but is likely to cause forgetting of the old stable knowledge, i.e., the model performs significantly worse in recognizing a large number of unchanged original explanations of noun while improving in recognizing new explanations of noun.

[0071] The method provided in the application includes step S40:

[0072] According to the explanation of noun data in the hierarchical explanation of noun rule library, multi-scenario simulation training corpora are collected, and the explanation of noun in each training corpus is labeled to obtain a sample content recognition set, wherein each sample content recognition includes the recognition content of all explanations of noun in the training corpus and corresponding version source information;

[0073] Based on machine learning, a network architecture of a file content recognition model is constructed;

[0074] The multi-scenario simulation training corpus and the sample content recognition set are used to supervise the training of the file content recognition model until the verification accuracy converges, and the construction of the file content recognition model is completed;

[0075] Determine the updateable parameter proportion dynamically based on the conflict types annotated in the multi-scenario simulation training corpus, wherein the highest update proportion corresponds to the complete replacement type conflict, the medium update proportion corresponds to the partial update type conflict, and the lowest update proportion corresponds to the semantic expansion type conflict;

[0076] Incrementally train the new parameter set using the multi-scenario simulation training corpus, select the corresponding proportion of parameters from the new parameter set for weight update according to the updateable parameter proportion, adjust the parameter weight through the small-batch gradient descent algorithm, and dynamically adjust the gradient clipping threshold during the training process using the adaptive gradient clipping.

[0077] Mark the optimized model parameters with versions and establish a mapping relationship with the corresponding explanatory rule library version to form a versioned file content recognition model.

[0078] In the embodiments of the present application, the multi-scenario simulation training corpus is used to optimize the file content recognition model through a hierarchical parameter update mechanism, and the optimized file content recognition model is applied to perform explanatory recognition on the input file text. Specifically, according to the explanatory data in the hierarchical explanatory rule library, multi-scenario simulation training corpus is collected, and the explanatory content in each training corpus is annotated to obtain a sample content recognition set, wherein each sample content recognition includes the recognition content of all explanatory in the training corpus and the corresponding version source information. For example, a sample content recognition can be [personal information, data that can identify a specific individual alone or in combination with other information, v2.0 version].

[0079] Further, based on machine learning, the network architecture of the file content recognition model is constructed. For example, the BERT model is used as the basic architecture, which contains 12 layers of Transformer encoder, each layer has 768 hidden units, and the number of attention heads is 12. A context-based text classification task is used as the training target, and a fully connected layer is added at the end of the model as a classifier, and the output dimension corresponds to the number of classification categories of the explanatory.

[0080] Further, the file content recognition model is supervised trained by using the multi-scene simulation training corpus and the sample content recognition set until the verification accuracy converges, and the construction of the file content recognition model is completed. Specifically, the multi-scene simulation training corpus and the sample content recognition set are divided into a training set and a verification set according to a ratio of 8:2, the multi-scene simulation training corpus in the training set is taken as an input, and the corresponding sample content recognition set is taken as a supervised label. The difference between the predicted result and the true label is calculated by using a cross-entropy loss function. The model parameters are updated by using a small batch gradient descent algorithm, the batch size is set to 32, the learning rate is set to 0.00001, the training is stopped when the accuracy on the verification set no longer significantly improves for a plurality of training periods, and the construction of the file content recognition model is completed.

[0081] Further, based on the conflict types labeled in the multi-scene simulation training corpus, the updateable parameter proportion is dynamically determined, wherein the completely replacement conflict corresponds to the highest update proportion, the partial update conflict corresponds to the medium update proportion, and the semantic expansion conflict corresponds to the lowest update proportion. Specifically, the conflict type labeled in each sample in the training corpus is analyzed. For the completely replacement conflict such as the definition of “natural person” updated from the biological concept to the legal concept, the highest updateable parameter proportion is set, for example, 80%; for the partial update conflict such as the definition of “personal data” adding the anonymization processing exception content, the medium updateable parameter proportion is set, for example, 50%; and for the semantic expansion conflict such as the definition of “personal information” expanding the recognition mode, the lowest updateable parameter proportion is set, for example, 20%. Through this process, the parameter update strategy dynamically adjusted according to different conflict types is obtained.

[0082] Further, the incremental training of the new parameter set is performed by using the multi-scene simulation training corpus, and the parameters of the corresponding proportion are selected from the new parameter set for weight update according to the updateable parameter proportion. The parameter weight is adjusted by using a small batch gradient descent algorithm, and the adaptive gradient clipping is used in the training process to dynamically adjust the gradient clipping threshold. Specifically, the parameters of the corresponding proportion are randomly selected from the new parameter set for weight update according to the preset updateable parameter proportion, and the remaining parameters are kept frozen. The parameter weight is adjusted by using a small batch gradient descent algorithm, and the cross-entropy loss is used as the loss function. The adaptive gradient clipping is used in the training process. Specifically, the ratio of the gradient norm to the weight norm of each parameter group is calculated in real time when the model weight is updated, and if the ratio exceeds the preset critical point, the gradient value is automatically reduced in proportion, so that the best learning step is maintained while avoiding unstable training.

[0083] Further, the optimized model parameters are version marked and mapped with the corresponding noun explanation rule base version to form a versioned file content recognition model. After the model training is completed, a unique version identifier such as "model_v2.1" is generated for the current model parameters, and the corresponding relationship between the model version and the dependent noun explanation rule base version such as "rulebase_v2.0" is recorded in the version mapping table, so as to have a versioned file content recognition model with complete version tracing capability.

[0084] Further, the file text to be recognized is input into the optimized file content recognition model, and the model interprets and recognizes the nouns in the text based on the learned parameters, and outputs the explanation content corresponding to each noun and its version source information.

[0085] By implementing the hierarchical parameter updating mechanism, precise and stable optimization of the file content recognition model is realized. According to the conflict types marked in the training corpus, the update proportion of different parameter sets in the model is dynamically determined, and a conservative update strategy is adopted for parameters related to core knowledge structure, while parameters responsible for adapting to new changes are adjusted. This differentiated updating method can effectively inject new noun explanation knowledge while maximizing the retention of the model's memory of stable noun explanation, significantly alleviating the catastrophic forgetting problem and ensuring smooth transition of the overall performance of the model. Finally, the optimized model not only improves the recognition ability of new noun explanation and conflict scenarios, but also maintains the recognition accuracy of the original knowledge, thereby achieving comprehensive and stable performance improvement.

[0086] S50: Based on the training effect feedback of the multi-scene simulation training corpus, when any of the accuracy, recall rate or F1 value is lower than a preset threshold, the hierarchical noun explanation rule base self-correction process is started, and the priority and correlation weight of noun explanation are readjusted.

[0087] The rule base itself may have configuration defects after initial construction or subsequent dynamic update, such as unreasonable priority setting of some noun explanations, or mismatched correlation weight with context features. These potential internal problems of the rule base will directly lead to poor quality of the generated training corpus, and ultimately result in the recognition performance indicators such as accuracy of the optimized model failing to meet expectations.

[0088] In the embodiments of the present application, based on the training effect feedback of the multi-scene simulation training corpus, when any one of the accuracy, recall rate or F1 value is lower than the preset threshold, the hierarchical noun explanation rule library self-correction process is started, and the priority and correlation weight of noun explanation are re-adjusted. Among them, the accuracy measures the proportion of correctly identified samples, the recall rate measures the proportion of correctly identified samples among the samples that should be identified, and the F1 value is the harmonic mean of the accuracy and the recall rate. In the model verification stage, when any one of the indicators, for example, the accuracy, is lower than the preset threshold, the self-correction process will be automatically triggered, and the priority and correlation weight of noun explanation will be re-adjusted. For example, samples with an identification confidence lower than a certain threshold are extracted from the training log, for example, the identification accuracy of the model for “independent individual” is continuously low, and through analysis it is found that the priority of the noun in the rule library is set too low, and the correlation weight with the key context such as “right subject” is insufficient. Further, the priority of “independent individual” in the basic rule layer is increased from a lower level to a medium level, and the correlation weight with the context features such as “right subject” and “natural person” is increased. The weight adjustment adopts a weighting algorithm based on co-occurrence frequency, and the new weight is equal to the original weight × (1+co-occurrence frequency), so as to obtain the optimized and adjusted noun explanation rule library. The updated rule library is used to regenerate the training corpus, and a new self-correction record is added in the version record table, including the correction version number, the correction time, the specific noun adjusted, the priority change and the weight adjustment basis.

[0089] Further, the method provided in the embodiments of the present application further comprises:

[0090] The rule library hot update strategy is implemented to complete version switching of the hierarchical noun explanation rule library under the premise of ensuring continuous operation, and version recoverability when updating fails is ensured through a version rollback mechanism.

[0091] The running state of the deployed noun explanation rule library is monitored based on the distributed monitoring node, and the consistency of the rule library version and the noun explanation data source is verified regularly.

[0092] The incremental update process is continuously triggered according to the node monitoring result, the dynamic synchronization of the hierarchical noun explanation rule library and the noun explanation data source is maintained, and a closed-loop update for continuous optimization is formed.

[0093] Specifically, in the embodiments of the present application, the rule library hot update strategy is implemented to complete version switching of the hierarchical noun explanation rule library under the premise of ensuring continuous operation, and version recoverability when updating fails is ensured through a version rollback mechanism. Specifically, a dual-database switching mechanism is adopted, and the updated rule library is deployed to a standby database during continuous operation, and the consistency is ensured through data synchronization, and then the new version is quickly switched to. When an update exception is detected, the last stable version is automatically rolled back.

[0094] Based on the distributed monitoring nodes, the running state of the deployed term explanation rule library is monitored, and the consistency of the rule library version and the term explanation data source is verified regularly. Specifically, through the deployed multiple monitoring nodes, whether the online rule library version is consistent with the central data source version is checked every fixed time, such as 24 hours. When it is found that the versions are inconsistent, an abnormal log is recorded and a warning is executed.

[0095] According to the node monitoring result, an incremental update process is continuously triggered to maintain the dynamic synchronization of the hierarchical term explanation rule library and the term explanation data source, and a continuous optimization closed loop is formed. Specifically, when the monitoring node detects that the data source has a new version published, the incremental update process is automatically triggered, the changed term explanation content is downloaded and updated, the timeliness of the rule library is maintained, and a continuous optimization closed loop is formed.

[0096] By establishing a self-correction process based on training effect feedback, the entire recognition method is given the ability of self-diagnosis and self-optimization, forming a continuous improvement closed loop. When the performance index of the optimized model improves, the rule library may have configuration problems. Locate the difficult samples and deduce the potential defects of the rule library in priority or correlation weight. Then dynamically adjust these configuration parameters, optimize the internal representation structure of knowledge, and put the adjusted rule library back into the training corpus generation cycle, so that the recognition method is no longer passively responding to external knowledge changes, but can actively optimize the rationality of internal knowledge organization, thereby fundamentally improving the quality of training data and the potential of subsequent model optimization.

[0097] Embodiment two, as shown in Figure 2 based on the same inventive concept of the term explanation update driven file content recognition optimization method provided in embodiment one, the present embodiment also provides a term explanation update driven file content recognition optimization system, comprising:

[0098] The rule library construction module 100 is configured to construct a hierarchical term explanation rule library, wherein the hierarchical term explanation rule library comprises a basic rule layer and a dynamic update layer.

[0099] The monitoring and updating module 200 is configured to deploy distributed monitoring nodes, periodically check the version consistency of the term explanation data source through a lightweight verification protocol, and trigger an incremental update process when detecting that the version difference exceeds a preset fault tolerance threshold.

[0100] The corpus generation module 300 is configured to generate multi-scenario simulation training corpora based on the conflict resolution mechanism in the hierarchical term explanation rule library, wherein the multi-scenario simulation training corpora contain mixed text instances of new and old term explanations, and are filtered in quality by a semantic consistency verifier.

[0101] The noun explanation identification module 400 is configured to adopt the multi-scene simulation training corpus, optimize the file content identification model through a hierarchical parameter updating mechanism, and apply the optimized file content identification model to perform noun explanation identification on the input file text.

[0102] The self-correction module 500 is configured to start a self-correction process of the hierarchical noun explanation rule library, readjust the priority and correlation weight of the noun explanation based on the training effect feedback of the multi-scene simulation training corpus when any index of the accuracy, recall rate or F1 value is lower than a preset threshold.

[0103] In one embodiment, the rule library construction module 100 is further configured to:

[0104] Based on the authoritative term library and the standard document, core noun explanations and corresponding correlation semantic features are extracted, and a normalized storage structure containing a synonym set, an antonym set, a context use mode and a domain-specific constraint rule is constructed as a basic rule layer.

[0105] Through a digital signature verification mechanism based on an asymmetric encryption algorithm, external newly added noun explanations and revised rules are received and verified, and the update source is authenticated and the data integrity is checked, and the verified content is stored as a dynamic update layer.

[0106] A query priority rule, a conflict resolution mechanism and a dynamic update strategy are established between the basic rule layer and the dynamic update layer.

[0107] The query priority rule, the conflict resolution mechanism and the dynamic update strategy are established between the basic rule layer and the dynamic update layer, including:

[0108] The query priority rule is set, and the core noun explanation of the basic rule layer is preferentially queried in the file content identification process, and the dynamic update layer is automatically triggered for query only when the corresponding explanation is not matched in the basic rule layer.

[0109] The conflict resolution mechanism is constructed, and the noun explanations in the basic rule layer and the dynamic update layer are compared for consistency based on a time stamp and a version number, and when the explanation content conflict is detected, the explanation with the latest version number and passing the digital signature verification is used as the reference;

[0110] The dynamic update strategy is configured, and for the noun explanation processed by the conflict resolution mechanism, the basic rule layer is updated in synchronization according to a preset rule, and the version iteration record is reserved.

[0111] In one embodiment, the monitoring update module 200 is further configured to:

[0112] A plurality of monitoring nodes are deployed in a distributed architecture, each monitoring node establishes a secure connection with a glossary data source through a lightweight authentication protocol, and performs version consistency verification at a preset time interval;

[0113] When the monitoring node detects that the glossary data source version number and the local rule library version difference exceeds a preset fault tolerance threshold, an incremental update process is automatically triggered, and only the glossary data of the version difference part is obtained;

[0114] In the incremental update process, the obtained glossary data is subjected to integrity check and digital signature verification, and after verification, the corresponding glossary data is temporarily stored in the waiting for review area, and after passing the semantic consistency check, it is synchronized to the dynamic update layer.

[0115] In one embodiment, the corpus generation module 300 is also used for:

[0116] Extracting glossary explanation pairs with version differences or semantic conflicts from the basic rule layer and the dynamic update layer, and constructing mixed text instances containing new and old glossary explanations based on conflict types, wherein the conflict types include complete replacement type conflict, partial update type conflict and semantic extension type conflict;

[0117] According to the feature distribution of the actual application scene, the mixed text instances are combined into a training corpus set according to a preset proportion, wherein each training corpus unit contains a basic version explanation, an updated version explanation and a corresponding context environment description;

[0118] The training corpus set is subjected to quality filtering by a pre-constructed semantic consistency verifier, and text instances with semantic contradictions or logical inconsistencies are removed, and training corpus meeting the semantic consistency requirements is retained;

[0119] In the training corpus filtered by quality, the version source and conflict resolution result of each glossary explanation are labeled to form a multi-scene simulation training corpus with version tracing capability.

[0120] In one embodiment, the glossary explanation identification module 400 is also used for:

[0121] According to the glossary explanation data in the hierarchical glossary explanation rule library, multi-scene simulation training corpus is collected, and the glossary explanation content in each training corpus is labeled to obtain a sample content identification set, wherein each sample content identification includes the identification content of all glossary explanations in the training corpus and the corresponding version source information;

[0122] Based on machine learning, a network architecture of a file content identification model is constructed;

[0123] The file content recognition model is supervised trained by using the multi-scene simulation training corpus and the sample content recognition set until the verification accuracy converges, and the construction of the file content recognition model is completed.

[0124] Based on the conflict types labeled in the multi-scene simulation training corpus, the updateable parameter proportion is dynamically determined, wherein the completely replacement type conflict corresponds to the highest update proportion, the partial update type conflict corresponds to the medium update proportion, and the semantic expansion type conflict corresponds to the lowest update proportion.

[0125] The incremental training is performed on the new parameter set by using the multi-scene simulation training corpus, and the corresponding proportion of parameters is selected from the new parameter set according to the updateable parameter proportion to update the weight, the parameter weight is adjusted by using the small batch gradient descent algorithm, and the adaptive gradient clipping is used to dynamically adjust the gradient clipping threshold during the training process.

[0126] The optimized model parameters are marked with a version, and a mapping relationship is established with the corresponding version of the noun explanation rule library, so as to form a versioned file content recognition model.

[0127] In one embodiment, the self-correcting module 500 is further configured to:

[0128] The rule library hot update strategy is implemented to complete the version switching of the hierarchical noun explanation rule library under the premise of ensuring continuous operation, and the version recoverability when the update fails is ensured by using the version rollback mechanism.

[0129] Based on the distributed monitoring node, the running state of the deployed noun explanation rule library is monitored, and the consistency of the rule library version and the noun explanation data source is verified regularly.

[0130] According to the node monitoring result, the incremental update process is continuously triggered to maintain the dynamic synchronization of the hierarchical noun explanation rule library and the noun explanation data source, and a closed loop update for continuous optimization is formed.

[0131] In summary, the embodiments of the present application have at least the following technical effects:

[0132] The application provides a file content recognition optimization method and system driven by term explanation update. By constructing a hierarchical term explanation rule library including a basic rule layer and a dynamic update layer, deploying distributed monitoring nodes based on a lightweight verification protocol, generating high-quality multi-scenario simulation training corpus by using a conflict resolution mechanism, optimizing the model by using a hierarchical parameter update mechanism based on conflict types, and introducing a self-correction process based on performance feedback, the accuracy, robustness and self-adaptive ability of the file content recognition system in the dynamic update environment of term explanation are significantly improved. Specifically, through the hierarchical rule library structure and strict verification mechanism, the normativity and traceability of term explanation data are ensured, and a safe and reliable entrance for dynamic update is provided, effectively reducing the risk caused by false or malicious updates to the system; through the distributed monitoring and incremental update process, near-real-time perception and synchronization of changes in the source of term explanation data are realized, greatly improving the timeliness of the rule library and avoiding recognition errors caused by outdated information; through the multi-scenario simulation training corpus generated based on the conflict resolution mechanism, the file content recognition model can fully learn and adapt to the complex language environment where new and old term explanations coexist and alternate, enhancing the model's generalization ability to semantic changes and reducing confusion caused by rule conflicts; through the hierarchical parameter update mechanism, fine management of model optimization is realized, different degrees of model adjustment are applied to different properties of update content, new knowledge is effectively absorbed while the memory of stable knowledge is maximized, and the overall performance of the model is ensured to transition smoothly and improve continuously; finally, through the integration of the adaptive correction closed loop of training effect feedback, the recognition method has the ability of self-perception and continuous optimization, and can automatically trigger the re-adjustment of the rule library and model parameters when the recognition performance fluctuates, thereby maintaining a high recognition accuracy in the long term. Compared with the traditional method, the technical scheme provided by the application significantly overcomes the problem of recognition performance decline caused by the static term explanation, update lag and version conflict, and builds an intelligent file content recognition method that can actively adapt to external knowledge changes, has internal conflict coordination ability, and can be continuously self-optimized, providing a complete, efficient and reliable solution to the technical problem of dynamic evolution of term explanation.

[0133] It should be noted that the above sequence of the embodiments of the application is only for description, and does not represent the advantages and disadvantages of the embodiments. Moreover, the above describes specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

[0134] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0135] The specification and drawings are only exemplary and illustrative of the present application and are considered to cover any and all modifications, variations, combinations or equivalents that are within the scope of the present application. Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the scope of the present application. Thus, it is intended that the present application cover the modifications and changes as they come within the scope of the application, and that the scope of the application be limited only by the claims.

Claims

1. A method for optimizing file content recognition driven by a terminology explanation update, characterized in that: The method includes: Construct a hierarchical noun definition rule base, wherein the hierarchical noun definition rule base includes a basic rule layer and a dynamically updated layer; Deploy distributed monitoring nodes and periodically check the consistency of the data source version through a lightweight verification protocol. When the version difference exceeds the preset fault tolerance threshold, trigger the incremental update process. Based on the conflict resolution mechanism in the hierarchical noun explanation rule base, a multi-scenario simulation training corpus is generated. The multi-scenario simulation training corpus contains mixed text instances of new and old noun explanations and is quality filtered by a semantic consistency verifier. Using the aforementioned multi-scenario simulated training corpus, the file content recognition model is optimized through a hierarchical parameter update mechanism. The optimized file content recognition model is then applied to perform term explanation recognition on the input file text. Based on the training effect feedback of the multi-scenario simulated training corpus, when any of the indicators such as accuracy, recall, or F1 score is lower than a preset threshold, the self-correction process of the hierarchical noun explanation rule base is initiated to readjust the priority and association weight of noun explanations.

2. The file content recognition optimization method driven by terminology explanation update according to claim 1, characterized in that, The steps for building a hierarchical noun definition rule base include: Based on authoritative terminology databases and standard documents, we extract core terminology explanations and corresponding related semantic features, and construct a standardized storage structure that includes a set of synonyms, a set of antonyms, contextual usage patterns, and domain-specific constraint rules, which serves as the basic rule layer. The digital signature verification mechanism based on asymmetric encryption algorithm receives and verifies newly added terminology definitions and revision rules from external sources. After verifying the identity and data integrity of the update source, the verified content is stored as a dynamic update layer. Establish query priority rules, conflict resolution mechanisms, and dynamic update strategies between the basic rule layer and the dynamic update layer.

3. The file content recognition optimization method driven by terminology explanation update according to claim 2, characterized in that, Establish query priority rules, conflict resolution mechanisms, and dynamic update strategies between the basic rule layer and the dynamic update layer, including: Set query priority rules to prioritize querying the core term explanations of the basic rule layer during the file content recognition process, and automatically trigger the dynamic update layer query only when the basic rule layer does not match the corresponding explanation; A conflict resolution mechanism is constructed, which compares the definitions of terms in the basic rules layer and the dynamic update layer based on timestamps and version numbers. When a conflict is detected, the definition with the latest version number and verified by digital signature shall prevail. Configure a dynamic update strategy. For the definitions of terms that have been processed by the conflict resolution mechanism, update them synchronously to the basic rule layer according to preset rules, while retaining version iteration records.

4. The file content recognition optimization method driven by terminology explanation update according to claim 1, characterized in that, Deploy distributed monitoring nodes to periodically check the consistency of the glossary data source version using a lightweight verification protocol. When a version difference exceeds a preset fault tolerance threshold, an incremental update process is triggered, including: In a distributed architecture, multiple monitoring nodes are deployed. Each monitoring node establishes a secure connection with the glossary data source through a lightweight verification protocol and performs version consistency checks at preset time intervals. When the monitoring node detects that the difference between the version number of the terminology data source and the version of the local rule base exceeds a preset fault tolerance threshold, it automatically triggers an incremental update process and only obtains the terminology data for the version difference portion. During the incremental update process, the acquired glossary data undergoes integrity verification and digital signature verification. After successful verification, the corresponding glossary data is temporarily stored in the review area. After passing the semantic consistency check, it is synchronized to the dynamic update layer.

5. The file content recognition optimization method driven by the terminology explanation update according to claim 1, characterized in that, Based on the conflict resolution mechanism in the hierarchical noun explanation rule base, multi-scenario simulated training corpus is generated, including: Extract noun explanation pairs with version differences or semantic conflicts from the basic rule layer and the dynamic update layer, and construct a mixed text instance containing old and new noun explanations based on the conflict type, wherein the conflict type includes complete substitution conflict, partial update conflict and semantic expansion conflict; Based on the feature distribution of actual application scenarios, the mixed text instances are combined into a training corpus according to a preset ratio. Each training corpus unit contains a basic version explanation, an updated version explanation, and a corresponding context description. The training corpus is quality filtered by a pre-built semantic consistency validator to remove text instances with semantic contradictions or logical inconsistencies, and retain the training corpus that meets the semantic consistency requirements. In the training corpus that has passed quality filtering, the version source and conflict resolution results of each term definition are labeled to form a multi-scenario simulation training corpus with version tracing capabilities.

6. The file content recognition optimization method driven by terminology explanation update according to claim 1, characterized in that, The steps for constructing the document content recognition model include: Based on the term explanation data in the hierarchical term explanation rule base, multi-scenario simulated training corpus is collected, and the term explanation content in each training corpus is labeled to obtain a sample content recognition set. Each sample content recognition includes the recognition content of all term explanations in the training corpus and the corresponding version source information. A network architecture for building a file content recognition model based on machine learning; The document content recognition model is trained under supervision using the multi-scenario simulation training corpus and sample content recognition set until the verification accuracy converges, thus completing the construction of the document content recognition model.

7. The file content recognition optimization method driven by terminology explanation update according to claim 1, characterized in that, Using the aforementioned multi-scenario simulated training corpus, the file content recognition model is optimized through a hierarchical parameter update mechanism, including: Based on the conflict types annotated in the multi-scenario simulation training corpus, the proportion of updatable parameters is dynamically determined. Among them, the highest update proportion corresponds to the complete substitution type conflict, the medium update proportion corresponds to the partial update type conflict, and the lowest update proportion corresponds to the semantic expansion type conflict. The new parameter set is incrementally trained using the multi-scenario simulation training corpus. Based on the proportion of updatable parameters, parameters of a corresponding proportion are selected from the new parameter set for weight update. The parameter weights are adjusted using a mini-batch gradient descent algorithm. During training, adaptive gradient clipping is used to dynamically adjust the gradient clipping threshold. Version tags are applied to the optimized model parameters, and a mapping relationship is established between them and the corresponding version of the glossary rule base, forming a versioned file content recognition model.

8. The file content recognition optimization method driven by the terminology explanation update according to claim 1 further includes: Implement a hot update strategy for the rule base to complete the version switching of the hierarchical terminology explanation rule base while ensuring continuous operation, and ensure the recoverability of the version in case of update failure through a version rollback mechanism; The distributed monitoring nodes are used to monitor the operational status of the deployed terminology definition rule base and periodically verify the consistency between the rule base version and the terminology definition data source. The incremental update process is continuously triggered based on the node monitoring results, maintaining dynamic synchronization between the hierarchical terminology explanation rule base and the terminology explanation data source, forming a continuously optimized closed-loop update.

9. A file content recognition and optimization system driven by terminology explanation updates, characterized in that: The system is used to implement the document content recognition optimization method driven by the terminology definition update as described in any one of claims 1-8, the system comprising: The rule base construction module is used to build a hierarchical term explanation rule base, wherein the hierarchical term explanation rule base includes a basic rule layer and a dynamic update layer; The monitoring and update module is used to deploy distributed monitoring nodes. It periodically checks the consistency of the data source version through a lightweight verification protocol. When the version difference exceeds the preset fault tolerance threshold, it triggers the incremental update process. The corpus generation module is used to generate multi-scenario simulation training corpus based on the conflict resolution mechanism in the hierarchical noun explanation rule base. The multi-scenario simulation training corpus contains mixed text instances of new and old noun explanations and is quality filtered by a semantic consistency verifier. The terminology definition recognition module is used to optimize the file content recognition model by using the multi-scenario simulated training corpus and a hierarchical parameter update mechanism, and then apply the optimized file content recognition model to perform terminology definition recognition on the input file text. The self-correction module is used to provide feedback on the training effect based on the multi-scenario simulated training corpus. When any of the indicators, such as accuracy, recall, or F1 score, is lower than a preset threshold, the self-correction process of the hierarchical noun explanation rule base is initiated to readjust the priority and association weight of the noun explanations.

Citation Information

Patent Citations

  • Intelligent processing system for enhanced training of legal text small samples

    CN120706434A

  • Streaming knowledge injection and adversarial self-optimization large language model training method and system

    CN121094006A