A thesaurus construction method and thesaurus for intelligent labor arbitration
Patent Information
- Application Number
- CN202610889162.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-25
AI Technical Summary
因而难以面向智慧劳动仲裁全业务场景应用需求,提供统一标准的关键词服务,给智慧劳动仲裁系统或平台的构建与应用均带来了困难
本发明针对智慧劳动仲裁系统或平台的设计需求,提出了“三层词库架构+双领域动态适配+统一审核发布”的词库构建方法,实现了智慧劳动仲裁领域词库的合理分类管理、动态智能迭代与标准化闭环应用。本发明构建静态刚性约束与动态柔性泛化相结合的分层协同词库体系,依托固化禁用词库筑牢合规底线、基础标准术语库统一法言法语基准、AI动态泛化词库实时适配新法新规,形成“底线拦截-规范约束-全面兜底”的梯度校验能力,有效解决传统系统术语不规范、隐性违规漏判、政策适配滞后等问题,提升仲裁文书的合法性与严谨性;通过法律、政策双领域独立适配机制,结合智能体专属领域约束与原文对标校验,搭配“AI柔性语义识别+Drools规则刚性终审”双闭环架构,可有效抑制大模型幻觉、释义偏差等问题,兼顾智能识别广度与仲裁业务的合规严肃性。同时,本发明针对动静两类词库设计差异化运维机制,静态词库依托审核、版本追溯、权限隔离的全生命周期管控模式,保障词条权威可审计,动态AI词库实现运行时实时生成、零常态化人工维护,大幅降低人工建库与迭代成本,提升系统可运维性。
Smart Images

Figure CN122817359A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent legal consultation and service technology, specifically to a method and lexicon for constructing a lexicon for intelligent labor arbitration. Background Technology
[0002] Currently, due to the heavy workload of grassroots labor arbitration and the shortage of senior arbitrators, there is an urgent need for intelligent labor arbitration, and intelligent labor arbitration application systems based on large models have emerged to meet this need.
[0003] The construction of intelligent labor arbitration application systems generally relies on the extraction, organization, and management of keywords. However, the knowledge level of labor arbitration service recipients is highly uncertain, and the diversity and uncertainty of their expressed demands and true intentions are significant. Furthermore, the systems need to adapt to a wide variety of laws, regulations, and policy documents, resulting in a long-standing lack of unified keyword planning and categorized management. Each application system typically relies on the personal understanding or preferences of designers or developers to define and select keywords, rarely involving the design (construction) of a standard keyword database (thesaurus), and lacking a unified and standardized thesaurus construction method. Consequently, it is difficult to provide a unified standard keyword service to meet the application needs of all intelligent labor arbitration business scenarios, posing challenges to the construction and application of intelligent labor arbitration systems or platforms. Summary of the Invention
[0004] The purpose of this invention is to overcome the deficiencies in the existing technology and provide a method and lexicon for constructing a lexicon for smart labor arbitration, so as to meet the development and application needs of smart labor arbitration platforms or systems in various places, and to achieve standardization and efficiency in the development of smart labor arbitration platforms or systems, as well as standardization and efficiency in smart arbitration services.
[0005] To achieve the above objectives, this invention proposes a thesaurus construction method for smart labor arbitration. The thesaurus provides core control over unified standard terminology, prohibited words, and semantic benchmarks to meet the needs of all business scenarios in smart labor arbitration. It provides a unified, authoritative, and controllable vocabulary benchmark for the entire process of smart labor arbitration services, including policy interpretation, legal structuring, content verification, question and answer generation, and large-scale model reasoning. The vocabulary benchmark includes unified standard terminology, prohibited words, verification rules, and semantic benchmarks. The construction method employs a "three-layer thesaurus architecture + dual-domain dynamic adaptation + unified review and release" approach to meet the needs of all scenarios, including legal structuring, policy document interpretation, content compliance verification, and large-scale model reasoning constraints. The three-layer lexicon architecture includes a fixed prohibited lexicon (blocking prohibited words), a basic standard terminology lexicon (high-frequency standardized terminology), and an AI dynamic generalized lexicon (a classification architecture that covers all commonly used semantic benchmarks without maintenance); the dual-domain dynamic adaptation includes dynamic adaptation of lexicons in the policy and legal domains. The fixed prohibited term library includes a reverse restricted keyword library. During the knowledge entry process of the intelligent labor arbitration platform or system (including conventional business databases, analytical databases, full-text search databases, and knowledge graph databases, etc.), if the parsed keywords match keywords in the fixed prohibited term library, knowledge entry is prohibited, and an alert or prompt for modification and confirmation is issued. The fixed prohibited term library serves as an interception component (bottom-line interception) of the compliance verification system. Its core function is to resolve fatal compliance issues in documents, such as unclear subject references, weakened legal force, non-neutral expression, and unauthorized expansion of authority. By accurately intercepting non-standard language, it ensures the seriousness, legality, and standardization of documents. The basic standard terminology database focuses on high-frequency core fixed or statutory terms in the fields of human resources and social security and the administrative legal system. As a standardized baseline terminology database, it is used to constrain high-frequency terms from being abbreviated, tampered with, or distorted, and to achieve unified standardization of common legal terminology. The AI dynamic generalization lexicon includes a policy lexicon and / or a legal lexicon, which is a lexicon of specialized terms or expressions formed after interpreting policy and / or legal documents (in real time or not), and is a fully covered, maintenance-free database. The unified review and release includes statically configured thesaurus such as the fixed prohibited thesaurus and the basic standard terminology library, providing full lifecycle management from entry editing, verification, review to release and effectiveness, ensuring that all online entries have undergone compliance review and preventing erroneous, conflicting, and non-standard words from entering the business operation process; the review and release module is a unified process module to ensure the authority, compliance, traceability, and controllability of the thesaurus across the entire platform.
[0006] Furthermore, adhering to the core design principles of "prioritizing the bottom line, following up with standardized procedures, providing backup, and final review of rules," the priority order of thelexicon calls is clearly defined to ensure that each verification step has a clear control objective, forming a standardized verification process; the thelexicon verification priority of the construction method includes: (a) Priority verification: Fixed prohibited word library (bottom-line hard blocking); (ii) Second verification: Basic standard terminology database (high-frequency specification constraints); (III) Last resort: AI dynamic generalization lexicon (full-coverage dynamic recognition); (iv) Final determination, which is made using the Drools rule engine.
[0007] Furthermore, the Drools rule engine includes a general Drools rule engine foundation built using a definition-as-code model, which is responsible for managing various knowledge governance rules, compliance verification rules, conflict resolution rules, and document generation constraint rules for the entire system. It enables unified rule management, dynamic configuration, rule updates without code, and unified adjudication, providing standardized rule support for knowledge base governance, knowledge fusion, element verification, business process scheduling, and agent reasoning constraints.
[0008] Furthermore, the fixed prohibited word library includes at least one of the following categories: vague pronouns, weakened constraint words, subjective evaluation words, and illegal expanded authority words; a verification method of fixed keyword precise matching + regular expression matching is adopted, and the Drools rule engine performs hard rule verification. If a violation is found, it is directly judged as a fatal violation and forcibly intercepted and added to the library, thus eliminating vague, weakened, and unauthorized expressions from the source (verification method).
[0009] Furthermore, the basic standard terminology database includes government and human resources standard terms, legal general standard terms, and prohibited abbreviations and variants; it adopts a combination of "whitelist + blacklist of illegal aliases" for verification, and the Drools rule engine combines the basic terminology database for accurate comparison to identify high-frequency term abbreviations, acronyms, and tampering issues, and allows, warns, or blocks (verification method) after a match is found.
[0010] Furthermore, the policy terminology database includes policy-specific terms, government jargon, hot topics, responsible entities, implementation measures, and policy concepts in the human resources and social security field (coverage). The verification method is based on policy-specific prompts, using the original policy text and higher-level statements as benchmarks. It is dynamically identified by a large model, with the Drools rule engine providing a rigid safety net for interception / early warning. Maintenance mode: zero manual maintenance, no static entries, dynamically generated at runtime, and automatically adapted to policy updates.
[0011] Furthermore, after the policy terminology database is completed through policy document interpretation and structured element extraction, the system relies on a comprehensive analytical agent (a public support agent at the agent layer) to identify non-standard expressions through exclusive constraints and textual benchmarking, ultimately forming a two-layer compliance verification closed loop of "AI flexible semantic recognition + rigid rule fallback".
[0012] Furthermore, the legal lexicon includes professional terms, niche terms (such as "bona fide acquisition" and "unjust enrichment" in law, regulations, and rules at all levels), statutory concepts, and newly added legal expressions, which are dynamically identified by a large model; the verification method is based on the law-specific Prompt constraint, using the original text of the legal provision and the legislative interpretation as the sole standard, dynamically identifying issues such as abbreviation, colloquialism, weakening, and misinterpretation, with Drools making the final rigid judgment; maintenance mode: zero manual maintenance, no static entries, automatic adaptation to new laws and regulations, and real-time dynamic identification; After the legal terminology database is decomposed into laws and regulations and structured elements are extracted, the system does not rely on a manual static massive terminology database. Instead, it uses a comprehensive parsing agent (a public support agent at the Agent layer) and relies on the domain understanding capabilities of a large model to complete the dynamic identification of all legal terms, comparison of standard interpretations, and screening of non-standard expressions. After the AI outputs standardized verification conclusions, they are sent back to the Drools rule engine for rigid review, risk classification, interception, or warning, thus realizing a two-layer compliance verification closed loop of "AI flexible semantic recognition + rule rigid fallback".
[0013] Furthermore, the unified review and release includes content or modules such as entry editing and pre-verification, hierarchical review, version management and traceability, release and activation strategies, and permission and security control.
[0014] On the other hand, a thesaurus for intelligent labor arbitration, the thesaurus including a thesaurus constructed based on any of the aforementioned thesaurus construction methods.
[0015] The advantages and beneficial effects of this invention are as follows: This invention addresses the design requirements of intelligent labor arbitration systems or platforms by proposing a "three-layer thesaurus architecture + dual-domain dynamic adaptation + unified review and release" thesaurus construction method. This achieves reasonable classification and management, dynamic intelligent iteration, and standardized closed-loop application of thesaurus in the field of intelligent labor arbitration. The invention constructs a layered collaborative thesaurus system combining static rigid constraints and dynamic flexible generalization. It relies on a fixed prohibited thesaurus to solidify compliance standards, a basic standard terminology library to unify legal terminology benchmarks, and an AI-driven dynamic generalized thesaurus to adapt to new laws and regulations in real time. This forms a gradient verification capability of "bottom-line interception - standardized constraints - comprehensive coverage," effectively solving problems such as non-standard terminology, hidden violations, and policy adaptation delays in traditional systems, thus improving the legality and rigor of arbitration documents. Through a dual-domain independent adaptation mechanism of law and policy, combined with agent-specific domain constraints and original text benchmarking verification, and a dual closed-loop architecture of "AI flexible semantic recognition + Drools rule rigid final review," it can effectively suppress problems such as large model illusion and interpretation bias, balancing the breadth of intelligent recognition with the compliance and seriousness of arbitration business. Meanwhile, this invention designs differentiated operation and maintenance mechanisms for static and dynamic thesauruses. Static thesauruses rely on a full lifecycle management model of review, version tracking, and permission isolation to ensure that the entries are authoritative and auditable. Dynamic AI thesauruses are generated in real time during runtime with zero routine manual maintenance, which greatly reduces the cost of manual database construction and iteration and improves the system's maintainability.
[0016] This invention achieves a unified, standardized, secure, and controllable thesaurus across the entire platform. It comprehensively covers core scenarios such as structured analysis of laws and regulations, policy interpretation, document compliance verification, and large-scale model reasoning constraints. It provides standardized thesaurus support for all business modules, including intelligent consultation, intelligent court hearings, and intelligent adjudication. It solves the problems of inconsistent keyword definitions, poor adaptability, and lack of compatibility among modules in traditional systems. It effectively reduces business errors and document compliance risks, and significantly improves the intelligence, standardization, and practical application level of the smart labor arbitration platform. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the dictionary construction method of the present invention; Figure 2 This is a diagram illustrating the principle of policy terminology generation. Figure 3 This is a diagram illustrating the principle of legal terminology generation. Detailed Implementation
[0018] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and examples. The following examples are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.
[0019] Example 1: This invention designs a lexicon construction method for intelligent labor arbitration, such as... Figure 1 As shown, the lexicon provides core control over unified standard terminology, prohibited words, and semantic benchmarks for the full range of smart labor arbitration business scenarios. It provides a unified, authoritative, and controllable lexicon benchmark for the entire process of smart labor arbitration services, including policy interpretation, legal structuring, content verification, question and answer generation, and large model reasoning. The lexicon benchmark includes unified standard terminology, prohibited words, verification rules, and semantic benchmarks. The construction method includes adopting a "three-layer lexicon architecture + dual-domain dynamic adaptation + unified review and release" construction method to meet the full-scenario needs of legal structuring, policy document interpretation, content compliance verification, and large model reasoning constraints. The three-layer lexicon architecture includes a fixed prohibited lexicon (blocking prohibited words), a basic standard terminology lexicon (high-frequency standardized terminology), and an AI dynamic generalized lexicon (a classification architecture that covers all commonly used semantic benchmarks without maintenance); the dual-domain dynamic adaptation includes dynamic adaptation of lexicons in the policy and legal domains. The fixed prohibited term library includes a reverse restricted keyword library. During the knowledge entry process of the intelligent labor arbitration platform or system (including conventional business databases, analytical databases, full-text search databases, and knowledge graph databases, etc.), if the parsed keywords match keywords in the fixed prohibited term library, knowledge entry is prohibited, and an alert or prompt for modification and confirmation is issued. The fixed prohibited term library serves as an interception component (bottom-line interception) of the compliance verification system. Its core function is to resolve fatal compliance issues in documents, such as unclear subject references, weakened legal force, non-neutral expression, and unauthorized expansion of authority. By accurately intercepting non-standard language, it ensures the seriousness, legality, and standardization of documents. The basic standard terminology database focuses on high-frequency core fixed or statutory terms in the fields of human resources and social security and the administrative legal system. As a standardized baseline terminology database, it is used to constrain high-frequency terms from being abbreviated, tampered with, or distorted, and to achieve unified standardization of common legal terminology. The AI dynamic generalization lexicon includes a policy lexicon and / or a legal lexicon, which is a lexicon of specialized terms or expressions formed after interpreting policy and / or legal documents (in real time or not), and is a fully covered, maintenance-free database. The unified review and release includes statically configured thesaurus such as the fixed prohibited thesaurus and the basic standard terminology library, providing full lifecycle management from entry editing, verification, review to release and effectiveness, ensuring that all online entries have undergone compliance review and preventing erroneous, conflicting, and non-standard words from entering the business operation process; the review and release module is a unified process module to ensure the authority, compliance, traceability, and controllability of the thesaurus across the entire platform.
[0020] Furthermore, the thesaurus can be divided into static thesaurus and dynamic thesaurus. The static thesaurus includes a fixed prohibited word library and a basic standard term library (manually maintained and subject to review and release). The dynamic thesaurus includes an AI dynamic generalized thesaurus (policy thesaurus + legal thesaurus, generated in real time during runtime and requiring no manual maintenance).
[0021] In other words, from another classification perspective, this invention divides the thesaurus into a static thesaurus and a dynamic thesaurus. The two types of thesaurus work together to build a full-process verification architecture of "bottom-line interception + standard constraints + comprehensive fallback" to support the efficient operation of the compliance verification system.
[0022] It should be noted that the lexicon construction method described in this invention mainly refers to how to plan and construct the lexicon structure. A reasonable structure and classification plan can make it more convenient to fill, improve, and use the lexicon. This does not specifically refer to how to obtain keywords in the lexicon or how to use the lexicon. Although the above content is mentioned in the specific method description of this invention, it is only to further illustrate the specific implementation and usage of the structure.
[0023] Preferably, this embodiment follows the core design principles of "bottom-line priority, standardized follow-up, fallback, and final rule review," clearly defining the priority order of the dictionary calls to ensure that each verification step has a clear control objective, forming a standardized verification process; the dictionary verification priority (or the order of verification) of the construction method includes: (a) Priority verification: Fixed prohibited word library (bottom-line hard blocking); (ii) Second verification: Basic standard terminology database (high-frequency specification constraints); (III) Last resort: AI dynamic generalization lexicon (full-coverage dynamic recognition); (iv) Final judgment.
[0024] Preferably, the final determination described in this embodiment utilizes Scheme 1: Drools business rule engine determination; as an equivalent alternative implementation, any one of the following can also be used to complete the final determination: QLExpress lightweight rule engine, database configuration rule table, hard-coded branch determination, or Flowable process node verification. All of these can achieve risk classification of the four-layer thesaurus matching results, issuance of handling actions, and solidification output of verification results.
[0025] Option 1: Drools business rule engine determination Drools (an open-source business rules engine based on Java, BRMS) rules are logical definitions used to describe "what action to perform under what conditions" when using Drools. It decouples business decisions from the code and expresses them using rules that are close to natural language, making it easier for business users to understand and for operations personnel to adjust.
[0026] Logic: Structured writing of if conditions - then actions, decoupling rules from code, maintainable by non-developers, supports batch and multi-level risk assessment, interception / early warning branch output, and adapts to complex government compliance clauses.
[0027] Option 2: MySQL / DM database configuration-based rule determination (lightweight and simple alternative) The interception threshold, violation handling methods, and consequences of keyword matching are all stored in a database configuration table. The program reads the configuration table to make branch decisions. Example: Create a table `rule_check_config` with fields including keyword type, matched keywords, risk level, and handling action (interception / warning / notification). After matching a keyword, the program automatically executes the corresponding action by looking up the table. Advantages: No middleware is required; suitable for small, lightweight systems. Disadvantages: Complex, multi-condition nested rules are cumbersome to write, and visual editing is weak.
[0028] Option 3: Hard-coded If-Else condition in Python (minimalist prototype solution) The system directly uses nested if statements within the business code to check the matching results from different thesauruses, directly blocking matches of prohibited words, issuing alerts for abnormal terminology, and performing secondary comparisons based on AI recognition results. Advantages: Simplest deployment, no additional components; Disadvantages: Rules are hardcoded into the code, requiring a release and code changes for every change in validation logic, making iteration cumbersome and unsuitable for long-term government operations.
[0029] Option 4: QLExpress, a lightweight script-based rules engine (an open-source alternative to BRMS). A lightweight JavaScript rule framework that allows you to write validation logic using custom expressions. It's similar in purpose to Drools but with a smaller package size. Example script: plaintext If (prohibited word detected == true), return "Fatal violation, forced blocking"; else if (non-standard terminology > 0), return "Warning: Correct terminology"; Advantages: Fast startup, low resource consumption, and adaptable to small and medium-sized platforms; its ability to link multiple legal provisions in complex arbitration is slightly weaker than Drools.
[0030] Option 5: Binding validation nodes to the Flowable workflow engine (integrated workflow solution) The thesaurus validation is made into an automated service node in the workflow. As the workflow progresses, four layers of thesaurus matching are triggered sequentially, with node configurations for approval, interception, and branch return. Applicable scenarios: Arbitration documents undergoing a full online approval process, where validation and business processes are integrated; pure thesaurus validation scenarios have high redundancy.
[0031] Drools is the optimal choice for this solution, with core advantages including: a large number of arbitration compliance rules and frequent changes in subsequent laws and policies; complete decoupling of rules and code; and the ability for business personnel to independently edit and adjust interception and warning standards without repeated program iterations and releases. The other options are only suitable for lightweight, low-change-frequency, and simple deployment scenarios. In addition to a powerful rule engine, it also includes rich features such as rule management, decision tables, decision trees, process orchestration (using jBPM), and complex event processing (CEP). Drools supports more complex rule structures and more comprehensive business scenarios. This embodiment introduces the Drools open-source business rule engine to decouple and standardize compliance verification logic. It independently extracts the complex compliance judgment logic in the labor arbitration field from the program code, defining the verification logic in a standardized "condition-action" business rule format. This generates rule files that can be independently parsed and executed in batches, overcoming the drawbacks of traditional hard-coded logic, which is rigid, cumbersome to modify, and has high iteration costs. Business and operations personnel can flexibly adjust arbitration compliance verification standards and update control rules based on a rule definition method that is close to natural language. Verification logic optimization and iteration can be completed without modifying the underlying code, significantly improving the flexibility, timeliness, and maintainability of system rule adjustments.
[0032] Preferably, the Drools rule engine includes a general Drools rule engine base built using a definition-as-code model, which is responsible for managing various knowledge governance rules, compliance verification rules, conflict resolution rules, and document generation constraint rules for the entire system. It enables unified rule management, dynamic configuration, rule updates without code, and unified adjudication, providing standardized rule support for knowledge base governance, knowledge fusion, element verification, business process scheduling, and agent reasoning constraints.
[0033] Preferably, the fixed prohibited word library includes at least one of four categories: vague pronouns, constraint weakened words, subjective evaluation words, and illegal expanded rights words. The verification method of the fixed prohibited word library generally includes fixed keyword precise matching, regular expression matching, etc. As an equivalent extended implementation method, one or more combinations of word segmentation semantic matching, vector semantic similarity matching, and function word removal preprocessing matching can also be superimposed. All verification results are uniformly sent to the Drools rule engine to perform hard rule verification. If a violation is found, it is directly judged as a fatal violation, and the violation is forcibly intercepted and archived to eliminate vague, weakened, and unauthorized expressions from the source.
[0034] The regular expression matching verification method refers to the program background pre-compiling regular expression matching templates for various types (such as the four types mentioned above) of prohibited words in batches. The template configuration is stored in the configuration file or the DM database. After the program reads the text to be verified, it calls the regular expression engine to perform full-text segment-by-segment scanning and matching. It can match the combination of function words, auxiliary words, conjunctions, and punctuation marks interspersed before and after the target prohibited words, and also supports limiting the syntactic position of the words. After the program captures the matching result, it outputs the matching position, the matched words, and the matched original text fragment, and pushes them to the Drools rule engine for risk assessment and handling judgment, making up for the defect that pure precise keywords can only match complete continuous text and cannot recognize modified and deformed sentences.
[0035] Two check branches of accurate keyword matching and regular matching are run in parallel by the program, and the two types of hit results are uniformly summarized and encapsulated into a check message and sent to Drools; as long as a disabled entry is hit by either method, Drools immediately triggers a hard rule for fatal violation, issues a mandatory interception instruction, blocks the circulation of the manuscript and stores the violation record in the database.
[0036] The word segmentation semantic matching includes relying on a Chinese word segmentation tool to segment and disassemble the whole sentence of the document, which is not limited to consecutive keywords, and identifies the illegally distributed illegal semantic combinations. For example, the two words "as appropriate" and "simplify approval" appear separately in a sentence, which cannot be identified by split keyword matching, while word segmentation semantic matching can determine the risk of overreach in combination.
[0037] The vector semantic similarity matching includes converting disabled entries and document sentences into vectors, calculating semantic similarity, and identifying synonymous paraphrasing and disguised replacement wording. For example, "flexibly adjust" is used to replace "adapt according to circumstances", the literal keywords are different, but the semantics are consistent, and vector matching can capture implicit illegal expressions.
[0038] The preprocessing matching for function word removal refers to accurate matching after stop word filtering, which includes filtering meaningless function words such as "de (the)", "le (past tense marker)", "he (and)", "ji (and)" first and then performing keyword comparison, so as to avoid continuous keyword breaking and matching failure caused by interspersed function words.
[0039] The coverage of the fixed disabled word library in this embodiment includes: The vague referring words include: vague subject expressions such as relevant departments, relevant units, competent authorities, relevant personnel and related units; The constraint weakening words include: words that weaken statutory force such as in principle, under normal circumstances, as appropriate, appropriate, according to the situation, as far as possible, to the best of one's ability; The subjective evaluation words include: non-neutral subjectively biased terms such as excellent, efficient, key, important, suggest, best; The illegal power-expanding words include: overreach and adaptation expressions such as in a disguised form, adaptation, relaxation, simplify approval, and relaxation as appropriate.
[0040] The maintenance mode of the fixed disabled word library adopts one-time initialization configuration, and only a small amount of iteration is performed when the superior document specifications and legal supervision requirements are significantly adjusted, and no manual maintenance is required daily.
[0041] As an interception component of the compliance check system, the core function of the fixed disabled word library is to solve fatal compliance problems such as unclear subject reference, weakened statutory force, non-neutral expression, and illegal power expansion in the manuscript. Through reasonable classification, omissions are avoided, so as to achieve accurate interception of non-standard expressions and ensure the seriousness, legality and standardization of the manuscript.
[0042] Preferably, the basic standard terminology library includes government and human resources standard terms, general legal standard terms, and contrast words for prohibited abbreviated alienation; a combination of "whitelist + illegal alias blacklist" is adopted for verification, and one or more combinations of character abbreviation deformation matching and edit distance similarity verification can also be used for collaborative verification. The Drools rule engine combined with the basic terminology library completes multi-dimensional accurate comparison, identifies problems such as tampering of high-frequency terms by abbreviations, short forms, and common names, and releases, gives early warning or intercepts when a match is found.
[0043] The character abbreviation deformation matching includes configuring an abbreviation mapping dictionary to identify Chinese character abbreviations, pinyin abbreviations, English abbreviations, and arbitrary abbreviated writing. For example, "Lao Zhong" is matched with the standard word "labor arbitration", and the program reads the mapping table for batch comparison, correction and recognition.
[0044] The edit distance similarity verification includes adopting an edit distance algorithm to calculate the character difference between text vocabulary and standard terms, setting a threshold, and marking non-standard vocabulary with similar character shapes and tampering by missing or adding extra characters as abnormal.
[0045] The coverage of the basic standard terminology library in this embodiment includes: Government and human resources standard terms: administrative counterpart, statutory duty, performing duties responsibly, labor supervision, arbitration jurisdiction, administrative adjudication, etc.; General legal standard terms: administrative penalty, administrative license, rights and obligations, prohibitive provisions, statutory time limit, legal liability, etc.; Prohibited abbreviated alienation contrast: defines the contrast relationship between standard terms and illegal abbreviations, common names, and colloquial alternative words.
[0046] Maintenance mode: the basic standard terminology library described in this embodiment is a lightweight static configuration library with small volume and concentrated coverage. The maintenance mode performs low-frequency incremental update by year or at the node when new regulations are issued, and does not require continuous large-scale manual maintenance.
[0047] Preferably, the policy terminology library includes policy-specific terms, government caliber, hot expressions, responsible subjects, implementation measures and policy concepts in the field of human resources and social security (coverage range). The verification method is constrained based on policy-specific Prompt, takes the original policy text and superior caliber as the benchmark, is dynamically identified by a general large model, and completes interception / warning through the rigid underpinning of the Drools rule engine; maintenance mode: zero manual maintenance, no static entries, dynamically generated at runtime, and automatically adapts to policy updates.
[0048] Preferably, after the policy terminology library completes interpretation and structured element extraction of policy documents, the system relies on the comprehensive analysis intelligent agent (Agent layer public support intelligent agent) to implement non-standard expression inspection through exclusive constraint and original text benchmarking, and finally forms a double-layer compliance verification closed loop of "AI flexible semantic recognition + rule rigid underpinning".
[0049] Specifically, such as Figure 2 As shown, the policy document interpretation module submits the structured policy text to the comprehensive analysis agent (the public support agent at the Agent layer). The agent loads the policy-specific prompt constraints and scans the text segment by segment using a general large model. It dynamically extracts policy terminology, official terminology, and implementation requirements, and compares the original policy text with official standard expressions to intelligently identify implicit non-standard issues such as colloquialisms, expansions, ambiguities, weakening, and overreach. The large model returns the verification results to the comprehensive analysis agent (the public support agent at the Agent layer) in JSON structured format. The agent then sends the results back to the policy document interpretation module and calls the Drools rule engine to perform rigid judgment. Drools performs cross-verification using a fixed rigid thesaurus, executes interception or warning according to risk level, and returns a unified verification report. After successful verification, the policy document interpretation module writes the structured data into the Dameng database and marks it as "approved".
[0050] Comprehensive analytical agent: This approach targets three types of documents: policy documents, laws and regulations, and case documents. It employs standardized structured parsing through prompt-guided techniques and RAG enhancement technology (see "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" by Patrick Lewis, Ethan Perez, & Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, & Mike Lewis†, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela, Facebook AI Research; or "A Survey on Large Language Model based Autonomous Agents" by Wang Lei, Ma Chen, Feng Xueyang, et al.; or "Structured Information Extraction from Government Policy Documents Based on Prompt Fine-tuning and RAG" by Wang L and Liu Q, published in the *Journal of Chinese Information Processing* in 2024; or "Automatic Construction of Knowledge Graphs from Text and Structured Data: A Preliminary Literature"). (Review, etc.) to achieve the following: convert policy texts into popular semantics, extract all-dimensional elements, automatically generate official interpretations, and construct question-and-answer datasets; achieve structured extraction of legal provisions, automatic generation of question-and-answer knowledge bases, and automated construction of legal knowledge graph entities and relationships; achieve core element analysis of arbitration case transcripts and awards, and automatic construction of corresponding knowledge graphs, providing raw data preprocessing support for upper-level business reasoning.
[0051] Preferably, the legal terminology database includes professional terms, niche terms (such as "bona fide acquisition" and "unjust enrichment" in law, regulations, and rules at all levels), statutory concepts, and newly added legal expressions, which are dynamically identified by a general large model; the verification method is based on the law-specific Prompt constraint, taking the original text of the legal provision and the legislative interpretation as the sole standard, dynamically identifying problems such as abbreviation, colloquialism, weakening, and misinterpretation, with Drools completing the final rigid judgment; maintenance mode: zero manual maintenance, no static entries, automatic adaptation to new laws and regulations, and real-time dynamic identification; After the legal terminology database is decomposed into laws and regulations and structured elements are extracted, the system does not rely on a manual static massive terminology database. Instead, it uses a comprehensive parsing agent (a public support agent at the Agent layer) and relies on the domain understanding capabilities of a large model to complete the dynamic identification of all legal terms, comparison of standard interpretations, and screening of non-standard expressions. After the AI outputs standardized verification conclusions, they are sent back to the Drools rule engine for rigid review, risk classification, interception, or warning, thus realizing a two-layer compliance verification closed loop of "AI flexible semantic recognition + rule rigid fallback".
[0052] Specifically, such as Figure 3 As shown, the legal regulations decomposition module submits structured elements to the comprehensive analysis agent. Relying on the comprehensive analysis agent, it dynamically identifies all legal terms in the text using legally specific constraint prompts. It compares the official original text of the legal provisions with the legislative interpretation to complete the standard interpretation comparison, and intelligently screens for implicit compliance issues such as abbreviations, colloquial modifications, weakened definitions, and misinterpretations of connotations. The large model outputs the term verification results in a structured format and simultaneously sends them back to the Drools rule engine for rigid review and risk classification control. Combined with the fixed prohibited word library and the basic standard term library, it forms a multi-layered verification linkage, which not only solves the problem of the high cost of manual maintenance of a large number of legal terms, but also avoids the risk of illusion by the large model through the rule engine, ensuring the uniformity of legal language, the integrity of legal definitions, and the integrity of legal constraints.
[0053] Example 2: The difference from Example 1 is that the fixed prohibited word library verification method described in this example, in addition to using fixed keyword precise matching and regular expression matching, also incorporates three other methods: word segmentation semantic matching, vector semantic similarity matching, and function word removal preprocessing matching.
[0054] Example 3: The difference from Example 1 is that the unified review and release described in this example includes not only the content modules of entry editing and pre-verification, hierarchical review, release and effectiveness strategy, but also content or modules such as version management and traceability, permissions and security control.
[0055] The thesaurus review and release is a unified process module that ensures the authority, compliance, traceability, and controllability of the thesaurus across the entire platform. It provides full lifecycle management for statically configured thesauruses such as fixed rigid thesauruses and basic standard terminology databases, from entry editing, verification, review to release and effectiveness. It ensures that all online entries have undergone compliance review and prevents erroneous, conflicting, and non-standard words from entering the business operation process.
[0056] (1) Entry editing and pre-validation It supports adding, modifying, deleting, enabling / disabling, batch importing and exporting of entries; it automatically performs format verification, duplicate verification and conflict verification before submission to ensure that entries are standardized, complete and without redundancy.
[0057] (2) Tiered review mechanism The system implements a three-tiered permission system for editing and submission, review and approval, and publication and effectiveness. Reviewers can view entry details, modification history, applicable scenarios and risk levels, and can perform operations such as approval, rejection, and return for modification.
[0058] (3) Version management and traceability It automatically records the content, operator, operation time, and review comments of each thesaurus change, forming a version history; it supports version rollback, version comparison, and historical query, meeting the audit and traceability requirements of government systems.
[0059] (4) Release and Implementation Strategy Once approved, the application can be published with one click. After publication, it will be synchronized in real time to all business scenarios, including the comprehensive analysis intelligent agent (the common support intelligent agent at the Agent layer), the Drools rule engine, the policy document interpretation module, and the legal decomposition module. It supports immediate and scheduled effects to ensure a smooth business transition.
[0060] (5) Access Control and Security Management The dictionary management permissions are assigned according to roles, distinguishing between administrators, editors, reviewers, and viewers, to ensure clear responsibilities and controllable operations; all key operations are logged to ensure the safe, compliant, and monitorable operation of the dictionary.
[0061] The scope of application for the thesaurus review and release: It only applies to the release and online management of fixed rigid thesaurus and basic standard terminology database; AI dynamic generalized thesaurus is dynamically identified in real time by a large model and is not included in the manual review and release process.
[0062] Error handling: When encountering situations such as validation failure, hitting a fixed rigid vocabulary, or non-standard terminology, the system will automatically intercept and return a clear error message, prohibiting access to the knowledge base and business processes.
[0063] This embodiment also includes monitoring and statistics, supporting keyword hit statistics, violation interception statistics, hot violation word statistics, full operation log retention, and multi-dimensional queries by module, time, and type.
[0064] Example 4: A thesaurus for intelligent labor arbitration, the thesaurus comprising a thesaurus constructed based on any of the thesaurus construction methods described above.
[0065] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, such as different dictionary verification methods, different classification methods for fixed prohibited words, different classification methods for basic standard terminology libraries, different models and training methods, different dynamic dictionary recognition and extraction methods, etc. These improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for constructing a thesaurus for intelligent labor arbitration, wherein the thesaurus provides a vocabulary benchmark for the full range of business scenarios in intelligent labor arbitration, characterized in that, The vocabulary benchmark includes unified standard terms, prohibited words, verification rules and semantic benchmarks, and the construction method includes: a construction method of "three-layer thesaurus architecture + dual-domain dynamic adaptation + unified review and release"; The three-layer thesaurus architecture includes a classification architecture of a fixed prohibited thesaurus, a basic standard terminology library, and an AI dynamic generalized thesaurus; the dual-domain dynamic adaptation includes dynamic adaptation of thesaurus in the policy and legal domains; The fixed prohibited word library includes a reverse restricted keyword library; The basic standard terminology database focuses on high-frequency core fixed or legal terms, serving as a standardized thesaurus. The AI dynamic generalization lexicon includes a policy lexicon and / or a legal lexicon, which is a lexicon of specialized terms or expressions formed after interpreting policy and / or legal documents; The unified review and release includes a fixed list of prohibited terms and a basic standard terminology database, providing full lifecycle management from term editing, verification, review to release and effectiveness.
2. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The dictionary verification priority of the construction method includes: (a) Priority verification: Fixed prohibited word list; (ii) Second verification: Basic standard terminology database; (III) Last resort: AI dynamic generalization lexicon; (iv) Final judgment.
3. The method for constructing a thesaurus for intelligent labor arbitration according to claim 2, characterized in that, The final determination is made using the Drools rule engine, which includes a general Drools rule engine base built using a definition-as-code model. This base is responsible for managing various knowledge governance rules, compliance verification rules, conflict resolution rules, and document generation constraint rules across the entire system.
4. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The fixed prohibited word library includes fuzzy pronouns, words with weakened constraints, subjective evaluation words, and words with illegally expanded authority; it adopts a verification method of "precise matching + regular expression matching" of fixed keywords.
5. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The basic standard terminology database includes government and human resources standard terms, legal general standard terms, and prohibited abbreviations and variant terms; it uses a combination of "whitelist + blacklist of illegal aliases" for verification.
6. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The policy terminology database includes policy-specific terms, government policy statements, hot topics, responsible entities, implementation measures, and policy concepts in the field of human resources and social security.
7. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The policy terminology database was created through the interpretation of policy documents and the extraction of structured elements.
8. The method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The legal terminology database includes professional terms, legal concepts, and newly added legal expressions from laws, regulations, and rules at all levels. The legal terminology database was created through the decomposition of laws and regulations and the extraction of structured elements.
9. A method for constructing a thesaurus for intelligent labor arbitration according to claim 1, characterized in that, The unified review and release includes entry editing and pre-verification, hierarchical review, version management and traceability, release and activation strategies, and permission and security control.
10. A thesaurus for intelligent labor arbitration, characterized in that, The lexicon includes a lexicon constructed based on any one of the lexicon construction methods of claims 1 to 9.