A coding and application method and system based on rule files and corpus construction

By encoding the rules file and chapter clauses, parsing morphemes, and constructing traceable encoding combinations, the problems of semantic ambiguity and high maintenance costs in text tag management are solved, achieving accurate semantic tag generation and traceability, and improving the interpretability and consistency of text processing.

CN121960503BActive Publication Date: 2026-07-24SSE INFORMATION NETWORK LTD
4 Cites 0 Cited by

Patent Information

Application Number
CN202610424977.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-07-24
Estimated Expiration
2046-04-02

AI Technical Summary

Technical Problem

Existing text tag management methods suffer from problems such as keywords not containing contextual semantic information, high maintenance costs, and insufficient interpretability of generated results, especially the lack of clear semantic tag evidence chains in the output of deep learning models.

Method used

By encoding the rule files, chapters and clauses, and parsing morphemes, corpus units of conditional and conclusion morphemes are constructed. Then, traceable encoding combinations are formed through semantic relation operators. Combined with semantic disambiguation mechanisms and vectorized embedding technology, accurate semantic labels are generated and traceability is achieved.

Benefits of technology

It achieves a complete chain of evidence from semantic tags directly to the original rule text, ensuring the certainty and consistency of the tag results, improving the interpretability of scenarios such as compliance review and legal reasoning, and achieving precise matching in intelligent question answering and fuzzy search.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application relates to the technical field of text label application management, and provides a coding and application method and system based on a rule file and a corpus, the method comprising the following steps: performing file coding on a rule file; performing section clause coding on section clauses in the rule file; constructing a corpus unit; and combining the file coding, the section clause coding and the corpus unit to form traceable coding combinations, which are applied to the standardized marking of key information of a business file, standard reference tracing is formed through coding, and a business corpus can be formed through generalization mapping combined with business practices to improve application effects, so that the original rule file, the clause position and the morpheme semantics can be traced in reverse, and a text meeting a condition can be searched in a forward direction through coding matching. The application realizes semantic disambiguation and accurate tracing of rule texts through a multilevel coding system, is compatible with the generalization understanding ability of AI technology, and improves the accuracy of text marking.
Need to check novelty before this filing date? Find Prior Art