An intelligent translation method and system based on patent literature structure analysis and field term alignment

CN122819271APending Publication Date: 2026-09-25QIANHEYI (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610900767.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

(1)结构解析能力不足:通用翻译系统无法准确识别专利文献的权利要求层级结构和引用关系,翻译后权利要求编号混乱、缩进丢失,导致译文失去法律效力;

Benefits of technology

[0014]与现有技术相比,本发明具有以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819271A_ABST
    Figure CN122819271A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent translation method and system based on patent document structure analysis and field term alignment, and relates to the technical field of natural language processing. The method comprises the following steps: obtaining a patent document electronic file to be translated, analyzing the structured mark information thereof, extracting independent claim texts and dependent claim texts in a claim section, and recording the citation relationship and hierarchical nested structure among the claims; according to the technical field of the patent document, calling a professional term comparison table of the corresponding technical field from a pre-constructed multi-field patent term library; inputting the extracted texts into a neural machine translation model for preliminary translation, and performing forced term alignment on the terms in the professional term comparison table during the translation process, so that the same terms remain consistent in the translation in the claims and the specification; according to the recorded citation relationship and hierarchical nested structure, back-filling the translated texts into the corresponding sections according to the format specification of the original patent document, and generating a translated patent document which maintains the original claim number, paragraph indentation and figure mark format. The application can effectively solve the problems of structure loss, term inconsistency and figure mark misplacement in patent document translation, and significantly improve the accuracy and standardization of patent translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing (NLP) technology, specifically to an intelligent translation method and system based on patent document structure parsing and domain terminology alignment, applicable to cross-language translation scenarios of patent documents in patent agencies, corporate intellectual property departments, and research institutes. Background Technology

[0002] With the acceleration of globalization, the demand for cross-border patent applications and examinations is increasing. Patent documents are highly structured, including sections such as claims, description, abstract, and drawings. The claims section contains independent claims and dependent claims, and there are complex referencing relationships and hierarchical nesting structures among the claims. In addition, patent documents make extensive use of technical terms and figure labels, which have specific meanings in different technical fields (such as machinery, electricity, chemistry, and biomedicine).

[0003] Existing machine translation systems (such as general translation engines based on neural machine translation) have the following shortcomings when processing patent documents: (1) Insufficient structural analysis capability: General translation systems cannot accurately identify the claim hierarchy and citation relationships of patent documents. After translation, the claim numbers are chaotic and the indentation is lost, resulting in the translation losing its legal effect; (2) Poor consistency of terminology: The same technical term may be translated into different expressions in the claims and the specification, especially in long patent documents, where the inconsistency of terminology is more prominent. (3) Improper handling of figure marks: Figure marks in patent documents (such as “10”, “100a”, etc.) are easily mistranslated or misplaced during the translation process, which affects the accurate expression of the technical solution; (4) Weak domain adaptability: The general translation model is not optimized for patent terms in specific technical fields, resulting in low accuracy in the translation of professional terms.

[0004] Therefore, there is an urgent need for an intelligent translation solution that can accurately parse the structure of patent documents, maintain terminology consistency, and support multi-domain adaptation.

[0005] The purpose of this invention is to provide an intelligent translation method and system based on patent document structure analysis and domain terminology alignment, so as to solve the technical problems of lost structure, inconsistent terminology, misaligned figure labels and weak domain adaptability in the prior art after patent document translation. Summary of the Invention

[0006] To achieve the above objectives, this invention provides an intelligent translation method based on patent document structure analysis and domain terminology alignment, comprising the following steps: Step 1: Obtain the electronic file of the patent document to be translated. The electronic file of the patent document contains structured markup information (such as XML, HTML tags or PDF structured data). The structured markup information at least identifies the claim section, the specification section and the hierarchical number within each section.

[0007] Step Two: Perform structural analysis on the electronic patent document, extract the independent and dependent claims text from the claim section, and record the citation relationships between claims (e.g., "as described in claim 1...") and the hierarchical nesting structure (e.g., the hierarchical depth of dependent claims citing independent claims). Simultaneously, identify the figure reference numbers in the patent document and mark their positions in the text.

[0008] Step 3: Based on the user-inputted target technical field or the system's automatic identification results based on IPC classification numbers, retrieve the corresponding technical field terminology lookup table from a pre-built multi-field patent terminology database. This multi-field patent terminology database covers mainstream patent technical fields such as machinery, electrical engineering, chemistry, biomedicine, computers, new materials, and new energy. Each field includes a two-way mapping relationship between Chinese and English patent legal terms (such as "independent claim," "dependent claim," and "implementation") and technical terms.

[0009] Step 4: Input the extracted claims and specification texts into a pre-trained neural machine translation model (such as the Transformer architecture) for initial translation. During the decoding phase, the system injects a terminology constraint mask, forcing the model to map source language terms to their corresponding target language terms in a terminology lookup table. After translation, a terminology consistency check is performed on the entire text: scan all occurrences of the same term in the claims and specification; if translation conflicts exist, a global replacement is performed based on the terminology lookup table.

[0010] Step 5: Based on the citation relationships and hierarchical nesting structure recorded in Step 2, backfill the translated text according to the format specifications of the original patent document. Specifically, this includes: maintaining the continuity of claim numbers (e.g., "1.", "2.", "3."), maintaining the indentation level of dependent claims (e.g., indenting a dependent claim referencing claim 1 by one level), and keeping the figure reference numerals in their original positions and consistent with their correspondence to the technical features. The final result is a translated patent document with the correct format.

[0011] Furthermore, this invention can also construct a translation memory, storing the translated claims of already translated patent family documents in the database. When translating a patent document to be translated, the system calculates its semantic similarity to existing paragraphs in the translation memory. If it exceeds a preset threshold (e.g., 0.85), the corresponding translation is directly reused, avoiding repeated translation and further ensuring terminology consistency.

[0012] The present invention also provides an intelligent translation system for implementing the above method, including a document parsing module, a terminology matching module, a translation processing module, and a format backfilling module.

[0013] The present invention also provides a computer device and a computer-readable storage medium for implementing the above method. Beneficial effects

[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) Ensure the integrity of the patent document structure: Through structured analysis and hierarchical citation records, the claim numbers, indentation format and paragraph structure of the translated patent document are completely consistent with the original document, ensuring that the legal effect of the translation is not affected; (2) Significantly improved consistency in terminology translation: Through terminology constraint decoding and full-text consistency verification, the translation of the same term in the claims and the specification remains consistent, and the terminology alignment accuracy rate can reach more than 95%; (3) Zero loss of figure labels: Through the figure label recognition and position locking mechanism, the figure label numbers are not mistranslated or shifted during the translation process, ensuring the integrity of the technical solution expression; (4) Flexible adaptation to multiple fields: By dynamically retrieving the professional terminology comparison table of different technical fields, the system can flexibly adapt to the translation needs of various patent documents such as mechanical, electrical, chemical, and biomedical fields; (5) Significantly improved translation efficiency: By reusing translations of family patents through translation memory, more than 30% of repetitive translation work can be reduced, which is particularly suitable for series patent applications and family patent management scenarios. Attached Figure Description

[0015] Figure 1 This is an overall flowchart of the intelligent translation method in an embodiment of the present invention; Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0017] like Figure 1 As shown, the intelligent translation method provided in this embodiment includes the following steps: S1: Obtain the patent document to be translated. Users upload the patent document file to be translated (supports PDF, Word, XML, HTML, and other formats) through the system front end. The system detects the file format; if it is in PDF format, it extracts the text content and hierarchical information through OCR and structural analysis engine; if it is in XML / HTML format, it directly parses the tag structure.

[0018] S2: Structural Analysis. The system performs deep structural analysis on patent documents, identifying and extracting the following information: ① Claims section: Distinguish between independent claims (e.g., "Claim 1: A...") and dependent claims (e.g., "Claim 2: According to claim 1..."), and record the reference relationship and nesting level of dependent claims to independent claims; ② Specification Section: This section includes subsections such as the technical field, background technology, invention content, description of drawings, and detailed embodiments; ③ Figure reference numerals: Identify the figure reference numerals (such as "10", "100a", "200") that appear in the text, and record their location and the technical features they modify.

[0019] S3: Domain Identification and Terminology Matching. Based on the IPC classification number of the patent document or user input, the system automatically determines the technical field to which the patent belongs (e.g., H01M for the battery field, G06F for the computer field), and retrieves a terminology lookup table for the corresponding field from a multi-domain patent terminology database. This lookup table contains bilingual mappings of core terms in the field, for example: Legal terms: Independent claim; Dependent claim; Embodiment; Novelty; Inventive step Technical terms (using the battery field as an example): Cathode (positive electrode); Anode (negative electrode); Electrolyte; Lithium-ion (lithium ion) S4: Neural Machine Translation and Terminology Constraint Alignment. The system inputs the extracted text content into a neural machine translation model based on the Transformer architecture. During the decoding phase, the system generates a terminology constraint mask matrix, forcing the model to output the corresponding target language vocabulary from the terminology lookup table when translating specific terms. Specifically, in the attention mechanism of the model decoder, constraints are applied to the output positions corresponding to the source words of the terms, causing their probability distribution to be biased towards the target terms in the terminology lookup table.

[0020] S5: Format Backfilling and Output. Based on the structural information recorded in S2, the system backfills the translated text according to the original format: claim numbers remain unchanged (e.g., "1.", "2.", "3."); dependent claims are indented according to the citation level; figure reference numerals remain in their original positions; and the headings of each section of the specification retain their original format. Finally, the system generates a translated patent document (PDF / Word format) for users to download and use.

[0021] Building upon Example 1, the system adds a Translation Memory (TM) module. When a user submits a new patent document translation task, the system first calculates the semantic similarity (based on cosine similarity or edit distance) between the text to be translated and existing translations in the Translation Memory. If the similarity between a paragraph and content in the Translation Memory exceeds a threshold (e.g., 0.85), the system directly reuses the existing translation, skipping translation model inference, thereby improving efficiency and ensuring consistency of terminology within the same patent family. The reused translation still undergoes terminology consistency verification to ensure quality.

[0022] This embodiment provides an intelligent translation system, including: Document parsing module: responsible for receiving and parsing uploaded patent document files, extracting structured information and text content; Terminology matching module: responsible for domain identification and retrieval of terminology lookup tables; Translation processing module: responsible for neural machine translation reasoning and term constraint alignment; Format backfilling module: responsible for assembling and outputting the translation results according to the original format; Translation memory: responsible for storing and retrieving historical translation results.

[0023] The above description is merely a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. Claim 1 A smart translation method based on patent document structure analysis and domain terminology alignment, characterized in that, The method includes the following steps: Obtain the electronic file of the patent document to be translated, wherein the electronic file of the patent document contains structured markup information, wherein the structured markup information at least identifies the claim section, the specification section and the hierarchical number within each section; The electronic patent document is structurally analyzed to extract the independent claims and dependent claims texts from the claims section, and the citation relationships and hierarchical nesting structure between the claims are recorded. Based on user input or automatic system identification, the technical field to which the patent document belongs is determined, and a professional terminology comparison table of the corresponding technical field is retrieved from a pre-constructed multi-field patent terminology database. The multi-field patent terminology database at least includes the mapping relationship between Chinese and foreign patent legal terms and professional technical terms in the fields of mechanics, electrical engineering, chemistry, biomedicine and computer science. The extracted claim text and specification text are input into a neural machine translation model for preliminary translation. During the translation process, the terms in the terminology lookup table are subject to forced term constraint alignment to ensure that the same term is translated consistently in the claim and specification. Based on the citation relationships and hierarchical nesting structure recorded, the translated text is backfilled into the corresponding sections according to the format specifications of the original patent document, generating a translated patent document that retains the original claim numbers, paragraph indentation, and figure mark formats.

2. Claim 2 The intelligent translation method according to claim 1 is characterized in that, The structural analysis of electronic patent documents also includes: identifying the figure mark numbers in the patent documents, and retaining the figure mark numbers from being translated and keeping their position relative to the technical features unchanged during the translation process.

3. Claim 3 The intelligent translation method according to claim 1 is characterized in that, The mandatory terminology constraint alignment specifically includes: injecting terminology constraint masks during the decoding stage of the neural machine translation model, so that the source words of the terms are forcibly mapped to the corresponding target language terms in the terminology lookup table, and performing terminology consistency verification on the translation output. If there are terminology translation conflicts, the terminology lookup table is used as the standard for post-processing replacement.

4. Claim 4 The intelligent translation method according to claim 1 is characterized in that, The method further includes: constructing a translation memory, storing the translations of the claims in the translated family of patent documents into the translation memory, and when translating the patent document to be translated, if a paragraph with a similarity to the content in the translation memory exceeds a preset threshold is detected, the corresponding translation is directly reused and the terminology consistency is updated synchronously.

5. Claim 5 The intelligent translation method according to claim 1 is characterized in that, The method also includes: performing format verification on the translated patent documents, checking the continuity of claim numbers, the completeness of figure mark citations, and the standardization of paragraph indentation; if any abnormalities are found, a manual review process is automatically triggered.

6. Claim 6 An intelligent translation system based on patent document structure analysis and domain terminology alignment, characterized in that: The system includes: The document parsing module is configured to acquire electronic files of patent documents to be translated and parse their structured markup information, extracting the text of the claims section and hierarchical citation relationships; The terminology matching module is configured to retrieve a corresponding professional terminology comparison table from a multi-field patent terminology database based on the technical field to which the patent document belongs; The translation processing module is configured to perform neural machine translation on the extracted text and apply terminology constraint alignment of the terminology lookup table during the decoding stage; The format backfilling module is configured to generate translated patent documents by backfilling the translated text based on the hierarchical nesting structure and format specifications of the original patent documents.

7. Claim 7 A computer device includes a memory and a processor, wherein the memory stores a computer program, characterized in that... When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.

8. Claim 8 A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.