Method for constructing multilingual parallel corpus of patent translation and patent translation system

By constructing a multilingual parallel corpus and training a large-scale patent translation model, the problems of insufficient translation accuracy and corpus adaptation in patent translation were solved, achieving efficient and accurate cross-language patent information dissemination and analysis.

CN122389891APending Publication Date: 2026-07-14BEIJING AUGUST MELON TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-17
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies in patent translation suffer from insufficient translation accuracy, lack of professionalism, non-standard corpus construction processes, and difficulty in adapting corpora and models to technological iterations. This leads to errors in cross-border technology layout decisions, low efficiency in cross-language technology intelligence analysis, and a large workload of repetitive examination.

Method used

By acquiring patent family identifiers for the same technical solution in different countries or regions, standardizing them, and forming a structured patent family data set, multi-level cross-parallel corpus alignment is performed by combining the inherent structural features and semantic information of the patent text. A multi-dimensional quality scoring model is used for automated evaluation, and parallel corpus pairs that meet the preset quality threshold are selected. A multilingual parallel corpus is constructed, and a large patent translation model is trained to adapt to patent scenarios.

Benefits of technology

The construction of a high-quality, multilingual parallel corpus has been achieved, which has improved the accuracy and adaptability of patent translation, solved the problem of insufficient accuracy of general translation software in patent scenarios, improved the accuracy and consistency of translation, and reduced the workload of repetitive examination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122389891A_ABST
    Figure CN122389891A_ABST
Patent Text Reader

Abstract

The application provides a multilingual parallel corpus construction method for patent translation, comprising: based on international patent application information, obtaining the same family patent document identifiers of the same technical solution in different countries or regions, forming a structured family patent data set; obtaining corresponding multilingual patent full-text texts, sequentially performing noise removal, structure analysis and language standardization processing, and obtaining a pair of multilingual structured family patent texts; combining the inherent structure characteristics and semantic information of the patent texts, performing multi-level cross-parallel corpus alignment from documents, chapters to sentences, generating a preliminary parallel corpus pair; using a quality scoring model containing multiple preset evaluation dimensions, performing automatic quality evaluation; according to the evaluation results, screening out parallel corpus pairs that meet the preset quality threshold to constitute a multilingual parallel corpus, so as to solve the problems of insufficient precision and professionalism of general translation software adapted to the patent scene, and non-standard construction process of patent multilingual parallel corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of patented machine translation technology in natural language processing, and in particular to a method for constructing a multilingual parallel corpus for patent translation and a patent translation system. Background Technology

[0002] With the deep integration of global technological innovation and intellectual property protection, patents, as the core carriers of technological, legal, and commercial value, face an increasingly urgent need for cross-language dissemination and utilization. Currently, global patent texts cover dozens of languages, and patent families for the same invention filed in different countries / regions provide a natural foundation for cross-language technical information association. Existing general-purpose translation software primarily uses major languages ​​such as Chinese and English as its core, achieving multilingual coverage through intermediary translation. Meanwhile, attempts have emerged in the industry to construct translation datasets based on patent texts. Some solutions, leveraging the textual correlation of patent families, have initially formed simple bilingual comparative corpora to support the training of patent translation models.

[0003] However, existing technologies have significant shortcomings in patent translation scenarios: First, general-purpose translation software lacks accuracy in translating patents in less common languages. Due to a lack of optimization for the characteristics of the patent field, the translation of legal and technical terms is inconsistent and difficult to adapt to the fixed structure of patent texts. Second, existing methods for constructing parallel patent corpora lack a systematic and standardized process. Cross-parallel corpus alignment only reaches the level of surface text matching and fails to achieve accurate alignment of both structure and semantics. At the same time, the quality assessment dimension is singular and cannot take into account both the accuracy and coverage of the corpus. Third, the rapid iteration of patent technologies makes it difficult to update existing corpora in real time, causing translation models to be unable to adapt to new technical terms and expressions, further reducing the technical adaptability of translation.

[0004] The aforementioned problems directly lead to many adverse consequences: when enterprises conduct transnational technology layout, they may misjudge technological barriers and competitors' layouts due to errors in patent translation in less common languages, resulting in decision-making errors; research institutions and patent examiners face problems such as low efficiency in cross-language technical intelligence analysis and a large workload of repetitive examinations; and patent translation models, due to the lack of high-quality and highly adaptable training data, are unable to meet the stringent requirements of translation accuracy and consistency in professional scenarios.

[0005] Therefore, there is an urgent need for a method for constructing a multilingual parallel corpus for patent translation, as well as a method for training a large-scale patent translation model, in order to solve the technical problems of insufficient accuracy and professionalism of general translation software in adapting to patent scenarios, non-standard construction process of existing multilingual parallel corpora for patents with lack of alignment and quality control, and difficulty in adapting corpora and models to technological iterations. Summary of the Invention

[0006] To overcome the problems existing in related technologies, this disclosure provides a method for constructing a multilingual parallel corpus for patent translation, in order to solve the technical problems in related technologies such as the lack of accuracy and professionalism of general translation software in adapting to patent scenarios, the non-standard construction process of existing patent parallel corpus and the lack of alignment and quality control, and the difficulty in adapting corpora and models to technological iterations.

[0007] This specification provides one or more embodiments of a method for constructing a multilingual parallel corpus for patent translation, including the following steps: Based on international patent application information, obtain the patent family document identifiers of the same technical solution in different countries or regions, and standardize the document identifiers to form a structured patent family data set. Obtain the full-text multilingual patents corresponding to the patent family data set, and sequentially perform noise removal, structural parsing and language normalization on the full-text multilingual patents to obtain multilingual structured patent family text pairs with unified chapters, paragraphs and sentences. Based on the aforementioned multilingual structured family of patent text pairs, and combining the inherent structural features and semantic information of the patent texts, multi-level cross-parallel corpus alignment from documents, chapters to sentences is performed to generate preliminary parallel corpus pairs. The multi-level cross-parallel corpus alignment process comprehensively utilizes word alignment methods based on statistical models and rule matching methods based on patent structural features. An automated quality assessment of the preliminary parallel corpus pairs is performed using a quality scoring model that includes multiple preset evaluation dimensions. The quality scoring model includes a structural consistency score for evaluating structural correspondence, a term density comparison score for evaluating terminology coverage, and a two-way translation consistency score for evaluating translation fidelity. Based on the evaluation results, parallel corpus pairs that meet the preset quality threshold are selected to form the final multilingual parallel corpus.

[0008] Preferably, the rule matching method based on patent structural features specifically includes the following steps: Using the claim numbers, formula numbers, and chapter titles in the patent text as alignment anchors, the corresponding paragraphs and sentences in the same patent text are initially matched and located using these alignment anchors. This clarifies the corresponding relationships between the texts, reduces initial alignment deviations, and is then calibrated in conjunction with sentence length ratios.

[0009] Preferably, after performing multi-level cross-parallel corpus alignment of sentences, semantic layer alignment is also included, specifically comprising the following steps: The semantic features of sentences are extracted using a pre-trained multilingual model, and adaptive feature fusion is performed by combining surface features. The semantic similarity between sentence pairs is calculated based on the fused features, thereby verifying and optimizing the sentence-level alignment results.

[0010] Preferably, the quality scoring model specifically includes the following seven evaluation dimensions: Structural consistency score, term density comparison score, length ratio score, bidirectional translation consistency score, semantic similarity score, numerical and symbol consistency score, and citation consistency score; The final quality score is calculated by weighted average, and the weights of each dimension can be dynamically adjusted according to different parallel corpus pairs or technical fields.

[0011] Preferably, it also includes corpus enhancement, specifically including the following steps: The corpus enhancement includes back-translating the selected parallel corpus pairs to generate semantically consistent variant corpora, and / or using a two-way cross-validation algorithm to further improve alignment confidence through multi-stage similarity calculation.

[0012] This specification provides one or more embodiments of a patented translation system that applies the multilingual parallel corpus construction method as described in claims 1-5, comprising: The corpus construction module is used to construct and maintain a multilingual parallel corpus based on the aforementioned multilingual parallel corpus construction method. The model training module is used to train a large-scale patent translation model using the aforementioned multilingual parallel corpus. The translation service module integrates the aforementioned patent translation model, which is used to receive patent text, images, or documents input by users, call the patent translation model to perform translation, and supports users to upload custom terminology lists to prioritize the use of custom terms in translation.

[0013] Preferably, the training method for the large-scale patent translation model includes the following steps: Using the aforementioned multilingual parallel corpus, the open-source large language model was fine-tuned with all parameters to learn the general language style, structure, and terminology of patent texts. Based on the full parameter fine-tuning, a low-rank adaptation technique is used to fine-tune the model using patent data from specific technical subfields; It integrates the DeepSeed ZeRO3 acceleration solution and Flash Attention tuning technology, adopts adaptive gradient optimization algorithm and fine-tuned hyperparameter configuration, and improves training efficiency through multi-machine and multi-card distributed training; The translation quality of the model is verified by combining automatic evaluation indicators with specific indicators and supplementing them with manual evaluation. Inference deployment is based on the vLLM framework, supports model quantization processing, adopts a hybrid deployment mode of edge computing and cloud services, and provides a user-specific terminology upload interface, which is used first during translation.

[0014] Preferably, the open-source large model is the Qwen3 base model, and the full parameters include the characteristic expressions of the patent claims and the complex technical principle explanations in the specification.

[0015] This specification provides one or more embodiments of a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described above.

[0016] This specification provides one or more embodiments of a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the method described above.

[0017] This disclosure provides a method for constructing a multilingual parallel corpus for patent translation, a training method, system, device, and medium for a large-scale patent translation model. Its advantages lie in that, based on international patent application information, it obtains patent document identifiers for the same technical solution in different countries or regions, and standardizes these identifiers to form a structured set of patent family data. This ensures the uniqueness and consistency of the document identifiers, providing a clear index for subsequent text acquisition and cross-parallel corpus alignment, thus improving the orderliness and efficiency of corpus construction. It obtains the multilingual full-text patent data corresponding to the patent family data set, and sequentially performs noise removal, structural analysis, and language standardization on these full-text patents to obtain multilingual structured patent family text pairs with unified chapters, paragraphs, and sentences. This effectively solves the problems of heterogeneous multilingual text formats and inconsistent expressions, reducing the difficulty and improving the accuracy of subsequent cross-parallel corpus alignment. Based on these multilingual structured patent family text pairs, and combined with the inherent structural features and semantic information of the patent text, it performs text-level alignment from document to chapter. The process involves multi-level cross-parallel corpus alignment from section to sentence level to generate preliminary parallel corpus pairs. This multi-level alignment process integrates word alignment methods based on statistical models and rule matching methods based on patent structural features, achieving precise matching of both structure and semantics. This overcomes the alignment challenges of complex patent structures and dense terminology, laying a foundation for high-quality corpus. A quality scoring model with multiple preset evaluation dimensions is used to automatically evaluate the preliminary parallel corpus pairs. This model includes a structural consistency score for evaluating structural correspondence, a terminology density comparison score for evaluating terminology coverage, and a two-way translation consistency score for evaluating translation fidelity. This avoids the bias of a single evaluation, improves the efficiency of quality review, and provides a quantitative and reliable basis for corpus selection. Based on the evaluation results, parallel corpus pairs that meet preset quality thresholds are selected to form the final multilingual parallel corpus. This corpus balances multilingual coverage and technical adaptability, providing core training data for a large-scale patent translation model and solving the problem of insufficient accuracy of general corpora in patent scenarios. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a method for constructing a multilingual parallel corpus for patent translation, provided for one or more embodiments of this specification; Figure 2 A schematic diagram of the structure of the patent translation system provided in one or more embodiments of this specification; Figure 3 Paragraph translation pages provided for one or more embodiments of this specification; Figure 4 A schematic diagram illustrating the recognition and translation of an image received from user input, provided in one or more embodiments of this specification; Figure 5 A schematic diagram illustrating the recognition and translation of a file received from user input, provided in one or more embodiments of this specification; Figure 6 Fine-tuning architecture diagrams provided for one or more embodiments of this specification; Figure 7 A schematic diagram illustrating the English-to-Chinese translation results of the patent translation big model provided for one or more embodiments of this specification; Figure 8 A schematic diagram illustrating the English-to-Chinese translation results of existing translation services provided for one or more embodiments of this specification; Figure 9 A schematic diagram illustrating the Chinese-to-English translation results of a large-scale patent translation model provided for one or more embodiments of this specification; Figure 10 A schematic diagram illustrating the Chinese-to-English translation results of existing translation services provided for one or more embodiments of this specification; Figure 11 This is a schematic diagram of the structure of a computer device provided for one or more embodiments of this specification. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this invention.

[0021] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings.

[0022] Method Implementation Examples According to embodiments of the present invention, a method for constructing a multilingual parallel corpus for patent translation is provided, such as... Figure 1The diagram shows a flowchart of the multilingual parallel corpus construction method for patent translation provided in this embodiment. The core value of the family-based parallel corpus (i.e., the corpus of patent texts of the same invention applied for in different countries / regions after alignment with cross-parallel corpora) lies in solving the precise association and reuse of cross-language technical information in the patent field. The family-based parallel corpus is a semantic bridge across languages, solving the problem of precise correspondence when different languages ​​describe the same thing; while the high-quality dataset of the basic patent corpus is a high-quality raw material for a single language, solving the problem of high-quality utilization of information within the same language. The two serve different scenarios but can work synergistically, such as first using the basic corpus to train a single-language patent LLM, and then using the parallel corpus to expand its multilingual capabilities. The multilingual parallel corpus construction method for patent translation according to this embodiment of the invention includes the following steps: S110. Based on international patent application information, obtain the patent family document identifiers of the same technical solution in different countries or regions, and standardize the document identifiers to form a structured patent family data set.

[0023] Specifically, this includes (i) obtaining information on patent families: 1) Extraction of PCT International Application Information Extract core information from the PCT international publication text, including the PCT application number (format: WO + year + serial number, e.g., WO2025123456), international filing date, and priority information (if there are multiple priorities, all priority numbers and corresponding countries / regions must be recorded). Simultaneously, analyze the family patent information fields recorded in the text, such as the "INPADOC family" or "EPO family" sections, to initially obtain possible family application country / region codes (e.g., US, EP, JP, etc.).

[0024] (ii) Patent database search Utilize a database containing 200 million professional patents for in-depth searching: Search by PCT application number: Enter the PCT application number in the database search box to trigger the patent family clustering function. The database will generate a global list of patent families corresponding to that PCT application based on algorithms such as technical similarity and priority association.

[0025] Extended search strategy: If the initial search results are incomplete, supplementary searches can be conducted by combining priority number and invention name keywords. Among them, invention name keywords need to be expanded by multilingual translation, such as French, German and other keywords corresponding to English keywords, to ensure that no important related members are missed.

[0026] (iii) Standardization of application numbers within the same family The retrieved application numbers from the same family were categorized and organized, mainly including: PCT system numbering: International application number (e.g., PCT / CN2023 / 012345), International publication number (WO2025123456).

[0027] Country / Region Phase Number: This includes the application number that has entered the national phase (e.g., US20250123456, corresponding to the US national phase), publication number (e.g., EP4567890A1, European patent publication text), and grant announcement number (e.g., JP6789012B2, Japanese granted patent).

[0028] Unified encoding format: The WIPO standard ST.3 patent document number format is used for standardization, removing prefixes and suffixes specific to different databases to ensure the uniqueness of the number. For example, “US-2025-0123456-A1” is unified as “US20250123456A1”.

[0029] (ii) Obtaining patent texts from multiple countries / regions 1) Text Acquisition Channels Official database interfaces and commercial database purchase channels, etc.

[0030] (ii) Text type filtering Select the target text based on the patent examination stage: (1) Application for public disclosure: contains a complete description of the technical solution, is suitable for early technology tracking, and has a relatively standardized text structure.

[0031] (2) Authorization Announcement Text: After review and revision, the claims are more accurate and suitable for legal status analysis.

[0032] (3) Revised text: If there is a revision of the patent content, the differences between the text before and after the revision must be obtained at the same time.

[0033] (iii) Deduplication and Integrity Verification (1) Text fingerprint generation: Calculate the MD5 hash value of each patent text. If the text hash values ​​corresponding to different application numbers are the same, they are determined to be duplicate texts, and only the latest version is retained.

[0034] (2) Structural integrity check: Verify whether the text contains complete sections, such as title, abstract, claims, technical field, background art, specification, invention summary, description of drawings, and detailed embodiments. Texts missing key sections need to be marked and retrieved again.

[0035] S120. Obtain the multilingual full-text patent data corresponding to the patent family data set, and sequentially perform noise removal, structural analysis and language normalization processing on the multilingual full-text patent data to obtain multilingual structured patent family text pairs with unified chapters, paragraphs and sentences.

[0036] Among them, noise removal specifically includes the following steps: (1) Fixed template cleaning: Remove page numbers, repeated application numbers, and legal declarations in the header and footer through regular expressions, such as the matching pattern \d{4}-\d{2}-\d{2}Page\d+ (date and page number combination).

[0037] (2) Non-text element filtering: Identify and delete drawing numbers, chemical structural formulas, and mathematical formulas, and retain pure text content.

[0038] (3) Redundant content processing: Merge consecutive repeated whitespace characters, merge paragraphs separated by multiple lines into continuous text, and judge based on punctuation and semantic coherence. For example, a line break after a period or semicolon is regarded as a separation within a paragraph.

[0039] [[ID=!2]]Structural analysis specifically includes the following steps: Adopt a hierarchical analysis method to decompose the patent text into a three-level structure: (1) Section level: Identify the main sections through title matching (such as "Specification", "Description", "発明の説明"), and establish a section index table (including section name, starting page number, and text offset).

[0040] (2) Paragraph level: Within a section, split paragraphs according to punctuation and indentation rules (for example, in Chinese, a period, exclamation mark, or question mark is used as the end mark of a paragraph; in English, a sentence starting with a capital letter after a period is used as the start of a paragraph), and generate a unique ID for each paragraph (format: section code - paragraph number, e.g., SPEC-001).

[0041] (3) Sentence level: Use tokenization tools such as NLTK (for English), jieba (for Chinese), and MeCab (for Japanese) to perform sentence segmentation, handle complex punctuation cases (such as ellipsis "..." in English and dash "——" in Chinese), and ensure accurate sentence boundaries (accuracy requirement ≥ 95%).

[0042] Language normalization processing, specifically language normalization: (1) Character encoding conversion: Unify to UTF-8 encoding, handle special characters (such as ß in German, é in French, and ぁ in Japanese), and use Python's unicodedata library for standardization.

[0043] (2) Format unification: Convert full-width symbols to half-width (such as converting full-width space U+3000 to half-width space U+0020), and unify the number format (such as converting Chinese numeral "一" to Arabic numeral "1", applicable to numerical descriptions in claims). It should be noted that in the original text, "结构解析" was misspelled as "结构解析" in line 12. It should be "结构解析". The above translation has been corrected accordingly.

[0044] (3) Terminology standardization: Establish a multilingual terminology database in the field of patents that includes technical terms, legal terms, and institutional names. Use regular expression substitution to achieve terminology unification, such as unifying "patent of invention" as "patent of invention".

[0045] S130. Based on the aforementioned multilingual structured family of patent text pairs, and combining the inherent structural features and semantic information of the patent texts, perform multi-level cross-parallel corpus alignment from document, chapter to sentence to generate preliminary parallel corpus pairs. The multi-level cross-parallel corpus alignment process comprehensively utilizes a word alignment method based on statistical models and a rule matching method based on patent structural features, specifically including the following steps: Using the claim numbers, formula numbers, and chapter titles in the patent text as alignment anchors, the corresponding paragraphs and sentences in the same patent text are initially matched and located using these alignment anchors. This clarifies the corresponding relationships between the texts, reduces initial alignment deviations, and is then calibrated in conjunction with sentence length ratios.

[0046] Specifically, at the structural level: the text alignment strategy employs a three-level alignment framework, combining statistical models with rule matching. (1) Document-level alignment: Based on the application number association of patent family, a multilingual text mapping table is established, with fields including: PCT application number, country code, language code, text storage path, and alignment status.

[0047] (2) Chapter-level alignment: Utilizing the standardized structure of the patent text, the Levenshtein distance algorithm is used to match the titles in multiple languages. The matching threshold is set to 0.8, that is, a similarity of ≥80% is used to determine the corresponding chapter, and a chapter-level alignment relationship is established, such as the “Instruction Manual” in the CN text corresponding to the “Specification” in the US text.

[0048] (3) Sentence-level alignment: ① Statistical model method: Use IBM Model 1-5 for word alignment. First, segment the parallel corpus into words, and Chinese texts need to be tagged with part-of-speech tags. Calculate the word alignment probability matrix and iteratively optimize it using the EM algorithm (Expectation-Maximization Algorithm). The number of iterations needs to be ≥10 to generate sentence-level alignment results.

[0049] ② Rule-based matching method: Utilizing the structured features of the patent text, the numbers in the claims ("1.", "(1)", "Item 1") and formula numbers (Equation 1) are used as anchor points for sentence alignment. Calibration is performed by combining punctuation marks and sentence length ratios (allowing a length difference of ±20%). Specifically, punctuation mark calibration refers to comparing the distribution of punctuation marks (such as the position and number of commas, periods, and semicolons) in the initially matched sentences to determine whether the sentence segmentation logic is consistent. If there is a large deviation in punctuation position or an inconsistency in punctuation type, the sentence matching range is adjusted to ensure that the matched sentences are complete and semantically corresponding. Sentence length ratio calibration refers to calculating the length ratio between the initially matched sentences. Taking the sentence length of one version as a benchmark, a length difference of ±20% is allowed. If the ratio exceeds this range, it indicates that there is a deviation in the initial matching. The anchor points are re-located, and the matched sentences are adjusted until the length ratio meets the requirements, ultimately achieving accurate sentence alignment.

[0050] It also includes semantic layer alignment, which specifically includes the following steps: The semantic features of sentences are extracted using a pre-trained multilingual model, and adaptive feature fusion is performed by combining surface features. The semantic similarity between sentence pairs is calculated based on the fused features, thereby verifying and optimizing the sentence-level alignment results.

[0051] Semantic layer: Combining cross-linguistic sentence vector models and a patent terminology database, sentence pair similarity is calculated. A domain-adaptive pre-trained model (fine-tuned on 10 million patent sentence pairs) is used to improve the accuracy of professional terminology recognition to 96%. Pre-trained multilingual models (such as mBERT and XLM-Roberta) are used to extract semantic features, while shallow networks are combined to extract surface information features (word frequency vectors and word vectors). The semantic feature vectors, word frequency vectors, and word vectors are concatenated according to preset rules through a feature fusion layer to form a fused feature vector.

[0052] To address the characteristics of patent texts, which are characterized by dense technical terms and complex structures, an adaptive feature fusion mechanism is designed to dynamically adjust the weights of different features across different parallel corpus pairs and technical domains. The system first extracts semantic features using XLM-R, but unlike traditional methods, it employs an attention-guided feature extraction strategy. By adding a patent structure-aware attention mechanism to each layer of XLM-R, the model focuses more on key technical terms and structural elements within the patent.

[0053] In specific implementation, in the 12-layer Transformer structure of XLM-R, structure-aware attention heads are added to layers 3, 6, 9, and 12, respectively, assigning different attention weights to different parts such as claims, specifications, and abstracts. For surface information features, a domain-adaptive word representation method is adopted, which fine-tunes the word vector model on patent corpora in specific technical fields to make the word vectors better capture the semantics of professional terms. At the same time, a term-aware word frequency calculation method is designed to give higher weight to professional terms. The word frequency calculation formula is tf(t)=α·frequency(t)^β, where α and β are parameters that are automatically adjusted according to the importance of the term, t refers to the term (term / word / vocabulary), frequency(t): the number of times this term / vocabulary t appears in the text, and tf(t): the word frequency of term t after weighting. In the feature fusion layer, an adaptive gating mechanism is introduced to dynamically adjust the fusion ratio of semantic features and surface features according to the characteristics of the input text: F_fusion = G·F_semantic + (1-G)·F_surface, where G is the gating value learned by the neural network and the value range is [0,1].

[0054] Experiments show that when processing patent texts from European and Asian language families, the system achieves parallel corpus accuracy of 97.2% and 96.5% respectively. Furthermore, while maintaining 95% accuracy, the corpus coverage is increased by 12%, effectively resolving the contradiction between accuracy and coverage.

[0055] It also includes corpus enhancement, specifically including the following steps: The selected parallel corpus pairs are back-translated to generate semantically consistent variant corpora, and / or a two-way cross-validation algorithm is used to further improve the alignment confidence through multi-stage similarity calculation.

[0056] Specifically, it includes the following steps: ① Back Translation: Translate the text of a minor language into English first, and then translate it back into the original language to generate variant corpora, which are used to expand the data of low-frequency languages.

[0057] ②Synonym replacement: Synonym replacement is performed on non-key terms, and the corpus diversity is enriched based on dictionaries such as WordNet and Chinese thesaurus.

[0058] ③ Two-way cross-validation: Improve alignment accuracy by verifying the consistency of translations from the source language to the target language and from the target language to the source language.

[0059] 1) Construct a bidirectional sentence vector model specific to the patent field, train encoders in two directions on each parallel corpus pair, and use contrastive learning methods to optimize the vector space so that similar sentence pairs are closer in the vector space; 2) Design a four-stage cross-validation algorithm: First, calculate the similarity score S1 between the source sentence and the target sentence. Then, translate the target sentence back into the source language using machine translation and calculate the similarity score S2 between the original source sentence and the translated sentence. Next, translate the source sentence back into the target language and calculate the similarity score S3 between the original target sentence and the translated sentence. Finally, use a weighted fusion formula: Q = α·S1 + β·S2 + γ·S3; Where α+β+γ=1, calculate the final mass fraction Q; 3) Introducing a patent structure-aware context enhancement mechanism, which improves the ability to process patent-specific long and complex sentences by considering the semantic coherence of sentences before and after paragraphs and correcting the quality scores of isolated sentence pairs.

[0060] Experiments show that while maintaining an accuracy of 96.2%, this method increases the effective corpus coverage from 83% to 87.5%, especially when dealing with complex paragraphs such as patent specifications and claims, where the alignment accuracy is improved by 4.8 percentage points.

[0061] S140. An automated quality assessment is performed on the preliminary parallel corpus pairs using a quality scoring model that includes multiple preset evaluation dimensions. The quality scoring model includes a structural consistency score for evaluating structural correspondence, a term density comparison score for evaluating terminology coverage, and a two-way translation consistency score for evaluating translation fidelity.

[0062] The quality scoring model specifically includes the following seven evaluation dimensions: 1) Structural consistency score (weight 0.25): Based on the node matching degree of the patent document structure tree, the degree of correspondence between two text segments in structural position is calculated. The Tree Edit Distance (TED) algorithm is used, and the threshold is set to 0.85. 2) Terminology density comparison score (weight 0.20): Calculate the ratio of the coverage of professional terms in bilingual texts, and use a terminology database specific to patent classification numbers for matching, with a threshold set to 0.75; 3) Length ratio score (weight 0.15): Based on the statistical characteristics of different parallel corpus pairs, set parallel corpus pair specific parameters (e.g., English-Chinese 1:0.6±0.15, English-Japanese 1:1.2±0.2) and calculate the deviation between the actual length ratio and the ideal ratio. 4) Bidirectional translation consistency score (weight 0.15): The BLEU score of the source language to target language back translation is calculated using a multi-head attention mechanism (12 heads, hidden layer dimension 768), with a threshold set to 0.7; 5) Semantic similarity score (weight 0.10): Cosine similarity is calculated using semantic vectors extracted by the cross-lingual Sentence-Transformers model, with a threshold of 0.8; 6) Numerical and symbol consistency score (weight 0.10): The matching degree of numerical values, units and special symbols in the text is extracted and compared by regular expressions, with a threshold of 0.9; 7) Citation consistency score (weight 0.05): Detects the consistency of patent numbers, document numbers and figure numbers cited in bilingual texts, with a threshold of 0.9.

[0063] The system uses a weighted average method to calculate the final quality score and dynamically adjusts the weights of each indicator (within ±0.05) according to different parallel corpus pairs and patent fields, achieving 97% parallel corpus accuracy and 88% corpus coverage, demonstrating excellent performance in both Eurasian and non-Eurasian language families.

[0064] S150. Based on the evaluation results, select parallel corpus pairs that meet the preset quality threshold to form the final multilingual parallel corpus.

[0065] The storage and management of multilingual parallel corpora includes the following steps: 1) Database Design A three-tier data model is built using a relational database (such as PostgreSQL): (1) Metadata layer: Stores basic patent information, including PCT application number (primary key), international application date, priority information (foreign key associated with priority table), and list of countries in the same family (array type, storing ISO3166-1alpha-2 code).

[0066] (2) Text layer: Stores patent texts in various languages. Fields include application number (foreign key), language code (ISO639-1 code, such as zh, en, ja), text content (CLOB type, supports large text storage), and text type (application publication / authorization announcement / amendment text).

[0067] (3) Alignment layer: Stores alignment relationships, including source language ID, target language ID, alignment unit type (document / chapter / sentence), source unit ID, target unit ID, and alignment confidence (0-1 value, automatically evaluated results).

[0068] (ii) Indexing and Query Optimization (1) Full-text search: An inverted index is built for the text content, and the full-text search module of PostgreSQL (tsvector / tsquery) is used to support multilingual word segmentation (by configuring parsers such as pg_catalog.zh_parser and pg_catalog.en_parser).

[0069] (2) Distributed storage: For corpora of TB or more, a distributed database (such as Elasticsearch) is used for sharded storage. Data is sharded by country code + year (e.g., index name patent_corpus_zh_2025) to improve query efficiency (response time ≤ 200ms).

[0070] (iii) Version Control and Updates (1) Version management mechanism: Generate a version number for each patent text (format: application number-year-version number, e.g. WO2025123456-2025-001), and record the modification history (including acquisition time, preprocessing operation, alignment correction record).

[0071] (2) Incremental update strategy: Scan the PCT bulletin (WIPO publishes it once a week) daily to obtain newly published PCT applications, and determine whether they are new members of the same family through priority matching to achieve dynamic expansion of the corpus (update delay ≤ 24 hours).

[0072] The method provided in this embodiment obtains patent family document identifiers for the same technical solution in different countries or regions based on international patent application information, and standardizes these document identifiers to form a structured patent family data set. This ensures the uniqueness and consistency of the document identifiers, providing a clear index for subsequent text acquisition and cross-parallel corpus alignment, and improving the orderliness and efficiency of corpus construction. The method then obtains the multilingual full-text patents corresponding to the patent family data set, and sequentially performs noise removal, structural analysis, and language standardization on these multilingual full-text patents to obtain multilingual structured patent family text pairs with unified chapters, paragraphs, and sentences. This effectively solves the problems of heterogeneous multilingual text formats and inconsistent expressions, reducing the difficulty and improving the accuracy of subsequent cross-parallel corpus alignment. Based on these multilingual structured patent family text pairs, and combining the inherent structural features and semantic information of the patent texts, multi-level cross-parallel corpus alignment from documents and chapters to sentences is performed to generate preliminary parallel corpus alignment. The corpus pairs, in particular, employ a multi-level cross-parallel corpus alignment process that integrates word alignment methods based on statistical models and rule matching methods based on patent structural features. This achieves precise matching of both structure and semantics, overcoming the alignment challenges of complex patent structures and dense terminology, and laying the foundation for high-quality corpora. A quality scoring model with multiple preset evaluation dimensions is used to automatically evaluate the quality of the initial parallel corpus pairs. This model includes a structural consistency score for evaluating structural correspondence, a terminology density comparison score for evaluating terminology coverage, and a two-way translation consistency score for evaluating translation fidelity. This avoids the bias of a single evaluation, improves the efficiency of quality review, and provides a quantitative and reliable basis for corpus selection. Based on the evaluation results, parallel corpus pairs that meet preset quality thresholds are selected to form the final multilingual parallel corpus. This corpus balances multilingual coverage and technical adaptability, providing core training data for the large-scale patent translation model and solving the problem of insufficient accuracy of general corpora in patent scenarios.

[0073] System Implementation Examples According to embodiments of the present invention, a patent translation system is provided, such as... Figure 2 The diagram shown is a structural schematic of the patent translation system provided in this embodiment. The patent translation system according to this embodiment includes: Corpus construction module 21 is used to execute the multilingual parallel corpus construction method for patent translation as described above, and to construct and maintain the multilingual parallel corpus.

[0074] The model training module 22 is used to train the patent translation large model using the multilingual parallel corpus and the training method of the patent translation large model as described above.

[0075] Translation service module 23 integrates the aforementioned patent translation model. It receives patent text, images, or documents input by the user, calls the patent translation model for translation, and supports user-uploaded custom terminology lists for priority use in translation. Users can upload their own terminology lists for mutual translation of professional terms; each user has their own exclusive terminology list, which will be prioritized when translating using the patent translation model. For example... Figure 3 As shown, this is a display of the paragraph translation page provided in this embodiment. Figure 4 The diagram shown is a schematic representation of the image recognition and translation provided in this embodiment. Figure 5 The diagram shown is a schematic of receiving and translating user-input files according to this embodiment.

[0076] The apparatus provided in this embodiment obtains family patents based on international patent application information and standardizes document identifiers through the corpus construction module 21 to form a structured dataset; it acquires multilingual patent texts, and after denoising, structural parsing, and language normalization processing, obtains hierarchically unified multilingual structured family patent text pairs; it generates preliminary corpus pairs through multi-level cross-parallel corpus alignment of documents, chapters, and sentences, combined with statistical models and structural rule matching; it uses a multi-dimensional quality scoring model for automated evaluation, and after screening, constructs a high-precision corpus that balances multilingual coverage and technical adaptability; the model training module 22 uses the constructed parallel corpus as training data to fine-tune all parameters of the open-source large model to adapt it to the specific needs. It leverages general style, structure, and terminology; then refines and adapts to specific sub-domain patent data using LoRA technology to reduce computational costs; integrates acceleration technologies and optimization algorithms, and improves efficiency through multi-machine, multi-GPU distributed training; ensures translation accuracy through collaborative verification of automatic evaluation, specific indicators, and manual evaluation; the translation service module 23 integrates the trained large model, supporting patent text, images, and various document formats for input; deployed based on the vLLM framework, it supports quantization processing and hybrid deployment modes, improving inference speed by 3-5 times, with a single node capable of processing 200+ requests per second; it provides a user-specific terminology upload interface, prioritizing the use of custom terms to balance translation accuracy, low latency, and personalized needs.

[0077] In one embodiment, the training method for the large-scale patent translation model includes the following steps: Using the aforementioned multilingual parallel corpus, a full-parameter fine-tuning technique is employed to fine-tune the open-source large language model, learning the general language style, structure, and terminology of patent texts. The open-source large model is based on the Qwen3 platform, and the full parameters include the signature expressions of patent claims and the complex technical principle explanations in the specification.

[0078] Specifically, full-scale fine-tuning technology achieves a comprehensive adjustment of the entire parameter space by directly updating all weight parameters of the model, thereby allowing the model to perform deep learning and optimization within a specific domain. While this method requires significant computational resources, such as multi-GPU distributed training environments, gradient accumulation strategies, and mixed-precision training to manage memory and computational overhead, it more thoroughly captures domain-specific subtle patterns, avoiding approximation biases that might be introduced by parameter-efficient methods, and ensuring overall performance improvement in patent translation tasks. In the patent translation scenario, the R&D team constructed a dedicated fine-tuning dataset based on the unique terminology, sentence structure, and logical features of patent texts. For example, for frequently occurring phrases such as "according to the claims" and "characterized by," as well as complex technical principle exposition sentences in the specification, full-scale fine-tuning guides the model to comprehensively learn and accurately generate translations that conform to patent specifications. This approach allows the model to retain its general language understanding capabilities while deeply adapting to the translation needs of the patent domain, significantly improving the accuracy and consistency of specialized terminology translation.

[0079] Building upon the comprehensive parameter fine-tuning, to further enhance expertise in specific sub-domains, a low-rank adaptation technique (LoRA) is employed to fine-tune the model using patent data specific to sub-domains such as biomedicine. LoRA adds a low-rank matrix to the model's weight matrix after comprehensive fine-tuning, concentrating training parameters in a small number of newly added low-rank matrices. This avoids a secondary comprehensive adjustment of the already optimized, large number of parameters, significantly reducing additional computational costs and training time, while preserving the generalized patent knowledge gained from comprehensive fine-tuning. In the context of biomedical patent translation, the R&D team focuses on unique professional elements in this field, such as molecular structure descriptions, clinical trial terminology, drug mechanisms of action, and biological sequence annotation, constructing targeted, refined datasets. For example, for proprietary expressions commonly found in biomedical patents such as "inhibitor," "receptor binding," and "gene expression regulation," as well as complex sentences in experimental methods and safety assessments, LoRA fine-tuning guides the model to learn and generate highly accurate, industry-standard translations. This two-stage optimization strategy—first laying the foundation for patents with full fine-tuning, and then finely adapting to sub-domains using LoRA—ensures that the model achieves more refined domain adaptation under the premise of controllable computing resources. This improves the professional accuracy, consistency, and contextual coherence of biomedical patent translation, while also being feasible for practical deployment. For example, completing the LoRA stage training on a standard cloud GPU cluster only takes a few hours to a few days.

[0080] To address the performance challenges of large-scale model training and inference, an integrated DeepSeed ZeRO3 acceleration solution and Flash Attention optimization technology are introduced. The integrated DeepSeed ZeRO3 acceleration solution starts with computation graph optimization, innovatively employing a hybrid parallel computing mode combining tensor parallelism, pipelined parallelism, and data parallelism. During the training phase, DeepSeed ZeRO3 uses an intelligent task scheduling algorithm to efficiently allocate computational tasks to multiple computing nodes, minimizing data transfer overhead between nodes and significantly improving computational resource utilization. During the inference phase, DeepSeed ZeRO3 merges multiple computational operators into one through operator fusion technology, reducing intermediate data read / write operations during computation; combined with memory optimization strategies, it dynamically manages GPU memory resources, reducing translation response time by more than 50%, meeting the stringent low-latency requirements of patent retrieval and real-time translation scenarios. Furthermore, DeepSeed ZeRO3 supports automatic mixed-precision computation, reducing GPU memory usage while ensuring translation quality, enabling the model to run efficiently with limited hardware resources. Flash Attention optimization technology optimizes the attention mechanism by reorganizing the computation process and memory access patterns, addressing the issues of high memory consumption and low computational efficiency in traditional attention computation when processing long sequences. In patent translation, patent specifications, claims, and other texts often have long sequences. Flash Attention optimization technology significantly reduces the time and space complexity of attention computation through strategies such as block computation and memory reuse, enabling the model to quickly process extremely long patent texts while reducing GPU memory consumption. This technology complements DeepSeed ZeRO3, jointly ensuring model performance during both training and inference phases.

[0081] Adaptive gradient optimization algorithms such as AdaGrad, RMSProp, and AdamW, along with fine-tuned hyperparameter configuration, are employed to improve training efficiency through distributed training on multiple machines and GPUs.

[0082] Among them, adaptive gradient optimization algorithms such as AdaGrad, RMSProp, and AdamW can dynamically adjust the learning rate based on historical gradient information of different parameters, avoiding the model from getting trapped in local optima and accelerating the convergence process. Addressing the long-tail distribution characteristics of patent translation data—specifically, the extremely small number of patents and low-frequency terminology in certain technical fields—the research team improved the gradient optimization algorithm by introducing a domain-adaptive weight adjustment mechanism. This mechanism analyzes the distribution of patent data across different technical fields and dynamically adjusts the update weights of parameters in each field, enhancing the model's ability to learn low-frequency patent terms and special sentence structures, ensuring stable and high-quality translation performance across various patent texts.

[0083] Precise setting of hyperparameters is a key factor determining model performance. Before training the large-scale translation model, the R&D team systematically explored the hyperparameter space using methods such as Bayesian optimization and random search. They focused on optimizing core parameters such as learning rate, batch size, number of training epochs, and number of attention heads. For example, through multiple rounds of experiments, they determined that the learning rate should be dynamically adjusted using a cosine annealing strategy, with an initial value of 4e-5, gradually reduced to 1e-9 in the later stages of training. This adjustment method effectively balances the model's convergence speed and generalization ability. The batch size was dynamically adapted based on the hardware's memory capacity, and reasonably set to 512 in a multi-machine, multi-GPU environment to maximize the utilization of computing resources. Furthermore, considering the generally long sequence characteristics of patent texts, the R&D team specifically adjusted the positional encoding length of the Transformer model to ensure that the model could accurately handle the semantic information and logical relationships in extremely long patent specifications.

[0084] The training of large language models relies on large-scale distributed computing clusters and adopts mature distributed training frameworks such as DeepSpeed ​​or Megatron-LM to achieve multi-machine and multi-GPU collaborative computing.

[0085] The DeepSpeed ​​framework utilizes a series of optimization techniques, such as the ZeRO optimizer stages (ZeRO-Offload, ZeRO-Infinity), to manage model parameters, gradients, and optimizer states in shards, distributing them across multiple computing nodes to effectively address the issue of insufficient GPU memory for extremely large models. During training, it supports a hybrid parallelism strategy encompassing data parallelism, pipeline parallelism, and model parallelism. Through dynamic task scheduling and communication optimization, it reduces data transfer overhead between nodes and improves the utilization of computing resources. For example, when processing complex long text training data in the patent field, DeepSpeed ​​can dynamically adjust its parallelism strategy based on the data scale and hardware resources, resulting in a several-fold increase in training efficiency.

[0086] The Megatron-LM framework focuses on model parallelism, particularly for large models with the Transformer architecture. It distributes different layers of the model across different computing nodes, with each node responsible for processing a portion of the model. This approach overcomes the limitations of single GPU memory capacity, enabling the training of ultra-large-scale language models. Furthermore, Megatron-LM employs efficient communication mechanisms and memory management strategies to reduce latency in inter-layer data transfer and improves overall training speed by optimizing core operations such as attention calculation. In the training of the Miaosuan Translation large-scale model, Megatron-LM can fully utilize multi-node computing resources to accelerate the model's learning process of patent text features.

[0087] During training, a data parallelism strategy divides the training data into multiple subsets and distributes them across different GPUs for parallel computation, using the AllReduce operation to synchronize gradients. A model parallelism strategy distributes different layers of the model across different computing nodes, effectively addressing the issue of insufficient GPU memory during the training of ultra-large models. To ensure the stability of the training process, a resilient training mechanism is introduced. When a computing node fails, the task automatically migrates to other nodes to continue execution, ensuring training continuity. Through multi-machine, multi-GPU distributed training, model training efficiency is improved by more than 3 times, significantly shortening the time required from data preparation to model convergence.

[0088] Throughout the training process, the system monitors key metrics such as loss function, accuracy, and perplexity in real time. Visualization tools like TensorBoard, Weights & Biases, and Wandb provide a clear view of the model's training dynamics, enabling developers to promptly identify issues such as overfitting and underfitting. To address potential issues with excessive GPU memory usage during training, gradient checkpointing reduces the storage of intermediate activation values, and model quantization techniques such as 4-bit quantization lower parameter storage overhead. Simultaneously, a dynamic memory management mechanism automatically releases unused computing resources, increasing single-GPU memory utilization to over 90%, ensuring stable and efficient model training even with limited hardware.

[0089] After model training, the translation quality is verified by combining automatic evaluation metrics with specialized metrics, along with human evaluation. The automatic evaluation uses classic machine translation metrics such as BLEU, ROUGE, and METEOR to compare the similarity between the model's translation and the human-referenced translation, quantifying the translation quality at the lexical and sentence levels. Specifically, the BLEU value calculates the n-gram matching degree (n=1-4) of aligned sentences, with a threshold set (BLEU-4 ≥ 0.4 is considered valid alignment). The TER value measures the Translation Error Rate, with a threshold ≤ 0.3 indicating acceptable alignment quality. Additionally, specific metrics such as patent terminology accuracy and claim logical consistency are introduced, focusing on the model's ability to handle the specialized content of patent texts. The human evaluation is conducted by a review team composed of experienced patent translation experts, who score and comprehensively evaluate the translation from multiple dimensions, including professionalism, grammatical accuracy, and logical coherence, ensuring that the model's output meets the strict standards and practical application requirements of the patent field. Specifically, 5% of the aligned data was manually reviewed, focusing on verifying: consistency of legal terminology, such as whether "claims" corresponds to "claims" or "scope of the claim"; accuracy of technical parameters, such as whether chemical formulas and mechanical unit dimensions correspond; and completeness of logical relationships, such as whether causal relationships and conditional statements are consistent in the alignment. For sentences with automatic alignment errors, a dynamic programming algorithm was used to re-search for the optimal alignment path. Based on sentence vector similarity, cosine similarity was used for calculation, with a threshold of ≥0.7, and the alignment model parameters were updated. Through multiple rounds of evaluation and iterative optimization, the model parameters were continuously adjusted, ultimately achieving a patent translation accuracy rate of over 95%.

[0090] During the model deployment phase, to meet the real-time and high-concurrency processing requirements of patent retrieval and translation scenarios, the vLLM framework, based on PagedAttention technology, further optimizes memory management for attention computation, achieving efficient inference acceleration. Two optimization strategies can be selected during deployment: one is to directly utilize the efficient inference capabilities of the vLLM framework to deploy the original model; the other is to first quantize the model, converting model parameters from high-precision data types to low-precision types, such as 8-bit or 4-bit, significantly reducing the computational load and memory consumption during model inference without sacrificing translation quality, and then deploying the model based on the vLLM framework. This supports model quantization and adopts a hybrid deployment mode of edge computing and cloud services, while also providing a user-specific terminology upload interface, which is prioritized during translation. Through operator fusion, model pruning, and quantized inference techniques, the model inference speed is improved by 3-5 times, with a single node capable of processing over 200 translation requests per second. Furthermore, the system employs a hybrid deployment mode of edge computing and cloud services, dynamically allocating computing resources based on user geographic location and request load to ensure that users worldwide can obtain low-latency, highly stable patent translation services. At the same time, a comprehensive model monitoring and automatic update mechanism is established to monitor translation quality and performance indicators in real time. Once a performance decline is detected or new patent data is accumulated, the model retraining and update process is automatically triggered to ensure that the model always maintains a leading translation level and technical competitiveness.

[0091] like Figure 6 The diagram shown is a fine-tuning architecture diagram provided in this embodiment.

[0092] The method provided in this embodiment uses the multilingual parallel corpus to perform full-parameter fine-tuning of an open-source large language model, learning the general language style, structure, and terminology of patent texts. This full-parameter fine-tuning allows the model to deeply learn the general style, structure, and terminology of patents, laying the foundation for patent translation and avoiding style adaptation bias and terminology mistranslation in general models. Based on this full-parameter fine-tuning, a low-rank adaptation technique (LoRA) is used to fine-tune the model with patent data from specific technical subdomains. LoRA fine-tuning focuses on patent data from specific subdomains, eliminating the need for full-parameter tuning, reducing computational costs, and improving the professionalism and contextual coherence of subdomain translations. DeepSeed is also integrated. The ZeRO3 acceleration solution, combined with FlashAttention optimization technology, employs adaptive gradient optimization algorithms and refined hyperparameter configuration. It enhances training efficiency through multi-machine, multi-GPU distributed training, integrating acceleration technologies and optimization algorithms to improve efficiency by over 300% through distributed training. It is compatible with patented long text processing and ensures training stability under limited hardware resources. Combining automatic and specialized evaluation metrics with manual evaluation to verify model translation quality, the solution comprehensively verifies translation quality, ensuring model output meets patent domain standards and practical application requirements. Based on the vLLM framework, it supports model quantization and adopts a hybrid deployment mode of edge computing and cloud services. It also provides a user-specific terminology upload interface, prioritizing the use of this terminology during translation. Deployed based on the vLLM framework, it supports quantization and hybrid deployment, improving inference speed by 3-5 times. Combined with the user-specific terminology upload function, it adapts to efficient translation needs across multiple scenarios.

[0093] The following specific implementation examples further illustrate this point: Case 1 - English to Chinese Translation Translate "The method of claim 6, further comprising causing statistics to be displayed along with the content, the statistics including at least one of offer acceptance or offer availability statistics for the at least one advertisement based on the feedback." like Figure 7 The image shown is a schematic diagram illustrating the English-to-Chinese translation results of the large-scale patent translation model provided in this embodiment. Figure 8 The image shown is a schematic diagram illustrating the English-to-Chinese translation result of the existing translation service provided in this embodiment.

[0094] Case 2 - Chinese to English Translation The translation describes an external logic device for a network interface controller that enables interrupt coupling. The network interface controller has a cause register for storing information about interrupt causes and drives an interrupt line. The external logic device is connectable to the cause register for reading its contents and is also connectable to the interrupt line of the network interface controller and an interrupt input of a processor for forwarding an interrupt from the interrupt line of the network interface controller to the processor. It also includes a timer that, when an interrupt is contained on the interrupt line, is initialized and configured to delay interrupt forwarding based on the current contents of the cause register until the timer times out. like Figure 9 As shown, this is a schematic diagram of the Chinese-to-English translation results of the patent translation big data model provided in this embodiment. Figure 10 The image shown is a schematic diagram of the Chinese-to-English translation result provided by the existing translation service in this embodiment.

[0095] Therefore, in English to Chinese translation: (1) The fine-tuned model expression is more in line with the patent expression style and has some patent terms (which did not appear in the original English), such as "according to the claim", "the", "characterized by", "a kind of", etc.

[0096] (2) The fine-tuned model is more in line with the patent expression habits. For example, when the fine-tuned model explains a noun, it will first list the noun and then explain it, while the existing translation service translates it in the form of modifier + noun.

[0097] (3) The fine-tuned model can better describe the claims description and can clearly and systematically express the methods protected by the claims.

[0098] Chinese to English: (1) The fine-tuned model will merge some modifiers to make the statements concise and clear.

[0099] (2) The fine-tuned model translates professional terms better.

[0100] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operations of each module processing step can be understood with reference to the description of the method embodiments, and will not be repeated here.

[0101] like Figure 11As shown, the present invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the method for constructing a multilingual parallel corpus for patent translation / training a large-scale patent translation model as described in the above embodiments, or when the computer program is executed by a processor, it implements the method for constructing a multilingual parallel corpus for patent translation / training a large-scale patent translation model as described in the above embodiments.

[0102] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0103] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are known to those skilled in the art.

Claims

1. A method for constructing a multilingual parallel corpus for patent translation, characterized in that, Includes the following steps: Based on international patent application information, obtain the patent family document identifiers of the same technical solution in different countries or regions, and standardize the document identifiers to form a structured patent family data set. Obtain the full-text multilingual patents corresponding to the patent family data set, and sequentially perform noise removal, structural parsing and language normalization on the full-text multilingual patents to obtain multilingual structured patent family text pairs with unified chapters, paragraphs and sentences. Based on the aforementioned multilingual structured family of patent text pairs, and combining the inherent structural features and semantic information of the patent texts, multi-level cross-parallel corpus alignment from documents, chapters to sentences is performed to generate preliminary parallel corpus pairs. The multi-level cross-parallel corpus alignment process comprehensively utilizes word alignment methods based on statistical models and rule matching methods based on patent structural features. An automated quality assessment of the preliminary parallel corpus pairs is performed using a quality scoring model that includes multiple preset evaluation dimensions. The quality scoring model includes a structural consistency score for evaluating the structure, a term density comparison score for evaluating the coverage of specialized terms, and a two-way translation consistency score for evaluating translation fidelity. Based on the evaluation results, parallel corpus pairs that meet the preset quality threshold are selected to form the final multilingual parallel corpus.

2. The method for constructing a multilingual parallel corpus for patent translation as described in claim 1, characterized in that, The rule matching method based on patent structural features specifically includes the following steps: Using the claim numbers, formula numbers, and chapter titles in the patent text as alignment anchors, the corresponding paragraphs and sentences in the same patent text are initially matched and located using these alignment anchors. This clarifies the corresponding relationships between the texts, reduces initial alignment deviations, and is then calibrated in conjunction with sentence length ratios.

3. The method for constructing a multilingual parallel corpus for patent translation as described in claim 1, characterized in that, After performing multi-level cross-parallel corpus alignment of sentences, semantic level alignment is also included, which specifically includes the following steps: The semantic features of sentences are extracted using a pre-trained multilingual model, and adaptive feature fusion is performed by combining surface features. The semantic similarity between sentence pairs is calculated based on the fused features, thereby verifying and optimizing the sentence-level alignment results.

4. The method for constructing a multilingual parallel corpus for patent translation as described in claim 1, characterized in that, The quality scoring model specifically includes the following seven evaluation dimensions: Structural consistency score, term density comparison score, length ratio score, bidirectional translation consistency score, semantic similarity score, numerical and symbol consistency score, and citation consistency score; The final quality score is calculated by weighted average, and the weights of each dimension can be dynamically adjusted according to different parallel corpus pairs or technical fields.

5. The method for constructing a multilingual parallel corpus for patent translation as described in claim 1, characterized in that, It also includes corpus enhancement, specifically including the following steps: The selected parallel corpus pairs are back-translated to generate semantically consistent variant corpora, and / or a two-way cross-validation algorithm is used to further improve the alignment confidence through multi-stage similarity calculation.

6. A patented translation system applying the multilingual parallel corpus construction method as described in claim 1, characterized in that, include: The corpus construction module is used to construct and maintain a multilingual parallel corpus based on the aforementioned multilingual parallel corpus construction method. The model training module is used to train a large-scale patent translation model using the aforementioned multilingual parallel corpus. The translation service module integrates the aforementioned patent translation model, which is used to receive patent text, images, or documents input by users, call the patent translation model to perform translation, and supports users to upload custom terminology lists to prioritize the use of custom terms in translation.

7. The patent translation system as described in claim 6, characterized in that, The training method for the patent translation large model includes the following steps: Using the aforementioned multilingual parallel corpus, the open-source large language model was fine-tuned with all parameters to learn the general language style, structure, and terminology of patent texts. Based on the full parameter fine-tuning, a low-rank adaptation technique is used to fine-tune the model using patent data from specific technical subfields; It integrates the DeepSeed ZeRO3 acceleration solution and Flash Attention tuning technology, adopts adaptive gradient optimization algorithm and fine-tuned hyperparameter configuration, and improves training efficiency through multi-machine and multi-card distributed training; The translation quality of the model is verified by combining automatic evaluation indicators with specific indicators and supplementing them with manual evaluation. Inference deployment is based on the vLLM framework, supports model quantization processing, adopts a hybrid deployment mode of edge computing and cloud services, and provides a user-specific terminology upload interface, which is used first during translation.

8. The patent translation system as described in claim 7, characterized in that, The open-source large model is the Qwen3 base model, and the full parameters include the characteristic expressions of the patent claims and the complex technical principle explanations in the specification.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • IC ESD protection with distributed silicon controlled rectifier

    EP4567890A1

  • Heat supply device

    JP6789012B2

  • Optical module

    US20250123456A1

  • Time sequence data prediction based on three-dimensional fully-connected fusion, and model training

    WO2025123456A1