Man-machine combination translation batch processing translation method containing artificial intelligence

This human-computer collaboration platform, which combines pre-translation analysis with a neural network translation engine using artificial intelligence technology, solves the problems of low efficiency and translation consistency in existing translation systems when processing large batches of documents, and achieves efficient and accurate translation results and rapid version management.

CN120952020APending Publication Date: 2025-11-14HARBIN ENG UNIV

Patent Information

Application Number
CN202510904397.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing translation systems are inefficient when processing large volumes of documents, lack context learning and language style adaptation capabilities, rely on human experience for post-translation quality control, and lack uniformity in terminology processing, making it difficult to guarantee the consistency and quality of translations.

Method used

Artificial intelligence technology is used for pre-translation semantic analysis to extract terminology and language style parameters. Combined with a neural network translation engine, the initial translation results are generated. Post-editing is performed through intelligent quality assessment and a human-computer collaboration platform to achieve full-text consistency verification and terminology consistency detection, and to support the knowledge accumulation and model optimization of the translation version library.

Benefits of technology

It significantly improves the efficiency and accuracy of translation task allocation, reduces manual coordination costs, ensures the semantic accuracy and terminological consistency of the translation, enhances translation speed and quality, and supports rapid reuse and version management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952020A_ABST
    Figure CN120952020A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine combination translation batch processing translation method containing artificial intelligence, and relates to the technical field of intelligent translation. In order to solve the defect that automatic batch scheduling, man-machine collaborative optimization and intelligent quality control cannot be realized in the prior art, the technical scheme provided by the invention comprises the following steps of: analyzing a manuscript to be translated to form a translation task package; term suggestions and text features are generated; performing machine translation on the to-be-translated manuscript original text to generate an initial translation result; performing intelligent quality evaluation on the initial translation result, and marking the initial translation result as an available translation, a light review task or a deep editing task; pushing the post-editing task to a man-machine cooperation platform, and performing manual modification and confirmation by a translator in combination with term suggestions and context prompts; and executing full-text consistency verification, detecting and prompting term inconsistency and style mutation problems, and outputting a final finalization translation. The method is suitable for intelligent translation service work with high requirements for translation quality, consistency and processing efficiency in multi-language mass translation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This relates to the field of intelligent translation technology, specifically a human-machine collaborative translation method incorporating artificial intelligence for batch translation processing. Background Technology

[0002] Human-machine collaborative batch translation is a translation model that integrates artificial intelligence technology with human expertise. Its core lies in combining the automation of the translation process with the precision of human intervention through technological means. Especially in the context of globalization, international cooperation in the maritime industry is becoming increasingly frequent, and effective communication has become the key to successful cooperation. By designing batch translation methods, users can overcome language barriers, ensure the accurate transmission of professional knowledge, and thus improve the efficiency and success rate of international cooperation projects.

[0003] A search revealed a Chinese patent publication number CN110837742A, which discloses a human-machine collaborative batch translation method incorporating artificial intelligence. This method utilizes techniques such as file parsing and sentence segmentation, original text translation / standardized original text writing, translation of repeated sentences, and simultaneous editing of multiple documents by multiple users to achieve batch processing of similar translated texts, thereby improving translation efficiency. Furthermore, by combining translation memory corpora, machine translation technology, and human translation, it significantly increases translation speed while ensuring quality. Finally, quality inspection technology is used to conduct in-depth analysis of machine and human translations, detecting problems in both, thus reducing workload while ensuring translation quality, lowering costs, and increasing efficiency. Besides the aforementioned technical documents, common translation workflows currently include: machine translation using online translation platforms (such as Google Translate and DeepL), followed by post-editing by human translators to correct terminology, sentence structure, and semantic errors. This human-machine collaborative translation model has become the mainstream method in the industry. For example, translation software platforms such as Trados and MemoQ widely utilize terminology memory, automatic recognition technology, and post-editing interfaces to assist translators in handling repetitive content and improving consistency. However, the following typical problems still exist in large-scale document processing, multilingual batch translation, terminology standardization, and post-translation quality control: Excessive human intervention and significant efficiency bottlenecks: Although current translation systems have certain terminology matching and automatic translation capabilities, batch processing still requires a large amount of manual distribution, scheduling, and post-editing, resulting in low efficiency.

[0004] Lack of context learning and language style adaptation capabilities: Many NMT systems are unable to fine-tune according to different clients and different writing styles, resulting in a lack of consistency in translations.

[0005] Post-translation quality control relies on human experience: Currently, mainstream platforms lack intelligent quality assessment mechanisms, and post-translation review still relies on experienced human quality inspectors, which involves a certain degree of subjectivity and uncertainty.

[0006] Terminology processing is isolated and lacks a unified control mechanism: Although terminology databases can be used for terminology unification, they cannot dynamically participate in the contextual decisions of the entire manuscript, which easily leads to inconsistencies in terminology.

[0007] In summary, existing technologies lack a batch translation processing method that combines artificial intelligence judgment capabilities with manual operation, and thus cannot achieve automatic batch scheduling, human-machine collaborative optimization, and intelligent quality control while ensuring semantic accuracy and terminological consistency of the translation. Summary of the Invention

[0008] To address the shortcomings of existing technologies in achieving automated batch scheduling, human-machine collaborative optimization, and intelligent quality control while ensuring semantic accuracy and terminological consistency in translations, the technical solution provided by this invention is as follows: A human-machine collaborative translation batch processing method incorporating artificial intelligence includes: The steps involved are: receiving and parsing the manuscript to be translated, extracting information on language, field, and urgency, and forming a translation task package; The steps include performing pre-translation semantic analysis on the translation task package, extracting terms, keywords and language style parameters, and generating term suggestions and text features; The steps involve calling a neural network translation engine to perform machine translation of the original text to be translated, and then fusing the results based on confidence and quality scoring models to generate an initial translation. Perform intelligent quality assessment on the initial translation results, and classify the tasks according to the scoring results, marking the steps as usable translations, light proofreading tasks, or deep editing tasks; After the task is pushed to the human-computer collaboration platform, the translator will manually modify and confirm it based on terminology suggestions and contextual prompts. The steps involve performing a full-text consistency check, detecting and alerting to inconsistencies in terminology and abrupt style changes, and outputting the final translated text.

[0009] Furthermore, a preferred implementation method is provided, which also includes the steps of storing the final translation in the translation version library and updating terminology resources to achieve knowledge accumulation and model optimization.

[0010] Furthermore, a preferred embodiment is provided in which the translation task package is formed by clustering according to the professional field to which the manuscript belongs and setting a priority queue according to the urgency of the task.

[0011] Furthermore, a preferred implementation is provided in which the pre-translation semantic analysis employs a natural language processing algorithm based on a pre-trained language model to identify proper nouns, terminology phrases, and sentimental statements, and annotates them in the original text.

[0012] Furthermore, a preferred embodiment is provided, wherein the neural network translation engine includes at least one self-developed translation engine and one external commercial engine, and the fusion method is a dynamic optimization strategy based on sentence pair quality scoring results.

[0013] Furthermore, a preferred implementation method is provided, wherein the intelligent quality assessment includes four scoring dimensions: semantic fidelity, grammatical integrity, terminology consistency, and sentence naturalness, and a multi-model integration is used to determine the translation quality level.

[0014] Based on the same inventive concept, the present invention also provides a human-machine combined translation batch processing device incorporating artificial intelligence, comprising: The module receives and parses the manuscript to be translated, extracts information on language, field, and urgency, and forms a translation task package; A module is used to perform pre-translation semantic analysis on the translation task package, extract terminology, keywords and language style parameters, and generate terminology suggestions and text features; This module calls a neural network translation engine to perform machine translation of the original text of the manuscript to be translated, and generates an initial translation result by fusing the confidence and quality scoring models. The initial translation results are subjected to intelligent quality assessment, and tasks are assigned based on the scoring results, marking modules as usable translations, light proofreading tasks, or deep editing tasks; After the task is pushed to the human-machine collaboration platform, the translator will manually modify and confirm it based on terminology suggestions and contextual prompts. This module performs full-text consistency checks, detects and alerts users to inconsistencies in terminology and abrupt style changes, and outputs the final translation.

[0015] Based on the same inventive concept, the present invention also provides a computer storage medium for storing a computer program, wherein when the computer program is read by a computer, the computer executes the method described thereon.

[0016] Based on the same inventive concept, the present invention also provides a computer, including a processor and a storage medium, wherein when the processor reads a computer program stored in the storage medium, the computer executes the method described thereon.

[0017] Based on the same inventive concept, the present invention also provides a computer program product, which, when executed, implements the method described.

[0018] Compared with the prior art, the advantages of the technical solution provided by the present invention are as follows: This solution introduces a batch scheduling module, which can intelligently classify and group pre-translation tasks based on task urgency, manuscript professional attributes, and target language, achieving structured breakdown and priority ranking of translation tasks. Unlike traditional scheduling based on file order or manual scheduling, this method significantly improves the efficiency and rationality of task allocation, reduces manual coordination costs, and effectively solves the problems of chaotic scheduling and uneven division of labor in large-scale tasks.

[0019] This solution introduces an AI-powered terminology recognition and translation suggestion module in the pre-translation processing stage. It automatically performs semantic analysis on the source text, marking proper nouns, terminological phrases, and sentence structures, and provides translation suggestions based on an existing terminology database. Compared to the traditional approach of relying on static terminology databases, this method improves the real-time performance and accuracy of terminology matching, and provides consistency assurance in subsequent translation, avoiding inconsistencies in terminology.

[0020] This solution introduces an AI quality assessment module, which intelligently scores machine translation results from multiple dimensions, including semantics, grammar, terminology, and sentence structure. Based on the scoring results, it automatically determines whether to submit the translation to a human translator for review or proceed to the next round of AI optimization. This strategy breaks away from the inefficient traditional process of "all translations requiring human review," enabling refined scheduling of human review resources and significantly improving processing speed while ensuring translation quality.

[0021] The human-machine collaboration platform designed in this solution supports seamless integration of AI intelligent suggestions and human post-editing. It captures the translator's editing behavior in real time during the revision process and continuously optimizes the subsequent AI translation model. Unlike the traditional "AI translation + human editing" separation model, this platform enables real-time feedback of translator experience to the AI ​​model, improving the quality of machine translation and giving the system self-iterative capabilities.

[0022] Furthermore, this solution supports a full-text-level terminology and style consistency verification module, which can perform contextual consistency analysis on the entire translation, identify inconsistencies in terminology, sentence structure, and style, and provide suggested modifications. This approach offers a more holistic perspective than traditional sentence-level evaluation-based review processes, effectively avoiding potential quality issues such as style inconsistencies and inconsistencies in expression within long texts.

[0023] Finally, this solution establishes a traceable translation version management and reuse mechanism. By recording each translation generation and modification process, it supports the rapid reuse of similar manuscripts and version backtracking. It is particularly effective for industry scenarios with stable terminology and repetitive templates (such as government documents, technical documents, etc.), and significantly improves the processing efficiency of subsequent translation tasks.

[0024] This intelligent translation service is suitable for multilingual, high-volume translation tasks where high quality, consistency, and processing efficiency are required. Attached Figure Description

[0025] Figure 1 A flowchart of a human-machine combined batch translation method. Detailed Implementation

[0026] To make the advantages and benefits of the technical solution provided by the present invention clearer, the technical solution provided by the present invention will now be described in further detail with reference to the accompanying drawings, specifically: Implementation Method 1: This implementation method provides a human-machine collaborative translation batch processing method incorporating artificial intelligence, including: The steps involved are: receiving and parsing the manuscript to be translated, extracting information on language, field, and urgency, and forming a translation task package; The steps include performing pre-translation semantic analysis on the translation task package, extracting terms, keywords and language style parameters, and generating term suggestions and text features; The steps involve calling a neural network translation engine to perform machine translation of the original text to be translated, and then fusing the results based on confidence and quality scoring models to generate an initial translation. Perform intelligent quality assessment on the initial translation results, and classify the tasks according to the scoring results, marking the steps as usable translations, light proofreading tasks, or deep editing tasks; After the task is pushed to the human-computer collaboration platform, the translator will manually modify and confirm it based on terminology suggestions and contextual prompts. The steps involve performing a full-text consistency check, detecting and alerting to inconsistencies in terminology and abrupt style changes, and outputting the final translated text.

[0027] It also includes steps such as storing the final translation in the translation version library and updating terminology resources to achieve knowledge accumulation and model optimization.

[0028] The process of forming the translation task package includes clustering according to the professional field of the manuscript and setting priority queues according to the urgency of the task.

[0029] The pre-translation semantic analysis employs a natural language processing algorithm based on a pre-trained language model to identify proper nouns, technical phrases, and sentimental statements, and annotates them in the original text.

[0030] The neural network translation engine includes at least one self-developed translation engine and one external commercial engine, and the fusion method is a dynamic optimization strategy based on sentence pair quality scoring results.

[0031] The intelligent quality assessment includes four scoring dimensions: semantic fidelity, grammatical integrity, terminology consistency, and sentence naturalness. It uses multi-model integration to determine the translation quality level.

[0032] The human-machine collaboration platform has real-time terminology prompts, bilingual comparisons, and editing behavior recording functions, and synchronously feeds back the translator's modification behavior to the AI ​​model for parameter fine-tuning.

[0033] The full-text consistency check is based on the construction of a terminology and style distribution map of the entire text using a contextual language model. Modification suggestions are generated for inconsistent areas and the translator makes the final confirmation and revision.

[0034] The term extraction process includes constructing term context word vectors and filtering redundant terms using semantic similarity thresholds, retaining only key terms that are semantically stable and appear multiple times.

[0035] When there are multiple candidate translations in the initial AI translation results, the system selects the candidate translation that is consistent with the target style as the final translation output according to the language style parameters obtained from the pre-translation analysis.

[0036] The translator post-editing platform has a terminology consistency prompt module, which automatically pops up a confirmation prompt when it detects that the translator has modified the terminology to be inconsistent with the recommended terminology.

[0037] The full-text consistency verification process employs a contextual semantic retrieval algorithm based on bidirectional language modeling to construct a style curve for the entire translation and detect style abrupt change points.

[0038] The version management of the translated text after delivery includes version number, editor identity identifier and modification record tracking, and supports full-text search and reuse by task number.

[0039] The grouping method of the translation task package further includes load balancing based on the length of the original text, the complexity of the mixed text and image structure, and the historical translation difficulty score, so as to improve resource utilization efficiency.

[0040] The AI ​​model fusion strategy includes a sentence-level confidence assessment module and a paragraph-level semantic consistency scoring module, which, combined with a voting mechanism, selects the fused translation.

[0041] The records of manual editing behavior include the location of each terminology modification, content changes, editing time, and operation frequency, which are used to construct a translator editing profile and behavior preference model.

[0042] The terminology consistency verification supports real-time integration with external industry standard terminology databases, automatically comparing whether the use of terms in the translation conforms to industry standards and providing replacement suggestions.

[0043] The system supports the automatic generation of translation quality assessment reports after the translation is finalized, including translation score distribution, manual revision rate, terminology consistency index, and style deviation prompts.

[0044] Implementation Method Two: This implementation method is a further detailed description of the technical solution provided in Implementation Method One, specifically: Step 1: Task Reception and Intelligent Scheduling The system is deployed on a cloud platform or a local translation server, supporting users to upload documents for translation via a web interface or submit documents in batches via API. Document formats include, but are not limited to, Word, PDF, TXT, and Excel. The system parses document metadata, identifying key parameters such as language pairs (e.g., Chinese to English, English to Spanish), subject area (e.g., legal, medical, technical specifications), and delivery deadline. The scheduling module constructs scheduling priorities based on preset rules, such as prioritizing delivery deadlines and using subject area complexity as a secondary weight, forming multiple task packages. Each task package contains several translation units (the smallest divisible text segment).

[0045] Step Two: Pre-translation Analysis and Terminology Extraction After the task package is generated, the system performs semantic analysis, including sentence segmentation, part-of-speech tagging, and named entity recognition. Combining existing terminology databases, translation memories, and industry knowledge graphs, it identifies high-value terms such as technical terms, proper nouns, and technical phrases in the original text. For example, for technical documents related to "photovoltaic inverter control strategy," the system identifies and tags terms such as "inverter" and "maximum power point tracking" as key items. The terminology extraction module generates a terminology list and provides recommended translations for subsequent machine translation and human review processes. It also analyzes sentence style (e.g., technical, colloquial, or policy-related language) and generates language style parameters.

[0046] Step 3: AI Initial Translation and Multi-Engine Integration: The translation engine invocation module automatically matches a suitable combination of machine translation engines based on the task type and language pair. For example, for highly specialized technical documents, the system prioritizes using its self-developed professional engine, while for general documents, it can simultaneously use external engines such as DeepL and Google Translate. After each sentence segment is translated by multiple engines, the system calculates the confidence level (such as BLEU score and semantic consistency score) based on the set quality scoring model, and finally selects the highest quality translation as the initial translation result, while retaining the second-best candidate translation for later use.

[0047] Step 4: Intelligent Translation Quality Assessment and Task Assignment: The system performs automatic quality assessment sentence by sentence based on the initial translation results. Evaluation metrics include: semantic fidelity (whether the source and translation are consistent), terminology matching rate (whether recommended terms are used), grammatical structure completeness, and language naturalness. The scoring model can combine semantic models such as BERT with proprietary translation quality prediction models for comprehensive scoring. The system categorizes the initial translation results into three levels: high quality (ready for direct delivery), medium quality (requires light manual review), and low quality (requires in-depth manual editing), and constructs a corresponding list of manual editing tasks.

[0048] Step 5: Post-human-computer collaborative editing and processing: Translators access the system through a human-machine collaboration platform. The left side of the interface displays a comparison between the original text and the AI ​​translation, while the right side shows terminology suggestions, contextual references, and suggestions for modification. The platform supports word association and replacement, quick terminology replacement, one-click switching of candidate translations, and records all translator editing actions. When a translator replaces terminology or adjusts sentence structure in a text, this action is recorded and input as annotation data into the AI ​​fine-tuning module for subsequent model parameter optimization. The platform also supports evaluation and feedback of the entire content and allows translators to annotate semantic ambiguity, unclear context, and other issues.

[0049] Step Six: Full Text Consistency Verification and Finalization: After editing, the system automatically performs a full-text consistency analysis, checking whether terminology is used consistently throughout the text, whether the style is consistent, and whether there are any logical conflicts between paragraphs. For example, if the term "inverter" is used earlier in the text and "converter" is used later, the system will mark it as inconsistent and suggest modifications. Similarly, if some sentences use English-style expressions while other paragraphs use Chinese-style structures, it will also be marked as a style abrupt change. The system will output a consistency verification report, which the translator will then confirm before submitting the final translation.

[0050] Step Seven: Translation Delivery and Knowledge Accumulation Once the final translation is confirmed, an output file is generated in the format requested by the client (Word, PDF, Excel, JSON, etc.) and delivered via email or API. Simultaneously, the system writes terminology updates, translation modifications, and scoring results from this translation task into the translation corpus and terminology database. This data is used to continuously train the translation engine, optimize the terminology suggestion system, and improve scheduling rules, enabling the translation service platform to achieve long-term self-learning capabilities.

[0051] Implementation Method 3: Combination Figure 1 This embodiment describes the technical solution provided above in further detail through specific examples. Specifically: To enhance terminology processing and contextual understanding capabilities, and effectively improve human-computer interaction efficiency, this embodiment provides the following technical solution: a human-computer collaborative translation batch processing method incorporating artificial intelligence, comprising the following steps: S1. System Architecture Setup: The core translation engine is built based on the Electron framework and the ChatGLM3 API. The semantic matching of terms is achieved by deploying the Embedding model and a real-time term update channel is built. S2. Batch document preprocessing: Parse the uploaded document format and extract the text, segment and mark the text into blocks, identify domain terms based on the BiLSTM-CRF model and match and mark them with the terminology database; S3. Human-computer collaborative translation: Matches with the local terminology database and performs real-time online term queries, performs contextualized translation based on the contextual information, and injects term explanations; S4. Quality Enhancement and Intervention: AI-assisted proofreading is used to detect differences in translation results, and a human collaboration mechanism is set up to revise the translation. S5. Dynamic optimization of terminology database: Set up a user rating system to provide feedback on terminology issues, and dynamically update the terminology database according to the automatic optimization mechanism.

[0052] The translation system includes a platform application as the front end, an integrated NLP library, ChatGLM model, and GPTAPI as the back end, and a terminology database and user feedback database as local databases. The steps for parsing uploaded document formats and extracting text in the translation system include: 1) Input documents of different formats, extract text using pdfminer and python-docx, and identify text format based on file type; 2) Deep text parsing processing based on layout analysis and table detection algorithms, including: a. A layout analysis algorithm based on pdfminer parses the format parameters of PDF text, including character merging threshold, line merging threshold, word spacing threshold, and text flow direction parameters; b. An OpenCV-based table detection algorithm parses text format, and its formula is expressed as: in Indicates the score in the table. Indicates linear density. Indicates cell alignment. , All are weighted parameters; After text parsing, a structured text object is output while preserving the original format metadata. A BERT-based domain classifier is constructed, taking document character samples as input and outputting text domain labels. The corresponding terminology library is loaded. The classification formula of the domain classification model architecture is expressed as: in The [CLS] label is used to label the hidden layer output of the BERT model. This is the classification layer weight matrix. This is the bias vector.

[0053] The steps for identifying domain terms based on the BiLSTM-CRF model and matching tags with a terminology database include: 1) Setting up a semantic embedding computation and adaptive block segmentation algorithm in a dynamic block segmentation based on semantic coherence, inputting the original text stream data, and outputting a sequence of text blocks with semantic tags, specifically including: a. Load the pre-trained model, generate sentence embeddings and batch process their size, and return a NumPy array to normalize the data vectors; b. For sentence segmentation, the sentence is divided into semantic similarity clusters based on its length, and otherwise treated as independent blocks; 2) Construct a BiLSTM-CRF term recognition model for pre-identification of domain terms. The term recognition formula includes: a. Calculation of LSTM hidden states, expressed as: ; b. For CRF decoding, the decoding formula is expressed as: in For the transition fraction matrix, The transmit fraction is the output of the LSTM. 3) Perform word segmentation and POS annotation on the text blocks. Based on the identification of compound nouns, they are divided into noun phrase blocks and single term detection. Term tags are generated by detecting term boundaries and matching them with the term database. 4) When performing global contextual analysis of the document, calculate the self-attention weights. When performing local contextual analysis of the document, set a sliding window to assign weights to the context. The assignment formula is expressed as follows: ; 5) Construct a FAISS index based on the approximate nearest neighbor search method of FAISS and embed dimensions and preloaded terminology library, and realize term matching tagging according to the set similarity threshold.

[0054] The steps to build a human-machine collaborative translation engine include: 1) Construct a dynamic Prompt that includes terminology, context, and professional requirements, specifically: a. Format the terminology list to generate a contextual summary, add index numbers to improve readability, and define field truncation based on the empty terminology list processing mechanism to achieve dynamic truncation of long terminology lists; b. Automatically add key entity annotations based on the generated context summary, while preserving paragraph structure information and intelligently processing the context; c. Dynamic control parameters, including configurable terminology display quantity, priority terminology switch domain and target language parameterization, automatic length control mechanism, explicit priority rules, added format specifications, and set output format requirements; 2) Implement a dual-engine translation scheduling system, intelligently allocating translation tasks to the optimal engine. The text complexity is calculated locally based on a scheduling algorithm, with the following formula: in Term density is the ratio of the number of terms to the total number of words. is the syntactic depth, representing the maximum depth of the dependency tree; The proportion of passive voice; , , All are weighted coefficients and The value is 0.4. The value is 0.3. The value is 0.3; 3) Based on the dynamic Prompt input of the GPT-4 translation core algorithm, a preliminary translation is output. Bundle search is then used to optimize the translation, with the following formula: After obtaining the initial sequence, expand the candidates to generate new candidates, set the detection termination condition and length penalty, calculate the comprehensive score and merge and sort the candidates; 4) Based on the ChatGLM local enhancement processing requirements for sensitive data and low latency, a ChatGLM enhancement model architecture is constructed, represented as follows: Output=TransformerDecoder(InputEmbedding+PositionalEncoding) Introducing an attention mechanism into the model is represented as: in The dimension of the key vector is represented, and computation is accelerated using FlashAttention-2. 5) Determine the confidence level based on the initial AI translation. If it fails to pass, trigger human intervention. Correct the translation using a bilingual editor and feed it back to the learning system, thus establishing a human-computer collaborative interaction mechanism. The confidence level calculation formula is as follows: The explanation generation model generates terminology explanations in real time, and the model's loss function is expressed as follows: in Represents the L2 regularization coefficient. Indicates the maximum generated length.

[0055] Based on AI-assisted proofreading and multi-dimensional comparison of semantics and syntax, significant differences between GPT-4 and ChatGLM translation results are automatically identified, including: 1) Calculate semantic similarity: in , , This represents the vector representation of the CLS markers from the last layer of the BERT model; 2) Calculate term overlap: in This represents the set of terms extracted from translated text 1. This represents the set of terms extracted from translated text 2; 3) Calculate the similarity of the syntactic tree kernels: in For syntactic analysis trees, , For a set of tree nodes, =1 indicates a node and The phrases must be of the same type and have the same subtree structure; otherwise, the value is 0. ∈(0,1) is the attenuation coefficient; 4) Calculate the overall weight: in , , All are weighting coefficients, and the dynamic threshold is set to... ,according to Make a judgment.

[0056] An automatic risk labeling system is constructed to identify potential problems in translation, including terminology risks, sentence complexity, passive voice, and logical ambiguity. Different risk problem types are labeled and processed based on detection algorithms, including: 1) The confidence level of the terminology database is used to calculate and mark terminology risk issues. The calculation formula is as follows: in This represents the probability score of term t in translation. The domain matching degree of term t is represented by σ, which represents the Sigmoid function that maps linear combinations to the interval (0,1). The formula for the terminology risk penalty is as follows: in For the set of terms in the translated text, Term confidence threshold; 4) The complexity of long sentences is calculated using syntactic tree depth detection. The calculation formula is as follows: in Let be the syntactic tree depth of sentence s, and let represent the longest path from the root node to the leaf node; Indicates the number of nested clauses; 5) Dependency analysis is used to detect and mark passive voice issues. The calculation formula is as follows: in It is a passive voice verb, identified through dependency relation tags. For all predicate verbs, the style penalty is as follows: , The passive voice tolerance threshold; 6) Using SRL missing detection calculations to label logically ambiguous problems, the calculation formula is as follows: Where R is the set of necessary semantic roles. This describes the filling situation for character r in sentence s.

[0057] A weighted composite score is calculated based on the different risk problem types, and the score is expressed as follows: in , , , The final quality judgment is expressed as ; After revision based on the human collaboration mechanism, the incremental fine-tuning based on LoRA transforms the human correction into system knowledge and realizes consensus memory learning, including: 1) configuring the LoRA adapter to realize the incremental training process; 2) updating parameters during training, updating the original weights and calculating gradients through LoRA.

[0058] The steps to collect user feedback on terminology issues include: 1) Provide explicit user feedback when terms are corrected or added; provide implicit feedback when the frequency of term hover queries; and enable AI self-checking for discrepancies by examining inconsistencies between the two engine terms. 2) The validity verification algorithm is used to verify the validity of the original feedback deduplication and merging, expressed as: in For user reputation score, For domain matching degree, This indicates semantic consistency with existing terminology databases, and , , ; Automatic terminology optimization is achieved through term vectorization and clustering, including: 1) Vectorization algorithm: 2) HDBSCAN hierarchical density clustering is used, and the minimum number of cluster terms, the number of neighbors required for the core point, and the merging distance threshold are set. Using term vectors as nodes and semantic relevance as edges, a semantic relationship network between terms is established by constructing similar edges. A large-scale terminology database is updated using incremental indexing based on FAISS, including term merging and disambiguation processing: 1) Merging algorithm: 2) Calculate the context matching degree based on context term disambiguation.

[0059] Compared with existing technologies, this embodiment provides a human-machine collaborative translation batch processing method incorporating artificial intelligence, which has the following beneficial effects: This AI-integrated human-machine batch translation method achieves efficient term recognition by combining FAISS approximate search and BiLSTM-CRF model, dynamically updates the terminology database through HDBSCAN clustering and concept graph construction, achieves context-aware processing through dynamic Prompt engineering and dual-engine scheduling, and achieves deep human-machine collaborative interaction by combining AI-assisted proofreading and consensus memory learning.

[0060] In the embodiment: Example 1 In this embodiment, the steps for initializing the translation system and setting up the environment include: 1) Server resources are provided using SSD and S3 terminology library storage, and a dedicated VPC and security group are set up to open ports 443 / 8000 for networking, and an operating system and Docker containers are configured. 2) Framework Integration a. Electron front-end architecture, including a multi-document upload area, a real-time terminology prompt floating window, and a translation progress dashboard module; b. Python backend services: GPT-4 API encapsulation, ChatGLM3 local deployment, terminology recognition model, Elasticsearch interface, MongoDB feedback collection, and RESTful entry point; 3) AI model deployment: Connect to the GPT API, deploy local models, and use FlashAttention-2 and vLLM to accelerate inference; 4) Build a terminology database system: Deploy an Elasticsearch cluster and import the initial terminology database, including the WIPO patent terminology database, UniProt medical terminology database, ISO terminology database, and IEC electrical terminology database; 5) Real-time system updates: Redis message queues include term update publishers and Electron client subscriptions. The incremental synchronization mechanism is set up as follows: the user terminal submits the updated terms to the API gateway, the API gateway's audit service adds the new terms to the audit queue, the NLP engine automatically verifies the consistency of the terms according to the audit service, and writes them to ES in the terminology database after passing the verification, the terminology database publishes the update event through the message bus, and the message bus pushes the incremental package to all clients. 6) Security and permission configuration: Configure the Electron login authentication module to use TLS 1.3 and two-way certificate authentication at the transport layer, AES-256-GCM encryption at the storage layer using a terminology library, and client-side local encryption before transmission of user documents.

[0061] Example 2 In this embodiment, the steps for outputting and integrating the translation include: 1) Reconstructing a tree structure based on location mapping: creating a copy of the document object, traversing all text nodes, preserving the original format tags, and handling tables specially to accurately inject the translation into the original document structure; 2) Create professional-grade bilingual reports using a dynamic layout engine, including: The original paragraph alignment is analyzed and dependency tree matching is performed to check if the alignment is correct. If the alignment is correct, the paragraph is compared precisely; otherwise, the sentence is compared. Then, the left and right layout is generated and navigation anchors are added. 3) Provide an enterprise-grade REST API encapsulation, including: The client sends a POST / translate message to the API gateway. The API gateway performs authentication via JWT verification. After authentication, a 200 / 401 error is sent to the API gateway. The API gateway then creates a Celery task in the task queue, and the task queue assigns translation tasks to worker nodes. The worker node then saves the progress in the database. At this point, the worker node returns the task ID to the client, which then sends a GET / result / {task_id} to the result endpoint.

[0062] The beneficial effects of this implementation are: efficient term recognition is achieved by combining FAISS approximate search and BiLSTM-CRF model; the terminology database is dynamically updated by HDBSCAN clustering and concept graph construction; context-aware processing is achieved by dynamic Prompt engineering and dual-engine scheduling; and deep human-machine collaborative interaction is achieved by combining AI-assisted proofreading and consensus memory learning.

[0063] The above description of several specific embodiments further details the technical solution provided by the present invention in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above-described specific embodiments are not intended to limit the present invention. Any reasonable modifications and improvements to the present invention, combinations of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A human-machine collaborative translation method incorporating artificial intelligence for batch processing of translated texts, characterized in that: include: The steps involved are: receiving and parsing the manuscript to be translated, extracting information on language, field, and urgency, and forming a translation task package; The steps include performing pre-translation semantic analysis on the translation task package, extracting terms, keywords and language style parameters, and generating term suggestions and text features; The steps involve calling a neural network translation engine to perform machine translation of the original text to be translated, and then fusing the results based on confidence and quality scoring models to generate an initial translation. Perform intelligent quality assessment on the initial translation results, and classify the tasks according to the scoring results, marking the steps as usable translations, light proofreading tasks, or deep editing tasks; After the task is pushed to the human-computer collaboration platform, the translator will manually modify and confirm it based on terminology suggestions and contextual prompts. The steps involve performing a full-text consistency check, detecting and alerting to inconsistencies in terminology and abrupt style changes, and outputting the final translated text.

2. The human-machine combined translation batch processing method incorporating artificial intelligence according to claim 1, characterized in that, It also includes steps such as storing the final translation in the translation version library and updating terminology resources to achieve knowledge accumulation and model optimization.

3. The human-machine combined translation batch processing method incorporating artificial intelligence according to claim 1, characterized in that, The process of forming the translation task package includes clustering according to the professional field of the manuscript and setting priority queues according to the urgency of the task.

4. The human-machine combined translation batch processing method incorporating artificial intelligence according to claim 1, characterized in that, The pre-translation semantic analysis employs a natural language processing algorithm based on a pre-trained language model to identify proper nouns, technical phrases, and sentimental statements, and annotates them in the original text.

5. The human-machine collaborative translation batch processing method incorporating artificial intelligence according to claim 1, characterized in that, The neural network translation engine includes at least one self-developed translation engine and one external commercial engine, and the fusion method is a dynamic optimization strategy based on sentence pair quality scoring results.

6. The human-machine combined translation batch processing method incorporating artificial intelligence according to claim 1, characterized in that, The intelligent quality assessment includes four scoring dimensions: semantic fidelity, grammatical integrity, terminology consistency, and sentence naturalness. It uses multi-model integration to determine the translation quality level.

7. A human-machine collaborative translation batch processing device incorporating artificial intelligence, characterized in that, include: The module receives and parses the manuscript to be translated, extracts information on language, field, and urgency, and forms a translation task package; A module is used to perform pre-translation semantic analysis on the translation task package, extract terminology, keywords and language style parameters, and generate terminology suggestions and text features; This module calls a neural network translation engine to perform machine translation of the original text of the manuscript to be translated, and generates an initial translation result by fusing the confidence and quality scoring models. The initial translation results are subjected to intelligent quality assessment, and tasks are assigned based on the scoring results, marking modules as usable translations, light proofreading tasks, or deep editing tasks; After the task is pushed to the human-machine collaboration platform, the translator will manually modify and confirm it based on terminology suggestions and contextual prompts. This module performs full-text consistency checks, detects and alerts users to inconsistencies in terminology and abrupt style changes, and outputs the final translation.

8. A computer storage medium for storing computer programs, characterized in that, When the computer program is read by the computer, the computer executes the method of claim 1.

9. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method of claim 1.

10. A computer program product, as a computer program, is characterized by: When the computer program is executed, it implements the method of claim 1.

Citation Information

Patent Citations

  • Man-machine combined translation batch processing translation method containing artificial intelligence

    CN110837742A

Cited By

  • Homework correction method and device based on human-AI efficient cooperation

    CN121391566A

  • Intelligent translation method and system

    CN121543608A

  • Intelligent translation method and system

    CN121543608B

  • Manual data translation working platform system and method

    CN121809498A

  • Interactive continuous writing translation method and system based on large language model

    CN122088521A