Method and device for adaptively translating multiple languages based on model cooperation
By employing a translation method that combines an OCR engine with a large model, the problem of traditional translation tools being unable to handle image documents is solved, enabling efficient and accurate multilingual translation that is suitable for translation needs across multiple fields and languages.
Patent Information
- Application Number
- CN202510960976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-12
- Publication Date
- 2025-10-21
AI Technical Summary
Traditional translation tools rely on manually generated rules or high-quality parallel corpora, have weak generalization ability, cannot handle image-type documents, and lack image recognition capabilities, making it impossible to parse the text content in images.
The OCR engine is used for layout analysis, the convolutional neural network is used to locate the coordinates of the rectangular box, and the large model engine is combined for text recognition and translation. The Transformer architecture is used to achieve cross-language mapping, and the translation process is optimized through the bidirectional text alignment algorithm and beam search algorithm.
It enables automated translation of image documents, improving the accuracy and efficiency of multi-domain and multi-language translation, reducing human intervention, and maintaining the consistency and accuracy of translated document format.
Smart Images

Figure CN120822528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of translation processing technology, and in particular to a method and device for translating multiple languages based on model collaboration and self-adaptation. Background Art
[0002] In cross-language communication scenarios, traditional translation tools (including rule-based engines and statistical machine translation) primarily process textual content, converting source language text into target language content through established rules or statistical models, thereby fulfilling the basic needs of cross-language information transfer. While these technologies have a certain application basis in plain text translation scenarios, their limitations are becoming increasingly apparent as document types diversify.
[0003] Traditional translation tools rely heavily on manually crafted rule systems or high-quality parallel corpora. For example, rule engines require pre-defined grammar and vocabulary conversion rules, while statistical machine translation relies on large amounts of aligned bilingual corpora for model training. This reliance significantly reduces translation accuracy and adaptability when faced with scenarios not covered by the rules or where corresponding corpora are lacking, making it unable to meet the generalized translation needs in complex scenarios.
[0004] Traditional translation tools suffer from fundamental technical flaws when it comes to image-based documents. Due to their lack of image recognition capabilities, they are unable to parse the textual content within images, much less complete the subsequent translation process. For example, traditional tools are unable to extract textual information from documents containing charts, scans, or mixed text and images, rendering translation impossible. This creates significant barriers to application in areas such as engineering drawings, electronic documents, and printed materials. Summary of the Invention
[0005] To this end, the present invention provides a method and device for self-adaptive translation of multiple languages based on model collaboration, which solves the problem that traditional translation tools rely on manual rules or high-quality parallel corpus, have weak generalization ability, and are unable to translate when encountering image-type documents due to the lack of image recognition ability.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solution: a method for model-based collaborative adaptive translation of multiple languages, comprising the following steps:
[0007] Performing layout analysis on the uploaded image document through the OCR engine, analyzing and locating the rectangular frame coordinate information of each content block in the image document, and generating rectangular blocks based on the rectangular frame coordinate information;
[0008] After the layout analysis is completed, the content of each rectangular block is recognized and converted into plain text by the OCR engine to obtain recognized text;
[0009] Assemble the recognition text and the translation prompt template, and use the large model engine to generate the target language translation content;
[0010] After setting the content in the corresponding rectangular box to empty, replace it with the content translated by the large model;
[0011] Repeat the steps of assembling, generating translations, and replacing to translate and replace all rectangular boxes obtained from the layout analysis;
[0012] Through the image document synthesis tool, the translated content is merged to generate a new document and then presented to the user.
[0013] As a preferred solution for the method of model-based collaborative adaptive translation of multiple languages, the OCR engine adopts a convolutional neural network in the model, and the forward propagation process is as follows:
[0014] y=f(Wx+b)
[0015] Where x is the input image feature, W is the weight matrix, b is the bias, and f is the activation function;
[0016] The back-propagation algorithm is used to minimize the loss function and optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle:
[0017]
[0018] Where n is the number of samples, y i is the model prediction value, is the true value.
[0019] As a preferred solution for the method of model-based collaborative adaptive translation of multiple languages, during the layout analysis process, the connected domain analysis algorithm is used to locate the boundaries of text lines and paragraphs, and the minimum enclosing rectangle of the connected pixel area is calculated:
[0020] R=(x,y,w,h)
[0021] Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle.
[0022] As a preferred solution for the method of model-based collaborative adaptive translation of multiple languages, the large model engine adopts the Transformer architecture and implements contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is:
[0023]
[0024] Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the correlation of the input sequence; V is the value matrix, which is used to extract semantic features; d k is the key vector dimension, used to scale the attention weights.
[0025] As a preferred solution for model-based collaborative adaptive translation of multiple languages, the content translated by the large model is used for replacement. A bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is:
[0026]
[0027] Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
[0028] As a preferred solution for model-based collaborative adaptive translation of multiple languages, a beam search algorithm is used to optimize the decoding process when generating target language translation content using a large model engine. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is:
[0029]
[0030] Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.
[0031] The present invention also provides a device for self-adaptive translation of multiple languages based on model collaboration, comprising:
[0032] A layout analysis module is used to perform layout analysis on the uploaded image document through an OCR engine, analyze and locate the rectangular frame coordinate information of each content block in the image document, and generate rectangular blocks according to the rectangular frame coordinate information;
[0033] A content recognition and conversion module is used to convert the content of each rectangular block into plain text through the OCR engine after the layout analysis is completed to obtain recognized text;
[0034] An assembly translation module is used to assemble the recognition text and the translation prompt template, and generate the target language translation content using a large model engine;
[0035] The content replacement module is used to first clear the content in the corresponding rectangular box and then replace it with the content translated by the large model;
[0036] Repeat processing module, used for repeated assembly, translation generation and replacement, and translation replacement of all rectangular boxes obtained from layout analysis;
[0037] The synthesis return module is used to merge the translated content through the image document synthesis tool to generate a new document and return it to the user.
[0038] As a preferred solution for the device of model-based collaborative adaptive translation of multiple languages, in the layout analysis module, the model adopted by the OCR engine introduces a convolutional neural network, and the forward propagation process is:
[0039] y=f(Wx+b)
[0040] Where x is the input image feature, W is the weight matrix, b is the bias, and f is the activation function;
[0041] The back-propagation algorithm is used to minimize the loss function and optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle:
[0042]
[0043] Where n is the number of samples, y i is the model prediction value, is the true value;
[0044] In the layout analysis module, the connected domain analysis algorithm is used to locate the boundaries of text lines and paragraphs by calculating the minimum enclosing rectangle of the connected pixel area:
[0045] R=(x,y,w,h)
[0046] Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle.
[0047] As a preferred solution for the device of model-based collaborative adaptive translation of multiple languages, in the assembly translation module, the large model engine adopts the Transformer architecture and realizes contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is:
[0048]
[0049] Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the correlation of the input sequence; V is the value matrix, which is used to extract semantic features; d k is the key vector dimension, used to scale the attention weight;
[0050] In the content replacement module, a bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is:
[0051]
[0052] Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
[0053] As a preferred solution for the device of model-based collaborative adaptive translation of multiple languages, the assembly translation module adopts a beam search algorithm to optimize the decoding process. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is:
[0054]
[0055] Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.
[0056] The present invention has the following advantages:
[0057] First, by uploading image documents to an OCR engine for layout analysis and combining OCR text recognition with large-scale model translation throughout the entire process, this completely overcomes the inability of traditional translation tools to process image documents due to their lack of image recognition capabilities. For example, for documents containing images, such as engineering drawings and scanned contracts, the text areas can be automatically analyzed and translated, filling a gap in traditional technologies for processing mixed text and image documents.
[0058] Second, the large-scale model engine, powered by massive amounts of pre-trained multilingual data, eliminates the need for manually formulated translation rules or high-quality parallel corpora for specific fields, enabling it to adaptively handle translation needs across multiple fields and languages. Compared to the limitations of traditional statistical machine translation in niche languages or specialized fields, it offers improved accuracy in translating terminology-intensive documents such as medicine and law, and can automatically adapt to vocabulary expressions in emerging fields.
[0059] Third, the OCR engine's fully automated processing of layout analysis, text recognition, large-scale model translation, and document synthesis significantly reduces manual intervention. For example, a 100-page document with mixed text and images would require two to three working days for traditional manual translation. However, this solution can complete the entire process, from image analysis to translation document generation, in just 30 minutes, increasing efficiency by over 90% while eliminating the cost of manual rule maintenance.
[0060] Fourth, by obtaining the coordinates of the rectangular frame of the content block, the text is replaced in its original position after translation. A bidirectional text alignment algorithm is then used to ensure that formatting attributes such as font and size remain consistent with the original document. For complex documents with tables and columns, the translated document maintains the original paragraph spacing and image-text relationships, eliminating the need for secondary typesetting and meeting the formatting requirements of formal documents.
[0061] Fifth, the system is not only suitable for general image and document translation, but can also be applied in scenarios such as archive digitization, cross-border e-commerce product image translation, and multinational engineering drawing collaboration. It can batch process image and text translation for product detail pages, supporting translation between over 20 languages, including Chinese, English, Japanese, and Korean. It also achieves an accuracy rate of over 95% for documents with complex layouts, such as posters and manuals. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.
[0063] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.
[0064] Figure 1 This is a flow chart of the method for model-based collaborative adaptive translation of multiple languages provided in Example 1 of the present invention;
[0065] Figure 2 This is a schematic diagram of the device architecture for model-based collaborative adaptive translation of multiple languages provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0066] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0067] Example 1
[0068] See also Figure 1 The embodiment of the present invention provides a method for translating multiple languages based on model collaboration and self-adaptation, including the following steps:
[0069] S1. Analyze the layout of the uploaded image document using an OCR engine, analyze and locate the rectangular frame coordinate information of each content block in the image document, and generate a rectangular block based on the rectangular frame coordinate information;
[0070] Specifically, layout analysis is based on computer vision technology. The OCR engine uses a convolutional neural network (CNN) to extract image features such as edges and textures, and combines it with a connected domain analysis algorithm (such as a four-neighborhood or eight-neighborhood search) to identify text regions. Connected domain analysis marks connected regions of pixels and calculates the minimum bounding rectangle of each region to determine the spatial position (x, y), width w, and height h of the text block, providing a coordinate basis for subsequent text recognition and position replacement.
[0071] S2. After the layout analysis is completed, the content of each rectangular block is recognized and converted into plain text by the OCR engine to obtain recognized text;
[0072] Specifically, the text recognition stage uses an OCR character recognition model (such as CRNN or Transformer OCR). After inputting a rectangular block image into the model, the model solves the problem of aligning the image and text sequence through feature extraction, sequence prediction, and the CTC (Connectionist Temporal Classification) loss function. Through training, the model learns the mapping relationship between the visual features and semantics of characters, converting the image pixel matrix into the corresponding character sequence, achieving image-to-text conversion.
[0073] S3, assembling the recognition text and the translation prompt template, and generating the target language translation content using a large model engine;
[0074] Specifically, the translation prompt template includes a source language identifier, target language instructions, and formatting controls (e.g., "Please translate the following text into English: {text}"). This prompt engineering optimizes the translation performance of the large model. The large model engine utilizes the Transformer architecture, which uses a self-attention mechanism to calculate semantic relevance within the input sequence and implements cross-language mapping through a multi-layer encoder-decoder structure. This attention mechanism allows the model to focus on key semantic information.
[0075] S4. After setting the content in the corresponding rectangular box to empty, replace it with the content translated by the large model;
[0076] Specifically, the replacement process is based on a bidirectional text alignment algorithm, which uses dynamic programming to calculate the optimal alignment path between the source text and the translated text. The alignment loss function is obtained by combining the character matching function δ and the position weight w i,j , ensuring that the length and format of the translated text are compatible with the original rectangle. Clearing the original content before replacing it can avoid text overlap or formatting errors. Combining the rectangle's coordinate information (x, y, w, h), the document rendering engine fills the translated text in its original position and font attributes to ensure layout consistency.
[0077] S5, repeat the steps of S3 assembling, translation generation and S4 replacing, and translate and replace all the rectangular boxes obtained by the layout analysis;
[0078] Specifically, the looping mechanism is suitable for complex documents with multiple content blocks. By iterating through all rectangular blocks, it ensures that every text area undergoes the "recognize-translate-replace" process. This iterative process is automated through programming logic, such as using a for loop to iterate through the list of rectangular blocks, executing steps S3 and S4 in sequence until all blocks are processed. This batch processing method ensures the integrity and efficiency of document translation and avoids omissions caused by manual intervention.
[0079] S6. Using the image-document synthesis tool, the translated content is combined to generate a new document and presented to the user.
[0080] Specifically, the document synthesis tool, based on image rendering technology (such as PDF.js or the Canvas drawing interface), recombines the replaced text blocks with the non-text elements of the original document. The synthesis process maintains the original document's page size, margins, pagination, and other attributes. Through coordinate mapping, the translated content blocks are rendered in their original positions, ultimately generating a new document with a uniform format. This process involves adapting the image synthesis algorithm to document format specifications to ensure that the output document is ready for reading and printing.
[0081] In this embodiment, in step S1, a convolutional neural network is introduced into the model adopted by the OCR engine, and the forward propagation process is:
[0082] y=f(Wx+b)
[0083] In the formula, x represents the input image features, W is the weight matrix, b is the bias, and f is the activation function. The CNN forward propagation performs a linear transformation on the input image features x using the weight matrix W and the bias b. The activation function f (such as ReLU or Sigmoid) then introduces a nonlinear mapping, enabling the model to learn complex features. This process extracts low-level features such as edges and corners from the image, which are then combined into high-level features such as text lines and characters, providing a foundation for layout parsing and text recognition. The activation function overcomes the expressive limitations of linear models, enabling the network to learn the nonlinear relationship between pixels and semantics.
[0084] Among them, the back propagation algorithm is used to minimize the loss function to optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle:
[0085]
[0086] Where n is the number of samples, y i is the model prediction value, The back propagation algorithm calculates the gradient of the loss function L to the parameters of each layer and uses the gradient descent method to update W and b so that the model predicts the value y i and the true value The loss function uses the mean squared error (MSE) to measure the deviation between the model output and the annotated data. During the optimization process, the model gradually improves the layout parsing accuracy through iterative training, making the generated rectangular box coordinates closer to the actual values annotated by humans.
[0087] In this embodiment, in step S1, during the layout analysis process, a connected domain analysis algorithm is used to locate text lines and paragraph boundaries, and the minimum enclosing rectangle of the connected pixel area is calculated:
[0088] R=(x,y,w,h)
[0089] Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle. Connected domain analysis is based on pixel neighborhood relationships and considers adjacent regions with similar pixel values in a grayscale image to be a connected domain. By scanning image pixels, each connected domain is marked and its minimum enclosing rectangle R is calculated, where (x, y) is the coordinate of the upper left corner, and w and h are the width and height. This rectangle serves as the spatial boundary of the text block and is used for subsequent text recognition and position replacement, enabling structured parsing of the document layout.
[0090] In this embodiment, in step S3, the large model engine adopts the Transformer architecture and implements contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is:
[0091]
[0092] Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the correlation of the input sequence; V is the value matrix, which is used to extract semantic features; d k is the key vector dimension, used to scale the attention weights.
[0093] Specifically, the Transformer's self-attention mechanism calculates the semantic relevance of each position in the input sequence to other positions through the operation of the Q, K, and V matrices. Q is used to query the target semantics, K is used to calculate the relevance, and V is used to extract features. T "Calculate the similarity between the query and the key, divided by It is used to scale gradients to avoid excessive values; softmax converts similarity into a probability distribution, which is used as a weighted sum over V to produce the focused feature output. This mechanism enables the model to capture long-range dependencies and understand contextual semantics, thereby achieving accurate cross-language translation.
[0094] In this embodiment, in step S4, during the replacement process using the content translated by the large model, a bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is:
[0095]
[0096] Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
[0097] Specifically, the bidirectional text alignment algorithm is used to solve the problem of inconsistent lengths between the source text s and the translated text t. The alignment matrix is constructed through dynamic programming to find the alignment path that minimizes the loss function L(align). i ,t j ) is a character matching function (match is 1, non-match is 0), w i,j is the position weight (e.g. the beginning and end of a sentence have a higher weight). This algorithm takes into account the semantic order and position importance of the text to ensure that the translated text is compatible with the original rectangle in terms of semantics and format.
[0098] In one possible embodiment, when using a large model engine to generate target language translation content, a beam search algorithm is used to optimize the decoding process. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is:
[0099]
[0100] Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.
[0101] Specifically, beam search is an optimization algorithm between greedy search and exhaustive search. At each decoding step, the k candidate paths with the highest probability (beam width k) are retained, and the other paths are discarded. The complete translation sequence is generated by recursively expanding these paths. In the formula, x is the input text, is the output sequence, is the conditional probability of generating the tth word given the input and the generated prefix. By maximizing the product of this joint probability, the most likely translation sequence is found, balancing translation quality and computational efficiency.
[0102] The application scenarios of the present invention are as follows:
[0103] Image document translation scenario
[0104] It can translate common office documents in image formats, such as scanned contracts, invoices, and reports. The OCR engine analyzes the image document layout, locates the coordinates of the content rectangle, and then uses a large model to translate and replace it, finally synthesizing a new document. This enables rapid translation of office image documents and improves cross-border office efficiency.
[0105] For e-books in picture format, the text content can be analyzed and translated to meet the reading needs of readers of different languages and promote cultural dissemination and exchange.
[0106] Professional document translation scenarios
[0107] In the engineering field, numerous drawings exist in the form of images. This invention can accurately interpret text, symbols, and other content within drawings, and generate target-language versions through large-scale model translation, ensuring accurate information transfer during engineering collaboration. Translating images of diagnostic reports accompanying medical imaging can facilitate international medical research and case exchange, providing patients with access to a wider range of medical advice.
[0108] Business and marketing scenarios
[0109] In cross-border e-commerce, the text on product images needs to be translated into the languages of different target markets. This invention can batch process product images and quickly generate multilingual product display images, helping merchants expand into international markets. When companies advertise in different countries, they can translate the text in promotional images into the local language, making the ads more relevant to the target audience and improving the effectiveness of the promotion.
[0110] Cultural and educational scenes
[0111] Translating ancient texts in pictorial form helps preserve and pass on cultural heritage, allowing more people to understand the history and culture of different countries and regions. Translating pictorial textbooks and supplementary materials into multiple languages meets the learning needs of students from different countries and promotes the international sharing of educational resources.
[0112] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the method of model-based collaborative adaptive translation of multiple languages.
[0113] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0114] Example 2
[0115] See also Figure 2 Embodiment 2 of the present invention further provides a device for model-based collaborative adaptive translation of multiple languages, comprising:
[0116] The layout analysis module 100 is used to perform layout analysis on the uploaded image document through the OCR engine, analyze and locate the rectangular frame coordinate information of each content block in the image document, and generate rectangular blocks according to the rectangular frame coordinate information;
[0117] The content recognition and conversion module 200 is used to convert the content of each rectangular block into plain text through the OCR engine after the layout analysis is completed to obtain recognized text;
[0118] The assembly translation module 300 is used to assemble the recognition text and the translation prompt template and generate the target language translation content using a large model engine;
[0119] The content replacement module 400 is used to first set the content in the corresponding rectangular box to empty and then replace it with the content translated by the large model;
[0120] Repeat processing module 500, for repeating assembly, translation generation and replacement, and performing translation replacement on all rectangular boxes obtained by parsing the layout;
[0121] The synthesis return module 600 is used to merge the translated content into a new document through an image document synthesis tool and return it to the user.
[0122] In this embodiment, in the layout analysis module 100, the OCR engine adopts a convolutional neural network model, and the forward propagation process is as follows:
[0123] y=f(Wx+b)
[0124] Where x is the input image feature, W is the weight matrix, b is the bias, and f is the activation function;
[0125] The back-propagation algorithm is used to minimize the loss function and optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle:
[0126]
[0127] Where n is the number of samples, y i is the model prediction value, is the true value;
[0128] In the layout analysis module 100, the connected domain analysis algorithm is used to locate the text line and paragraph boundaries by calculating the minimum enclosing rectangle of the connected pixel area:
[0129] R=(x,y,w,h)
[0130] Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle.
[0131] In this embodiment, in the assembly translation module 300, the large model engine adopts the Transformer architecture and implements contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is:
[0132]
[0133] Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the correlation of the input sequence; V is the value matrix, which is used to extract semantic features; d k is the key vector dimension, used to scale the attention weight;
[0134] In the content replacement module 400, a bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is:
[0135]
[0136] Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
[0137] In this embodiment, the assembly translation module 300 uses a beam search algorithm to optimize the decoding process. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is:
[0138]
[0139] Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.
[0140] It should be noted that the information interaction, execution process and other contents between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of this application, and the technical effects they bring are the same as those of the method embodiment of this application. For specific contents, please refer to the description in the method embodiment shown above in this application, and no further details will be given here.
[0141] Example 3
[0142] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which the program code of the method for model-based collaborative self-adaptive translation of multiple languages is stored, and the program code includes instructions for executing the method for model-based collaborative self-adaptive translation of multiple languages of embodiment 1 or any possible implementation thereof.
[0143] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0144] Example 4
[0145] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0146] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method of model-based collaborative self-adaptive translation of multiple languages in embodiment 1 or any possible implementation thereof.
[0147] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.
[0148] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.
[0149] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0150] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A method for self-adaptive translation of multiple languages based on model collaboration, characterized in that: The following steps are involved: Performing layout analysis on the uploaded image document through the OCR engine, analyzing and locating the rectangular frame coordinate information of each content block in the image document, and generating rectangular blocks based on the rectangular frame coordinate information; After the layout analysis is completed, the content of each rectangular block is recognized and converted into plain text by the OCR engine to obtain recognized text; Assemble the recognition text and the translation prompt template, and use the large model engine to generate the target language translation content; After setting the content in the corresponding rectangular box to empty, replace it with the content translated by the large model; Repeat the steps of assembling, generating translations, and replacing to translate and replace all rectangular boxes obtained from the layout analysis; Through the image document synthesis tool, the translated content is merged to generate a new document and then presented to the user.
2. The method for model-based collaborative adaptive translation of multiple languages according to claim 1, characterized in that: The OCR engine uses a convolutional neural network model, and the forward propagation process is as follows: y=f(Wx+b) Where x is the input image feature, W is the weight matrix, b is the bias, and f is the activation function; The back-propagation algorithm is used to minimize the loss function and optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle: Where n is the number of samples, y i is the model prediction value, is the true value.
3. The method for model-based collaborative adaptive translation of multiple languages according to claim 2, characterized in that: During the layout analysis process, the connected domain analysis algorithm is used to locate the boundaries of text lines and paragraphs, and the minimum enclosing rectangle of the connected pixel area is calculated: R=(x,y,w,h) Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle.
4. The method for model-based collaborative adaptive translation of multiple languages according to claim 1, characterized in that: The large model engine adopts the Transformer architecture and implements contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is: Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the relevance of the input sequence; V is the value matrix used to extract semantic features; d k is the key vector dimension, used to scale the attention weights.
5. The method for model-based collaborative adaptive translation of multiple languages according to claim 4, characterized in that: In the process of replacing the content translated by the large model, a bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is: Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
6. The method for model-based collaborative adaptive translation of multiple languages according to claim 1, characterized in that: When using a large model engine to generate target language translation content, a beam search algorithm is used to optimize the decoding process. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is: Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.
7. A device for self-adaptive translation of multiple languages based on model collaboration, characterized in that: include: A layout analysis module is used to perform layout analysis on the uploaded image document through an OCR engine, analyze and locate the rectangular frame coordinate information of each content block in the image document, and generate rectangular blocks according to the rectangular frame coordinate information; A content recognition and conversion module is used to convert the content of each rectangular block into plain text through the OCR engine after the layout analysis is completed to obtain recognized text; An assembly translation module is used to assemble the recognition text and the translation prompt template, and generate the target language translation content using a large model engine; The content replacement module is used to first clear the content in the corresponding rectangular box and then replace it with the content translated by the large model; Repeat processing module, used for repeated assembly, translation generation and replacement, and translation replacement of all rectangular boxes obtained from layout analysis; The synthesis return module is used to merge the translated content through the image document synthesis tool to generate a new document and return it to the user.
8. The device for model-based collaborative adaptive translation of multiple languages according to claim 7, characterized in that: In the layout analysis module, the OCR engine uses a convolutional neural network model, and the forward propagation process is as follows: y=f(Wx+b) Where x is the input image feature, W is the weight matrix, b is the bias, and f is the activation function; The back-propagation algorithm is used to minimize the loss function and optimize the model parameters, analyze the layout of the image document, and obtain the coordinate information of the content block rectangle: Where n is the number of samples, y i is the model prediction value, is the true value; In the layout analysis module, the connected domain analysis algorithm is used to locate the boundaries of text lines and paragraphs by calculating the minimum enclosing rectangle of the connected pixel area: R=(x,y,w,h) Where (x, y) is the coordinate of the upper left corner, w is the width of the rectangle, and h is the height of the rectangle.
9. The device for model-based collaborative adaptive translation of multiple languages according to claim 7, characterized in that: In the assembly translation module, the large model engine adopts the Transformer architecture and implements contextual semantic understanding and cross-language mapping based on the self-attention mechanism. The attention calculation formula is: Where Q is the query matrix, which is used to obtain the target semantic association information; K is the key matrix, which is used to calculate the relevance of the input sequence; V is the value matrix used to extract semantic features; d k is the key vector dimension, used to scale the attention weight; In the content replacement module, a bidirectional text alignment algorithm is used to calculate the optimal alignment path between the source text and the translated text through dynamic programming. The alignment loss function is: Where s i is the i-th character of the source text, t j is the jth character of the translated text, δ is the character matching function, which takes 1 when it matches and 0 when it does not match, w i,j is the position weight, which represents the influence of the position within the text line on the alignment accuracy.
10. The device for model-based collaborative adaptive translation of multiple languages according to claim 7, characterized in that: In the assembly translation module, a beam search algorithm is used to optimize the decoding process. The beam width is k, and the k candidate paths with the highest probability are retained. The calculation formula is: Where x is the input text, is the output sequence, The conditional probability of generating the tth word given the input text and the generated sequence prefix.