Artificial intelligence-based foreign language bill image translation method, device and medium

By combining image recognition, a dedicated model for invoice translation, and a large language model, the problems of identifying and translating specialized terms in foreign language invoice translation have been solved. This has enabled efficient and accurate extraction and translation of foreign language invoice information, improving the processing efficiency and user experience of financial software.

CN120913229BActive Publication Date: 2026-07-24INSPUR GENERSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR GENERSOFT CO LTD
Filing Date
2025-07-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing accounting software struggles to accurately identify and translate specialized terminology in foreign language documents, resulting in inaccurate translations and a poor user experience, which in turn affects the smoothness and accuracy of accounting work.

Method used

An artificial intelligence-based approach is adopted, which uses an image recognition model to obtain text information and location coordinates of the invoice, uses a dedicated invoice translation model for accurate translation, combines a large language model to optimize language fluency, and uses a key information extraction model to extract key information from the invoice to generate an invoice image that conforms to the Chinese context.

Benefits of technology

It enables efficient and accurate translation of foreign language invoices, ensuring the clear presentation of key information and the integrity of financial information, thereby improving the efficiency of the foreign language invoice reimbursement process and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913229B_ABST
    Figure CN120913229B_ABST
Patent Text Reader

Abstract

The application provides a foreign language bill image translation method, device and medium based on artificial intelligence, and belongs to the technical field of data information. The foreign language bill image translation method comprises the following steps: acquiring a foreign language bill image, inputting the foreign language bill image into an image recognition model, and outputting bill text information and position coordinate information; inputting the bill text information into a pre-trained bill translation special model to obtain Chinese bill information; inputting the information into a large language model, optimizing the Chinese bill information according to a preset bill question and answer prompt word; using the position coordinate information, arranging and combining the optimized Chinese bill information according to a Chinese bill format, and drawing a Chinese bill image; using a key information extraction model to extract bill key information from the Chinese bill information; and displaying the Chinese bill image and the bill key information. The application can solve the problem that the prior art cannot accurately translate bill data sets, thereby causing difficulty in training accurate and effective bill information extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data information technology, and specifically relates to a method, device and medium for translating foreign language invoice images based on artificial intelligence. Background Technology

[0002] As a crucial component of modern enterprise data management, financial software not only provides businesses with efficient and standardized financial accounting management but also comprehensively monitors their cash accounts and income and expenditure. To enhance the usability and user experience of financial software, existing systems are beginning to integrate automatic document recognition and entry functions. This feature automatically recognizes and enters financial documents, eliminating the tedious steps of traditional manual input and greatly improving the intuitiveness of financial information. Through this intelligent approach, financial software can provide more accurate and efficient financial information support, thereby driving enterprise financial management towards a more intelligent and automated direction.

[0003] In the context of economic globalization, many enterprises inevitably encounter foreign language invoices when conducting overseas business. As enterprises continue to expand into overseas markets, existing financial software reveals several significant shortcomings when faced with massive amounts of foreign language invoices: First, the diverse formats of foreign language invoices and the multiple forms in which key information can be presented increase the complexity of information extraction. When attempting to extract key information from invoices, existing natural language processing often produces errors, such as incorrect amount recognition, failed date format conversion, or incomplete extraction of project descriptions. Furthermore, foreign language invoices commonly contain a large number of specialized terms, which often have specific industry meanings or even multiple meanings, further affecting translation accuracy. Most existing translation models only translate text images. For example, existing patent CN118552965A provides a text image translation model training method that uses a pre-trained model and training data to perform feature encoding and decoding on text images and source text to generate translation results; it also calculates text image translation loss and multi-level knowledge transfer loss, fusing them as the training loss to update the pre-trained model parameters.

[0004] However, foreign language documents are often highly specialized, such as... Figure 1 The data not only includes polysemous words but also involves the differentiation of numerous similar information points. Because existing translation models are general-purpose models and mostly only capable of translating text and images, they cannot specifically translate invoice data, resulting in poor translation accuracy. For example... Figure 2 As shown, the existing M2M-100 translation model directly addresses... Figure 1 The document shown will be translated. Figure 1The word "tip" in the original text is directly translated as "hint" instead of the more common "small tip" used on Chinese receipts. This issue makes it difficult for existing technology to quickly and accurately extract valid receipt information, resulting in shortcomings in practicality and user experience, and impacting the smoothness and accuracy of financial work. Summary of the Invention

[0005] This application aims to provide an AI-based foreign language invoice image translation solution that can efficiently and accurately translate foreign language invoices and extract valid invoice information, thereby improving the processing efficiency, fluency, and accuracy of financial software when processing foreign language invoices.

[0006] According to a first aspect of this application, this application provides an artificial intelligence-based method for translating foreign language invoice images, including: Acquire images of foreign language invoices, input the images into an image recognition model, and output the text information and location coordinates of the invoices. The text information of the invoice is input into a pre-trained invoice translation model to translate it into Chinese invoice information; The Chinese invoice information is input into the large language model, and the fluency of the Chinese invoice information is optimized according to the preset question and answer prompts for invoices. Using location coordinate information, the optimized Chinese invoice information is arranged and combined according to the Chinese invoice format to draw a Chinese invoice image in the Chinese invoice format. A key information extraction model is used to extract key information from Chinese invoice information. The key information extraction model is pre-trained with keyword tags. Displays images of Chinese invoices and key information about them.

[0007] Preferably, in the above-mentioned foreign language document image translation method, the steps of inputting the foreign language document image into the image recognition model and outputting the document text information and location coordinate information include: Obtain the training set of the recognition model, and classify the foreign language invoice images in the training set according to the language type and format type of the foreign language invoice images; Based on the distribution location and type of key information in the foreign language invoice images, the classified foreign language invoice images are labeled with data, and clustering algorithms are used to generate clustering anchor boxes corresponding to the key information in the foreign language invoice images. The image recognition model is trained using data-annotated and foreign language invoice images containing clustered anchor boxes, combined with an attention mechanism. Once the image recognition model has been trained, the acquired foreign language invoice image will be input into the image recognition model. Image recognition models are used to extract text information, location coordinates, and confidence levels from images of foreign language invoices.

[0008] Preferably, in the above-mentioned foreign language invoice image translation method, the step of inputting the invoice text information into a pre-trained invoice translation-specific model to translate it into Chinese invoice information includes: Based on the language and format of the foreign language document image, and combined with the location coordinate information, determine the data type of the document's text information; By classifying the textual information on invoices, a training set for translation models containing textual information on invoices of various data types is obtained; The translation model training set is input into the preset translation model, and the preset Chinese invoice corpus is used to train the preset translation model to obtain a special model for invoice translation; the preset Chinese invoice corpus contains professional Chinese invoice corpus corresponding to the data types; Input the text information of the invoice into the invoice translation model, and output the corresponding Chinese invoice information and Boolean value.

[0009] Preferably, in the above-mentioned foreign language document image translation method, the step of training a preset translation model using a preset Chinese document corpus to obtain a dedicated document translation model includes: A dedicated semantic space is constructed using a Chinese invoice corpus, and a pre-defined translation model is trained using the Chinese invoice corpus so that the pre-defined translation model learns the invoice translation rules corresponding to the Chinese invoice corpus; The system controls a preset translation model to map the textual information of a document into specialized word vectors in a dedicated semantic space according to document translation rules. Combine all the professional word vectors corresponding to the text information of the invoice, convert all the professional word vectors into Chinese text, and obtain the Chinese invoice information.

[0010] Preferably, in the above-mentioned foreign language invoice image translation method, the step of inputting Chinese invoice information into a large language model and optimizing the fluency of the Chinese invoice information according to preset invoice question-and-answer prompts includes: Compile multiple question-and-answer prompts for invoices according to Chinese expression habits, and use these prompts to construct an invoice prompt template; The Chinese invoice information is input into a large language model, and the output of the large language model is optimized using question-and-answer prompts in the invoice prompt template to obtain Chinese invoice information with improved fluency.

[0011] Preferably, in the above-mentioned method for translating foreign language document images, the step of using position coordinate information to arrange and optimize the Chinese document information according to the Chinese document format to draw a Chinese document image includes: Generate a Chinese invoice framework according to the Chinese invoice format. The Chinese invoice framework includes the data type corresponding to the Chinese invoice information and the Chinese input area. Based on the correspondence between location coordinates and Chinese text fields, the Chinese document information is mapped to the data type field in the Chinese document frame.

[0012] Preferably, in the above-mentioned foreign language invoice image translation method, the step of extracting key information from Chinese invoice information using a key information extraction model includes: The key information of Chinese invoices was labeled using a labeling tool to obtain a training set. Train the key information extraction model using the training set until the loss function of the key information extraction model converges. The optimized Chinese invoice information is input into the key information extraction model; Use the key information extraction module to extract key information from Chinese invoice information.

[0013] Preferably, in the above-mentioned foreign language invoice image translation method, the step of extracting key information from Chinese invoice information using a key information extraction model includes: Use the Docano annotation tool to annotate key information on Chinese invoice samples to obtain keyword tags; Set the initial learning rate and loss function for the key information extraction model, and input Chinese ticket samples containing keyword tags into the key information extraction model for multiple rounds of training until the loss function converges. The Chinese invoice information is input into the key information extraction model to extract the key information of the invoice.

[0014] According to a second aspect of this application, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the artificial intelligence-based foreign language document image translation method provided by any of the above technical solutions.

[0015] According to a third aspect of this application, a computer storage medium is also provided, on which a computer program is stored, which, when executed, implements the artificial intelligence-based foreign language document image translation method provided by any of the above technical solutions.

[0016] The technical solution of this application has at least the following technical effects: The foreign language invoice image translation solution provided in this application acquires a foreign language invoice image and inputs it into an image recognition model to obtain the invoice text information and positional coordinates. Then, the text information is input into a pre-trained invoice translation model, which is trained on invoices and specifically designed for translating invoice text information, accurately translating specialized terminology within the foreign language invoice text. Compared to existing technologies that use large models or general translation rules to translate foreign language invoices, this solution reduces semantic bias and awkward phrasing in the translation results. Furthermore, to further improve the accuracy and readability of the translated content, the Chinese invoice information is input into a large language model. Leveraging the general language capabilities of this model, the fluency of the Chinese invoice information is optimized according to preset question-and-answer prompts, allowing for detailed adjustments and refinement of the translation. This ensures accurate expression of the invoice information while conforming to the Chinese context. Finally, the translated Chinese invoice information is arranged according to the obtained positional coordinates to create a Chinese invoice image in Chinese invoice format, resulting in a complete, accurate, and original layout-compliant Chinese invoice image, providing users with an intuitive visual experience. Finally, for subsequent financial work, a key information extraction model is used to efficiently and accurately extract key information from Chinese invoice information, ensuring the accuracy and completeness of the invoice information. This key information extraction model is pre-trained with keyword tags related to key information, thus accurately extracting and identifying key information presented in various formats such as amount, date, and project description. Through this method, core financial information can be accurately extracted while accurately translating foreign language invoice images, resulting in highly accurate Chinese invoice information. This method not only ensures the clear presentation of key invoice information but also provides complete Chinese invoice images, greatly improving the efficiency of the foreign language invoice reimbursement process and optimizing the user experience. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A structural schematic diagram of an original foreign language invoice image provided for existing technology; Figure 2 For existing translation models Figure 1 A schematic diagram of the structure of a Chinese invoice image directly translated from the invoice image shown; Figure 3 A flowchart illustrating the first method for translating foreign language invoice images based on artificial intelligence, provided in this application embodiment; Figure 4 A flowchart illustrating the second method for translating foreign language invoice images based on artificial intelligence, provided in this application embodiment; Figure 5 This application provides an illustration of the effect of extracting text information, location coordinate information, and confidence level from a foreign language invoice image. Figure 6 This is a schematic diagram illustrating the effect of a translation model used to translate Chinese invoice information, as provided in an embodiment of this application. Figure 7 This application provides an example of the effect of optimizing Chinese invoice information using a large language model. Figure 8 This application provides an illustration of the effect of optimizing a large language model using prompt word templates, as shown in the embodiments of this application. Figure 9 This application provides a schematic diagram of the structure of a redrawn Chinese invoice image; Figure 10 This is a schematic diagram of the structure of a key information extraction image provided in an embodiment of this application; Figure 11 This application provides a schematic diagram of the structure of an original English invoice image. Figure 12 This application provides a schematic diagram of the structure of a redrawn Chinese invoice image; Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.

[0019] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.

[0020] In this application, unless otherwise expressly specified and limited, the terms "above" and "below" the second feature can refer to direct contact between the first and second features, or indirect contact between the first and second features through an intermediate medium. In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

[0021] The existing technology has the following drawbacks: Foreign language invoices are often highly specialized, containing not only polysemous words but also requiring the differentiation of numerous similar information points. A major challenge in foreign language invoice recognition and translation applications is processing invoices containing many domain-specific terms. Existing methods mostly employ natural language processing (NLP) models. However, NLP models tend to use general translation rules for displaying and optimizing results when performing translation tasks. This process can lead to semantic biases or awkward phrasing in the translation. Specifically, the model may incorrectly translate specialized terms into non-specialized or uncommon equivalents, resulting in decreased accuracy and even semantic errors.

[0022] like Figure 1 The original text of the foreign language invoice image shown is... Figure 2 To utilize the M2M-100 translation model for Figure 1 The image shown is a direct translation of a foreign language receipt into Chinese. The word "tip" on the foreign language receipt usually refers to a tip, not a monetary reward. Figure 2 The "hints" following the model's interpretation. In this case, the accuracy and readability of the translation are severely affected, which in turn impacts the efficiency and reliability of financial data processing. Furthermore, invoices often contain various formats for amounts, invoice dates, and delivery dates, making it easy for mis-ordered matching issues to occur when extracting this crucial information.

[0023] To overcome the aforementioned technical challenges, the following embodiments of this application provide an AI-based foreign language invoice image translation solution that can accurately translate foreign language invoices, extract core financial information, and finally obtain highly accurate Chinese translated invoices. This technology not only ensures the clear presentation of core financial information but also provides complete Chinese invoice images, greatly improving the efficiency of the foreign language invoice reimbursement process and optimizing the user experience.

[0024] To achieve the above objectives, see [link to relevant documentation]. Figure 3 , Figure 3 A flowchart illustrating the first artificial intelligence-based foreign language document image translation method provided in this application embodiment is shown below. Figure 3 As shown, this AI-based method for translating foreign language invoice images includes: S110: Acquire the image of the foreign language invoice, input the image of the foreign language invoice into the image recognition model, and output the text information and location coordinate information of the invoice.

[0025] The image recognition model used here can be the Paddle-OCR model. Paddle-OCR is an open-source OCR (Optical Character Recognition) toolkit within a deep learning framework, capable of supporting the detection, recognition, and orientation classification of multilingual text. This embodiment integrates a text recognition model and a coordinate judgment model (e.g., a "coordinate system transfer interface HEP") on top of the traditional Paddle-OCR model. This allows for the acquisition of foreign language document images through scanning or photographing, which are then input into the image recognition model to accurately identify the document's text information and location coordinates.

[0026] Specifically, in a preferred embodiment, in the above-described foreign language document image translation method, step S110: inputting the foreign language document image into an image recognition model and outputting document text information and location coordinate information includes: S111: Obtain the training set of the recognition model, and classify the foreign language invoice images in the training set according to their language type and format type.

[0027] First, foreign language document images are classified according to language and format. This allows the image recognition model to be trained using these images based on language (e.g., English, French, and Russian) and format (e.g., invoices, receipts, and forms). Different languages ​​often correspond to different countries or document usage habits, and different formats often have unique text information and location coordinates. Therefore, classifying foreign language document images according to language and format enables the image recognition model to accurately understand foreign language documents from different countries, with different usage habits, and in different formats, thus quickly and accurately extracting the text information and location coordinates.

[0028] S112: According to the distribution location and type of key information in the foreign language invoice images, the classified foreign language invoice images are labeled with data, and clustering algorithms are used to generate clustering anchor boxes corresponding to the key information in the foreign language invoice images.

[0029] Here, foreign language document images are labeled according to the distribution and type of key information (e.g., time type, symbol type, and text type) within the image. This makes the image recognition model more sensitive to key information, reducing false positives and false negatives. By labeling key information in the classified foreign language document images and using clustering algorithms to generate cluster anchor boxes, the image recognition model can accurately locate the key information during training, thereby improving the model's accuracy in recognizing key information and accurately extracting its coordinates. The key information here is in foreign language format.

[0030] S113: Use data-annotated images of foreign language invoices containing clustered anchor boxes to train an image recognition model using an attention mechanism.

[0031] This application utilizes the aforementioned data annotation (such as labels of annotation type) and clustering anchor boxes to enable the image recognition model to accurately locate text information on foreign language invoices, such as the aforementioned key information. Furthermore, by combining this with an attention mechanism for training, the image recognition model can enhance its sensitivity to key information, thereby accurately extracting this key information from the foreign language invoice image and improving the recognition accuracy of the image recognition model.

[0032] This application embodiment uses a recognition model training set containing a large number of foreign language invoice images to train the image recognition model. After training, the image recognition model can accurately identify the text information and location coordinate information of the foreign language invoice images.

[0033] S114: When the image recognition model has been trained, the obtained foreign language invoice image will be input into the image recognition model.

[0034] S115: Use an image recognition model to extract text information, location coordinates, and confidence level from foreign language document images. The delivery text information includes foreign language content.

[0035] Specifically, such as Figure 5 As shown, the foreign language information extracted from the foreign language document image by this image recognition model includes location coordinates, foreign language content, and confidence level.

[0036] This embodiment first inputs the foreign language document image into an image recognition model, which can accurately identify the text content in the foreign language document image and obtain key information such as the specific location coordinates of each character. This step is the foundation of the entire foreign language document image translation method, ensuring the accuracy and efficiency of subsequent result processing.

[0037] Figure 3 The technical solution provided in the illustrated embodiment, after the step of inputting the foreign language document image into the image recognition model and outputting the document text information and location coordinate information, further includes: S120: Input the text information of the invoice into a pre-trained invoice translation model to translate it into Chinese invoice information.

[0038] By inputting the text information on invoices output by the image recognition model into a pre-trained invoice translation model, foreign language content can be accurately translated into Chinese. Because this invoice translation model differs from general language translation models—it is pre-trained using invoice text information—it can accurately translate invoice text. Furthermore, this invoice translation model incorporates a pre-set Chinese invoice corpus and uses specialized translation rules to translate the invoice text. This effectively addresses the problem in existing natural language processing models that tend to use general translation rules for foreign language invoices containing numerous domain-specific terms, potentially leading to semantic biases or awkward phrasing in the translation results. Through this method, specialized terms can be translated into their corresponding professional or specialized Chinese vocabulary, thereby improving the accuracy of the translation results and reducing semantic errors.

[0039] In addition, as a preferred embodiment, in the above-described foreign language invoice image translation method, step S120: inputting the invoice text information into a pre-trained invoice translation-specific model to translate it into Chinese invoice information includes: S121: Determine the data type of the text information on the foreign language document based on the language type and format type of the image, combined with the location coordinate information.

[0040] The data types of text information on these documents include time, name, address, label, and amount. Using language and format types allows us to determine the arrangement and format of each piece of text information in a foreign language document image. Combining this with position coordinate information further clarifies the data type of the text information, enabling us to train a dedicated document translation model specifically for that data type. Furthermore, by combining the aforementioned data type and position coordinate information, we can reconstruct the position coordinates in the Chinese format.

[0041] S122: Classify the textual information of the invoices to obtain a training set for translation models containing textual information of invoices of various data types.

[0042] By classifying the textual information on the invoices according to the above data types, the resulting training set for the translation model can be used to train the translation model, enabling the model to accurately locate textual information on invoices of different data types and coordinate positions.

[0043] S123: Input the translation model training set into the preset translation model, and train the preset translation model in combination with the preset Chinese invoice corpus to obtain a special model for invoice translation; wherein, the preset Chinese invoice corpus contains professional Chinese invoice corpus corresponding to the data types.

[0044] Because the document text information in the translation model training set is categorized according to data type, the preset translation model can be trained separately for document text information of different data types, thus obtaining a dedicated document translation model. This model can accurately locate document text information of different data types and coordinate positions, improving the translation accuracy of document text information. The untrained preset translation model here can be the M2M-100 translation model.

[0045] In one preferred embodiment, step S123 above involves training a preset translation model using a preset Chinese document corpus to obtain a document translation-specific model, including: S1231: Construct a dedicated semantic space using a Chinese invoice corpus, and train a pre-defined translation model using the Chinese invoice corpus so that the pre-defined translation model learns the invoice translation rules corresponding to the Chinese invoice corpus.

[0046] S1232: Controls the preset translation model to map the textual information of the invoice into professional word vectors in a dedicated semantic space according to the invoice translation rules.

[0047] S1233: Combine all the professional word vectors corresponding to the text information of the invoice, convert all the professional word vectors into Chinese text, and obtain the Chinese invoice information.

[0048] The technical solution provided in this application, targeting specialized terminology in the field of negotiable instruments, uses a preset translation model to learn the corresponding negotiable instrument translation rules in the dedicated semantic space of this field. By mapping the specialized terminology in the negotiable instrument text information to specialized word vectors in the aforementioned dedicated semantic space according to these translation rules, it can accurately understand the Chinese meaning of the specialized terminology in the negotiable instrument text, thereby improving the accuracy of the translation of specialized terminology. In summary, the technical solution provided in this application, by combining a Chinese negotiable instrument corpus and negotiable instrument translation rules, accurately understands the corresponding Chinese text meaning based on the context and linguistic information of the negotiable instrument text information, achieving accurate translation of specialized negotiable instrument text information.

[0049] S124: Input the text information of the invoice into the dedicated invoice translation model, and output the corresponding Chinese invoice information and Boolean value. After using the dedicated invoice translation model to output the Chinese invoice information and Boolean value, the model can combine the aforementioned positional coordinate information and confidence level to further redraw and restore the Chinese invoice image in Chinese format, thereby improving the accuracy of the model's translation.

[0050] By optimizing the translation model using a large-scale model, a dedicated model for invoice translation is obtained. Inputting the invoice's text information will yield the desired result. Figure 6 The translated information includes the location coordinates, Chinese document information, confidence level, and Boolean value indicating whether translation was performed. Based on this translation, it is possible to redraw... Figure 9 The Chinese document image shown ensures accurate information delivery and visual consistency. In summary, the technical solution provided in this application, by inputting the text content and location information output by the image recognition model into the translation model according to the language type and format type of the foreign language document image, can accurately translate foreign language content into Chinese.

[0051] Furthermore, the Chinese information on the invoice translated using a specialized invoice translation model may appear somewhat stiff and literal, failing to fit the Chinese context. For example... Figure 6 The translated phrase "100% of the money goes to the driver" conveys the correct meaning, but the expression feels awkward, resulting in poor readability and accuracy. To further improve the accuracy and readability of the translation, after obtaining the Chinese invoice information, it is necessary to input the translated Chinese invoice information into a large language model for optimization, using the large language model to refine the fluency of the translated content.

[0052] After step S120 above, where the text information of the invoice is input into a pre-trained invoice translation model to obtain the Chinese invoice information, the following steps are also included: S130: Input the Chinese invoice information into the large language model, and optimize the language fluency of the Chinese invoice information according to the preset question-and-answer prompts.

[0053] In order to further improve the accuracy and readability of the translated content, this application embodiment improves the translation of Chinese invoice information by writing question-and-answer prompts for invoices when inputting Chinese invoice information into a large language model. This successfully optimizes the translation effect and perfectly adapts to the expression habits of the Chinese context.

[0054] Specifically, in a preferred embodiment, step S130, which involves inputting Chinese bill information into a large language model and optimizing the fluency of the Chinese bill information according to preset bill question-and-answer prompts, includes: S131: Compile multiple question-and-answer prompts for negotiable instruments according to Chinese expression habits, and use these prompts to construct a negotiable instrument prompt template. These prompts should include professional terminology, noun concepts, Chinese language features, and special considerations for negotiable instruments.

[0055] S132: Input the Chinese invoice information into the large language model, and use the question-and-answer prompts in the invoice prompt template to optimize the output of the large language model, resulting in Chinese invoice information with improved fluency. Here, the large language model combined with question-and-answer prompts is used to optimize the Chinese invoice information. Specifically, it can optimize the semantic accuracy of polysemous professional terms and correct the use of professional invoice terms in foreign language abbreviations, thereby reducing the challenges posed by the frequent use of professional abbreviations in foreign language invoices for translation.

[0056] See details Figure 7 ,like Figure 7 The data shown includes the original translations, such as "Terms and Conditions," "Cambridge," "News about football weapons," "Front car," and "You can send a check to East Repair Company." After being polished with question-and-answer prompts, the translated content is as follows: East Repair Company, 1912 Harvest Lane, New York City, New York, 12210. Bill to: John Smith… The translated text is completely consistent with the usage habits of the Chinese context.

[0057] To further improve the accuracy and readability of the translated content, this embodiment of the application feeds the translated Chinese information into a large-scale language model for optimization. Through carefully designed question-and-answer prompts for invoices, the large model can meticulously adjust and refine the translated content, ensuring that the expression of the invoice information is both accurate and consistent with the Chinese context. Specifically, the large language model is optimized using prompt templates, and the optimization results are as follows: Figure 8 As shown, large language models demonstrate significant advantages in language processing, particularly excelling in text understanding and optimization. This application leverages the powerful capabilities of large language models to refine translation results. Question-and-answer prompts are used to standardize responses and polish translated answers. Prompt templates successfully optimize translation quality.

[0058] Furthermore, to further improve the accuracy and readability of the translated content, this embodiment combines a large language model and a dedicated invoice translation model to construct an invoice translation-large language model. The Chinese invoice information output by the dedicated invoice translation model is used as the predicted value, while the optimized Chinese invoice information output by the large language model is used as the ground truth. This ground truth is then rolled back to the dedicated invoice translation model to further optimize its output. Simultaneously, a loss function for the dedicated invoice translation model is designed, which is the error function between the ground truth and the predicted value. When the loss function converges, the dedicated invoice translation model is considered successfully trained. This method of using the output of the large language model to roll back to the dedicated invoice translation model further improves the accuracy and readability of the dedicated invoice translation model's output.

[0059] Figure 3The technical solution provided in the illustrated embodiment, in step S130: after inputting the Chinese invoice information into the large language model and optimizing the language fluency of the Chinese invoice information according to preset invoice question-and-answer prompts, further includes: S140: Using location coordinate information, the optimized Chinese invoice information is arranged and combined according to the Chinese invoice format to draw a Chinese invoice image.

[0060] This application embodiment utilizes the position coordinate information obtained during the image recognition stage to redraw the Chinese invoice image according to the Chinese invoice format; thus, a complete, accurate Chinese invoice image that conforms to the original layout can be obtained, providing users with an intuitive visual experience.

[0061] Specifically, in a preferred embodiment, step S140 involves: using location coordinate information to arrange and optimize the Chinese invoice information according to the Chinese invoice format, and then drawing a Chinese invoice image; specifically including: S141: Generate a Chinese invoice framework according to the Chinese invoice format, wherein the Chinese invoice framework includes the data type corresponding to the Chinese invoice information and the Chinese input area.

[0062] S142: Based on the correspondence between the location coordinate information and the Chinese text input area, map the Chinese document information to the data type in the Chinese document frame.

[0063] Foreign language invoices come in various formats, and key information such as amounts, dates, and item descriptions may be presented in multiple ways, increasing the complexity of information extraction. Existing general-purpose natural language processing models often encounter errors when attempting to extract key information from invoices, such as misidentifying amounts, failing to convert date formats, or incompletely extracting item descriptions.

[0064] Finally, to facilitate subsequent financial work, key information was extracted from the optimized translation. The key information extraction model efficiently extracted crucial information such as date, amount, and purpose from the invoices, ensuring the accuracy and completeness of the information and providing strong data support for the enterprise. This series of processes not only enabled rapid translation and image conversion of foreign language invoices but also provided convenient and efficient services for the enterprise's financial work.

[0065] Figure 3 The technical solution provided in the illustrated embodiment further includes, after drawing the Chinese invoice image: S150: Extract key information from Chinese invoice information using a key information extraction model. This model is pre-trained using keyword tags. The UIE model within the PaddleNLP framework can be used for training to obtain a specialized information extraction model for extracting key information from invoices.

[0066] The technical solution provided in this application generates a Chinese invoice frame according to the Chinese invoice format, then maps the Chinese invoice to the Chinese writing area according to the obtained position coordinate information and the Chinese writing area, and combines it with data types such as time, fee, and visa information to accurately reconstruct the Chinese invoice image. In summary, Figure 11 The original foreign language invoice image shown can be used to generate an image rendering model after the optimized Chinese invoice text is input into the model. Figure 12 The image shown is a fluent and accurate translated document image.

[0067] Specifically, in a preferred embodiment, in the above-described foreign language invoice image translation method, step S150: extracting key information of the invoice from the Chinese invoice information using a key information extraction model, specifically includes: S151: Use a labeling tool to label the key information of Chinese invoices to obtain the training set for the extraction model.

[0068] S152: Use the extraction model training set to train the key information extraction model until the loss function of the key information extraction model converges.

[0069] S153: Input the optimized Chinese invoice information into the key information extraction model.

[0070] S154: Use the key information extraction module to extract key information from Chinese invoice information.

[0071] The technical solution provided in this application uses a labeling tool to label key information in Chinese invoices, obtaining a training set for the extraction model. This training of the key information extraction model ensures the accuracy and completeness of the invoice information, providing strong data support for enterprises. Through this series of processes, not only is rapid translation and image conversion of foreign language invoices achieved, but also convenient and efficient services are provided for enterprise financial work.

[0072] In a preferred embodiment, the step S150 of the above-mentioned foreign language invoice image translation method, which involves extracting key information from Chinese invoice information using a key information extraction model, includes: S155: Use the Docano annotation tool to annotate key information on Chinese invoice samples to obtain keyword tags; S156: Set the initial learning rate and loss function of the key information extraction model, input Chinese ticket samples containing keyword tags into the key information extraction model for multiple rounds of training until the loss function converges; S157: Input the Chinese invoice information into the key information extraction model to extract the key information of the invoice.

[0073] In the technical solution provided by this application embodiment, in the field of invoice extraction, key information such as amount and date needs to be extracted for archiving and approval, such as... Figure 10 Therefore, after obtaining the translated invoice results, it is necessary to extract the Chinese invoice information using a key information extraction model. This embodiment uses the UIE model under the PaddleNLP framework for extraction. However, the basic extraction model's extraction content is incomplete; while the date content can be extracted completely, complex monetary information is difficult to identify. Therefore, this embodiment first annotates and trains the invoice data. Specifically, using the Docano annotation tool, the dataset is annotated for the "date" and "amount" information in the invoice, resulting in labeled data. Then, an initial learning rate is set; the number of input data samples is 32 each time the neural network is trained; the maximum number of tokens in the text input to the information extraction model is 512; after 100 rounds of training, the resulting information extraction model is applied in the key information extraction process, obtaining readable and ordered key information, and can extract complete dates and total amounts.

[0074] Figure 3 The technical solution provided in the illustrated embodiment further includes, after extracting key information from the invoice: S160: Displays Chinese image of the invoice and key information about the invoice.

[0075] The foreign language invoice image translation method based on artificial intelligence provided in this application involves acquiring a foreign language invoice image, inputting the image into an image recognition model to obtain the invoice text information and position coordinate information, and then inputting the text information into a pre-trained invoice translation-specific model. This model, trained on invoices, is specifically designed for translating invoice text information and accurately translates specialized terminology within the foreign language invoice information. Compared to existing methods that use large models or general translation rules to translate foreign language invoices, this method reduces semantic bias and awkward phrasing in the translation results. Furthermore, to further improve the accuracy and readability of the translated content, the Chinese invoice information is input into a large language model. Leveraging the general language capabilities of this model, the fluency of the Chinese invoice information is optimized according to preset question-and-answer prompts, allowing for detailed adjustments and refinement of the translated content. This ensures accurate expression of the invoice information while conforming to the Chinese context. Finally, the translated Chinese invoice information is arranged according to the obtained position coordinate information to create a Chinese invoice image in Chinese invoice format. This results in a complete, accurate Chinese invoice image that conforms to the original layout, providing users with an intuitive visual experience. Finally, for subsequent financial work, a key information extraction model is used to efficiently and accurately extract key information from Chinese invoice information, ensuring the accuracy and completeness of the invoice information. This key information extraction model is pre-trained with keyword tags related to key information, thus accurately extracting and identifying key information presented in various formats such as amount, date, and project description. Through this method, core financial information can be accurately extracted while accurately translating foreign language invoice images, resulting in highly accurate Chinese invoice information. This method not only ensures the clear presentation of key invoice information but also provides complete Chinese invoice images, greatly improving the efficiency of the foreign language invoice reimbursement process and optimizing the user experience.

[0076] Additionally, see Figure 4 , Figure 4 This is a flowchart illustrating the second artificial intelligence-based foreign language document image translation method provided in this application embodiment. Figure 4 As shown, the method for translating foreign language document images includes: S210: Foreign language invoice image input.

[0077] S220: Image content recognition and information location tagging.

[0078] S230: Image recognition content is fed into the translation model to obtain Chinese information.

[0079] S240: Import Chinese information into a large model and use prompt words to optimize the translation content.

[0080] S250: Extract key information from the optimized translation content.

[0081] S260: Image display of translated Chinese invoice.

[0082] Furthermore, the beneficial effects of the product embodiments provided in the following embodiments of this application are the same as the beneficial effects of the foreign language invoice image translation method based on artificial intelligence provided in the above embodiments, and other technical features in the product embodiments are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0083] The following is for reference. Figure 13 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals and / or fixed terminals. Figure 13 The device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0084] The following is for reference. Figure 13 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals and / or fixed terminals. Figure 13 The device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0085] like Figure 13 As shown, the electronic device can include a processing unit 1001, such as a central processing unit and / or a graphics processing unit, which can perform various appropriate actions and processes according to a program stored in ROM 1002 or a program loaded from storage device 1003 into RAM 1004. RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output interface 1006 is also connected to bus 1005. Typically, the following systems can be connected to input / output interface 1006: input devices 1007, such as touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, and / or gyroscopes; output devices 1008, such as liquid crystal displays (LCDs), speakers, and / or vibrators; storage devices 1003, such as magnetic tape and / or hard disks; and communication devices 1009. Communication device 1009 is capable of enabling the electronic device to exchange data with other devices wirelessly or via wired communication. Although the diagram shows a model building device with various systems, it should be understood that it is not required to implement or have all of the systems shown. It is possible to implement or have more or fewer systems alternatively.

[0086] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the AI-based foreign language document image translation method shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the AI-based foreign language document image translation method of the embodiments disclosed in this application.

[0087] This application provides a computer-readable storage medium having computer-readable program instructions stored thereon, namely the computer program described above, which is used to execute the foreign language invoice image translation method based on artificial intelligence in the above embodiments.

[0088] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the model-building device, can be written in one or more programming languages ​​or combinations thereof to perform the operations of this application. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, or C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer, for example, via the Internet using an Internet service provider.

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0090] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0091] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0092] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for translating foreign language invoice images based on artificial intelligence, characterized in that, include: Acquire an image of a foreign language invoice, input the image of the foreign language invoice into an image recognition model, and output the invoice text information and location coordinate information; The text information of the invoice is input into a pre-trained invoice translation model to translate it into Chinese invoice information; The Chinese invoice information is input into a large language model, and the fluency of the Chinese invoice information is optimized according to preset question-and-answer prompts for invoices. Using the location coordinate information, the optimized Chinese invoice information is arranged and combined according to the Chinese invoice format to draw a Chinese invoice image in the Chinese invoice format. The key information extraction model is used to extract key information from the Chinese invoice information, wherein the key information extraction model is pre-trained with keyword tags. Display the Chinese image of the invoice and its key information; The step of inputting the document text information into a pre-trained document translation model to translate it into Chinese document information includes: determining the data type of the document text information according to the language type and format type of the foreign language document image, combined with the position coordinate information; classifying the document text information to obtain a translation model training set containing multiple data types; inputting the translation model training set into a preset translation model, and training the preset translation model in combination with a preset Chinese document corpus to obtain the document translation model; wherein the preset Chinese document corpus contains professional Chinese document corpus corresponding to the data types; inputting the document text information into the document translation model, and outputting the corresponding Chinese document information and Boolean value; The step of training the preset translation model using a preset Chinese invoice corpus to obtain the invoice translation-specific model includes: constructing a dedicated semantic space using the Chinese invoice corpus, and training the preset translation model using the Chinese invoice corpus so that the preset translation model learns the invoice translation rules corresponding to the Chinese invoice corpus; controlling the preset translation model to map the invoice text information into professional word vectors in the dedicated semantic space according to the invoice translation rules; and combining all professional word vectors corresponding to the invoice text information and converting all professional word vectors into Chinese text to obtain the Chinese invoice information. The step of extracting key information from the Chinese invoice information using a key information extraction model includes: labeling the key information in the Chinese invoice information using a labeling tool to obtain an extraction model training set; training the key information extraction model using the extraction model training set until the loss function of the key information extraction model converges; inputting the optimized Chinese invoice information into the key information extraction model; and using the key information extraction module to extract the key information from the Chinese invoice information.

2. The method as described in claim 1, characterized in that, The step of inputting the foreign language invoice image into an image recognition model and outputting invoice text information and location coordinate information includes: Obtain the recognition model training set, and classify the foreign language invoice images in the recognition model training set according to the language type and format type of the foreign language invoice images; According to the distribution location and type of key information in the foreign language invoice image, the classified foreign language invoice image is labeled with data, and a clustering algorithm is used to generate clustering anchor boxes corresponding to the key information in the foreign language invoice image; The image recognition model is trained using data-annotated images of foreign language invoices containing the clustered anchor boxes, combined with an attention mechanism. When the image recognition model has been trained, the acquired foreign language invoice image will be input into the image recognition model. The image recognition model is used to extract text information, location coordinates, and confidence level from the foreign language invoice image.

3. The method as described in claim 1, characterized in that, The step of inputting the Chinese invoice information into a large language model and optimizing the fluency of the Chinese invoice information according to preset question-and-answer prompts includes: Compile multiple question-and-answer prompts for invoices according to Chinese expression habits, and use these multiple question-and-answer prompts to construct an invoice prompt template; The Chinese invoice information is input into the large language model, and the output of the large language model is optimized using the question-and-answer prompts in the invoice prompt template to obtain the Chinese invoice information with optimized language fluency.

4. The method as described in claim 1, characterized in that, The step of using the location coordinate information to draw a Chinese invoice image by arranging and optimizing the Chinese invoice information according to the Chinese invoice format includes: A Chinese invoice framework is generated according to the Chinese invoice format, wherein the Chinese invoice framework includes the data type corresponding to the Chinese invoice information and the Chinese input area; Based on the correspondence between the location coordinate information and the Chinese text input area, the Chinese document information is mapped to the data type in the Chinese document frame.

5. The method as described in claim 1, characterized in that, The step of extracting key information from the Chinese invoice information using a key information extraction model includes: Use the Docano annotation tool to annotate key information on Chinese invoice samples to obtain keyword tags; Set the initial learning rate and loss function of the key information extraction model, and input Chinese ticket samples containing the keyword tags into the key information extraction model for multiple rounds of training until the loss function converges; The Chinese invoice information is input into the key information extraction model to extract the key information of the invoice.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the artificial intelligence-based foreign language invoice image translation method as described in any one of claims 1 to 5.

7. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the artificial intelligence-based foreign language invoice image translation method as described in any one of claims 1 to 5.