An AI large model-based multilingual menu translation method

By using an AI-powered large-scale model for menu translation, combined with OCR enhancement and a catering knowledge base, the problems of accuracy, structure, and aesthetics in menu translation in the catering industry have been solved, achieving efficient and low-cost multilingual menu translation.

CN121562640BActive Publication Date: 2026-06-16XINHE LIANSHENG (CHENGDU) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XINHE LIANSHENG (CHENGDU) INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-11-27
Publication Date
2026-06-16

Smart Images

  • Figure CN121562640B_ABST
    Figure CN121562640B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multilingual menu translation methods based on AI big model, the present application is from two aspects of OCR identification enhancement, layout analysis and commodity area segmentation, to reduce the interference of picture blur and complex layout to image OCR identification;At the same time, with the aid of catering knowledge base and the longest match priority rule, the words of the structured field of commodity are recalled, and the three scenarios of complete hit, partial hit and miss are distinguished, to give knowledge base enhancement translation strategy under the three scenarios, to carry out secondary translation enhancement and completion based on this, and then obtain the translation text containing structured field and commodity label of commodity, in this way, the foregoing operation not only ensures that translation result is reliable, but also can output structured translation text with explanation;Finally, the present application also supports direct rendering on translation text, and can output menu translation file in multiple formats, based on this, the aesthetic degree and readability of translation result can be improved while reducing cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and natural language processing application technology, specifically relating to a multilingual menu translation method based on a large AI model. Background Technology

[0002] Currently, intelligent translation of menus in the catering industry mainly involves manually copying text or uploading menu images by taking photos, and then using general machine translation tools to translate the menus. This approach has the following shortcomings: (1) Low OCR recognition accuracy. When the images are blurry or the layout is complex, the translation error rate is high. Furthermore, the general model lacks a dedicated knowledge base for catering menus. In tests, it was found that low-quality translation results accounted for as much as 20% when translating between Chinese and English. (2) Lack of structured output. The translation results are mostly reference patches with messy information that cannot be automatically categorized. Information such as product names, prices, specifications, descriptions, and classifications are scattered, which is not conducive to sharing. Secondary use; (3) Lack of interpretability of results: Traditional technology only translates word by word without combining the background, ingredients and characteristics of the dishes, which makes it difficult for users to understand; (4) The results are mostly PDF or image comparison documents, which are less aesthetically pleasing and readable, and are difficult to use directly as a customer menu. If they are handed over to a design company for beautification, the price is often several thousand yuan per menu, which is costly. Therefore, based on the above shortcomings, how to provide a multilingual menu translation method based on AI large model with high translation accuracy, low cost, multi-form delivery, and explanatory translation and structured output has become an urgent problem to be solved. Summary of the Invention

[0003] The purpose of this invention is to provide a multilingual menu translation method based on a large AI model, in order to solve the problems of low translation accuracy, lack of structured output, lack of interpretability of results, high cost, and poor aesthetics and readability of existing technologies.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] Firstly, a multilingual menu translation method based on a large AI model is provided, including:

[0006] Obtain the menu image and perform OCR recognition enhancement processing on the menu image to obtain the enhanced image;

[0007] The enhanced image is subjected to layout analysis and product region segmentation to obtain product text blocks;

[0008] Multilingual recognition and script conversion are performed on the product text block to obtain the product structured fields;

[0009] The longest match priority rule based on fuzzy matching mechanism is adopted, and the catering knowledge base is used to recall terms for the product structured fields to obtain the candidate term set corresponding to the product structured fields and the recall hit type of each candidate term in the candidate term set. The recall hit type includes complete hit, no hit and partial hit.

[0010] Based on the recall hit type of each candidate term in the candidate term set, a knowledge base enhancement translation strategy is determined for each candidate term. Among them, the knowledge base enhancement translation strategy includes an enhancement translation strategy based on the catering knowledge base, an enhancement translation strategy based on the large language model, and an enhancement translation strategy based on NER and the catering knowledge base.

[0011] The translation strategy is enhanced by the knowledge base of each candidate term. Each candidate term is subjected to knowledge base enhanced translation processing to obtain the translated text of the menu image. The translated text includes the complete product name, complete product description, complete product label, product category information, price and specifications.

[0012] Based on the rendering template, the translated text is rendered using a menu template to obtain an HTML rendered menu;

[0013] Render menus using HTML to generate translated menu files in multiple formats.

[0014] Based on the above-disclosed content, after acquiring the menu image, this invention first performs OCR recognition enhancement processing to reduce the interference of image quality issues such as blurriness on OCR recognition. Then, the enhanced image undergoes layout analysis and product region segmentation (i.e., dish region segmentation) to obtain product text blocks. This layout analysis ensures the accuracy of product region recognition even in complex layouts, thereby improving the reliability of subsequent OCR recognition. Simultaneously, after obtaining the product text blocks, multilingual recognition and script conversion processing are performed to obtain product structured fields. Then, based on the longest match priority rule of the fuzzy matching mechanism and using a dedicated catering knowledge base, term recall is performed on the product structured fields, resulting in a candidate term set and the recall hit type of each candidate term. Next, based on the recall hit type of each candidate term, knowledge base-enhanced translation of the candidate terms can be performed. That is, this invention provides different knowledge base-enhanced translation strategies for different candidate term recall hit types, thereby ensuring that the translation results are both stable and reliable, and can be intelligently generated when knowledge gaps exist.

[0015] Furthermore, compared to traditional technologies that only provide a literal translation, the translated text obtained by this invention includes the complete product name, complete product description, complete product label, product category information, price, and specifications. This achieves a structured translation output, and the complete product label in the translation can be used to explain the product's (dish's) background, ingredients, and characteristics, thus making the translation interpretable and easy for users to understand. Finally, this invention also utilizes a rendering template to render the translated text, obtaining an HTML rendered menu, and based on this, generating multi-format translated menu files. This not only improves the aesthetics and readability of the translation results but also enables the translation results to support multi-format delivery, thereby reducing costs while increasing ease of use.

[0016] Through the above design, this invention reduces the interference of image blur and complex layout on OCR recognition of menu images from two aspects: OCR recognition enhancement, layout analysis, and product region segmentation, thereby improving the accuracy of OCR recognition. Simultaneously, after completing the initial product recognition, it uses a catering knowledge base and the longest match priority rule to recall words in the product's structured fields. Furthermore, it differentiates between three scenarios: complete hit, partial hit, and no hit, and provides targeted knowledge base enhancement translation strategies for each scenario. Based on this, it performs secondary translation enhancement and completion, resulting in translated text containing product structured fields and product tags. Thus, the aforementioned operations not only ensure that the translation results are stable and reliable but also output structured translated text with interpretability. Finally, this invention also supports direct rendering of the translated text and can output menu translation files in multiple formats. Based on this, it can improve the aesthetics and readability of the translation results while reducing costs and improving ease of use. Therefore, this invention provides a multilingual menu translation technology with high translation accuracy, low cost, multi-format delivery, and the ability to perform interpretable translation and structured output, making it highly suitable for large-scale application and promotion.

[0017] In one possible design, the enhanced image undergoes layout analysis and product area segmentation to obtain product text blocks, including:

[0018] The enhanced image is subjected to text region detection processing to obtain several polygonal text image blocks;

[0019] Several polygonal text image blocks are merged to obtain multiple merged image blocks, and the merged image blocks are then segmented to obtain segmented text blocks.

[0020] A clustering algorithm based on the proximity of y-centers is used to cluster multiple segmented text blocks to obtain several text block groups. Then, the segmented text blocks in each text block group are grouped by line to obtain several line groups.

[0021] Perform in-row sorting and merging on each row group to obtain several menu rows;

[0022] The layout semantics of several menu rows are processed to obtain several menu fields, including product name field, price field, specification field, product description field and product tag field;

[0023] Using the product name field as an anchor point, the specified fields in the same row and the next row of the anchor point are aggregated to obtain the product text block after aggregation. The specified fields include the price field, specification field, product tag field, and product description field.

[0024] In one possible design, a text block clustering algorithm based on the proximity of y-centers is used to cluster multiple segmented text blocks, resulting in several text block groups, including:

[0025] Select one segmented text block from multiple segmented text blocks as the specified text block;

[0026] The y-centers of the specified text block and each target text block are determined, and the similarity between the specified text block and each target text block is calculated based on the y-centers of the specified text block and each target text block. Each target text block is a segmented text block other than the specified text block among multiple segmented text blocks.

[0027] From each target text block, select target text blocks with a similarity less than a similarity threshold, and use the selected target text blocks and the specified text block to form a text block group;

[0028] The specified text block is deleted from multiple segmented text blocks, and a new segmented text block is selected from the multiple segmented text blocks until each segmented text block has been polled, resulting in several text block groups.

[0029] In one possible design, the product text block undergoes multilingual recognition and script conversion processing to obtain the product's structured fields, including:

[0030] The product text block is segmented into multiple span segments, including product name segment, price segment, specification segment, product label segment, product description segment, and product category segment.

[0031] For any span fragment among multiple span fragments, dictionary hit matching processing and character script prior language recognition processing are performed on each span fragment to obtain dictionary hit results and first language recognition results;

[0032] Using a language recognition model, language recognition processing is performed on any of the span segments to obtain a second language recognition result;

[0033] Based on the dictionary hit results, the first language recognition results, and the second language recognition results, the language type of any span fragment is determined, and after all span fragments have been polled, the language type of each span fragment is obtained.

[0034] Each span segment is normalized to obtain several normalized span segments;

[0035] Each standardized span fragment is converted from traditional Chinese to simplified Chinese characters to obtain several standardized span fragments.

[0036] The product structured fields are composed of several standardized span fragments and the language type of each standardized span fragment.

[0037] In one possible design, a longest match priority rule based on fuzzy matching is adopted, and a catering knowledge base is used to retrieve terms from the structured fields of the products to obtain a set of candidate terms corresponding to the structured fields of the products and the recall hit type of each candidate term in the candidate term set, including:

[0038] The product structured fields are processed to generate a product n-gram set and a KB query string;

[0039] Using the KB query string and the product n-gram set, term retrieval and matching processing is performed in the catering knowledge base to obtain a candidate term set;

[0040] Based on the longest match priority rule, the recall match score of each candidate term in the candidate term set is calculated;

[0041] Based on the recall matching score of each candidate term, the recall hit type of each candidate term is determined.

[0042] In one possible design, the structured fields of the products are processed to generate a product n-gram set and a KB query string, including:

[0043] The structured fields of the product are cleaned to obtain cleaned structured fields;

[0044] The cleaned structured fields are then processed by a script to obtain the KB query string;

[0045] From the cleaned structured fields, filter out the structured fields corresponding to the product names, and use the structured fields corresponding to the product names as anchor words;

[0046] Based on the KB query string and the anchor words, the product n-gram set is generated.

[0047] In one possible design, the KB query string and the product n-gram set are used to perform term retrieval and matching processing in the food and beverage knowledge base to obtain a candidate term set, including:

[0048] Using the KB query string and according to the preset hit rules, the term is matched in the catering knowledge base to obtain the first candidate term;

[0049] In the food and beverage knowledge base, each word in the n-gram set is retrieved by backward indexing to obtain several initial candidate words, and the number of times each initial candidate word hits the n-gram set is counted.

[0050] The initial candidate terms are deduplicated, and then sorted in descending order of the number of hits. The first k initial candidate terms are selected as the second candidate terms, where k is a positive integer.

[0051] The candidate term set is formed by using the first candidate term and the second candidate term.

[0052] In one possible design, the product structured fields include: product tags, where, based on the longest match priority rule, the recall match score for each candidate term in the candidate term set is calculated, including:

[0053] For any candidate term in the candidate term set, the number of target words contained in any candidate term is determined, and the text overlap is obtained based on the determined number, wherein the target words are words in the n-gram set;

[0054] Determine the anchor word overlap for any of the candidate terms;

[0055] Calculate the semantic similarity between the term tag of any candidate term and the product tag in the product structure field;

[0056] Determine the longest possible inclusion length of any candidate term in the KB query string, and calculate the ratio between the longest possible inclusion length and the total length of the KB query string;

[0057] We calculate the weighted sum of text overlap, anchor word overlap, and semantic similarity to obtain the weighted sum result.

[0058] Obtain the longest coverage weight and calculate the product between the longest coverage weight and the ratio;

[0059] The sum of the product and the weighted sum is used as the recall matching score of any candidate term. After all candidate terms in the candidate term set have been polled, the recall matching score of each candidate term is obtained.

[0060] In one possible design, a knowledge-based augmentation translation strategy is used for each candidate term, performing knowledge-based augmentation translation processing on each candidate term, including:

[0061] If the knowledge base enhancement translation strategy for any candidate term is an enhancement translation strategy based on NER and the catering knowledge base, then the NER model is used to detect the unmatched word fragments in any candidate term.

[0062] Identify the matching words of any candidate term in the catering knowledge base, and obtain the standard translation of the matching words from the catering knowledge base;

[0063] Replace the matched words with the standard translations from the aforementioned food and beverage knowledge base;

[0064] By combining the standard translation and the missing word fragments, the complete product name is obtained;

[0065] Perform tag and description completion processing on any of the candidate terms to obtain complete product tags and complete product descriptions;

[0066] From the product structured fields, obtain the product category information, price, and specifications corresponding to the complete product name;

[0067] Using the complete product name, complete product label, complete product description, and the product category information, price, and specifications corresponding to the complete product name, a complete translation of any candidate term is constructed. After all candidate terms have been polled, the translated text of the menu image is generated based on the complete translations of each candidate term.

[0068] In one possible design, the method is applied to a multilingual menu translation system based on a large AI model, wherein, after obtaining the translated text of the menu image, the method further includes:

[0069] The translated text is proofread and optimized to obtain an optimized translated text. The proofreading and optimization process is as follows: when responding to human-computer editing interaction, the modified fields in the translated text are identified, and a local retranslation model is called to retranslate the modified fields so as to fill the retranslated fields back into the translated text to obtain the optimized translated text.

[0070] The quality of the optimized translation text is assessed, and it is determined whether the quality assessment passes.

[0071] If so, then the translated text will be rendered using a menu template based on the rendering template;

[0072] Accordingly, after generating the multi-format translated menu file, the method further includes:

[0073] Obtain the review results of the translated menu file;

[0074] Determine whether the review result is a problematic file;

[0075] If so, add the translated menu file to the problem sample pool;

[0076] By utilizing a problem sample pool, feedback optimization is performed on the multilingual menu translation system based on the AI ​​large model.

[0077] Secondly, a multilingual menu translation system based on a large AI model is provided, including:

[0078] The input and preprocessing module is used to acquire the menu image and perform OCR recognition enhancement processing on the menu image to obtain the enhanced image;

[0079] The product segmentation module is used to perform layout analysis and product region segmentation on the enhanced image to obtain product text blocks;

[0080] The language recognition module is used to perform multilingual recognition and script conversion on product text blocks to obtain product structured fields.

[0081] The term recall module is used to recall terms based on the longest match priority rule of fuzzy matching mechanism and the catering knowledge base to obtain the candidate term set corresponding to the product structured field and the recall hit type of each candidate term in the candidate term set. The recall hit type includes complete hit, no hit and partial hit.

[0082] The knowledge base retrieval and translation enhancement module is used to determine the knowledge base enhancement translation strategy for each candidate term based on the recall hit type of each candidate term in the candidate term set. The knowledge base enhancement translation strategies include enhancement translation strategies based on the catering knowledge base, enhancement translation strategies based on the large language model, and enhancement translation strategies based on NER and the catering knowledge base.

[0083] The knowledge base retrieval and translation enhancement module is also used to perform knowledge base enhancement translation processing on each candidate term using the knowledge base enhancement translation strategy for each candidate term, so as to obtain the translated text of the menu image, wherein the translated text includes the complete product name, complete product description, complete product label, product category information, price and specifications;

[0084] The rendering module is used to render menu templates on the translated text based on the rendering template, resulting in an HTML rendered menu;

[0085] The rendering module is also used to render menus using HTML and generate translated menu files in multiple formats.

[0086] Thirdly, a multilingual menu translation device based on an AI large model is provided. Taking the device as an electronic device as an example, it includes a memory, a processor, and a transceiver that are connected in sequence. The memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the multilingual menu translation method based on an AI large model as described in the first aspect or any possible design in the first aspect.

[0087] Fourthly, a storage medium is provided, on which instructions are stored, which, when executed on a computer, perform the multilingual menu translation method based on the AI ​​large model as described in the first aspect or any possible design of the first aspect.

[0088] Fifthly, a computer program product containing instructions is provided, which, when executed on a computer, causes the computer to perform the multilingual menu translation method based on an AI large model as described in the first aspect or any possible design of the first aspect.

[0089] Beneficial effects:

[0090] (1) This invention reduces the interference of image blur and complex layout on OCR recognition of menu images from two aspects: OCR recognition enhancement, layout analysis and product region segmentation, thereby improving the accuracy of OCR recognition. At the same time, after completing the initial recognition of products, it uses the catering knowledge base and the longest matching priority rule to recall words in the structured fields of products. It also distinguishes between three scenarios: complete hit, partial hit and no hit, and provides targeted knowledge base enhancement translation strategies for each scenario. Based on this, it performs secondary translation enhancement and completion, and then obtains translated text containing product structured fields and product tags. In this way, the above operations not only ensure that the translation results are stable and reliable, but also output structured translated text and make it interpretable. Finally, this invention also supports direct rendering of translated text and can output menu translation files in multiple formats. Based on this, it can improve the aesthetics and readability of translation results while reducing costs and improving ease of use. Thus, this invention provides a multilingual menu translation technology with high translation accuracy, low cost, multi-form delivery, and the ability to perform interpretable translation and structured output, which is very suitable for large-scale application and promotion.

[0091] (2) This invention realizes a user-edit-driven partial retranslation + learning closed loop. In traditional machine translation, once the result is wrong, it can only be retranslated as a whole or manually modified, and cannot be fed back into the system. However, this invention allows users to edit the result, and after editing, the partial retranslation model can be called to perform partial retranslation and automatically fill in the error to ensure style consistency. At the same time, this invention can also set up feedback optimization operation, that is, when the translation file fails the review, it is added to the problem sample pool, thereby optimizing the multilingual menu translation system of the AI ​​large model. In this way, the learning closed loop of the menu translation system can be realized, and the accuracy of its translation can be continuously improved during use. Attached Figure Description

[0092] Figure 1 This is a flowchart illustrating the steps of a multilingual menu translation method based on an AI large model provided in an embodiment of the present invention.

[0093] Figure 2 This is a schematic diagram of the structure of a multilingual menu translation system based on an AI large model provided in an embodiment of the present invention;

[0094] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0095] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0096] It should be understood that although the terms first, second, etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit, without departing from the scope of the exemplary embodiments of the invention.

[0097] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0098] Example:

[0099] See Figure 1 As shown, the multilingual menu translation method based on an AI large model provided in this embodiment reduces the interference of image blur and complex layout on OCR recognition of menu images from two aspects: OCR recognition enhancement, layout analysis, and product region segmentation, thereby improving the accuracy of OCR recognition. Simultaneously, after completing the initial product (dish) recognition, it uses a catering knowledge base and the longest match priority rule to recall words in the product's structured fields. Furthermore, it distinguishes between three scenarios: complete hit, partial hit, and no hit, and provides targeted knowledge base enhancement translation strategies for each scenario. Based on these strategies, secondary translation enhancement and completion are performed to obtain a translation containing the product's structured words. The method translates the text of the product label and other related text. Thus, the aforementioned operations not only ensure that the translation results are stable and reliable, but also output structured translated text that is interpretable. Finally, this method also supports direct rendering of the translated text and can output menu translation files in multiple formats. Based on this, it can improve the aesthetics and readability of the translation results while reducing costs and increasing ease of use. Therefore, this method is very suitable for large-scale application and promotion. For example, this method can be run on the product translation end of a multilingual menu translation system based on an AI large model. Optionally, the menu translation end can be, but is not limited to, a personal computer (PC), a smartphone, or a server. It is understood that the aforementioned execution entity does not constitute a limitation on the embodiments of this application. Accordingly, the operation steps of this method can be, but are not limited to, the steps S1 to S8 below.

[0100] S1. Obtain the menu image and perform OCR recognition enhancement processing on the menu image to obtain the enhanced image. In this embodiment, the menu file is obtained first, and then the target language for translation is selected, such as translating to Simplified Chinese or English. Then, the file type of the menu file is identified, i.e., whether it is a menu image. If it is, the menu image is loaded directly; otherwise, a non-menu image prompt is given to prompt the user to re-enter the information. After obtaining the menu image, OCR enhancement processing can be performed, and the process can be, but is not limited to, the steps S11 to S13 below.

[0101] S11. Perform image denoising and image enhancement processing on the menu image to obtain a denoised and enhanced image; in this embodiment, the image is denoised and contrast stretched to obtain a denoised and enhanced image; then, image super-resolution processing can be performed for different resolutions, as shown in step S12 below.

[0102] S12. Perform image super-resolution processing on the denoised and enhanced image to obtain a resolution-enhanced image. In this embodiment, it is first determined whether the resolution of the denoised and enhanced image is less than 200 dpi. If so, image super-resolution processing is performed to enhance it to 300-400 dpi. Otherwise, the current resolution is maintained and tilt correction processing is performed directly.

[0103] Furthermore, image super-resolution is a technique that amplifies the resolution of an image, converting a low-resolution image into a high-resolution super-resolution image. It is a common method in image processing, and its principle will not be elaborated here.

[0104] After resolution enhancement processing, tilt correction can be performed, as shown in step S13 below.

[0105] S13. Perform tilt correction processing on the resolution-enhanced image to obtain the enhanced image after tilt correction processing; in this embodiment, the purpose of image tilt correction is to correct image distortion caused by shooting angle or scanning error, and ensure the accuracy of the image in subsequent text recognition; of course, tilt correction is a common technique in image processing, and its principle will not be elaborated here.

[0106] By performing image enhancement on the menu image through the aforementioned steps S11 to S13, the interference of the blurred image on subsequent OCR recognition can be reduced, thereby improving the accuracy of OCR recognition.

[0107] After image enhancement is completed, layout analysis and product region segmentation can be performed to extract product text blocks (products include dishes on the menu (such as specific dishes), staple foods, beverages (such as coffee, tea, and alcohol)) from complex layout images; the segmentation process is shown in step S2 below.

[0108] S2. Perform layout analysis and product region segmentation on the enhanced image to obtain product text blocks; in specific implementation, for example, but not limited to, the following steps S21 to S26 can be used to complete the layout analysis and product region segmentation.

[0109] S21. Perform text region detection processing on the enhanced image to obtain several polygonal text image blocks. In this embodiment, for example, but not limited to, the DBNet model or the EAST model can be used to perform text region detection, that is, to perform OCR recognition and output image blocks corresponding to single words, image blocks corresponding to single characters, or irregular polygons. Among them, the aforementioned DBNet and ESAT are both segmentation-based text detection models that can be pre-trained and deployed on the menu translation terminal and called when needed. After obtaining the polygonal text image blocks, the image blocks can be merged, and the process is shown in step S22 below.

[0110] S22. Perform image merging processing on several polygonal text image blocks to obtain multiple merged image blocks, and perform column segmentation processing on the merged image blocks to obtain segmented text blocks; in specific implementation, the merging processing is performed on text regions with a spacing < 1.2 times the character height, and the process is as shown in the following steps S22a to S22c.

[0111] S22a. For any two adjacent polygonal text image blocks among a plurality of polygonal text image blocks.

[0112] S22b. Determine whether the image spacing between any two adjacent polygonal text image blocks is less than a preset threshold, wherein the preset threshold is calculated based on the average character height of the two adjacent polygonal text image blocks; in this embodiment, the preset threshold is obtained by multiplying the average character height of the two adjacent polygonal text image blocks by 1.2 (the average character height is the average of the image heights of the two polygonal text image blocks); at the same time, if the image spacing is less than the preset threshold, the image blocks can be merged, otherwise, they are not merged, as shown in step S22c below.

[0113] S22c. If so, merge any two adjacent polygonal text image blocks to obtain a merged image block, and after all polygonal text image blocks have been polled, obtain several merged image blocks.

[0114] After merging the polygonal text image blocks through the aforementioned steps S22a to S22c, column segmentation can then be performed.

[0115] This involves calculating the column gaps between each merged image block to divide the column structure and distinguish between two-column and multi-column text blocks. Specifically, it involves calculating the distribution of x_center, and if multiple peaks and valleys are detected, it is identified as a multi-column layout. Furthermore, the x_center of a merged image block is the horizontal center of the merged image block, x_center=(x_left + x_right) / 2, where x_left and x_right are the horizontal coordinates of the leftmost and rightmost pixels of the merged image block.

[0116] Thus, when multiple values ​​of x_center are detected to be less than the preset value, column splitting is performed. Then, the split text blocks obtained from column splitting can be clustered to achieve row merging of text blocks of the same class. The process is shown in step S23 below.

[0117] S23. The y-center proximity text block clustering algorithm is used to cluster multiple segmented text blocks to obtain several text block groups, and the segmented text blocks in each text block group are grouped by row to obtain several row groups; in this embodiment, the following steps S23a to S23d can be used, but are not limited to, to perform the segmented text block clustering process.

[0118] S23a. Select one segmented text block from multiple segmented text blocks as the designated text block.

[0119] After selecting a segmented text block from multiple segmented text blocks, its similarity to the remaining segmented text blocks can be calculated, so that image clustering can be performed based on the similarity.

[0120] The similarity calculation process is shown in step S23b below.

[0121] S23b. Determine the y-center of the specified text block and each target text block, and calculate the similarity between the specified text block and each target text block based on the y-center of the specified text block and each target text block, wherein each target text block is a segmented text block other than the specified text block among multiple segmented text blocks.

[0122] In practical applications, a specified text block essentially corresponds to an image block. Furthermore, a specified text block corresponds to (x_left, y_top, x_right, y_bottom), where y_top is the upper boundary coordinate (the y-coordinate of the topmost pixel of the text block), and y_bottom is the lower boundary coordinate (the y-coordinate of the bottommost pixel of the text block). w and h represent the width and height of the specified text block. Thus, the y-center of this specified text block is (y_top + y_bottom) / 2. Of course, the calculation method for the y-center of each target text block is the same, and will not be elaborated further here.

[0123] After determining the y-center of the specified text block and the y-centers of each target text block, similarity calculation can be performed. Specifically, for a specified text block and any target text block, firstly, the absolute value of the difference between the y-center of the specified text block and the y-center of the target text block is calculated as the similarity between the two. Thus, after obtaining the similarity between the specified text block and each target text block, similarity judgment can be performed, as shown in step S23c below.

[0124] S23c. From each target text block, select target text blocks with a similarity less than a similarity threshold, and use the selected target text blocks and the specified text block to form a text block group. In specific applications, for example, but not limited to, first calculate the average height of the specified text block and any target text block; then, multiply the average height by 1.2 to obtain the similarity threshold; thus, when the similarity is less than the similarity threshold, the specified text block is considered to be similar to any target text block, and the two can be added to a text block group; based on this, after completing the selection of the remaining target text blocks in the aforementioned manner and adding the target text blocks with a similarity less than the similarity threshold to the same text block group, one clustering can be completed; then, the next segmented text block can be used as the specified text block for clustering operations until all segmented text blocks have been polled; wherein, the polling clustering process is shown in step S23d below.

[0125] S23d. Delete the specified text block from multiple segmented text blocks, and reselect a segmented text block from multiple segmented text blocks until each segmented text block has been polled, resulting in several text block groups.

[0126] Thus, through the aforementioned steps S23a to S23d, the clustering of each segmented text block can be completed, thereby obtaining several text block groups; then, the segmented text blocks in the same text block group can be grouped by line, as shown in the following steps S23e to S23g.

[0127] S23e. For any segmented text block in any text block group, calculate the y-center of any segmented text block and the distance between it and the y-centers of the remaining segmented text blocks in the same text block group; in this embodiment, the calculation process of the y-center has been described above and will not be repeated here.

[0128] After calculating the distance between the y-center of any segmented text block and the y-center of each of the other segmented text blocks in the same text block group, the proximity threshold can be calculated, as shown in step S23f below.

[0129] S23f. Determine the proximity threshold between any segmented text block and all other segmented text blocks in that group of text blocks; in specific implementation, the calculation process of the proximity threshold between any segmented text block and all other segmented text blocks is: 1.2 × average character height; if we assume that the other segmented text blocks include text block A, then the proximity threshold between any segmented text block and text block A is: the average character height of the characters in any segmented text block and text block A multiplied by 1.2, and the average character height is the average of the heights of the two aforementioned text blocks.

[0130] After calculating the nearest threshold, row grouping can be performed, as shown in step S23g below.

[0131] S23g. Based on the distance between the y-center of any segmented text block and the y-center of each other segmented text block in the same text block group, and the proximity threshold between any segmented text block and each other segmented text block in the same text block group, the segmented text blocks in the same text block group are grouped into rows, and after all text block groups have been polled, several row groups are obtained.

[0132] In practical applications, lines of segmented text blocks whose distance is less than the corresponding proximity threshold are grouped together to obtain a line group. For example, assuming that the remaining segmented text blocks include text blocks A, B, and C, and the distance between the y-center of any of the segmented text blocks and the y-center of text block A is less than the corresponding proximity threshold, and the distance between the y-center of any of the segmented text blocks and the y-center of text block C is also less than the corresponding proximity threshold, then text blocks A and C, along with the text lines in any of the segmented text blocks, are merged into a line group.

[0133] Thus, after completing the line grouping of each text block group through the aforementioned steps S23e to S23g, the text lines within the line group can be sorted and merged, as shown in step S24 below.

[0134] S24. Perform in-line sorting and merging processing on each row group to obtain several menu rows. In this embodiment, within each row group, sort by x_left (i.e., the distance between the end string and the start string in a text line) from smallest to largest. Then, merge blocks whose distance to adjacent bboxes (i.e., the text boxes containing adjacent text lines) is less than a threshold into a single string. That is, for any row group, determine the number of character intervals between the end string and the start string in each text line of the row group, and sort each text line in the row group in ascending order of the number of character intervals to obtain a sorting sequence. Then, calculate the text box distance between adjacent text lines in the sorting sequence, and merge adjacent text lines whose text box distance is less than a distance threshold into a single line. After merging, obtain the menu row corresponding to any row group, and obtain several menu rows after iterating through all row groups.

[0135] After sorting and merging the text lines within the line group, the page semantic blocks can be divided, as shown in step S25 below.

[0136] S25. Perform semantic block processing on several menu rows to obtain several menu fields, including product name, price, specifications, product description, and product tag fields. In this embodiment, examples, but not limited to, can be used to perform semantic block processing on the layout to distinguish fields such as title, product name, price, and specifications.

[0137] Specifically, the aforementioned large model and pre-trained classification model are used to perform semantic recognition and attribute recognition on the menu fields to distinguish the field types.

[0138] After completing the semantic segmentation of the layout, the product area can be constructed, as shown in step S26 below.

[0139] S26. Using the product name field as an anchor point, aggregate the specified fields in the same row and the next row of the anchor point to obtain the product text block after aggregation. The specified fields include the price field, specification field, product tag field, and product description field.

[0140] In this embodiment, the product name field is used as the anchor point, and the price, specifications, description block and tags of the same / downward row of the anchor point are aggregated to form a dish_item area (product area, that is, product text block, i.e. dish text block).

[0141] Furthermore, the following examples illustrate the use of product text blocks:

[0142] Product text block: dish_item = {name: "Wenchang Chicken Rice",price: "18 RMB",spec: "Large",desc: "Clear Broth Boiled Chicken with Dipping Sauce",bbox: [..]}.

[0143] Thus, through the aforementioned steps S21 to S26, the layout analysis and product area segmentation of the enhanced image can be completed; then, the obtained product text blocks can be processed for multilingual recognition and script conversion to obtain structured product fields, as shown in step S3 below.

[0144] S3. Perform multilingual recognition and script conversion on the product text block to obtain the product structured fields. In this embodiment, before performing multilingual recognition, it is also necessary to detect the number of products contained in the product text block. When the number of products is less than the quantity threshold, multilingual recognition and script conversion can be performed directly. When the number of products is greater than or equal to the quantity threshold, an allocation processing method is used to perform multilingual recognition and script conversion. In this way, the problem of excessively long single input can be solved, thereby improving the processing stability.

[0145] Furthermore, for example, but not limited to, the following steps S31 to S37 can be used to generate product structured fields.

[0146] S31. Perform field segmentation on the product text block to obtain multiple span fragments, wherein the multiple span fragments include product name fragment, price fragment, specification fragment, product label fragment, product description fragment, and product category fragment.

[0147] In practical implementation, this is equivalent to segmenting the product text block into fields to obtain multiple continuous character fragments, which are then used as span fragments. After obtaining multiple span fragments, multilingual recognition processing can be performed. In this embodiment, each span fragment is individually subjected to three-level language recognition to generate language labels and confidence scores. The process is shown in steps S32 to S34 below.

[0148] S32. For any span fragment among multiple span fragments, dictionary hit matching processing and character script prior language recognition processing are performed on each span fragment to obtain dictionary hit results and first language recognition results. In specific implementation, the characters of different languages ​​usually have a clear Unicode script range (Unicode script range refers to a set of continuous code points allocated for a specific writing system or script in the Unicode encoding standard), which is used to quickly filter out impossible languages ​​and narrow down the candidate language range. Therefore, by performing character script prior processing on any span fragment, the first language recognition result can be obtained, that is, the probability of the corresponding language type can be obtained.

[0149] Similarly, by matching the span with entries in the knowledge base dictionary, when an entry is matched, the candidate language of any span segment is directly labeled, thereby improving the recognition accuracy of professional entries. In other words, the probability of matching the language of an entry is used as the language recognition result of any span segment.

[0150] After completing the aforementioned level two language recognition, the language recognition model can be used to perform level three language recognition, as shown in step S33 below.

[0151] S33. Using a language recognition model, perform language recognition processing on any of the span segments to obtain a second language recognition result; in this embodiment, for example, but not limited to, using the LID language recognition model (LanguageIdentification) to perform language recognition on any of the span segments, that is: the model outputs the probability distribution of each language on the input span text, and takes the language with the highest probability as the prediction result, focusing on solving the boundary cases that cannot be covered by scripts and dictionaries, especially multilingual mixing and transliterated words.

[0152] After obtaining the Level 3 language recognition result through the aforementioned steps, the language type of any span fragment can be obtained based on this, as shown in step S34 below.

[0153] S34. Based on the dictionary hit results, the first language recognition results, and the second language recognition results, determine the language type of any span fragment, and obtain the language type of each span fragment after all span fragments have been polled; in this embodiment, the one with the highest probability among the three language recognition results is taken as the final language type of any span fragment; then, the three are weighted and summed to obtain the confidence score of the final language type.

[0154] The following is an example of the aforementioned multi-level language recognition:

[0155] span = "Wenchang Chicken Rice"; Character script prior: Chinese characters → conf_script=0.9; Dictionary hit: "chicken rice" hit → conf_dict=0.8; LID model: fastText predicts zh probability 0.92 → conf_model=0.92.

[0156] Confidence score calculation (comprehensive score): conf = w1 × conf_script + w2 × conf_dict + w3 × conf_model, that is, conf = 0.3×0.9 + 0.4×0.8 + 0.3×0.92 ≈ 0.87; output (lang=zh,conf=0.87).

[0157] After obtaining the language type of each span fragment, script conversion can be performed, as shown in steps S35 and S36 below.

[0158] S35. Normalize each span fragment to obtain several normalized span fragments; in this embodiment, the aforementioned normalization process may include, but is not limited to, NFKC processing (which is a Unicode normalization form used in text processing to convert character sequences into a unified standard form for comparison and processing, i.e., full-width and half-width characters, compatible characters, superscript / subscript standardization processing), currency, unit, and other format standardization processing.

[0159] After obtaining the normalized span fragment, the simplified / traditional Chinese conversion can be performed, as shown in step S36 below.

[0160] S36. Perform traditional and simplified Chinese character conversion on each standardized span fragment to obtain several standardized span fragments; in this embodiment, for example, but not limited to, using a simplified-traditional dictionary mapping table (OpenCC, HanDian Data) to internally convert the input text into another form and perform traditional-simplified conversion to avoid knowledge base matching failure due to traditional-simplified differences.

[0161] After completing the conversion between traditional and simplified Chinese characters, the product structured fields can be generated based on this, as shown in step S37 below.

[0162] S37. Based on several standardized span fragments and the language type of each standardized span fragment, a product structured field is formed; in this embodiment, the product structured field includes the aforementioned product name, price, specifications, product label (i.e., dish label), product description (dish description), product category (dish category), language type, and confidence score.

[0163] After completing the initial OCR recognition through the aforementioned steps S31 to S17 and obtaining the product structured fields, knowledge base-enhanced translation can be performed, as shown in steps S4 to S6 below.

[0164] S4. The longest match priority rule based on fuzzy matching mechanism is adopted, and the catering knowledge base is used to recall terms for the product structured fields to obtain the candidate term set corresponding to the product structured fields and the recall hit type of each candidate term in the candidate term set. The recall hit type includes complete hit, no hit and partial hit.

[0165] In this embodiment, for example, but not limited to, the following steps S41 to S44 can be used to recall terms and determine the recall hit type of each candidate term.

[0166] S41. Perform query construction processing on the product structured fields to obtain the product n-gram set and KB query string; in specific implementation, the query construction is to convert the product structured fields into a normalized query key and n-gram set for KB retrieval to improve the stability of recall; the process can be, but is not limited to, as shown in steps S41a to S41d below.

[0167] S41a. Perform text cleaning on the product structured fields to obtain cleaned structured fields. In specific implementation, the aforementioned text cleaning process may include, but is not limited to, NFKC standardization (full-width and half-width characters, compatible characters, superscripts / subscripts), dotted lines, case conversion, and whitespace removal. After completing the text cleaning and obtaining the cleaned structured fields, script conversion processing can be performed, as shown in step S41b below.

[0168] S41b. Perform script conversion processing on the cleaned structured fields to obtain the KB query string; in specific applications, the script conversion is used for matching. If lang (language type) ∈ {zh-Hant, zh-Hans} (representing Traditional Chinese), then use OpenCC to perform Traditional-Simplified conversion to the KB-side main script to generate the query_norm, i.e., the KB query string. Then, anchor words can be selected, as shown in step S41c below.

[0169] S41c. From the cleaned structured fields, filter out the structured fields corresponding to the product names, and use the structured fields corresponding to the product names as anchor words; in this embodiment, in addition to using the structured fields corresponding to the product names as anchor words, proper nouns (brands / place names) will also be retained for "partial hits" later.

[0170] After obtaining the anchor words, the product n-gram set can be generated by combining them with the aforementioned KB query string, as shown in step S41d below.

[0171] S41d. Based on the KB query string and the anchor words, generate the product n-gram set; in this embodiment, for example, but not limited to, the Longest-First algorithm can be used to generate the aforementioned product n-gram set; at the same time, corresponding prompt words can be set to generate a standardized query key and n-gram set for knowledge base retrieval.

[0172] Thus, by following the steps S41a to S41d, the query construction of the product structured fields can be completed; then, term retrieval can be performed, as shown in step S42 below.

[0173] S42. Using the KB query string and the product n-gram set, perform term recall and matching processing in the catering knowledge base to obtain a candidate term set; in this embodiment, for example, but not limited to, the following steps S42a to S42d can be used to generate candidate terms.

[0174] S42a. Using the KB query string and according to the preset matching rules, perform term matching in the catering knowledge base to obtain the first candidate term; in this embodiment, use "query_norm" to search for name_zh / name_en in the KB → if a match is found, it is taken as the first candidate term, that is, search for name_zh (Chinese name field) and name_en (English name field) in the catering knowledge base. If a name_zh or name_en that completely matches the KB query string is found in the knowledge base, then the knowledge term corresponding to name_zh or name_en is taken as the first candidate term.

[0175] After completing the direct lookup of the precise product name (i.e., the aforementioned step S42a), an n-gram inverted index recall can be performed, as shown in step S42b below.

[0176] S42b. In the catering knowledge base, each word in the n-gram set is used for inverted index retrieval to obtain several initial candidate words, and the number of times each initial candidate word hits the n-gram set is counted. In this embodiment, the n-grams set is traversed to obtain candidate KBs (initial candidate words), the number of hits and the longest hit length of each candidate KB are counted, and a context blocking strategy is used to match words during each inverted index retrieval.

[0177] The context blocking strategy is as follows: if `lang_hint` (language type) is valid, entries in the same language are prioritized; if `category_hint` (category information) is valid, KBs with inconsistent categories are filtered out. Specifically, if the language detection result of the query string is reliable (i.e., `lang_hint` is valid), for each candidate KB, its language label is checked to see if it is the same as `lang_hint`. If they are the same, the KB is retained; otherwise, it is discarded. Similarly, if the category information of the query string is reliable (i.e., `category_hint` is valid), for each candidate KB, its category label is checked to see if it is the same as `category_hint`. If they are the same, the KB is retained; otherwise, it is discarded.

[0178] After the inverted index recall is completed, deduplication and sorting can be performed, as shown in step S42c below.

[0179] S42c. Perform deduplication on a number of initial candidate terms, and sort the deduplicated initial candidate terms in descending order of hit count, so that the first k initial candidate terms in the sorting are used as the second candidate terms, where k is a positive integer; in this embodiment, k is 3 for example.

[0180] After obtaining the second candidate term, it can be combined with the first candidate term to generate a candidate term set, as shown in step S42d below.

[0181] S42d. Using the first candidate term and the second candidate term, form the candidate term set.

[0182] Through the aforementioned steps S42a to S42d, word recall can be completed based on the catering knowledge base; then, candidate words can be scored to determine the subsequent knowledge base enhancement translation strategy for each candidate word.

[0183] The recall matching score calculation process is shown in step S43 below.

[0184] S43. Based on the longest match priority rule, calculate the recall match score for each candidate term in the candidate term set; in specific applications, for example, but not limited to, the following steps S43a to S43g can be used to calculate the recall match score.

[0185] S43a. For any candidate term in the candidate term set, determine the number of target words contained in the candidate term, and derive the text overlap based on the determined number, wherein the target words are words in the n-gram set; in this embodiment, the number of words in the n-gram set contained in any candidate term is counted, and then the number of target words contained is divided by the total number of words in the n-gram set to obtain the text overlap.

[0186] After obtaining the text overlap, the anchor word overlap can be calculated, as shown in step S43b below.

[0187] S43b. Determine the anchor word overlap of any candidate term; in this embodiment, it is to determine whether any candidate term contains a specified anchor word, wherein the specified anchor word is a keyword containing cooking method or category. If the specified anchor word is contained, the anchor word overlap is 1, otherwise it is 0.

[0188] After obtaining the anchor word overlap, the semantic similarity of the tags can be calculated, as shown in step S43c below.

[0189] S43c. Calculate the semantic similarity between the term tag of any candidate term and the product tag in the product structured field; in this embodiment, the term tag and product tag can be semantically vectorized (e.g., using the BERT model for vectorization), and then the vector similarity can be calculated to serve as the semantic similarity; at the same time, the intersection ratio of the term tags extracted from the text and the knowledge base tags can also be calculated to serve as the semantic similarity.

[0190] After obtaining the semantic similarity, the longest inclusion length can be determined, as shown in step S43d below.

[0191] S43d. Determine the longest inclusion length of any candidate term in the KB query string and calculate the ratio between the longest inclusion length and the total length of the KB query string. In specific implementation, find the longest length of consecutive characters contained in query_norm among the matching candidate names (i.e., any candidate term). For example, if the KB query string is "Hainan Wenchang Chicken Rice" and the candidate term is "Wenchang Chicken Rice", then the longest inclusion length is 4. After obtaining the longest inclusion length, the total length of the candidate term and the KB query string can be calculated. Then, the recall matching score can be calculated by combining the aforementioned semantic similarity, hit ratio, and text overlap. The process is shown in steps S43e to S43g below.

[0192] S43e. The text overlap, anchor word overlap and semantic similarity are weighted and summed to obtain the weighted sum result; in this embodiment, the weights of the aforementioned text overlap, anchor word overlap and semantic similarity can be set to 0.25, 0.2 and 0.1 respectively, but are not limited to.

[0193] After obtaining the weighted summation result, the longest coverage weight can be obtained, as shown in step S434f below.

[0194] S43f. Obtain the longest coverage weight and calculate the product between the longest coverage weight and the ratio; in this embodiment, the longest coverage weight is between [0.05, 0.1].

[0195] Thus, after obtaining the product between the longest coverage weight and the ratio, the summation is performed to obtain the recall matching score of any candidate term, as shown in step S43g below.

[0196] S43g. The sum of the product and the weighted sum is used as the recall matching score of any candidate term, and after all candidate terms in the candidate term set have been polled, the recall matching score of each candidate term is obtained.

[0197] Therefore, after obtaining the recall matching score of each candidate term through the aforementioned steps S43a to S43g, the recall hit type of each candidate term can be determined, as shown in step S44 below.

[0198] S44. Based on the recall matching score of each candidate term, determine the recall hit type of each candidate term; in this embodiment, if the recall matching score of any candidate term is greater than or equal to the first threshold, the recall hit type of any candidate term can be determined to be a complete hit; if the recall matching score of any candidate term is greater than the second threshold and less than the first threshold, the recall hit type of any candidate term is determined to be a partial hit; if the recall matching score of any candidate term is less than or equal to the second threshold, the recall hit type of any candidate term is determined to be a miss.

[0199] Furthermore, the first threshold can be set to 0.9, and the second threshold can be set to 0; of course, the specific values ​​of the two thresholds can be set according to actual use, and no specific limitation is made here.

[0200] Therefore, after obtaining the recall hit type of each candidate term through the aforementioned steps S41 to S44 and their corresponding sub-steps, the knowledge base enhancement translation strategy for each candidate term can be determined, as shown in step S5 below.

[0201] S5. Based on the recall hit type of each candidate term in the candidate term set, determine the knowledge base enhancement translation strategy for each candidate term. The knowledge base enhancement translation strategy includes an enhancement translation strategy based on the catering knowledge base, an enhancement translation strategy based on the large language model, and an enhancement translation strategy based on NER and the catering knowledge base. In this embodiment, a complete hit corresponds to the enhancement translation strategy based on the catering knowledge base, a no hit corresponds to the enhancement translation strategy based on the large language model, and a partial hit corresponds to the enhancement translation strategy based on NER and the catering knowledge base.

[0202] Once the knowledge base enhancement translation strategy for each candidate term is obtained, knowledge base enhancement translation can be performed, as shown in step S6 below.

[0203] S6. Utilize the knowledge base enhancement translation strategy for each candidate term to perform knowledge base enhancement translation processing on each candidate term to obtain the translated text of the menu image, wherein the translated text includes the complete product name, complete product description, complete product label, product category information, price, and specifications.

[0204] In practical applications, let's take any candidate term as an example to illustrate its knowledge base enhancement translation process:

[0205] Specifically, when the recall hit type of any candidate term is partial hit, an enhanced translation strategy based on NER (Named Entity Recognition) and a catering knowledge base is adopted to perform knowledge base enhanced translation. Specifically, through the catering knowledge base alias detection and proper name recognition hit method, semantic type determination is performed on the non-hit text, and information such as brand name / place name / proper noun / personal name is marked; the target language of the hit part, that is, the general target language name / description / tag, is concatenated, and secondary integration is performed to generate and structure the output (the personalized parts such as brand, place name, and proper noun are retained, and only the general parts are supplemented).

[0206] Optionally, the enhanced translation processing procedure based on the aforementioned NER and catering knowledge base enhanced translation strategy is shown in steps S61 to S67 below.

[0207] S61. If the knowledge base enhancement translation strategy for any candidate term is an enhancement translation strategy based on NER and the catering knowledge base, then the NER model is used to detect the missing word fragments in any candidate term. In this embodiment, the named entity model is used to detect the missing fragments. These fragments are usually not translated or are only transliterated and are directly retained in the target text, such as brand names (HaiXlao, XX Cola), place names, personal names (Zhang San Chicken Rice, Zhang San) and other proper nouns.

[0208] For example, suppose we use "Hainan Wenchang Chicken Rice" from the n-gram set to perform a KB query. The corresponding candidate term is "chicken rice". That is, any candidate term is chicken rice. Then, the frequency of the missing terms is Hainan and Wenchang. Of course, the above example is just an example and this embodiment is not limited to it.

[0209] In this embodiment, the NER model is used to detect the aforementioned proper nouns, which is a common technique in natural language processing, and its principle will not be elaborated here. After obtaining the unmatched word fragments in any candidate word, the enhanced translation of the matched words can be performed, and the process is shown in step S62 below.

[0210] S62. Determine the matching word in the catering knowledge base for any candidate term, and obtain the standard translation of the matching word from the catering knowledge base; after obtaining the standard translation of the matching word from the catering knowledge base, word replacement can be performed, as shown in step S63 below.

[0211] S63. Replace the matched words with the standard translations in the catering knowledge base; after replacing the matched words with their corresponding standard translations, word splicing can be performed, as shown in step S64 below.

[0212] S64. Combine the standard translation and the missing word fragments to obtain the complete product name; in this embodiment, after obtaining the complete product name, the label and description can be completed, as shown in step S65 below.

[0213] S65. Perform tag and description completion processing on any candidate term to obtain complete product tags and complete product descriptions; in specific applications, for example, but not limited to, custom tags can be used, that is, the product tags corresponding to the complete product name are matched from the tag database to serve as complete product tags; wherein, the tag database pre-stores product tags corresponding to different product names and is uploaded by users.

[0214] Meanwhile, if the number of product tags retrieved from the tag database is less than 3, then the LLM model can be used to refer to the tags in the catering knowledge base to complete the tags. That is, the complete product name and corresponding tag information in the catering knowledge base are used as prompt words, thereby using LLM to complete the tags.

[0215] Similarly, the process for completing product descriptions is the same as that for completing tags, prioritizing custom descriptions (i.e., searching the product description database based on the complete product name); if a product description corresponding to the product name cannot be found, the LLM model is used to generate 1-2 sentences of description based on the key information of the KB, with a length of less than or equal to 120 characters.

[0216] After obtaining the complete product label and complete product description, the complete translation corresponding to any candidate term can be formed by combining the product category, price and specifications information in the product structured fields, as shown in steps S66 and S67 below.

[0217] S66. Obtain the product category information, price, and specifications corresponding to the complete product name from the product structured fields.

[0218] S67. Using the complete product name, complete product label, complete product description, and the product category information, price, and specifications corresponding to the complete product name, construct the complete translation corresponding to any candidate term. After polling all candidate terms, generate the translated text of the menu image based on the complete translation of each candidate term.

[0219] The following example illustrates the steps S61 to S67 mentioned above.

[0220] Example: Source: Hainan Wenchang Chicken Rice, signboard; KB contains an entry (i.e., any candidate term): "Chicken Rice"; Match result: Hit part = "Chicken Rice", Missing part = "Hainan Wenchang".

[0221] Processing: Hit part → "Chicken Rice"; Missed part (place name) → "Hainan Wenchang".

[0222] Output: Product Name: "Hainan Wenchang Chicken Rice"; Tags: [Signature, RiceDish, Chicken]; Description: A regional specialty chicken rice from Wenchang, Hainan.

[0223] Of course, the product category information, price and specifications corresponding to the above examples will not be listed again; only the main product name, label and description will be listed.

[0224] By following the steps S61 to S67 described above, the knowledge base enhancement translation of the candidate terms that were not matched can be completed.

[0225] Furthermore, when the recall hit type of any candidate term is a complete hit, an enhanced translation strategy based on the catering knowledge base is adopted to perform knowledge base enhanced translation. Specifically, the standard translation (which represents the complete product name, i.e., the complete dish name) and the corresponding description (i.e., the complete product description), ingredients, cooking methods, and flavor tags (i.e., the complete product tags) that are already available in the catering knowledge base are directly taken, and the tag structure template is referenced to perform knowledge base enhanced translation processing, thereby obtaining the complete translation corresponding to any candidate term. At the same time, the product category information, price, and specifications corresponding to the complete product name are obtained from the product structured fields and recorded together in the tag structure template, thereby obtaining the complete translation.

[0226] Furthermore, the label structure template can be: [Main Ingredients / Preparation Method] + [Flavor / Taste] + [Optional Information (Spiciness / Vegetarian / Allergen Tips)], 1–2 sentences, and the number of characters should be less than or equal to 120.

[0227] Furthermore, when the recall hit type of any candidate term is no hit, an enhanced translation strategy based on the large language model is adopted to perform knowledge base enhancement translation. The process is as follows: call the large language model (LLM) to perform translation / transliteration, and automatically generate a short description and label.

[0228] Therefore, through the aforementioned steps S4 to S6, this embodiment addresses the problem that in multilingual menu translation scenarios, product names are often composed of a general part (such as "chicken rice") and a proper part (such as "Hainan Wenchang"). If only the shortest or partial match is relied upon, key information may be missing, leading to generalized or distorted translations. This embodiment innovatively proposes and adopts a fuzzy matching optimization mechanism to enhance translation. That is, by combining the longest match priority principle with knowledge base-enhanced translation, the accuracy and professionalism of the translation can be improved.

[0229] The optimal matching priority rule is:

[0230] Knowledge base entries may have hierarchical relationships: for example, short entries → general concepts (such as "chicken rice"); long entries → specific professional terms (such as "Hainanese Wenchang Chicken Rice"). Therefore, when recalling and ranking candidates, a length weighting factor (i.e., the longest coverage weight) is introduced. This strategy ensures that long entries that cover more characters / words are output first, thereby avoiding the system only hitting "chicken rice" and missing key information such as "Hainanese Wenchang".

[0231] After candidate recall and ranking are completed, knowledge base-enhanced translation can be introduced. That is, after determining the hit type, the system executes a differentiated enhancement strategy. For example, for a complete hit, the standard translation, description and tags in the knowledge base are directly used to ensure terminology consistency and controllability. For a partial hit, NER + rules are combined to identify brand names, place names and proper nouns, while retaining personalization. At the same time, the knowledge base is used to complete the common parts, and LLM is called to complete the tags and descriptions when necessary. Finally, for no hit, LLM translation or transliteration is called to automatically generate a short description and tags, and the results are regularized and quality evaluated.

[0232] Thus, the translation enhancement mechanism relying on knowledge base and rules ensures the preservation of personalized information and the standardized translation of general parts, achieving a structured splicing of proper names and general information.

[0233] After obtaining the translated text of the menu image based on the aforementioned step S6, the menu template can be rendered, as shown in step S7 below.

[0234] S7. Based on the rendering template, the translated text is rendered using a menu template to obtain an HTML rendered menu. In specific applications, before executing the menu template rendering, this embodiment also sets up a proofreading and secondary optimization step, the process of which is as follows: (1) The translated text is proofread and optimized to obtain an optimized translated text. The proofreading and secondary optimization process is as follows: when responding to the human-computer editing interaction operation, the modified fields in the translated text are determined, and the local retranslation model is called to retranslate the modified fields so as to fill the retranslated fields back into the translated text to obtain the optimized translated text; (2) The optimized translated text is evaluated for quality and it is determined whether the quality evaluation is passed; (3) If so, the translated text is rendered using a menu template based on the rendering template.

[0235] In practice, the system outputs the translated text and displays it to the user, who can then compare the original text with the translated text to determine if the translation is correct. If modifications are needed, the user selects the field to be modified. In response to the human-computer editing interaction, the system can locate the modified field and generate a minimal patch. Then, LLM partial retranslation (preserving style / terminology / length constraints) can be invoked to retranslate the modified field. Finally, the field is backfilled to obtain the optimized translated text.

[0236] At the same time, all modification records are stored for use in optimizing the version model of subsequent systems.

[0237] Thus, after completing the proofreading and secondary optimization steps, a quality assessment can be conducted to generate a quality analysis report.

[0238] Specifically, the dimensions of quality assessment include: accuracy (whether the main ingredients, methods, and units are fully covered, and whether the units / specifications / prices are maintained or correctly converted); authenticity (whether internationally recognized translations are used, such as Mapo Tofu instead of "Pockmarked Granny's Bean Curd," and different internationally recognized translations can be preset); expressiveness (whether specified types of descriptions have been added, and specific types of descriptions to increase attractiveness can also be preset, and type matching can be performed when using them); and clarity (whether the fields and tags are complete, i.e., whether the tags are less than or equal to 3).

[0239] Different scores can be assigned to the aforementioned accuracy. For example, if all accuracy assessment dimensions are met—that is, the main ingredients, methods, and units are fully covered, and the units / specifications / prices are maintained or correctly converted—the accuracy score is 5 points. Conversely, if the main ingredients, methods, and units are not covered, and the units, specifications, and prices are all incorrect, the score is 0 points. Excluding the two cases mentioned above, the score is 3 points. Similarly, for authenticity, if internationally accepted translations are used, the score is 5 points; otherwise, it is 0 points. For expressiveness, if a specified type of description is matched, the score is 5 points; otherwise, it is 0 points. Finally, for clarity, if all fields and tags are complete, the score is 5 points; if all fields and tags are completely missing, the score is 0 points; excluding the two cases mentioned above, the score is 3 points.

[0240] Finally, the quality assessment result is obtained by weighted summation of the aforementioned four dimensions, with the weights of the four dimensions set to 0.35, 0.25, 0.2, and 0.2, respectively.

[0241] Meanwhile, when the total score is greater than or equal to 4.5, template rendering can be performed directly without manual verification; when the total score is greater than or equal to 3 and less than 4.5, template rendering can be performed directly, but a prompt message for manual spot check needs to be output; when the total score is less than 3, the quality assessment fails and manual review is required, i.e., the result is output to the review terminal.

[0242] Furthermore, when the quality assessment result is passed (i.e., the total score is greater than or equal to 3), menu rendering can then be performed. That is, based on the rendering template, the translated text is rendered using the menu template to obtain an HTML rendered menu. The rendering template can be customized and uploaded by the user. Thus, this embodiment has multiple rendering templates (i.e., it supports customization of different brand styles). When using it, you can select the one that suits your needs.

[0243] After obtaining the HTML rendered menu, the menu can be exported, as shown in step S8 below.

[0244] S8. Render the menu using HTML to generate translated menu files in multiple formats; in this embodiment, the translated menu files may include, but are not limited to, A4-sized PNG / JPG files to meet printing and digital display requirements, as well as electronic links that support QR code access; thus, this embodiment can support multi-format delivery and application expansion.

[0245] Furthermore, in this embodiment, after generating multi-format translated menu files, a review step is also set, namely: obtaining the review result of the translated menu file; then, determining whether the review result is a problem file; wherein, if so, the translated menu file is added to the problem sample pool, so as to use the problem sample pool to complete the feedback optimization of the multilingual menu translation system based on the AI ​​big model.

[0246] Specifically, the operations staff reviews the generated results. Samples that pass the review are automatically added to the knowledge base, while samples that fail are added to the problem database. By optimizing the constructed prompts and improving the tag trimming logic, the accuracy of the system is improved in the next round. In this way, a closed-loop system of "recognition → translation → output → feedback → optimization" is formed, thereby continuously improving the accuracy of the entire translation system during use.

[0247] Therefore, the multilingual menu translation method based on an AI large model, as described in detail in steps S1 to S8 above, has the following significant advantages compared with the prior art:

[0248] (1) Value to the industry

[0249] Cost reduction: Taking a menu with 30 SKUs as an example, the traditional method would cost several hundred to several thousand yuan, while this solution costs only about 0.3 yuan.

[0250] Improved efficiency: The entire translation and generation process is shortened to a few minutes, making it suitable for large-scale use by small and medium-sized catering businesses.

[0251] Multilingual adaptation: The system supports multiple languages ​​for recognition, and the target language output supports Simplified Chinese, Traditional Chinese, and English, meeting the rigid demand for cross-language menus both domestically and internationally.

[0252] Communication and Brand Value: The generated results are aesthetically pleasing and standardized, helping catering businesses enhance their brand image and the experience of international customers.

[0253] (2) Value of technology

[0254] Knowledge Base Enhanced LLM (RAG): Through a mechanism of "fuzzy matching + knowledge base priority + LLM completion", it ensures that the translation results are both accurate and idiomatic. At the same time, it distinguishes between three scenarios: complete hit, partial hit, and no hit, and calls the knowledge base or LLM accordingly. This industry-specific RAG mechanism ensures that the translation results are both stable and controllable, and can be intelligently generated when there are knowledge gaps.

[0255] Sustainable learning: The system has self-learning capabilities. As users use it, the knowledge base expands continuously, and the model accuracy gradually improves. If the results of traditional machine translation are wrong, they can only be retranslated as a whole or manually modified, and cannot be fed back into the system. This invention allows users to edit single / batch results. The system will call LLM for partial retranslation and automatically fill in the missing information to ensure style consistency. It will also store user modification records to form a learning loop of "knowledge base supplementation + prompt word optimization".

[0256] Structured output: Unlike traditional translations that only provide text, this invention provides complete structured fields, facilitating subsequent recommendations, filtering, and visualization. Furthermore, the translation results can be exported as PNG / JPG (A4 printable) or generated as an online electronic menu link with a single click, supporting QR code sharing. This multi-format delivery method greatly enhances practical application value, making it particularly suitable for low-cost promotion by small and medium-sized businesses.

[0257] Modular and scalable: This process can be adapted to different categories such as tea drinks, wine lists, and desserts, and can also be used for cross-border e-commerce product catalog translation.

[0258] like Figure 2 As shown, the second aspect of this embodiment provides a hardware system for implementing the multilingual menu translation method based on an AI large model as described in the first aspect of the embodiment, comprising:

[0259] The input and preprocessing module is used to acquire the menu image and perform OCR recognition enhancement processing on the menu image to obtain the enhanced image.

[0260] The product segmentation module is used to perform layout analysis and product region segmentation on the enhanced image to obtain product text blocks.

[0261] The language recognition module is used to perform multilingual recognition and script conversion on product text blocks to obtain product structured fields.

[0262] The term recall module is used to recall terms based on the longest match priority rule of fuzzy matching mechanism and the catering knowledge base to obtain the candidate term set corresponding to the product structured field and the recall hit type of each candidate term in the candidate term set. The recall hit type includes complete hit, no hit and partial hit.

[0263] The knowledge base retrieval and translation enhancement module is used to determine the knowledge base enhancement translation strategy for each candidate term based on the recall hit type of each candidate term in the candidate term set. The knowledge base enhancement translation strategies include enhancement translation strategies based on the catering knowledge base, enhancement translation strategies based on the large language model, and enhancement translation strategies based on NER and the catering knowledge base.

[0264] The knowledge base retrieval and translation enhancement module is also used to utilize the knowledge base enhancement translation strategy for each candidate term to perform knowledge base enhancement translation processing on each candidate term to obtain the translated text of the menu image, wherein the translated text includes the complete product name, complete product description, complete product label, product category information, price and specifications.

[0265] The rendering module is used to render menu templates on the translated text based on the rendering template, resulting in an HTML rendered menu.

[0266] The rendering module is also used to render menus using HTML and generate translated menu files in multiple formats.

[0267] The working process, working details and technical effects of the system provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0268] like Figure 3 As shown, the third aspect of this embodiment provides a multilingual menu translation device based on an AI large model. Taking the device as an electronic device as an example, it includes: a memory, a processor, and a transceiver that are connected in sequence. The memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the multilingual menu translation method based on an AI large model as described in the first aspect of the embodiment.

[0269] The working process, working details and technical effects of the electronic device provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0270] The fourth aspect of this embodiment provides a storage medium that stores instructions containing the multilingual menu translation method based on an AI large model as described in the first aspect of the embodiment. That is, the storage medium stores instructions that, when executed on a computer, perform the multilingual menu translation method based on an AI large model as described in the first aspect of the embodiment.

[0271] The storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or memory sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0272] The working process, working details and technical effects of the storage medium provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.

[0273] The fifth aspect of this embodiment provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the multilingual menu translation method based on an AI large model as described in the first aspect of this embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0274] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multilingual menu translation method based on an AI large model, characterized in that, The method includes: Obtain the menu image and perform OCR recognition enhancement processing on the menu image to obtain the enhanced image; The enhanced image is subjected to layout analysis and product region segmentation to obtain product text blocks; Multilingual recognition and script conversion are performed on the product text block to obtain the product structured fields; The longest match priority rule based on fuzzy matching mechanism is adopted, and the catering knowledge base is used to recall terms for the product structured fields to obtain the candidate term set corresponding to the product structured fields and the recall hit type of each candidate term in the candidate term set. The recall hit type includes complete hit, no hit and partial hit. Based on the recall hit type of each candidate term in the candidate term set, a knowledge base enhancement translation strategy is determined for each candidate term. Among them, the knowledge base enhancement translation strategy includes an enhancement translation strategy based on the catering knowledge base, an enhancement translation strategy based on the large language model, and an enhancement translation strategy based on NER and the catering knowledge base. The translation strategy is enhanced by the knowledge base of each candidate term. Each candidate term is subjected to knowledge base enhanced translation processing to obtain the translated text of the menu image. The translated text includes the complete product name, complete product description, complete product label, product category information, price and specifications. Based on the rendering template, the translated text is rendered using a menu template to obtain an HTML rendered menu; Render menus using HTML to generate translated menu files in multiple formats.

2. The method according to claim 1, characterized in that, The enhanced image is subjected to layout analysis and product region segmentation to obtain product text blocks, including: The enhanced image is subjected to text region detection processing to obtain several polygonal text image blocks; Several polygonal text image blocks are merged to obtain multiple merged image blocks, and the merged image blocks are then segmented to obtain segmented text blocks. A clustering algorithm based on the proximity of y-centers is used to cluster multiple segmented text blocks to obtain several text block groups. Then, the segmented text blocks in each text block group are grouped by line to obtain several line groups. Perform in-row sorting and merging on each row group to obtain several menu rows; The layout semantics of several menu rows are processed to obtain several menu fields, including product name field, price field, specification field, product description field and product tag field; Using the product name field as an anchor point, the specified fields in the same row and the next row of the anchor point are aggregated to obtain the product text block after aggregation. The specified fields include the price field, specification field, product tag field, and product description field.

3. The method according to claim 2, characterized in that, The text block clustering algorithm based on the similarity of y-centers is used to cluster multiple segmented text blocks, resulting in several text block groups, including: Select one segmented text block from multiple segmented text blocks as the specified text block; The y-centers of the specified text block and each target text block are determined, and the similarity between the specified text block and each target text block is calculated based on the y-centers of the specified text block and each target text block. Each target text block is a segmented text block other than the specified text block among multiple segmented text blocks. From each target text block, select target text blocks with a similarity less than a similarity threshold, and use the selected target text blocks and the specified text block to form a text block group; The specified text block is deleted from multiple segmented text blocks, and a new segmented text block is selected from the multiple segmented text blocks until each segmented text block has been polled, resulting in several text block groups.

4. The method according to claim 1, characterized in that, Multilingual recognition and script conversion are performed on the product text block to obtain the product structured fields, including: The product text block is segmented into multiple span segments, including product name segment, price segment, specification segment, product label segment, product description segment, and product category segment. For any span fragment among multiple span fragments, dictionary hit matching processing and character script prior language recognition processing are performed on each span fragment to obtain dictionary hit results and first language recognition results; Using a language recognition model, language recognition processing is performed on any of the span segments to obtain a second language recognition result; Based on the dictionary hit results, the first language recognition results, and the second language recognition results, the language type of any span fragment is determined, and after all span fragments have been polled, the language type of each span fragment is obtained. Each span segment is normalized to obtain several normalized span segments; Each standardized span fragment is converted from traditional Chinese to simplified Chinese characters to obtain several standardized span fragments. The product structured fields are composed of several standardized span fragments and the language type of each standardized span fragment.

5. The method according to claim 1, characterized in that, The longest match priority rule based on fuzzy matching is adopted, and a catering knowledge base is used to retrieve terms from the structured fields of products, so as to obtain a set of candidate terms corresponding to the structured fields of products and the recall hit type of each candidate term in the candidate term set, including: The product structured fields are processed to generate a product n-gram set and a KB query string; Using the KB query string and the product n-gram set, term retrieval and matching processing is performed in the catering knowledge base to obtain a candidate term set; Based on the longest match priority rule, the recall match score of each candidate term in the candidate term set is calculated; Based on the recall matching score of each candidate term, the recall hit type of each candidate term is determined.

6. The method according to claim 5, characterized in that, The product structured fields are processed to construct a query, resulting in a product n-gram set and a KB query string, including: The structured fields of the product are cleaned to obtain cleaned structured fields; The cleaned structured fields are then processed by a script to obtain the KB query string; From the cleaned structured fields, filter out the structured fields corresponding to the product names, and use the structured fields corresponding to the product names as anchor words; Based on the KB query string and the anchor words, the product n-gram set is generated.

7. The method according to claim 5, characterized in that, Using the KB query string and the product n-gram set, term retrieval and matching processing is performed in the catering knowledge base to obtain a candidate term set, including: Using the KB query string and according to the preset hit rules, the term is matched in the catering knowledge base to obtain the first candidate term; In the food and beverage knowledge base, each word in the n-gram set is retrieved by backward indexing to obtain several initial candidate words, and the number of times each initial candidate word hits the n-gram set is counted. The initial candidate terms are deduplicated, and then sorted in descending order of the number of hits. The first k initial candidate terms are selected as the second candidate terms, where k is a positive integer. The candidate term set is formed by using the first candidate term and the second candidate term.

8. The method according to claim 5, characterized in that, The product structured fields include: product tags. Based on the longest match priority rule, the recall match score for each candidate term in the candidate term set is calculated, including: For any candidate term in the candidate term set, the number of target words contained in any candidate term is determined, and the text overlap is obtained based on the determined number, wherein the target words are words in the n-gram set; Determine the anchor word overlap for any of the candidate terms; Calculate the semantic similarity between the term tag of any candidate term and the product tag in the product structure field; Determine the longest possible inclusion length of any candidate term in the KB query string, and calculate the ratio between the longest possible inclusion length and the total length of the KB query string; We calculate the weighted sum of text overlap, anchor word overlap, and semantic similarity to obtain the weighted sum result. Obtain the longest coverage weight and calculate the product between the longest coverage weight and the ratio; The sum of the product and the weighted sum is used as the recall matching score of any candidate term. After all candidate terms in the candidate term set have been polled, the recall matching score of each candidate term is obtained.

9. The method according to claim 1, characterized in that, The translation strategy leverages the knowledge base of each candidate term, performing knowledge base-enhanced translation processing on each candidate term, including: If the knowledge base enhancement translation strategy for any candidate term is an enhancement translation strategy based on NER and the catering knowledge base, then the NER model is used to detect the unmatched word fragments in any candidate term. Identify the matching words of any candidate term in the catering knowledge base, and obtain the standard translation of the matching words from the catering knowledge base; Replace the matched words with the standard translations from the aforementioned food and beverage knowledge base; By combining the standard translation and the missing word fragments, the complete product name is obtained; Perform tag and description completion processing on any of the candidate terms to obtain complete product tags and complete product descriptions; From the product structured fields, obtain the product category information, price, and specifications corresponding to the complete product name; Using the complete product name, complete product label, complete product description, and the product category information, price, and specifications corresponding to the complete product name, a complete translation of any candidate term is constructed. After all candidate terms have been polled, the translated text of the menu image is generated based on the complete translations of each candidate term.

10. The method according to claim 1, characterized in that, The method is applied to a multilingual menu translation system based on a large AI model. After obtaining the translated text of the menu image, the method further includes: The translated text is proofread and optimized to obtain an optimized translated text. The proofreading and optimization process is as follows: when responding to human-computer editing interaction, the modified fields in the translated text are identified, and a local retranslation model is called to retranslate the modified fields so as to fill the retranslated fields back into the translated text to obtain the optimized translated text. The quality of the optimized translation text is assessed, and it is determined whether the quality assessment passes. If so, then the translated text will be rendered using a menu template based on the rendering template; Accordingly, after generating the multi-format translated menu file, the method further includes: Obtain the review results of the translated menu file; Determine whether the review result is a problematic file; If so, add the translated menu file to the problem sample pool; By utilizing a problem sample pool, feedback optimization is performed on the multilingual menu translation system based on the AI ​​large model.

Citation Information

Patent Citations

  • Intelligent translation calibration method based on RAG knowledge base

    CN120805944A

  • PDF optimization translation method and system based on intelligent text detection

    CN120893451A