Apparatus and methods for training a machine learning model to extract text items from text photographs

By using a generative artificial intelligence model to determine the correspondence between text items and image regions, and combining it with image extraction add-ons, the problem of low menu registration efficiency in existing technologies is solved, and an efficient and automated menu update and registration process is achieved.

CN122374737APending Publication Date: 2026-07-10GRABTAXI HOLDINGS PTE LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GRABTAXI HOLDINGS PTE LTD
Filing Date
2023-12-18
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, machine learning models for extracting text items from text photos require proper training to provide accurate results, but this involves a large amount of manual labor and is inefficient, especially when dealing with a large number of restaurant menus.

Method used

Generative artificial intelligence models are used to determine the correspondence between text items and image regions. Machine learning models are adjusted to improve accuracy. Generative AI models are used for data correction and feedback loop systems. Combined with image extraction add-ons, automated menu registration and updating are achieved.

Benefits of technology

It enables an efficient and automated menu registration and update process, reducing the cost of manual operations, improving the speed and accuracy of menu registration, and saving merchants time and effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122374737A_ABST
    Figure CN122374737A_ABST
Patent Text Reader

Abstract

Various aspects relate to a method for training a machine learning model to extract text items from a text photograph, including: extracting text items from a text photograph using the machine learning model; determining a correspondence between the extracted text items and an image of a display text from which the machine learning model has extracted the text items; and adjusting the machine learning model based on the determined correspondence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various aspects of this disclosure relate to apparatus and methods for training machine learning models that extract text items from text photographs. Background Technology

[0002] Many texts are easier to obtain as images (i.e., photographs) than as editable text, such as restaurant menus. However, for further processing, the text needs to be converted to another form (editable text), which can result in a significant amount of manual labor, such as converting multiple restaurant menus into editable formats so they can be published on food delivery service websites. Machine learning models can be applied to such tasks, but they need to be properly trained to provide accurate results.

[0003] Therefore, an efficient method is needed to train a machine learning model that extracts text items from text photos. Summary of the Invention

[0004] Various implementations relate to a method for training a machine learning model to extract text items from a text photograph, comprising: extracting text items from a text photograph using a machine learning model; determining, by means of a generative artificial intelligence model, a correspondence between the extracted text items and an image comprising a region of display text in the text photograph, the machine learning model having extracted the text items from the region of display text in the text photograph; and adjusting the machine learning model based on the determined correspondence.

[0005] According to one implementation, adjusting a machine learning model based on a determined correspondence includes adjusting the machine learning model's processing of text-based photos based on the determined correspondence.

[0006] According to one implementation, adjusting the machine learning model based on the determined correspondence includes: if the correspondence is below a predetermined threshold, adjusting the machine learning model to improve the correspondence.

[0007] According to one implementation, determining correspondence includes providing an image of the extracted text item and an image of the region from which the machine learning model, from which the text item has been extracted, to a generative artificial intelligence model.

[0008] According to one implementation, at least the extracted text items are provided to the generative artificial intelligence model in a text prompt for the generative artificial intelligence model.

[0009] According to one implementation, a generative artificial intelligence model is a large language model that supports image input.

[0010] According to one embodiment, the method includes: for each of a plurality of text items in one or more text photos, generating a plurality of training data elements by: extracting the text item from the corresponding text photo in the one or more text photos using a machine learning model; determining, by a generative artificial intelligence model, a correspondence between the extracted text item and a corresponding image of a region containing the displayed text of the corresponding text photo, the machine learning model having extracted the text item from the region containing the displayed text of the text photo; generating training data elements with the extracted text item as the true value of the corresponding text photo in response to the determined correspondence being higher than a predetermined threshold; and training the machine learning model using the generated training data elements.

[0011] According to one implementation, the method includes: in response to a determined correspondence being lower than a predetermined threshold, correcting an extracted text item by a generative artificial intelligence model, and generating training data elements with the corrected text item as the true value of the text photograph.

[0012] According to one implementation, the method includes: rejecting the use of an extracted text item for one of the training data elements in response to a determined correspondence being lower than a predetermined threshold.

[0013] According to one implementation, the method includes: for each of a plurality of text items in one or more text photos, obtaining user feedback regarding the correspondence between the extracted text item and the corresponding image; determining whether the user feedback confirms the determined correspondence; and in response to the user feedback confirming the determined correspondence, generating training data elements with the extracted text item as the ground truth value of the corresponding text photo.

[0014] According to one implementation, the method includes: in response to user feedback that the determined correspondence has not been confirmed, rejecting the use of an extracted text item for one of the training data elements.

[0015] In one implementation, the text item is a text description of the dish.

[0016] In one implementation, the text photo is a photo of the menu.

[0017] According to one embodiment, a method for extracting text items from multiple text photos is provided, comprising: training a machine learning model according to any of the methods described in the above embodiments; and extracting text items from the multiple text photos by providing the text photos to the trained machine learning model.

[0018] According to one embodiment, a server computer is provided, including an input interface, a memory interface, and a processing unit configured to perform the methods of any of the above embodiments.

[0019] According to one embodiment, a computer program element is provided, including program instructions that, when executed by one or more processors, cause one or more processors to perform any of the methods described above.

[0020] According to one embodiment, a computer-readable medium is provided, including program instructions that, when executed by one or more processors, cause one or more processors to perform the methods of any of the above embodiments. Attached Figure Description

[0021] The invention will be better understood with reference to the detailed description, in which the non-limiting examples and accompanying drawings are taken into consideration: - Figure 1 The communication setup of the market system is shown, including smartphones and servers (computers).

[0022] - Figure 2 The processing of input data provided by the merchant is shown.

[0023] - Figure 3 An example of viewing the page is shown.

[0024] - Figure 4 The diagram illustrates the menu image route, which is the main process for extracting text from the menu image.

[0025] - Figure 5 The generative AI (artificial intelligence) add-on shows the menu image route.

[0026] - Figure 6 An image extraction add-in is shown, illustrating the menu image route.

[0027] - Figure 7 A flowchart illustrating a method for training a machine learning model to extract text items from a text photograph is shown.

[0028] - Figure 8 A server computer according to an embodiment is shown. Detailed Implementation

[0029] The following detailed description refers to the accompanying drawings, which illustrate specific details and embodiments in which this disclosure can be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice this disclosure. Other embodiments may be utilized, and structural and logical changes may be made, without departing from the scope of this disclosure. The various embodiments are not necessarily mutually exclusive, as some embodiments may be combined with one or more other embodiments to form new embodiments.

[0030] The embodiments described in the context of one apparatus or method are similarly effective for other apparatuses or methods. Similarly, the embodiments described in the context of an apparatus are similarly effective for a vehicle or method, and vice versa.

[0031] Features described in the context of an implementation may be correspondingly applied to the same or similar features in other implementations. Features described in the context of an implementation may be correspondingly applied to other implementations, even if not explicitly described in those other implementations. Furthermore, additions and / or combinations and / or substitutions described for features in the context of an implementation may be correspondingly applied to the same or similar features in other implementations.

[0032] In the context of various implementations, the articles “a,” “an,” and “the” used with respect to a feature or element include references to one or more of the features or elements.

[0033] As used in this article, the term “and / or” includes any and all combinations of one or more associated listed items.

[0034] The implementation method will be described in detail below.

[0035] Figure 1 The communication setup of the market system is shown, including a smartphone 100 and a server (computer) 106.

[0036] The smartphone 100 has a screen that displays a graphical user interface (GUI) for using one or more of a variety of services (such as ordering food), which the smartphone user has previously installed on the smartphone and opened (i.e. launched) to use the service, such as ordering food.

[0037] GUI 101 includes graphical user interface elements 102 and 103 that help users use the service, such as a map of the user's location, available food nearby (which the application can determine based on location services, such as GPS-based location services), buttons for placing orders, etc.

[0038] When a user has made a selection for a service (e.g., choosing a restaurant or online supermarket and / or selecting food or groceries to order), the application communicates via a radio connection with the server 106 of the corresponding service. The server 106 (executing the corresponding server program via processor 107) can query memory 109 or data storage device 108 containing information about the service (e.g., price, availability, estimated delivery time, etc.). The server transmits any data relevant to the user or requested by the user (such as the estimated delivery time) back to the smartphone 100, and the smartphone 100 displays this information on the GUI 101. The user can then accept the service, such as ordering food. In this case, the server 106 accordingly notifies the service provider 104, such as the restaurant or online supermarket. The server 106 may also communicate with the service provider 104 earlier, for example, to determine the estimated delivery time.

[0039] It should be noted that although server 106 is described as a single server, its functions (such as providing food delivery services) are typically provided by an arrangement of multiple server computers (i.e., a server computer system) in practical applications (e.g., to implement cloud services). Therefore, the functions provided by a server (e.g., server 106) as described below can be understood as being provided by an arrangement of servers or server computers.

[0040] In order to provide a food ordering service, server 106 specifically needs to have information about the dishes offered by each restaurant where food can be ordered. Therefore, this information (such as merchant menus) needs to be incorporated into server 106 (e.g., into memory 109 or data storage device 108) and needs to be kept up-to-date.

[0041] Various implementations provide methods for efficiently incorporating (i.e., registering or enrolling) and updating merchant menu information, avoiding tedious, repetitive, and time-consuming manual processes, especially when there are a large number of merchants (e.g., restaurants). This allows for reduced menu registration and update costs, faster merchant registration, and increased self-service rates for merchant menus.

[0042] According to various implementation methods, the menu registration and update process is automated, allowing merchants to easily scan or record their menus using a (first) machine learning (ML) model and a large language model (LLM) (i.e., a second ML model, which is an LLM) to structure menu data, thereby improving efficiency and reducing the costs associated with manual updates.

[0043] According to various implementation methods, in addition to applying LLM to structure the output of the (first) ML model, one or more of the following may also be provided: ● An end-to-end system for capturing feedback from merchant editors to enable iterative improvements to the ML model (feedback loop system). ● A third ML model (which is a generative model) is used to assist in data augmentation / expansion to improve the first ML model. ● Photo-based food extraction (from menu photos containing one or more food photos) and linking to the first and third ML models (generative models). Photo-based food extraction can be performed by a fourth ML model.

[0044] According to various implementations, server 106 allows merchants (e.g., service provider 104) to scan their menus or record audio of their menu items. This input is then processed using a first ML model and a second model (LLM). The first ML model can be, for example, an OpenAI Whisper-based model for speech transcription, PSENet (Progressive Scale Extension Network) for text detection and recognition, and / or SVTR (Scene Text Recognition with a Single Visual Model). The LLM classifies, maps, and structures the data (generated by the first ML model) into a publishable menu format.

[0045] Figure 2 The processing of input data 201 provided by the merchant is shown, which includes menu images (i.e., images (e.g., photos) including text descriptions of the dishes) 202 and / or audio recordings of the merchant's menu 203 (i.e., spoken versions of the text descriptions of the dishes).

[0046] The first ML model applies text extraction 204 and / or speech transcription 205 to the input data to provide menu output data. For this purpose, the first ML model may include two sub-models, one for text extraction and one for speech transcription. The menu output data generated by the first ML model is fed to an LLM 206, which aligns the menu output data with a predetermined menu data structure. The LLM 206 can be viewed by a user (specifically a merchant) and published by a server 106 (menu publishing 208), for example, when it has been approved by the merchant.

[0047] Therefore, according to various implementation methods, the automated menu registration and / or update process has two routes: one is through the menu image 202, and the other is through the voice recording of the menu 203.

[0048] For using voice recording of routes

[0049] 1) To generate input 203, the merchant can, for example, use their mobile phone to record their voice. For example, server 106 can provide instructions on a registration webpage or application for the merchant to speak their menu items one by one.

[0050] 2) After the server 106 obtains the audio (i.e., voice recording, which may include multiple items) uploaded by the merchant, the server performs speech transcription through a first ML model, which includes, for example, an automatic speech recognition (ASR) model to transcribe the speech into text.

[0051] 3) The output of the first ML model (in this case, the ASR model) is menu output data, which is in raw text form, ideally such as text that the merchant has already read into their mobile device. LLM 206 helps extract food items, categories, and prices from the raw text to generate structured menu data. It should be noted that people speak and construct their content differently, so rule-based extraction may fail, and traditional ML models (which learn from observed patterns) may also fail, but LLMs are generally able to perform this task. Example: For example, the following text is (original text) menu output data (output by the first ML 204): "The first category is pasta, and below are all items. Item name is Beef Pasta, priced at $5. Item name is Black Pepper Pasta, priced at $7. Item name is Tomato Chicken Pasta, priced at $8." For this example, the LLM output (structured menu data) would be something like this: [ {'itemName': 'beef spaghetti', 'price': 5, 'categoryName': 'spaghetti', 'description': "}, {'itemName': 'black pepper pasta', 'price': 7, 'categoryName': 'spaghetti', 'description': "}, {'itemName': 'tomato chicken pasta', 'price': 8, 'categoryName': 'spaghetti', 'description': "} ] By providing hints, LLM 206 can output JSON-formatted data, as shown in the example above, allowing downstream backend services to easily parse the data and populate the corresponding fields (in the webpage or application GUI 101).

[0052] Figure 3 An example of viewing page 300 is shown, which presents the merchant with the processing result of input 203 in the example above and gives the merchant the opportunity to edit the result.

[0053] The food item 301 on the view page is generated by server 106 by parsing the structured menu data generated by LLM 206 for menu view 207. The item name, category, and price are given on this view page 300. View page 300 closes the loop because the merchant—the menu owner—can view the processing output and edit, delete, or add further items. In this way, the results of the automated menu registration and / or update process, acting as a merchant assistant, can be confirmed by the merchant. Therefore, a merchant menu assistant (driven by AI) is provided, which can help save time and effort for merchants or other users.

[0054] Similar to the voice recording route described above, the menu image route also includes the process input 202 -> first ML model (text extraction 204) -> LLM 206 -> menu viewing. Unlike the voice recording route, the menu image route can include additional components to make the entire system a better loop, which can continuously improve performance, especially the performance of the first ML model (i.e., text extraction 204). The high quality of the text extracted by the first ML model is crucial for the menu image route to work, and it also has greater room for improvement when retrieving data from the real world.

[0055] According to various implementations, the menu image route includes three functional components (e.g., implemented by three subsystems respectively), a main flow, a generative AI add-on, and an image extraction add-on, which will be described below.

[0056] Figure 4 The main flow 400, which shows the menu image route, is shown.

[0057] The main process 400 includes an end-user feedback loop process for iteratively improving the text extraction (OCR (Optical Character Recognition)) model 401 (in this case, the first ML model or a sub-model of the first ML model).

[0058] The main process includes the following items: 1) Merchant 402 inputs menu photo 403, for example, uploads menu photo 403 (including one or more food items) to server 106. 2) Text extraction model 401 processes the menu photo 403. Text extraction model 401 includes a text detection model and a text recognition model. The text extraction result 404 is the recognition of the original text 405 and the corresponding position of the text in the image 406 for each food item (positional information, i.e., in the form of one or more bounding boxes). For example, the text recognition result 404 (e.g., for a food item) is {"Fried Rice": [12, 25, 31, 67]} 3) The text recognition result 404 is formatted (e.g., to fit the LLM text prompt) and then input as part of the prompt into the LLM 407. Here, the purpose of the LLM 407 is to understand the menu photo 401, but instead of sending the menu photo 403 directly to the LLM 407 (which could be more time-consuming and less accurate), the text extraction model 401 is used to first extract the text and location, and then the LLM 407 processes the result 404 based on its understanding of the menu.

[0059] 4) The output 408 of LLM 407 is, for example, similar to a voice recording route and can be parsed to generate a viewing page as described above for a voice recording route. Figure 3 (For three food items), it is used to present to merchants for review in order to form a merchant feedback loop.

[0060] 5) When merchant 402 views output 408 (in the form of a page view), if the merchant may find that the food item has been correctly extracted from the menu photo, the merchant can approve the food item, and approval 409 is recorded in log 410. Sometimes, merchant 401 may find that a food item has been extracted incorrectly. For example, menu photo 401 shows "fried rice," but output 408 indicates "fire rice." In this case, merchant 401 can correct the food item to the correct one. Such a correction 411 is also recorded in log 410.

[0061] The approvals 409 and modifications 411 from merchant 401 are concatenated back to the text recognition result 404; that is, the text extraction result 404 is mapped to merchant behavior. For example, the text extraction result 404 for a food item is {“Hot Rice”: [12, 25, 31, 67]}, while it should actually be {“Fried Rice”: [12, 25, 31, 67]}. This feedback data can then be used to improve the text extraction model 401. For example, updated data (modified food items) and approved data (approved food items) are collected in a “golden dataset” (i.e., training dataset) 412 as ground truth data (associated with the corresponding menu photos), and the golden dataset 412 is used to further fine-tune the text extraction model 401 to make the text extraction model better and improve the accuracy of the menu image route guide.

[0062] Figure 5 The diagram illustrates the generative AI add-in 500 for the menu image route.

[0063] Figure 5 The process shown is via Figure 4 and Figure 5 Point A in the middle is connected to Figure 4 The main process.

[0064] Generative AI add-on 501 serves two purposes: first, as a way to augment additional data for training text extraction model 401; and second, as a filter to select high-quality data from the merchant feedback loop.

[0065] For the purpose of generative AI add-on 500, a set of menu photos 501 were collected, such as those uploaded by merchants, open source menus, etc.

[0066] The menu photo 501 is then processed by the current version of the text extraction model 502 (corresponding to the text extraction model 401) to output the text of the food items and their corresponding positions in the menu photo 501 as the text extraction result 503.

[0067] For each food item, based on location (e.g., in the form of a bounding box), menu photo 501 was cropped to text and paired with the extracted text, see [link to relevant documentation]. Figure 5 The diagram illustrates output 504, which is a pair of text and a cropped image, including the OCR text: Grilled Eggplant and an image displaying "Grilled Eggplant Salad (Grilled EggplantZaalouk)".

[0068] In this example, an error exists in the OCR text: the word "salad (Zaalouk)" is missing. To detect this error and utilize it to improve text extraction models 401 and 502, the output 504 is sent to a multimodal generative AI model 505 (such as GPT 4V), which can accept both text and images as input. The generative AI model 505 is used to score the correspondence between the cropped image and the OCR text and / or (if applicable) correct the provided OCR text. Therefore, its output 506 is (possibly) a corrected pair of text and the cropped image. In effect, the generative AI model 505 is thus used to automate the data verification and review process, rather than relying on manual labor (thus saving effort and cost).

[0069] In this example, the generative AI model 505 corrects “baked eggplant” to “baked eggplant salad” to generate a ground truth 507 for the food item (which can be added to the Golden Dataset 412 for further training and fine-tuning of the text extraction models 401 and 502), or its output “poor correspondence”, which can also be used as feedback for further training and fine-tuning of the text extraction models 401 and 502.

[0070] For example, training data elements of the Golden Dataset 412 can be used for (re)training rounds of machine text extraction models 401 and 502, for example, by calculating the loss between their output and the true value for the corresponding text photo, and adjusting their parameters (e.g., neural network weights) in the direction of reducing the loss.

[0071] For the purpose of filtering high-quality data from the merchant feedback loop, the merchant outputs (i.e., modified or approved food items) are compared with the outputs from generative AI model 505. Only data that has good overlap with the outputs of the generative AI model, or data whose outputs are rated as good by generative AI model 505, are retained for further model improvement (e.g., put into the Golden Dataset 412).

[0072] Figure 6 The image extraction add-in 600 for the menu image route is shown.

[0073] Figure 6 The process shown is via Figure 4 and Figure 6 Point B in the middle is connected to Figure 4 The main process.

[0074] The purpose of the image extraction add-on 600 is to detect and extract images of dishes from the menu photo 401 when the menu photo 401 is a menu with photos of the dishes (i.e., showing the appearance of the dishes), and to match them with the correct dish names and their corresponding prices extracted by the text extraction model.

[0075] Therefore, the input 601 to the image extraction add-on 600 is again a menu photo uploaded by the merchant, but is assumed to include images of the dishes. The object detection model 602 is used to obtain the bounding boxes (location information) of the dish images in the menu photo 601 (or each one), that is, the object detection model 602 outputs one or more pairs 603 <dish image-location>.

[0076] Then, the text extraction result 604 from the main process 400 is obtained. This result includes a <text-location> pair for each food item. The extracted text and dish images can be mapped to each other, for example, through a multimodal generative AI model 605 (e.g., corresponding to multimodal generative AI model 505 or LLM 407) and / or based on the food item text (i.e., name) and the location of the dish image. The dish image can be assigned to the corresponding food item to form a dish image and food item text pair 606. This result 606 of the image extraction add-on 600 can be sent back to the merchant 401 for review, and when approved by the merchant 401, the dish image can be included in the menu presented to the customer (e.g., GUI 101). Therefore, the image extraction add-on 600 helps merchants register their dish photos more efficiently.

[0077] In summary, based on various implementation methods, the following are provided: Figure 7 The method shown.

[0078] Figure 7 A flowchart 700 illustrating a method for training a machine learning model to extract text items from a text photograph is shown.

[0079] In 701, a machine learning model is used to extract text items from text photos.

[0080] In 702, a generative artificial intelligence model is used to determine the correspondence between the extracted text items and the image of the region displaying the text in the text photo, which has extracted the text items from the region displaying the text in the text photo.

[0081] In 703, the machine learning model is adjusted based on the determined correspondence.

[0082] In various implementations, in other words, training data or feedback for a text extraction model is provided through a generative artificial intelligence model. The trained machine learning model (text extraction model) can be applied, for example, to extract food items (in editable text) from menu photos, i.e., to provide a menu assistant, but it can also be applied to other (e.g., industrial) use cases that require cataloging and structuring varied and extensive data from images (e.g., in addition to voice), such as restaurant reservation platforms, retail inventory management systems, or any other platform that processes product listings and updates.

[0083] As described above, for example, a system is provided that includes: ● Main process: Implement an end-to-end system to capture feedback from merchant editors, thereby enabling iterative ML model improvement (feedback loop system). ● Generative AI add-on, using generative artificial intelligence to assist in data augmentation / enlargement to improve ML models. ● Image extraction add-on, which performs photo extraction and dish name mapping through generative AI or ML.

[0084] Figure 7 Methods such as those derived from Figure 8 The server computer shown is executing.

[0085] Figure 8 A server computer 800 according to an embodiment is shown.

[0086] Server computer 800 includes an input interface 801 (e.g., configured to receive text images). Server computer 800 also includes a processing unit 802 and a memory 803. Processing unit 802 can use memory 803 to store, for example, data to be processed, such as information about demand and supply. The server computer is configured to perform... Figure 7 The method.

[0087] The methods described herein can be performed, and the various processing or computing units and devices described herein, as well as computing entities (e.g., server computers), can be implemented by one or more circuits. In implementations, "circuit" can be understood as any type of logical implementation entity, which can be hardware, software, firmware, or any combination thereof. Therefore, in implementations, "circuit" can be hardwired logic circuitry or programmable logic circuitry, such as a programmable processor, e.g., a microprocessor. "Circuit" can also be software implemented or executed by a processor, e.g., any kind of computer program, e.g., a computer program using virtual machine code. Any other kind of implementation of the corresponding functions described herein can also be understood as a "circuit" according to alternative implementations.

[0088] While this disclosure has been specifically shown and described with reference to particular embodiments, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims. Therefore, the scope of the invention is indicated by the appended claims and is thus intended to cover all changes falling within the meaning and scope of equivalents of the claims.

Claims

1. A method for training a machine learning model to extract text items from a text photograph, comprising: Use the machine learning model to extract text items from text photos; The corresponding relationship between the extracted text item and the image of the display text area from which the machine learning model has extracted the text item is determined by a generative artificial intelligence model; as well as The machine learning model is adjusted based on the determined correspondence.

2. The method according to claim 1, wherein, Adjusting the machine learning model based on the determined correspondence includes: adjusting the processing of the text photo by the machine learning model based on the determined correspondence.

3. The method according to claim 1 or 2, wherein, Adjusting the machine learning model based on the determined correspondence includes: if the correspondence is below a predetermined threshold, adjusting the machine learning model to improve the correspondence.

4. The method according to any one of claims 1 to 3, wherein, Determining the correspondence includes providing the extracted text item and an image of the region from which the machine learning model has extracted the text item, including the text photo, to the generative artificial intelligence model.

5. The method according to claim 4, wherein, At least the extracted text items are provided to the generative artificial intelligence model in the text prompts used in the generative artificial intelligence model.

6. The method according to any one of claims 1 to 5, wherein, The generative artificial intelligence model is a large language model that supports image input.

7. The method according to any one of claims 1 to 6, comprising: For each text item in one or more text images, generate multiple training data elements in the following manner: The machine learning model is used to extract the text item from the corresponding text photo in the one or more text photos; The generative artificial intelligence model determines the correspondence between the extracted text item and the corresponding image of the area from which the machine learning model has extracted the display text of the text item; In response to a determined correspondence exceeding a predetermined threshold, training data elements are generated using the extracted text items as the true values ​​of the corresponding text photos. as well as The machine learning model is trained using the generated training data elements.

8. The method of claim 7, comprising: In response to a determined correspondence falling below a predetermined threshold, the extracted text item is corrected by the generative artificial intelligence model, and training data elements are generated using the corrected text item as the true value of the text photograph.

9. The method of claim 7, comprising: In response to a determined correspondence being below the predetermined threshold, the extracted text item is rejected for use in one of the training data elements.

10. The method according to any one of claims 7 to 9, comprising: For each of the plurality of text items in the one or more text photos, Obtain user feedback regarding the correspondence between the extracted text items and the corresponding images; Determine whether the user feedback regarding the correspondence confirms the determined correspondence; as well as In response to the user feedback confirming the determined correspondence, training data elements are generated using the extracted text items as the true values ​​of the corresponding text photos.

11. The method of claim 10, comprising: In response to the user feedback that the determined correspondence has not been confirmed, the use of the extracted text item for one of the training data elements is rejected.

12. The method according to any one of claims 1 to 11, wherein, The text item is a text description of the dish.

13. The method according to any one of claims 1 to 12, wherein, The text photo is a photo of the menu.

14. A method for extracting text items from multiple text photographs, comprising: The machine learning model is trained using the method according to any one of claims 1 to 13; as well as Text items are extracted from the multiple text photos by feeding them into a trained machine learning model.

15. A server computer, comprising an input interface, a memory interface, and a processing unit configured to perform the method according to any one of claims 1 to 14.

16. A computer program element comprising program instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 14.

17. A computer-readable medium comprising program instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 14.