Catering image generation method and system based on AI

By using a restaurant dish image dataset to train an image generation model and combining it with LoRA model fine-tuning, we built a ComfyUI workflow, which solved the problems of authenticity and style limitations in restaurant image generation and achieved efficient and low-cost generation of diversified restaurant images.

CN120672886APending Publication Date: 2025-09-19BEIJING GUANGYAN TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510730693.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing AI image generation tools have problems with dish images' lack of authenticity, targeting, and style limitations when used in the catering industry, making it difficult to meet the catering companies' needs for rapid iteration and diversified visual content.

Method used

By using a dataset of restaurant images to train an image generation model, fine-tuning it with the LoRA model, and building a workflow based on ComfyUI, we can preprocess and parse user input information and generate high-quality and diverse restaurant images.

Benefits of technology

It enables catering companies to quickly generate high-quality promotional images without relying on professional photographers, reducing costs by more than 70%. The generated images are close to or reach the level of professional photography, and support one-click generation of multiple style variants, greatly improving work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672886A_ABST
    Figure CN120672886A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image generation, and discloses an AI-based catering image generation method, which comprises the steps of training an image generation model by utilizing a catering dish image data set to obtain a catering image generation large model and a LoRA model; constructing a workflow based on a ComfyUI, and preprocessing and analyzing the obtained user input information by using a workflow node to obtain a plurality of workflow node parameters; scheduling a LoRA model by using the plurality of workflow node parameters, and performing fine tuning on the catering image generation large model to obtain a fine-tuned catering image generation large model; generating a large model according to the catering image after fine adjustment, and generating a preliminary catering image in combination with the obtained user input information; and performing post-processing on the preliminary catering image by using the workflow node to obtain a catering image. The method has the advantages of cost reduction, efficiency increase, high authenticity and diversity, realization of end-to-end automation, great improvement of the working efficiency, and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation technology, and in particular to an AI-based catering image generation method and system. Background Art

[0002] In today's fiercely competitive restaurant industry, the production of dish images is facing significant change. Traditionally, restaurants have relied on professional photographers. While this approach ensures image quality, its high costs and lengthy production cycles are significant drawbacks. Furthermore, with rapidly evolving market demands and the prevalence of multi-scenario marketing strategies, this traditional approach is increasingly unable to meet the restaurant industry's demand for rapidly iterating and diverse visual content.

[0003] At the same time, although existing AI image generation tools have demonstrated strong capabilities in general image generation, their application in the catering industry has exposed many shortcomings: Lack of specificity. These general models lack dedicated datasets for restaurant scenarios, resulting in unsatisfactory representation of details in the generated dish images. For example, key features such as the color and texture of the dish are often not faithfully reproduced, affecting the visual appeal and professionalism of the image.

[0004] It is difficult to control the output image. When using these tools, it is difficult for users to achieve precise control of the image through simple interaction.

[0005] Limited style adaptation. The images generated by AI models, whether in lighting effects or background selection, cannot reach the level of professional photographers. This results in a single style of generated images that is difficult to meet the needs of different brand positioning and marketing.

[0006] Fragmented processes. Existing AI tools often require integration with other software to complete complex operations like image cutouts, style transfer, and background synthesis. This fragmented process not only reduces work efficiency but also increases user learning costs and technical barriers to entry.

[0007] Therefore, in response to the unique needs of the catering industry, it is necessary to develop a dedicated image generation solution to address the shortcomings of general models in terms of the authenticity of dish images. Summary of the Invention

[0008] The embodiments of the present invention provide an AI-based restaurant image generation method to solve the problems of insufficient authenticity, insufficient pertinence and style limitations of dish images in the prior art.

[0009] To provide a basic understanding of some aspects of the disclosed embodiments, the following is a brief summary. This summary is not intended to be an extensive review, identify key or critical elements, or delineate the scope of these embodiments. Its sole purpose is to present some concepts in a simplified form as a prelude to the detailed description that follows.

[0010] According to a first aspect of an embodiment of the present invention, an AI-based method for generating restaurant images is provided.

[0011] In one embodiment, an AI-based method for generating restaurant images includes: Using the restaurant image dataset, we trained the image generation model to obtain the restaurant image generation model and the LoRA model. Build a workflow based on ComfyUI and use workflow nodes to preprocess and parse the user input information to obtain several workflow node parameters; Using several workflow node parameters, the LoRA model is scheduled to fine-tune the large model for generating restaurant images, and the fine-tuned large model for generating restaurant images is obtained. A large model is generated based on the fine-tuned catering image, and a preliminary catering image is generated in combination with the obtained user input information; and the preliminary catering image is post-processed using the workflow node to obtain the catering image.

[0012] In one embodiment, using a catering dish image dataset to train an image generation model to obtain a catering image generation model and a LoRA model, using the LoRA model to train the image generation model to obtain a catering image generation model includes the following steps: Acquire catering dish image data, and pre-process the catering dish image data to obtain a catering dish image dataset; Select an image generation model, initialize the training parameters, and introduce a low-rank matrix into the image generation model in combination with the LoRA model; Based on the catering dish image dataset and combined with the training parameters, the image generation model is trained to obtain the trained image generation model and the LoRA model image generation model; Use the loss function to evaluate the trained image generation model and obtain the evaluation results; According to the evaluation results, the output effect of the workflow composed of the trained image generation model and the LoRA model is obtained, and the training parameters are adjusted in combination with the evaluation results to obtain a large catering image generation model and a LoRA model. The training parameters are adjusted in combination with the LoRA model to obtain a large catering image generation model.

[0013] In one embodiment, the training parameters include learning rate, batch size, number of training rounds, and network dimension; Among them, the learning rate is used to control the step size of model parameter update; Batch size, which determines the number of images processed simultaneously during training; The number of training rounds is used to determine the number of times the restaurant image dataset is fully trained; The network dimension determines the fineness of the generated image.

[0014] In one embodiment, a workflow is constructed based on ComfyUI, and workflow nodes are used to pre-process and parse the acquired user input information to obtain several workflow node parameters, including the following steps: Build workflow based on ComfyUI and configure workflow nodes; The image information input by the user, the prompt words input by the user and the options input by the user are preprocessed and parsed according to the workflow nodes to obtain a number of workflow node parameters.

[0015] In one embodiment, the workflow nodes include an image input parsing node, a style coding matching node, a LoRA model calling node, an additional control network node, an image migration node, a light effect optimization node, and a result output node.

[0016] In one embodiment, the image input parsing node is used to perform size compression and image segmentation on the image information input by the user to obtain the first workflow node parameter; The style code matching node is used to parse the prompt word and the option input by the user respectively to obtain the second workflow node parameter and the third workflow node parameter; LoRA model calling node, used to call the LoRA model and adjust the LoRA model weight; Additional control network nodes are used to strictly constrain the structure and details of the image during the restaurant image generation process; Image migration node, used to implement partial image migration of restaurant images; Lighting optimization node, used to enhance the lighting and shadow effects of restaurant images and improve the visual quality of restaurant images; The result output node is used to output the restaurant image.

[0017] In one embodiment, preprocessing and parsing the image information input by the user, the prompt word input by the user, and the option input by the user according to the workflow node to obtain several workflow node parameters includes the following steps: Obtaining image information input by the user, and using the image input parsing node to perform size compression and image segmentation on the image information to obtain first workflow node parameters; Obtain the prompt word input by the user, and use the style coding to match the node, extract the keywords of the prompt word, and perform mapping analysis to obtain the second workflow node parameters; The options input by the user are obtained, and the style coding matching node is used to parse the user options to obtain the third workflow node parameters.

[0018] In one embodiment, the LoRA model is scheduled using several workflow node parameters to fine-tune the large model for generating restaurant images. The fine-tuned large model for generating restaurant images includes the following steps: Load restaurant images to generate large models; Based on several workflow node parameters, configure the LoRA model to call the node to schedule the LoRA model and adjust the weight of the LoRA model; According to the weights of the LoRA model, the large model for generating restaurant images is fine-tuned to obtain a fine-tuned large model for generating restaurant images.

[0019] In one embodiment, generating a large model based on the fine-tuned restaurant image and generating a preliminary restaurant image in combination with the acquired user input information; and post-processing the preliminary restaurant image using a workflow node to obtain the restaurant image includes the following steps: Inputting user input information into the fine-tuned restaurant image generation model to obtain a preliminary restaurant image; Using additional control network nodes, image migration nodes, and light effect optimization nodes, the preliminary restaurant image is post-processed to obtain a restaurant image; Use the result output node to output the restaurant image.

[0020] According to a second aspect of an embodiment of the present invention, an AI-based restaurant image generation system is provided.

[0021] In one embodiment, an AI-based restaurant image generation system includes: The catering image generation model training module is used to train the image generation model using the catering dish image dataset to obtain the catering image generation model and the LoRA model; The workflow node parameter parsing module is used to build workflows based on ComfyUI and use workflow nodes to preprocess and parse the acquired user input information to obtain several workflow node parameters. The catering image generation model fine-tuning module is used to use several workflow node parameters to schedule the LoRA model and fine-tune the catering image generation model to obtain the fine-tuned catering image generation model; The catering image generation module is used to generate a large model based on the fine-tuned catering image, and generate a preliminary catering image based on the obtained user input information; and use the workflow node to post-process the preliminary catering image to obtain the catering image.

[0022] According to a third aspect of an embodiment of the present invention, a computer device is provided.

[0023] In some embodiments, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0024] According to a fourth aspect of embodiments of the present invention, a computer-readable storage medium is provided.

[0025] In one embodiment, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0026] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: 1. The present invention reduces costs and increases efficiency, allowing catering companies to quickly generate high-quality promotional pictures without relying on professional photographers, reducing costs by more than 70%.

[0027] 2. This invention has high authenticity and diversity. By combining a dedicated dataset with LoRA technology, the generated images are close to or reach the level of professional photography in terms of food texture, lighting effects, etc., while providing a variety of style options.

[0028] 3. This invention realizes end-to-end automation. It only takes 1-5 minutes for users to upload pictures and obtain finished products. It supports one-click generation of multiple style variants, which greatly improves work efficiency.

[0029] 4. The present invention is scalable, and its modular design facilitates the subsequent addition of style templates (such as holiday themes) and functions (such as short video generation), laying a solid foundation for future development.

[0030] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0032] Figure 1 This is a flowchart of a method for generating restaurant images based on AI according to an exemplary embodiment; Figure 2This is a principle block diagram of an AI-based restaurant image generation system according to an exemplary embodiment; Figure 3 This is a schematic diagram of workflow nodes in an AI-based restaurant image generation method according to an exemplary embodiment; Figure 4 1 is a schematic diagram of a user interface in an AI-based method for generating restaurant images according to an exemplary embodiment; Figure 5 This is an example image of a restaurant with a white background in an AI-based restaurant image generation method according to an exemplary embodiment; Figure 6 This is an example diagram of a restaurant in a traditional Chinese style in an AI-based restaurant image generation method according to an exemplary embodiment; Figure 7 This is an example diagram of an elegant Chinese-style restaurant in an AI-based restaurant image generation method according to an exemplary embodiment; Figure 8 This is an example diagram of an elegant Chinese restaurant style and upper left corner lighting in an AI-based restaurant image generation method according to an exemplary embodiment; Figure 9 The figure is a schematic structural diagram of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION

[0033] The following description and accompanying drawings sufficiently illustrate the specific embodiments herein to enable those skilled in the art to practice them. Portions and features of some embodiments may be included in or substituted for portions and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims, including all available equivalents thereof. Herein, the terms "first," "second," and the like are used solely to distinguish one element from another and do not require or imply any actual relationship or order between these elements. In practice, the first element can also be referred to as the second element, and vice versa. Furthermore, the terms "comprise," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a structure, device, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such structure, device, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the structure, device, or apparatus comprising the element. The various embodiments herein are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Similar or identical parts between the various embodiments can be referenced to each other.

[0034] The terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like in this document indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this document and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, they should not be understood as limitations on the present invention. In the description of this document, unless otherwise specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, they can be mechanical or electrical connections, or they can be internal connections between two elements. They can be directly connected or indirectly connected through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0035] As used herein, unless otherwise specified, the term "plurality" means two or more.

[0036] In this document, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.

[0037] In this article, the term "and / or" is used to describe the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0038] It should be understood that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0039] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software so that the processor can call and execute the operations corresponding to the above modules.

[0040] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0041] Figure 1An embodiment of an AI-based restaurant image generation method of the present invention is shown.

[0042] In this optional embodiment, the AI-based method for generating restaurant images includes: S101. Use the catering dish image dataset to train the image generation model to obtain the catering image generation model and the LoRA model.

[0043] In this optional embodiment, the image generation model is trained using a catering dish image dataset to obtain a catering image generation model and a LoRA model. The image generation model is trained using the LoRA model to obtain a catering image generation model, which includes the following steps: Acquire catering dish image data, and pre-process the catering dish image data to obtain a catering dish image dataset; Select an image generation model, initialize the training parameters, and introduce a low-rank matrix into the image generation model in combination with the LoRA model; Based on the catering dish image dataset and combined with the training parameters, the image generation model is trained to obtain the trained image generation model and the LoRA model image generation model; Use the loss function to evaluate the trained image generation model and obtain the evaluation results; According to the evaluation results, the output effect of the workflow composed of the trained image generation model and the LoRA model is obtained, and the training parameters are adjusted in combination with the evaluation results to obtain a large catering image generation model and a LoRA model. The training parameters are adjusted in combination with the LoRA model to obtain a large catering image generation model.

[0044] In this optional embodiment, the training parameters include learning rate, batch size, number of training rounds, and network dimension; Wherein, the learning rate is used to control the step size of model parameter update; The batch size determines the number of images processed simultaneously during training; The number of training rounds is used to determine the number of times the restaurant dish image dataset is fully trained; The network dimension is used to determine the fineness of the generated image.

[0045] It should be explained that the image generation model uses the Stable Diffusion series model and the FLUX series model as the basic models, and is trained using a self-built catering dish image dataset.

[0046] The restaurant dish image dataset contains more than 80,000 professionally photographed dish images and corresponding annotations, covering multiple cuisines, styles, dishes, angles, scenes, and lighting conditions, providing rich learning materials for image generation models.

[0047] During training, models such as Stable Diffusion 1.5 (Stable Diffusion Model 1.5) and Flux fp8 (FluxFP8 quantized model) were used to fine-tune the base model. A dedicated dataset covering professional photographic images of various restaurant dishes was constructed using our own high-quality restaurant image data.

[0048] The restaurant image dataset is derived from high-quality, real-world dish photos, covering a wide range of Chinese and Western cuisine, as well as diverse plating, background, and lighting scenarios. To ensure the model learns rich, detailed features, each image in the dataset is accompanied by a corresponding text description, including label information such as the dish name, appearance characteristics, and scene style (e.g., "wooden tabletop + warm lighting"), to annotate the image material for training. The dataset was also cleaned and enhanced: low-quality or blurry images were removed, resolution and size were standardized, and some data was enhanced (such as rotation, scaling, and brightness adjustment) to expand the data volume and improve model generalization.

[0049] LoRA (Low-Rank Adaptation) is an efficient model fine-tuning technique that introduces low-rank matrices into pre-trained models, enabling them to quickly adapt to new tasks and reducing the computational resources and time required for training. We build lightweight LoRA adaptation models for different photographic styles (such as side-lit and texture overlays), enabling rapid fine-tuning of the base model to ensure that generated images precisely match the user-specified style.

[0050] Through training with a self-built catering image dataset (including multi-angle shots of dishes, scene background library, and lighting parameter labels), the basic model is given a powerful ability to restore dish details.

[0051] LoRA technology is used to perform lightweight fine-tuning on the model to achieve rapid switching and combination of styles, providing users with a more flexible, efficient and accurate image generation experience.

[0052] The specific process of using the LoRA model to train the image generation model and obtain the large model for restaurant image generation is as follows: 1. Prepare training data.

[0053] Collect a dataset of restaurant dish images: Based on the training objective (usually a specific cuisine style), collect high-quality, relevant images, ensuring that the dataset covers different angles, lighting conditions, and backgrounds. The number of images should be between 50 and 500.

[0054] Data preprocessing: crop, scale, and other processing of food images to unify the size and resolution (usually in arrive to improve the training effect.

[0055] Label generation: Use automated labeling tools (such as the WD14 label generator) to generate initial labels for food and beverage images, which are then manually adjusted and supplemented by professional food and beverage photographers to ensure accuracy and completeness.

[0056] 2. Configure the training environment.

[0057] Select a basic model: Select an appropriate image generation model according to your needs, such as Stable Diffusion 1.5.

[0058] Install training tools: Use LoRA-scripts (LoRA training scripts) or other tools suitable for LoRA training to configure the training environment.

[0059] Set the training parameters: Learning Rate: controls the step size of model parameter updates and is usually set between 1e-4 and 1e-5.

[0060] Batch Size: The number of food image files processed simultaneously during each training session, typically 1-8.

[0061] Epochs: The number of times the entire restaurant image dataset is fully trained. This number is determined by the size of the restaurant image dataset and the training objectives, and is typically between 10 and 20 epochs.

[0062] Network Dimension: determines the fineness of the generated image and is usually set to 128.

[0063] 3. Start training.

[0064] Start the training process using the configured training parameters, and the image generation model will learn based on the restaurant image dataset and labels.

[0065] 4. Evaluation and adjustment.

[0066] During the training process, monitor the changes in the loss function to ensure that the image generation model is effectively learned.

[0067] The preview images generated during the training process can also monitor the accuracy of model training.

[0068] Based on the training results, it may be necessary to adjust the training parameters such as learning rate, batch size, and number of training rounds to achieve the best results and obtain a large model for restaurant image generation.

[0069] Through the above process, the LoRA model can efficiently fine-tune the image generation model to obtain a large model for catering image generation, so that the large model for catering image generation can adapt to specific tasks and generate high-quality images.

[0070] S102: Build a workflow based on ComfyUI, and use workflow nodes to pre-process and parse the acquired user input information to obtain several workflow node parameters.

[0071] In this optional embodiment, the steps of constructing a workflow based on ComfyUI and preprocessing and parsing the acquired user input information using workflow nodes to obtain several workflow node parameters include the following steps: Build workflow based on ComfyUI and configure workflow nodes; The image information input by the user, the prompt words input by the user and the options input by the user are preprocessed and parsed according to the workflow nodes to obtain a number of workflow node parameters.

[0072] In this optional embodiment, the workflow nodes include an image input parsing node, a style encoding matching node, a LoRA model calling node, an additional control network node, an image migration node, a lighting effect optimization node, and a result output node.

[0073] In this optional embodiment, the image input parsing node is used to perform size compression and image segmentation on the image information input by the user to obtain the first workflow node parameter; The style code matching node is used to parse the prompt word and the option input by the user respectively to obtain the second workflow node parameter and the third workflow node parameter; LoRA model calling node, used to call the LoRA model and adjust the LoRA model weight; Additional control network nodes are used to strictly constrain the structure and details of the image during the restaurant image generation process; Image migration node, used to implement partial image migration of restaurant images; Lighting optimization node, used to enhance the lighting and shadow effects of restaurant images and improve the visual quality of restaurant images; The result output node is used to output the restaurant image.

[0074] In this optional embodiment, preprocessing and parsing the image information input by the user, the prompt word input by the user, and the option input by the user according to the workflow node to obtain several workflow node parameters includes the following steps: Obtaining image information input by the user, and using the image input parsing node to perform size compression and image segmentation on the image information to obtain first workflow node parameters; Obtain the prompt word input by the user, and use the style coding to match the node, extract the keywords of the prompt word, and perform mapping analysis to obtain the second workflow node parameters; The options input by the user are obtained, and the style coding matching node is used to parse the user options to obtain the third workflow node parameters.

[0075] It needs to be explained that, Figure 3 The figure shows a workflow node diagram for building a workflow based on ComfyUI. ComfyUI (Comfortable User Interface) is used to build a workflow that parses user input (image, prompt, and options) into workflow node parameters (such as style code and lighting intensity), automatically triggering image-to-image or text-to-image tasks. This invention typically provides style options, allowing users to directly trigger specific nodes, models, or workflows by selecting a style.

[0076] Optimize resource allocation through multi-threaded scheduling to ensure fast response in high-concurrency scenarios.

[0077] To improve generation speed and response efficiency, this invention employs multi-threaded optimization and parallel scheduling technology in its software architecture. Multi-threading is used throughout all stages, from data preprocessing to image generation, achieving efficient resource utilization and pipelined processing. Key points include: Parallel preprocessing: After a user submits a request, a separate thread is created to handle image preprocessing tasks, such as reading, compressing, resizing, cropping, and edge detection. These operations primarily utilize the CPU (central processing unit), and executing them in a separate thread can fully utilize the power of multi-core processors. Simultaneously, another thread initiates model loading, which primarily utilizes GPU (graphics processing unit) and I / O (input / output) resources. By running CPU preprocessing and GPU model loading in parallel, overall waiting time is reduced.

[0078] Asynchronous model loading and inference: Stable Diffusion, the Flux base model, and various LoRA / ControlNet sub-models are loaded using asynchronous calls. While waiting for disk I / O to read model weights, other threads can continue executing logic such as prompt parsing and parameter configuration. Once the model is ready, the inference thread immediately takes over and performs image generation calculations. Multithreading is also adopted for decoupled sub-processes at different stages of the generation process. For example, CLIP text encoding can be executed in a separate thread (or GPU stream) in parallel with the main generation process. This design avoids hardware idleness caused by a single thread executing all steps serially.

[0079] Multi-threaded concurrent generation: When generating multiple style variants simultaneously, multiple generation threads can be launched, with each thread responsible for generating images of a specific style or parameter combination. Within the constraints of hardware resources (e.g., batch inference on multiple GPUs or a single GPU), multi-threaded concurrent generation can significantly improve throughput. Even when generating only a single image at a time, thread pool management allows for pre-loading the model and caching results for subsequent requests, accelerating responses to consecutive requests.

[0080] Resource Scheduling and Synchronization: To prevent resource conflicts caused by multi-threaded competition, a thread scheduling management module was designed. For example, this ensures that only one thread is executing diffusion sampling on the GPU at a time, while other threads wait. Alternatively, GPU tasks for different threads can be scheduled on separate streams to achieve true parallelism. Regarding memory, data is transferred between threads via shared memory or message queues, avoiding unnecessary copying overhead. Threads synchronize at key points (for example, waiting for all pre-processing to complete before starting synthesis) to ensure an orderly process.

[0081] Combining the above measures, this invention breaks down the generation process into several parallelizable subtasks, optimizing them through multithreading and asynchronous scheduling. In actual testing, the multithreaded architecture significantly reduces average processing time. For example, a process that originally took 10 seconds can be reduced to 6-7 seconds after optimization. This speed increase is of great value to the actual deployment of user experience, enabling real-time interaction and one-click image generation.

[0082] The image input parsing node intelligently preprocesses the image information input by the user, including steps such as automatic clipping and region segmentation. The uploaded original image is first compressed, typically to a length and width of 1024 pixels or less. The SegmentAnything (SAM) image segmentation algorithm is then used to detect the main dish and background areas in the image and generate an accurate segmentation mask. This step extracts the image's foreground (removing clutter or unnecessary background). After preprocessing, the image is compressed to a size suitable for computational operations, while the main dish is accurately clipped out and the background area is marked, facilitating subsequent operations such as replacement, blurring, or adding special effects. This image compression and automated clipping process requires no human intervention, ensuring processing efficiency and accuracy, and laying the foundation for generating high-quality, uniformly styled composite images.

[0083] The style code matching node is used to parse the prompt word and the option entered by the user respectively.

[0084] If the user selects a preset style (such as "Appetite - Side Backlight") or enters a text description (such as "Wooden Tabletop + Warm Tones"), the user's selection is converted into a prompt word or node configuration. Furthermore, after receiving the style description entered by the user, the prompt word module converts it into a prompt word or workflow node parameter configuration that the AI ​​model can interpret, thereby guiding the image generation process.

[0085] This prompt conversion process utilizes a pre-built style dictionary and template library. These libraries store mappings between common style descriptions and model prompts, as well as corresponding parameter combination templates. When a user enters free text to describe a style, such as "wooden tabletop + warm tones," the input text is first segmented and keyword extracted to identify the stylistic elements contained therein ("wooden tabletop" represents the background material, "warm tones" represents the lighting and color style). The style dictionary is then searched to replace these elements with specialized prompts or parameter settings understood by the model. For example, "wooden tabletop" might be converted into a background description fragment within the prompt (e.g., adding the English description "on a wooden table" to the prompt to adapt to the model training corpus), while the node configuration selects a background LoRA related to wood texture or enables a wood grain background node. "Warm tones" might be converted into a lighting description within the prompt (e.g., "warm lighting") and the rendering node's lighting parameters are adjusted (increasing the intensity of warm lighting). In this way, the user's natural language requirements are precisely translated into a set of instructions that are meaningful to the model. For Chinese prompt words, we will also combine the labels used during model training to map the Chinese style descriptions to corresponding English or multilingual labels to match the model's training language domain and ensure that the prompt words can be correctly parsed by the model.

[0086] If the user selects a predefined style option, the prompt word conversion module will directly call the predefined template. For example, the preset style "Appetite - Side Backlight" may correspond to a set of internal configurations: the prompt word template contains a description such as "Food photography under soft side backlighting", the LoRA module selects the dish texture enhancement model associated with "Appetite", and sets the light effect node to simulate the side and rear light source. The workflow will automatically apply this template to generate specific prompt word strings and parameters. When the generation process is started, these prompt words will be combined with the dish name or specific content uploaded by the user to form a complete text prompt input model, and the workflow node configuration will be ready. This prompt word and parameter conversion process is completed in real time and is transparent to the user, but it ensures that the user can drive the complex AI generation process using everyday language, lowering the operation threshold. In short, through the prompt word conversion module, the present invention realizes the bridge from the user's high-level intention to the underlying model instructions, so that the model can accurately understand the image effect expected by the user.

[0087] The LoRA model calling node loads the LoRA model of a specific style and the corresponding control parameters based on the user's selection and input, and performs style fine-tuning on the large model generated by the restaurant image.

[0088] The ControlNet node is an additional control network used to strictly constrain the image structure and details during the diffusion model generation process. Through ControlNet, key information (such as edge contours and segmented regions) is extracted from the input dish image and passed as conditions to the generative model, thereby preserving the original shape structure.

[0089] For example, the present invention uses pre-processing operators such as Canny edge detection to extract the contour lines and edge features of the dish from the original dish image uploaded by the user. This edge information serves as the input condition of ControlNet, so that the generative model strictly follows the outline and shape of the dish when drawing, and does not deviate from the basic structure of the original dish. At the same time, ControlNet can also control the generation details, such as maintaining the detailed texture inside the dish and ensuring that the plating structure is not distorted. Through the collaboration of the ControlNet module, fine control of AI-generated content is achieved. It not only ensures the introduction of new backgrounds and styles, but also ensures that the main shape and key features of the dish are consistent with the original input.

[0090] The image migration node is used to implement object migration, that is, to seamlessly migrate the original dish as an independent foreground object into the newly generated background scene. Its core principle is to introduce the feature representation of the original object (dish) into the generation process to ensure that the dish in the final image is highly consistent with the original image. In terms of specific implementation, the original dish image is first cut out and segmented to obtain the dish foreground and its mask. Then the Redux algorithm is used to perform stylized migration on the dish. Redux will combine the Flux style model and CLIP visual encoding to extract the visual features of the original dish image and embed them into the generation process. The dish position area is reserved in the generated background image, and Redux performs a targeted image transformation on this area, keeping the shape and texture details of the dish unchanged, and only adjusting its color matching and style consistency with the new background.

[0091] This approach, similar to the inpainting and restyling process, ensures that the dish retains its original texture and characteristics after being "migrated" to the new image. From a practical perspective, the image migration node ensures the authenticity of the main dish. Even with changes to the background and photography style, the dish still looks exactly like the one provided by the user, without being arbitrarily redrawn by AI. This is crucial for restaurant applications, as it ensures that brands can meet their requirements for consistent dish appearance.

[0092] The lighting optimization node is used for lighting optimization, that is, to enhance the light and shadow effects of the image to improve the visual quality. IC Light (short for "Imposing Consistent Light") is an AI lighting reshaping technology. The present invention integrates IC Light in the later stage of the workflow to relight the initially generated image. According to the lighting style selected by the user (such as "warm lighting in the upper left corner"), the lighting optimization node accepts corresponding text prompts or parameters to adjust the light source direction, brightness intensity and shadow projection of the image. In implementation, IC Light will input the generated image and its foreground / background information into its trained lighting model, and change the lighting distribution through diffusion model reasoning.

[0093] For example, when ambient lighting requires warm tones, IC Light enhances the warm highlights in the image and creates realistic highlights and shadows in appropriate locations on the dish, making the overall light and shadow contrast of the image more intense and richer in depth. Through IC Light processing, the final output image is closer to professional photography in terms of lighting effects—consistent light sources, natural shadows, and an atmosphere that matches the intended style, significantly enhancing the realism and visual impact of the finished image.

[0094] The result output node uses an SD upscaling algorithm to generate high-definition images, supporting resolutions up to 8K.

[0095] S103. Utilize several workflow node parameters to schedule the LoRA model, fine-tune the large model for generating restaurant images, and obtain a fine-tuned large model for generating restaurant images.

[0096] In this optional embodiment, the LoRA model is scheduled using several workflow node parameters to fine-tune the large model for generating restaurant images. The fine-tuned large model for generating restaurant images includes the following steps: Load restaurant images to generate large models; Based on several workflow node parameters, configure the LoRA model to call the node to schedule the LoRA model and adjust the weight of the LoRA model; According to the weights of the LoRA model, the large model for generating restaurant images is fine-tuned to obtain a fine-tuned large model for generating restaurant images.

[0097] S104: Generate a large model based on the fine-tuned restaurant image, and generate a preliminary restaurant image in combination with the acquired user input information; and use workflow nodes to post-process the preliminary restaurant image to obtain the restaurant image.

[0098] In this optional embodiment, a large model is generated based on the fine-tuned restaurant image, and a preliminary restaurant image is generated in combination with the acquired user input information; and the preliminary restaurant image is post-processed using a workflow node to obtain the restaurant image, including the following steps: Inputting user input information into the fine-tuned restaurant image generation model to obtain a preliminary restaurant image; Using additional control network nodes, image migration nodes, and light effect optimization nodes, the preliminary restaurant image is post-processed to obtain a restaurant image; Use the result output node to output the restaurant image.

[0099] It needs to be explained that, Figure 4 The user interface is shown in FIG. The specific process of a user obtaining a restaurant image is as follows.

[0100] 1. User uploaded images: Users upload original images of dishes (image file size not exceeding 20MB) through the mini program or web client as image input for the workflow.

[0101] 2. Style selection and parameter configuration: Users select a preset style or enter a text description, which is converted into prompt words or workflow node configuration based on the user's selection.

[0102] 3. Workflow operation: 1) Image processing and input parsing: resize and automatically crop images uploaded by users, parse prompts and options entered by users, and select the corresponding workflow, model, and node.

[0103] 2) Initialize the model: Load the self-trained restaurant images to generate a large model.

[0104] 3) Loading LoRA model: Based on the user's selection and input, load the LoRA model of a specific style and the corresponding control parameters to fine-tune the style of the large model generated by the restaurant image.

[0105] 4) Input text prompt (Prompt): Generate prompt words based on user input.

[0106] 5) Image Generation (Stable Diffusion Generation Process): Image generation begins using the input image, text prompt, and loaded model. The large restaurant image generation model works with the LoRA model to generate a preliminary image based on the prompt.

[0107] 6) Image post-processing: Optional post-processing steps, such as changing the lighting, image enhancement, high-definition, denoising, sharpening, etc.

[0108] 7) Export image: Export the final generated image to the user interface.

[0109] 8) Result delivery and iteration: Users can view and save generated images and decide whether to make further adjustments to them. Real-time preview and batch export are supported.

[0110] If the user is not satisfied with the generated image, or needs to make further adjustments to the image, he can click the "Regenerate" button below to re-enter the previous AI image generation process (consistent with the above process); or click "Continue Editing" below to enter the image editing page, and call the above-mentioned local image editing function to perform post-editing on the image.

[0111] like Figure 5-8 The figure below shows images generated in multiple styles for the same dish. Users can define the background by entering a prompt word and selecting specific options. For example, if they enter the name of a dish and select "Chinese Food" + "Minimalist Background" + "Dark Lighting", the corresponding LoRA module and algorithm nodes will be called to generate the target image.

[0112] The present invention also supports a variety of local editing functions, allowing post-processing of specific areas of the generated image to meet the user's personalized adjustment needs. Typical functions include background blur and adding special effects (such as flame effects). The technical implementation is as follows: Background blur: To highlight the main subject of a dish, it's often necessary to moderately blur the background (simulating a depth of field effect). Using the previous image segmentation results, the dish's foreground and background regions have already been separated. With a background mask, the blurring process is relatively straightforward: a blur filter algorithm (such as a Gaussian blur) is applied to the background region in the resulting high-definition image. Specifically, the Image Input Parsing node is called to gradually blur pixels outside the masked region. The blur radius and intensity can be adjusted based on user requirements or a preset style. Because the mask precisely defines the dish's location, the blurring operation doesn't affect the dish itself, resulting in a clear main subject and a blurred background. This method is efficient and secure, introducing no new AI generation uncertainties and ensuring smooth transitions between dish edges without any harsh cuts or pasting. For more advanced depth of field simulation, depth estimation can also be used to generate a depth map. Different blur levels are then applied to the near and far backgrounds based on depth, creating a gradual caustic effect. However, in most cases, a simple masked background blur is sufficient to achieve the desired visual quality of a professional photograph.

[0113] Adding special effects (using flames as an example): To address user needs for adding special effects to specific locations within an image, a localized special effect synthesis function is provided. For example, a flame effect can be added to an image of a hot pot dish to enhance its visual appeal. Implementation requires first defining the location and range of the special effect: this can be achieved through user interaction in the front-end interface (e.g., specifying the stove location) or through intelligent detection (e.g., detecting the location of the hot pot base). Once the area is determined, there are two ways to generate the flame effect: one is to use a pre-prepared special effects library, selecting a semi-transparent PNG image of flames with matching angles and colors, and overlaying it onto the corresponding location in the image. Alpha blending is used to adjust transparency and brightness during overlay, making the flames appear to blend into the scene. The other is to leverage the AI ​​generation model itself to generate the special effect through localized diffusion. Specifically, a flame-themed prompt is constructed, such as "Flame effect, orange halo." Using the original image as a reference, a localized Stable Diffusion (similar to a partial image repaint) is applied to the selected area. Thanks to the integration of ControlNet and other technologies, the flames can be restricted to that area. After generation, the area is then merged back into the full image. Each approach has its advantages and disadvantages: the library approach is fast and controllable, while the AI-generated approach is flexible but computationally expensive. You can choose the appropriate method to implement special effects based on your application needs.

[0114] In addition to background blur and flames, local editing can also be extended to other effects, such as adjusting the color temperature of a certain area, adding embellishments to dishes, etc. The common technical points are area demarcation (determined by masks or coordinates) and area-by-area processing (using image processing algorithms or generative models to act separately on the area). Since the entire workflow is based on the ComfyUI node graph architecture, we can attach various post-processing nodes to the generated image branch to implement these functions. For example, in ComfyUI, a new "local blur" node is added to process the background area, and a "local glow" node is added to enhance the brightness or add effects in a specific area. These local editing functions further enhance flexibility, allowing users to fine-tune local details after obtaining preliminary results to obtain a more satisfactory final image.

[0115] Take "generating a promotional image for hot pot dishes on the WeChat mini program" as an example: User-uploaded images: Users upload the original photos of hot pot dishes they collected and photographed through the WeChat mini-program. The photos are automatically uploaded to the AI ​​workflow, the image size is automatically adjusted from 1080x1920 to 576x1024, and uploaded to the image input node of the AI ​​workflow.

[0116] Style selection and parameter configuration: The user selects the preset style "Chinese food style", the background style "complex background", the lighting effect "warm colors", the image quality "4K", the number of raw images "1", and enters the text description "lively and festive". Based on the user's selection, the Chinese food style large model, the complex background LoRA model, and the warm color light algorithm node are automatically selected, and the text description is uploaded to the prompt word node as the prompt word input.

[0117] AI workflow runs: Initialize the model: Load the self-trained Chinese food style restaurant photography model.

[0118] Loading LoRA model: Based on the user's selection and input, the complex background LoRA is loaded, and the LoRA weight is set to 0.8.

[0119] Input text prompt (Prompt): Generate the prompt word "lively and festive" based on user input.

[0120] Image generation (Stable Diffusion generation process): Use the input resized image, text prompt and loaded Chinese style restaurant model and complex background LoRA to start generating images.

[0121] Image post-processing: Based on the user's selection, the warm light node and the 4K zoom node are called to perform post-processing of the generated image.

[0122] Export image: Export the final generated professional-level hot pot photography image to the user interface.

[0123] Result delivery and iteration: Users can view and save the generated hot pot photography image.

[0124] Figure 2 An embodiment of an AI-based restaurant image generation system of the present invention is shown.

[0125] In this optional embodiment, the AI-based restaurant image generation system includes: A catering image generation model training module 201 is used to train an image generation model using a catering dish image dataset to obtain a catering image generation model and a LoRA model; The workflow node parameter parsing module 202 is used to build a workflow based on ComfyUI and use workflow nodes to pre-process and parse the acquired user input information to obtain a number of workflow node parameters; The catering image generation large model fine-tuning module 203 is used to use a number of workflow node parameters to schedule the LoRA model and fine-tune the catering image generation large model to obtain a fine-tuned catering image generation large model; The catering image generation module 204 is used to generate a large model based on the fine-tuned catering image and generate a preliminary catering image in combination with the obtained user input information; and use workflow nodes to post-process the preliminary catering image to obtain the catering image.

[0126] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps of the above-mentioned method embodiment are implemented.

[0127] Those skilled in the art will understand that Figure 9 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0128] In addition, the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiment when executing the computer program.

[0129] In addition, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0130] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0131] The present invention is not limited to the structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A method for generating restaurant images based on AI, characterized in that: The method includes: Using the restaurant image dataset, we trained the image generation model to obtain the restaurant image generation model and the LoRA model. Build a workflow based on ComfyUI and use workflow nodes to preprocess and parse the user input information to obtain several workflow node parameters; Using several workflow node parameters, the LoRA model is scheduled to fine-tune the large model for generating restaurant images, and the fine-tuned large model for generating restaurant images is obtained. A large model is generated based on the fine-tuned catering image, and a preliminary catering image is generated in combination with the obtained user input information; and the preliminary catering image is post-processed using the workflow node to obtain the catering image.

2. The AI-based restaurant image generation method according to claim 1, characterized in that: The method of using the catering dish image dataset to train the image generation model to obtain the catering image generation model and the LoRA model includes the following steps: Acquire catering dish image data, and pre-process the catering dish image data to obtain a catering dish image dataset; Select an image generation model, initialize the training parameters, and introduce a low-rank matrix into the image generation model in combination with the LoRA model; Based on the catering dish image dataset and combined with the training parameters, the image generation model is trained to obtain the trained image generation model and LoRA model; Use the loss function to evaluate the trained image generation model and obtain the evaluation results; The output effect of the workflow consisting of the trained image generation model and the LoRA model is obtained, and the training parameters are adjusted based on the evaluation results to obtain the large catering image generation model and the LoRA model.

3. The AI-based restaurant image generation method according to claim 2, characterized in that: The training parameters include learning rate, batch size, number of training rounds and network dimension; Wherein, the learning rate is used to control the step size of model parameter update; The batch size determines the number of images processed simultaneously during training; The number of training rounds is used to determine the number of times the restaurant dish image dataset is fully trained; The network dimension is used to determine the fineness of the generated image.

4. The AI-based restaurant image generation method according to claim 1, characterized in that: The process of building a workflow based on ComfyUI and preprocessing and parsing the acquired user input information using workflow nodes to obtain several workflow node parameters includes the following steps: Build workflow based on ComfyUI and configure workflow nodes; The image information input by the user, the prompt words input by the user and the options input by the user are preprocessed and parsed according to the workflow nodes to obtain a number of workflow node parameters.

5. The AI-based restaurant image generation method according to claim 4, characterized in that: The workflow nodes include image input parsing node, style coding matching node, LoRA model calling node, additional control network node, picture migration node, light effect optimization node, and result output node.

6. The AI-based restaurant image generation method according to claim 5, characterized in that: The image input parsing node is used to perform size compression and image segmentation on the image information input by the user to obtain the first workflow node parameters; The style code matching node is used to parse the prompt word and the option input by the user respectively to obtain the second workflow node parameter and the third workflow node parameter; The LoRA model calling node is used to call the LoRA model and adjust the LoRA model weight; The additional control network node is used to strictly constrain the structure and details of the image during the restaurant image generation process; The image migration node is used to implement partial image migration of restaurant images; The light effect optimization node is used to enhance the light and shadow effects of the restaurant image and improve the visual quality of the restaurant image; The result output node is used to output the restaurant image.

7. The AI-based restaurant image generation method according to claim 5, characterized in that: The method of preprocessing and parsing the image information, prompt words and options input by the user according to the workflow nodes to obtain several workflow node parameters includes the following steps: Obtaining image information input by the user, and using the image input parsing node to perform size compression and image segmentation on the image information to obtain first workflow node parameters; Obtain the prompt word input by the user, and use the style coding to match the node, extract the keywords of the prompt word, and perform mapping analysis to obtain the second workflow node parameters; The options input by the user are obtained, and the style coding matching node is used to parse the user options to obtain the third workflow node parameters.

8. The AI-based restaurant image generation method according to claim 1, characterized in that: The method utilizes several workflow node parameters to schedule the LoRA model and fine-tune the large model for generating restaurant images to obtain the fine-tuned large model for generating restaurant images, including the following steps: Load restaurant images to generate large models; Based on several workflow node parameters, configure the LoRA model to call the node to schedule the LoRA model and adjust the weight of the LoRA model; According to the weight of the LoRA model, the large model for generating restaurant images is fine-tuned to obtain a fine-tuned large model for generating restaurant images.

9. The AI-based restaurant image generation method according to claim 1, characterized in that: Generating a large model based on the fine-tuned restaurant image and generating a preliminary restaurant image in combination with the acquired user input information; and post-processing the preliminary restaurant image using a workflow node to obtain the restaurant image includes the following steps: Inputting user input information into the fine-tuned restaurant image generation model to obtain a preliminary restaurant image; Using additional control network nodes, image migration nodes, and light effect optimization nodes, the preliminary restaurant image is post-processed to obtain a restaurant image; Use the result output node to output the restaurant image.

10. An AI-based restaurant image generation system, characterized in that: The system includes: The catering image generation model training module is used to train the image generation model using the catering dish image dataset to obtain the catering image generation model and the LoRA model; The workflow node parameter parsing module is used to build workflows based on ComfyUI and use workflow nodes to preprocess and parse the acquired user input information to obtain several workflow node parameters. The catering image generation model fine-tuning module is used to use several workflow node parameters to schedule the LoRA model and fine-tune the catering image generation model to obtain the fine-tuned catering image generation model; The catering image generation module is used to generate a large model based on the fine-tuned catering image, and generate a preliminary catering image based on the obtained user input information; and use the workflow node to post-process the preliminary catering image to obtain the catering image.

Citation Information

Patent Citations

  • Picture processing method and system

    CN119131202A

  • Material attribute acquisition method and system based on large language model and BIM

    CN119516287A

  • Paper-cut image generation method and system based on Stable Diffusion, medium and equipment

    CN119991867A