Image processing method, moving image processing method, and image processing system

By parsing image information, generating target prompt information, and using a reference information library to generate an initial image, the problems of image generation results deviating from expectations and compliance in existing technologies are solved, and an efficient and usable image generation process is achieved.

CN122115619APending Publication Date: 2026-05-29ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG TMALL TECH CO LTD
Filing Date
2026-01-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing image generation methods rely on vague or incomplete text prompts provided by users, which leads to the generated results deviating from expectations. Furthermore, the generated images require manual review or multiple iterations to meet compliance requirements, making the process cumbersome and inefficient.

Method used

By receiving image processing requests, parsing image information, identifying image elements across multiple visual attribute dimensions, generating target prompts, generating an initial image using a reference information library, and performing inspections to ensure compliance.

Benefits of technology

It improves the efficiency and usability of image generation, ensuring that the generated images meet user needs in terms of visual appeal and compliance, and reduces the need for manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115619A_ABST
    Figure CN122115619A_ABST
Patent Text Reader

Abstract

The embodiment of the present specification provides an image processing method, a moving image processing method and an image processing system, wherein the image processing method comprises: receiving an image processing request, and analyzing the image processing request to obtain image information; determining an image element matched with the image information in multiple visual attribute dimensions, and generating target prompt information based on the image element and initial prompt information; inputting the image information and the target prompt information into an image generation module, and generating an initial image by the image generation module through querying a reference information library; and detecting the initial image to obtain image detection information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of image processing technology, and in particular to image processing methods, moving image processing methods, and image processing systems. Background Technology

[0002] In the field of AI-driven image generation, existing methods typically rely on user-provided text prompts. However, these prompts are often vague, incomplete, or lack visual detail, leading to results that deviate from expectations, with missing elements or illogical layouts. Some solutions attempt to incorporate external knowledge to aid generation, but these often rely on static retrieval or general databases, limiting the expressiveness of the generated images. The generated images frequently require manual review or multiple iterations to meet compliance requirements such as content security, copyright, or task specifications, resulting in a cumbersome and inefficient process. Therefore, a more effective image processing method is urgently needed to address these issues. Summary of the Invention

[0003] In view of the above, embodiments of this specification provide an image processing method. One or more embodiments of this specification also relate to a moving image processing method, an image processing system, an image processing apparatus, a moving image processing device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising: Receive an image processing request and parse the image processing request to obtain image information; Image elements matching the image information are determined across multiple visual attribute dimensions, and target prompt information is generated based on the image elements and initial prompt information. The image information and the target prompt information are input into the image generation module, and the image generation module generates an initial image by querying the reference information database. The initial image is then detected to obtain image detection information.

[0005] According to a second aspect of the embodiments of this specification, a moving image processing method is provided, comprising: Receive an activity image processing request submitted by a user for a target activity, and parse the activity image processing request to obtain activity image information; The activity image element that matches the activity image information is determined by identifying multiple visual attribute dimensions associated with the target activity, and target prompt information is generated based on the activity image element and initial prompt information; The activity image information and the target prompt information are input into the activity image generation module, and an initial activity image containing the target activity information is generated by querying the activity element database. Perform a compliance check on the initial moving image to obtain compliance check information.

[0006] According to a third aspect of the embodiments of this specification, an image processing system is provided, including a client and a server, comprising: The client is configured to generate an image processing request in response to a user's touch operation on the image processing page, and send the image processing request to the server. The server is configured to parse the image processing request to obtain image information; determine image elements matching the image information across multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information; input the image information and the target prompt information into the image generation module, and generate an initial image by querying a reference information database; detect the initial image to obtain image detection information; and feed back the image detection information and the initial image to the client.

[0007] According to a fourth aspect of the embodiments of this specification, another image processing system is provided, including a request processing module, an image generation module, an image detection module, and a reference information library; The request processing module is configured to receive an image processing request, parse the image processing request to obtain image information; determine image elements that match the image information in multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information; and send the image information and the target prompt information to the image generation module. The image generation module is used to generate an initial image based on the image information and the target prompt information by querying the reference information database; and send the initial image to the image detection module. The image detection module is used to detect the initial image and obtain image detection information.

[0008] According to a fifth aspect of the embodiments of this specification, an image processing apparatus is provided, comprising: The parsing unit is configured to receive an image processing request and parse the image processing request to obtain image information; The generation unit is configured to determine image elements that match the image information in multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information; The input unit is configured to input the image information and the target prompt information into the image generation module, and generate an initial image by querying a reference information database. The detection unit is configured to detect the initial image and obtain image detection information.

[0009] According to a sixth aspect of the embodiments of this specification, a moving image processing apparatus is provided, comprising: The parsing unit is configured to receive an activity image processing request submitted by a user for a target activity, and to parse the activity image processing request to obtain activity image information; The generation unit is configured to determine an active image element that matches the active image information across multiple visual attribute dimensions associated with the target activity, and to generate target prompt information based on the active image element and initial prompt information. The input unit is configured to input the moving image information and the target prompt information into the moving image generation module, and generate an initial moving image containing target activity information by querying the activity element database. The detection unit is configured to perform activity content compliance detection on the initial activity image and obtain compliance detection information.

[0010] According to a seventh aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above method.

[0011] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the above-described method.

[0012] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0013] This specification provides an embodiment of an image processing method that receives an image processing request, parses the request, and obtains image information. It determines image elements matching the image information across multiple visual attribute dimensions and generates target prompt information based on these image elements and initial prompt information. This target prompt information is used to prompt the image generation module in the visual dimensions of the image. The image information and target prompt information are input into the image generation module, which then generates an initial image by querying a reference information library. By accessing the reference information library, elements that assist the image generation module in generating the initial image are obtained, resulting in a richly detailed initial image. Detection of the initial image and obtaining image detection information ensures compliance while improving the efficiency and usability of initial image generation. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating an image processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of an image processing method provided in one embodiment of this specification. Figure 3 This is a schematic diagram of an image generation page provided in one embodiment of the image processing method described in this specification; Figure 4 This is a schematic diagram of image elements of an image processing method provided in one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of an image processing system provided in one embodiment of this specification; Figure 6 This is an interactive schematic diagram of an image processing method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the structure of another image processing system provided in one embodiment of this specification; Figure 8 This is a schematic diagram of the data flow of another image processing system provided in one embodiment of this specification; Figure 9 This is a flowchart illustrating an embodiment of a moving image processing method provided in this specification; Figure 10 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification; Figure 11 This is a schematic diagram of the structure of an active image processing apparatus provided in one embodiment of this specification; Figure 12 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0015] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0016] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0017] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0018] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0019] The technical solutions provided in this application can employ deep learning models with relatively large parameter scales. However, this large model is merely an example; this application does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application can be artificial intelligence-based language models (LM) or multimodal models (MM).

[0020] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0021] Prompt words: A structured AI prompt template automatically generated based on the selected task type, including task fields (such as distribution channels, activity type, logo, main and sub-titles) and style guidance words.

[0022] Style selector: Visual style options that are dynamically loaded according to the task type (e.g., "surrealism" is recommended for "splash screen image" and "realistic photography" is recommended for "product detail image").

[0023] Logo Selector: A collection of logo components that can be retrieved from Taobao's internal resource library by brand / activity category and inserted with one click.

[0024] To address the aforementioned technical problems, this specification provides an image processing method, an image processing system, an image processing apparatus, a moving image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0025] See Figure 1 , Figure 1 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0026] Step 102: Receive an image processing request and parse the image processing request to obtain image information.

[0027] An image processing request can be a computer instruction submitted by a client to a server, requesting image processing. The client can be a mobile terminal capable of running applications or web pages with image processing capabilities. The server can be a service provider offering data processing, logical operations, and image generation services to the client; the server can be a cloud server. The image processing request carries image information, which is image description data determined based on user-provided interaction information. This user-provided interaction information includes, but is not limited to, image type information, the image itself, image description text, and requirement description text. This interaction information can be determined by the image generation page displayed on the client, such as the user entering text content, uploading an image, or selecting an image type. Image attribute data can represent the user's image generation requirements.

[0028] Based on this, the server receives image processing requests submitted by the client, parses the requests, and obtains the image information. Image processing requests can be generated based on content selected, input, and / or uploaded by the user on the image generation page provided on the client, representing the user's image generation needs.

[0029] Furthermore, considering that users can express their image generation needs through image information, these needs can be communicated to the client either through input or by selecting options provided on the image generation page. The specific implementation is as follows: The image processing request is parsed to obtain image type information and input information. The image type information corresponds to the image type selected by the user on the image generation page, and the input information is the information entered by the user through the image generation page. The image type information and the input information are used as the image information.

[0030] Specifically, the image generation page is an editable page provided by the client to users, offering an information input interface. The image generation page provides users with a variety of selectable image types, representing the user's image requirements. Image types include, but are not limited to, promotional posters, application splash screen images, marketing images, and product images. Users can select an image type from the available options by interacting with the touch elements on the interactive page, thus confirming the image type information. Users can also input image type information on the image generation page, including but not limited to text input and voice input. The multimodal information, such as text and images, entered by users on the image generation page constitutes the input information.

[0031] Based on this, the image processing request is parsed to obtain image type information and input information. The image type information corresponds to the image type selected by the user on the image generation page, and the image type represents the user's image requirement type. The input information is the information entered by the user through the image generation page. The image type information and input information are used as image information. The input information can be information entered by the user through touch elements provided on the image generation page, and the information input methods include, but are not limited to, text input and voice data.

[0032] For example, in image processing scenarios, image processing can be applied to the generation of various images in e-commerce, such as marketing images, promotional posters, product detail images, and application splash screen images. To meet users' customized needs, an image generation page can be provided. This page can be a user-facing application or webpage. Users can specify the image type information corresponding to their image generation needs through input or selection, and interact with the client via the interactive webpage, inputting text and uploading image materials, using the user-inputted text and uploaded images as input information. For example... Figure 2The page shown is the image generation page. The area containing the "Image Description (Required)" field is where users can provide text information from their input. The area containing the "Upload Image" control is where users can provide image information from their input. The image generation page includes various selectable image types, such as free mode, splash screen, promotional poster, channel header, product main image, and image processing. When the user selects "Splash Screen," predefined information linked to the splash screen, such as "Generate a splash screen...", will be displayed. This predefined information is included in the input information.

[0033] In summary, parsing image processing requests yields image type and input information, which are then used as the basis for image information, thus enriching the overall image data. This provides users with different methods for determining information, enabling them to more clearly express their image generation needs.

[0034] Step 104: Determine image elements that match the image information across multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information.

[0035] Specifically, after receiving and parsing the image processing request to obtain image information, image elements matching the image information can be determined across multiple visual attribute dimensions. Target prompt information is then generated based on these image elements and initial prompt information. The visual attribute dimensions can be the visual attributes of the image, including but not limited to image style, image size, image title, and keywords. Image elements are the visual attribute elements corresponding to these multiple visual attribute dimensions. Each visual attribute dimension can correspond to at least one image element, which can be text, image, or WordArt. Image elements include, but are not limited to, keywords, WordArt, style information, and titles. The initial prompt information can be a prompt template, i.e., a prompt template. The initial prompt information contains at least one prompt item to be filled, including but not limited to image requirement items, image element items, and output constraint items. The target prompt information is the prompt information obtained after filling in at least one prompt item to be filled in the initial prompt information. The content filled in the initial prompt information can be text or an image.

[0036] Based on this, after receiving and parsing the image processing request to obtain image information, image elements matching the image information are determined across multiple visual attribute dimensions, including image style, image size, image title, and keywords within the image. At least one image element can be determined. By filling in at least one unfilled prompt item in the initial prompt information based on at least one image element, the target prompt information can be obtained.

[0037] The target prompt message could be: "Generate a splash screen image, style [surrealism], logo [..." [Holiday], search box [ [Search box], Main title: Future Fashion Week, Subtitle: Live Stream Limited Time Reservation, Size: 750×1624px” where the fields in [] are dynamic placeholders and will be automatically replaced with the default values ​​corresponding to the current image type information.

[0038] Furthermore, considering the inseparable relationship between image generation and image element selection, image elements can be determined based on multiple visual attribute dimensions, as specifically implemented below: The process involves determining editable and associated attribute dimensions within multiple visual attribute dimensions; identifying editable image elements matching the image information within the editable attribute dimensions; and identifying associated image elements matching the editable image elements within the associated attribute dimensions; and using the editable image elements and the associated image elements as the image elements.

[0039] Specifically, the visual attribute dimension refers to the visual element types of an image, while the editable attribute dimension provides users with an interface to select image attribute types. Correspondingly, the associated attribute dimension is the attribute dimension that does not provide users with a direct editing interface for image attribute types. The associated attribute dimension is linked to both image type information and the editable attribute dimension. After the editable image elements are determined in the editable attribute dimension, the associated image elements are the predefined image elements associated with the editable image elements. After the editable elements are determined, the associated image elements and the editable image elements are displayed in tandem. The element types of editable elements include, but are not limited to, keywords, theme image elements, style image elements, and search box text elements or image elements. The element types of associated image elements can be image elements, keywords, or image-generated descriptive text.

[0040] Based on this, editable attribute dimensions and associated attribute dimensions are identified within multiple visual attribute dimensions. The associated attribute dimensions are linked to the editable attribute dimensions, and also to image type information. Editable image elements matching the image information are determined through input or control selection within the editable attribute dimensions, and associated image elements matching the editable image elements are determined within the associated attribute dimensions. The associated image elements and editable image elements are displayed in tandem; the editable image elements are displayed simultaneously with the associated image elements. Both editable and associated image elements are treated as image elements.

[0041] Continuing with the previous example, such as Figure 2 As shown, when the user selects "splash image" as the image type, the system automatically displays image descriptions such as "Generate a splash image..." in the input area containing "Upload Image". The style, logo, and search box in the image description are editable attribute dimensions, while the search terms, main title, subtitle, and time are associated attribute dimensions. When determining surrealism, Festivals and When the search box is an editable image element, it will display related information such as fashion shows, future fashion weeks, and limited-time live stream reservations. Related image elements. Editable image elements and related image elements are used as the image elements corresponding to the image type information.

[0042] In summary, editable image elements matching image information are determined at the editable attribute level, and associated image elements matching the editable image elements are determined at the associated attribute level. The associated image elements and editable image elements are displayed in tandem, showing the associated image elements simultaneously with the editable image elements. This simplifies the method of determining image elements and improves the efficiency of subsequent image generation.

[0043] Furthermore, considering that users have different image generation needs and require different image elements in image generation scenarios, an element selection page containing at least one candidate image element can be provided. Users can select editable image elements visually, as implemented below: A reference image element matching the image information is determined in the editable attribute dimension; in response to an element update request submitted for the reference image element, an element selection page containing at least one candidate image element is generated; in response to a confirmation request submitted for a target element contained in the at least one candidate image element in the element selection page, the editable image element corresponding to the target element is determined.

[0044] Specifically, the reference image element can be the initialization element corresponding to the image type information. The element update request can be a computer instruction submitted via an element selection control associated with the reference image element, used to display an element selection page associated with the editable attribute dimension. The element selection page contains at least one candidate image element, which is an optional image element matching the image type information. Candidate image elements can be displayed on the element selection page in the form of WordArt or images. The target element is the image element selected by the user from at least one candidate image element.

[0045] Based on this, a reference image element matching the image information is determined in the editable attribute dimension. This reference image element can be an initial image element corresponding to the editable attribute dimension. In response to an element update request submitted for the reference image element, an element selection page containing at least one candidate image element is generated. The element selection page may also contain the reference image element; if the user does not select a target image element, the reference image element in the element selection page is selected. In response to a confirmation request submitted for a target element contained in at least one candidate image element in the element selection page, an editable image element corresponding to the target element is determined. After the user selects the target element, the target element becomes an editable image element.

[0046] Continuing with the previous example, such as Figure 3 As shown in (a), when the editable attribute dimension is style, the initialized style is surrealist, and the surrealist subject element is the reference element style. Touching the element selection control (the triangle behind the reference image element) submits an element update request. The client responds to the element update request by displaying an element selection page containing at least one candidate image element (style 1, style 2, and multiple selectable styles). This element selection page corresponds to the linked style selector. When the user selects style 2 as the target element, element 2 is displayed in the area corresponding to that style. Style 2 is the editable image element. Accordingly, as... Figure 3 As shown in (b), when the editable attribute dimension is the logo dimension, the initialized logo is... festival, The festival element is the reference element logo. Touching the triangle control following the reference image element submits an element update request. The client responds to the element update request by displaying an element selection page containing at least one candidate image element (logo1, logo2, and multiple selectable logos). This element selection page corresponds to the linked logo selector. When the user selects logo2 as the target element, element2 will be displayed in the corresponding area. logo2 is an editable image element. Similarly, the search box can also be manually modified by the user in this way.

[0047] In summary, for an element selection page containing at least one candidate image element, the editable image element corresponding to the target element can be determined through visual element selection, thereby realizing the visual selection of editable image elements, improving the selection efficiency of editable image elements, and enhancing the user experience of selecting editable image elements.

[0048] Step 106: Input the image information and the target prompt information into the image generation module, and generate an initial image by querying the reference information database.

[0049] Specifically, after determining the image elements matching the image information across multiple visual attribute dimensions and generating the target prompt information based on the image elements and initial prompt information, the image information and target prompt information can be input into the image generation module. The image generation module then generates an initial image by querying a reference information library. This image generation module can be an image generation engine integrating multiple content processing models, including but not limited to text processing models, image processing models, and layout planning models. The reference information library can contain text, images, artistic fonts, image layout templates, and other information that can be used as model references. The initial image is the image output by the image processing module that corresponds to the image processing request.

[0050] Based on this, after determining the image elements that match the image information in multiple visual attribute dimensions and generating the target prompt information based on the image elements and the initial prompt information, the image information and the target prompt information are input into the image generation module. By querying the reference information library, the image generation module generates an initial image. Subsequently, the initial image can be further inspected for compliance to ensure the compliance of the image displayed to the user.

[0051] Furthermore, considering that the initial image to be generated may contain not only image content but also text content, the layout of the image and text in the initial image needs to be taken into account during the initial image generation. Therefore, in order to improve the usability of the initial image, an image generation module containing both text processing and image processing models can be used to generate the initial image. The specific implementation is as follows: The image information and the target prompt information are input into the image generation module, and the text processing model and image processing model included in the image generation module after input are determined. The text processing model and the image processing model are used to query the reference information library to obtain reference text and reference image. The text processing model is used to process the reference text, the image information and the target prompt information to obtain image text. The image processing model is used to process the reference image, the image information and the target prompt information to obtain a layout image. The initial image is generated based on the image text and the layout image.

[0052] Specifically, both the text processing model and the image processing model are machine learning models included in the image generation module. These models can also be large-scale models. The text processing model generates text content, while the image processing model generates images containing layout formatting. Reference information can be obtained by querying information in the reference information library. The types of reference information include, but are not limited to, text, image, and artistic font types. Text content is the reference text, and image content is the reference image. Image text can be copywriting with information delivery, promotional, and / or marketing functions. Layout images can be images containing layout information.

[0053] Based on this, image information and target prompt information are input into the image generation module. The text processing model and image processing model included in the image generation module after input are determined. The text processing model generates text content, and the image processing model generates an image containing a layout format. The text processing model and image processing model are used to query a reference information library to identify reference text that can be used for subsequent text generation, and reference images that can be used for subsequent image generation. The text processing model processes the reference text, image information, and target prompt information to obtain image text, and the image processing model processes the reference image, image information, and target prompt information to obtain a layout image. An initial image is generated based on the image text and layout image. The image text is added to the layout image, and the image text and layout image are then merged or overlaid to obtain the initial image.

[0054] In summary, a text processing model is used to process reference text, image information, and target prompt information to obtain image text, and an image processing model is used to process reference image, image information, and target prompt information to obtain a layout image. An initial image is generated based on the image text and layout image, improving the generation efficiency of the initial image and ensuring its quality in both the image and layout dimensions, thus enhancing its usability.

[0055] Furthermore, considering that the text processing model processes text and its output is primarily text content, text prompts can be identified within the target prompt information to assist the text processing model in text processing. These prompts are then input into the text processing model to process the reference text and image information. The specific implementation is as follows: The text prompt information is determined from the target prompt information; the text processing model is used to process the reference text, the image information and the text prompt information to obtain image text.

[0056] Based on this, text prompts are determined from the target prompt information. These text prompts are used to generate image text when the text processing model processes the reference text and image information. The text processing model processes the reference text, image information, and text prompts to obtain image text, ensuring the usability of the image text and its match with the user's image generation needs.

[0057] Continuing with the previous example, the text processing model is used to generate image text, which can be marketing copy. By processing the reference text, image information, and text prompts using the text processing model, copy that meets the user's image generation needs can be obtained.

[0058] In summary, by using a text processing model to process reference text, image information, and text prompts, image text can be obtained, thereby improving the efficiency and accuracy of image text determination.

[0059] Furthermore, considering that the image processing model primarily processes images and its output mainly consists of image content, image prompts can be identified within the target prompt information to assist the image processing model in image processing. These image prompts are then input into the image processing model to process the reference image and the image information. The specific implementation is as follows: Image prompt information is determined from the target prompt information; the image processing model is used to process the reference image, the image information, and the image prompt information to obtain a layout image.

[0060] Based on this, image prompts are determined from the target prompt information. These image prompts are used to generate the layout image according to their content when the image processing model processes the reference image and image information. The image processing model processes the reference image, image information, and image prompts to obtain the layout image, ensuring its usability and its match with the user's image generation requirements.

[0061] Continuing with the previous example, the image processing model is used to generate a layout image, which can contain a marketing image with a specific layout format. By processing the reference image, image information, and image prompts using the image processing model, a layout image that meets the user's image generation requirements can be obtained.

[0062] In summary, by using an image processing model to process the reference image, image information, and image prompts, a layout image can be obtained, thereby improving the efficiency and accuracy of layout image determination.

[0063] Step 108: Detect the initial image to obtain image detection information.

[0064] Specifically, after the image information and target prompt information are input into the image generation module and the initial image is generated by the image generation module by querying the reference information library, the initial image can be detected to obtain image detection information. The purpose of detecting the initial image is to perform compliance detection on the initial image. The image detection information includes the detection judgment result of whether it is compliant. If the detection judgment result is non-compliant, the detection result can also include the reason for non-compliance and modification suggestions.

[0065] Based on this, after inputting the image information and target prompt information into the image generation module and generating an initial image by querying the reference information database, the initial image is subjected to compliance detection to obtain image detection information, thereby realizing compliance detection of the initial image and preventing non-compliant initial images from being put into use.

[0066] Furthermore, considering the potential compliance risks of the generated initial image, after obtaining the initial image, it can be detected in at least one detection dimension, as specifically implemented as follows: At least one detection dimension is determined, and the initial image is detected using the at least one detection dimension to obtain the image detection information; if the initial image fails the detection based on the image detection information, the initial image is updated to an intermediate image based on the image detection information; the target image is generated based on the image template and the image information corresponding to the intermediate image.

[0067] Specifically, at least one detection dimension can include an image content dimension and an image attribute dimension. The image attribute dimension can include a size dimension, while the image content dimension includes, but is not limited to, version consistency, font compliance, copyright, and portrait rights. Image detection information can be either a pass / fail detection of the initial image across detection dimensions, or a fail / fail detection of the initial image across detection dimensions, providing the reason for the failure and suggested modifications. The intermediate image is the image that passes the detection after modifying the initial image in at least one detection dimension based on the detection results. The image template defines the size and position of at least one region in the target image. The target image is the compliant image obtained by performing compliance checks on the initial image and adjusting the initial image based on the detection results.

[0068] Based on this, at least one detection dimension is determined, and the initial image is detected using this dimension to obtain image detection information. This information includes detection conclusions indicating whether the detection passed or failed, as well as reasons for failure and suggested modifications for those failures. If the initial image fails detection based on the image detection information, it is updated to an intermediate image. A target image is generated based on the image template and the corresponding image information of the intermediate image. The intermediate image is adjusted according to the region size and position corresponding to the image template to obtain the target image.

[0069] Continuing with the previous example, at least one detection dimension includes, but is not limited to, size, version consistency, font compliance, copyright, and portrait rights. Detecting the initial image at the size dimension allows for the generation of multiple intermediate images in various aspect ratios (1:1, 4:5, 9:16, banner size, custom size, etc.) based on the image's intended use. The version consistency dimension checks for missing main images, obscured logos, and consistent white space in the title area. The font compliance dimension checks for compliant fonts in the initial image, i.e., whether they are licensed fonts. The copyright dimension checks for brand logos, people, cartoon characters, famous buildings, etc., in the initial image. The portrait rights dimension checks for clearly identifiable portraits of any person in the initial image.

[0070] In summary, by detecting the initial image in at least one detection dimension and generating the target image based on the image template and the image information corresponding to the intermediate image, the compliance and usability of the target image can be improved.

[0071] Furthermore, considering that the target image consists of text content and image content, the text and images in the target image need to be arranged according to a certain layout, as specifically implemented as follows: Text information and image content information are determined from the image information corresponding to the intermediate image, and text region and image region are determined from the image template; text content is generated in the text region based on the text information, and image content is generated in the image region based on the image information; the target image is generated based on the text content and the image content.

[0072] Specifically, the image information corresponding to the intermediate image is used to generate the target image. The text information contained in the image information is the text content, including but not limited to main titles, subtitles, and artistic fonts. The image content information is the image data corresponding to the image. The text area is used to display text information, and the image area is used to display image content information.

[0073] Based on this, text information and image content information are determined from the image information corresponding to the intermediate image, and text and image regions are determined from the image template. Text content is generated in the text region based on the text information, and image content is generated in the image region based on the image information. The text region can include multiple areas such as a main title area, subtitle area, and slogan area. The image region can include a background area, main image area, and icon area. Image rendering is performed based on the text and image content to generate the target image.

[0074] In summary, text content is generated from text areas based on text information, and image content is generated from image areas based on image information. Image rendering is then performed based on the text and image content to generate the target image, ensuring that the layout of the target image meets the user's image generation requirements.

[0075] Furthermore, considering that the target image contains various content such as text, images, and backgrounds, when the user needs to adjust the target image, the adjustment can be performed according to at least one image processing type, as specifically implemented as follows: Determine at least one image processing type, and determine a target image processing type among the at least one image processing type; process the target image according to the target image processing type to obtain an image to be displayed.

[0076] Specifically, image processing types include, but are not limited to, image expansion, local adjustment, and full-image fine-tuning. Image expansion can be the expansion of image content, local adjustment can be the partial redrawing of the target image, and full-image adjustment can be the adjustment of the background and tone of the target image.

[0077] Based on this, at least one image processing type is determined, and a target image processing type is selected from at least one image processing type by touching a type control. That is, the target image processing type can be selected by touching the type control corresponding to the target image processing type. The target image is then processed according to the target image processing type to obtain the image to be displayed.

[0078] Continuing with the previous example, such as Figure 4 As shown, once the target image is determined, it can be displayed on the preview page. Figure 1 , Figure 2 , Figure 3 and Figure 4 The target image is provided. Each target image provides adjustment controls, including sub-controls for various image processing types such as image expansion, local adjustment, and full-image fine-tuning. Select the sub-control corresponding to local adjustment and perform a touch operation to start redrawing the selected image 1 locally. After the redrawing is completed, the image to be displayed is obtained.

[0079] In summary, by processing the target image according to the target image processing type, the image to be displayed is obtained, and fine-tuning of the target image is achieved, ensuring that users can adjust the target image according to their needs and improving user satisfaction with image generation.

[0080] This specification provides an embodiment of an image processing method that receives an image processing request, parses the request, and obtains image information. It determines image elements matching the image information across multiple visual attribute dimensions and generates target prompt information based on these image elements and initial prompt information. This target prompt information is used to prompt the image generation module in the visual dimensions of the image. The image information and target prompt information are input into the image generation module, which then generates an initial image by querying a reference information library. By accessing the reference information library, elements that assist the image generation module in generating the initial image are obtained, resulting in a richly detailed initial image. Detection of the initial image and obtaining image detection information ensures compliance while improving the efficiency and usability of initial image generation.

[0081] Figure 5 This specification shows a schematic diagram of the structure of an image processing system according to one embodiment. Figure 5As shown, the image processing system 500 includes a request processing module 510, an image generation module 520, an image detection module 530, and a reference information database 540. The request processing module 510 receives image processing requests, parses the requests to obtain image information, determines image elements matching the image information across multiple visual attribute dimensions, and generates target prompt information based on the image elements and initial prompt information. It then sends the image information and the target prompt information to the image generation module 520. The image generation module 520 generates an initial image based on the image information and the target prompt information by querying the reference information database 540. The initial image is then sent to the image detection module 530. The image detection module 530 detects the initial image to obtain image detection information.

[0082] This specification provides an image processing system according to one embodiment, including a request processing module, an image generation module, an image detection module, and a reference information library. The request processing module handles image generation requests submitted by users through a client. The image generation module generates an initial image based on image information and target prompt information. The image detection module performs compliance checks on the initial image generated by the image generation module. The reference information library stores image elements used for image generation. Upon receiving an image processing request, the request processing module parses the request to obtain image information; determines image elements matching the image information across multiple visual attribute dimensions; and generates target prompt information based on the image elements and the initial prompt information. The image information and target prompt information are then sent to the image generation module. The image generation module queries the reference information library to generate the initial image based on the image information and target prompt information; the initial image is then sent to the image detection module. The image detection module detects the initial image and obtains image detection information. Detecting the initial image and obtaining image detection information ensures compliance of the initial image while improving the generation efficiency and usability of the initial image.

[0083] The following is in conjunction with the appendix Figure 6 Taking the image processing method provided in this specification in the generation of product promotional images as an example, the image processing method will be further explained. Among other things, Figure 6 An interactive schematic diagram of an image processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0084] Step 602: The user selects the task tool at the top of the image generation page and determines the image type.

[0085] Step 604: The front-end module retrieves the default configuration and placeholders based on the image type and sends them to the gateway module.

[0086] Step 606: The gateway module establishes or locates a session through the session module.

[0087] Step 608: The front-end module generates a structured prompt template that matches the image type and automatically fills in the default values ​​to generate image information.

[0088] Step 610: The front-end module displays the image information to the user, including auto-suggestions and style, size, and logo selectors.

[0089] Step 612: The user interacts with the image generation control via touch to inform the front-end module to generate an image with one click.

[0090] Step 614: The front-end module submits image information such as structured prompts and system prompts to the orchestration module.

[0091] Step 616: The orchestration module performs pre-compliance injection through the compliance module.

[0092] Step 618: The orchestration module submits the image generation task to the queue module.

[0093] Step 620: The queue module routes the image generation task to the large model engine for image generation.

[0094] Step 622: The adapter pool generates an initial graph through the merging graph module and performs merging and orchestration.

[0095] Step 624: The composite drawing module performs compliance verification on the initial drawing through the compliance module.

[0096] Step 626: The composite image module receives the verification pass information.

[0097] Step 628: The composite drawing module generates the final drawing and packages it.

[0098] Step 630: The output module stores the finished product image and its data.

[0099] Step 632: The output module returns a preview and download link for the finished product image to the front-end module.

[0100] Step 634: The front-end module displays the image generation result to the user.

[0101] The image processing method provided in one or more embodiments of this specification achieves efficient collaboration between its modules through message queues. It improves generation accuracy by using prompt word templates. The sequential invocation and fusion of text generation, screen layout, and visual model further enhances image generation efficiency.

[0102] Figure 7 This specification illustrates a schematic diagram of another image processing system provided in one embodiment, as shown below. Figure 7As shown, the image processing system 700 includes a client 710 and a server 720. The client 710 is configured to generate an image processing request in response to a user's touch operation on an image processing page and send the image processing request to the server 720. The server 720 is configured to parse the image processing request to obtain image information; determine image elements matching the image information across multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information; input the image information and the target prompt information into an image generation module, and generate an initial image by querying a reference information database; detect the initial image to obtain image detection information; and feed back the image detection information and the initial image to the client 710.

[0103] This specification provides an image processing system in one embodiment. The client responds to a user's touch operation on an image processing page, generates an image processing request, and sends the request to the server. The server receives the image processing request, parses it, and obtains image information. It determines image elements matching the image information across multiple visual attribute dimensions and generates target prompt information based on these image elements and initial prompt information. This target prompt information is used to guide the image generation module in the visual dimensions of the image. The image information and target prompt information are input into the image generation module, which then generates an initial image by querying a reference information library. By accessing the reference information library, elements that assist the image generation module in generating the initial image are obtained, resulting in a richly detailed initial image. Detection of the initial image and obtaining image detection information ensures compliance while improving the efficiency and usability of initial image generation.

[0104] In practical applications, such as Figure 8 As shown, the image processing system comprises four layers: a first layer (user and entry layer), a second layer (conversation and AI platform layer), a third layer (engine and execution layer), and a fourth layer (data and output layer). The user and entry layer includes a front-end interaction layer 110 and an API gateway processing layer 120. The conversation and AI platform layer includes conversation and chat 131, task semantic parsing 132, prompt parameter repository 145, resource library linkage 140, AI orchestration and fusion 133, and compliance and rule engine 134. The engine and execution layer includes task scheduling and queues, container orchestration controllers, brand input services, model adapters, font and layout services, masking and matting services, layout and compositing engine, quality enhancement services, formatting and packaging services, and object storage. The data and output layer includes a data storage layer 160 and an output and distribution layer 180.

[0105] The first layer is the user and entry point layer. The front-end input is: when a user clicks the "Task Tools" tab (e.g., splash screen / product main image / external poster), it displays selectable styles / sizes / logos, natural language supplementary descriptions, and image upload controls. The front-end retrieves the default configuration and placeholder templates for the task, automatically fills in a structured prompt draft in the input boxes, and presents style / size / logo selectors. It then assembles a "Generate Request" draft (including task, style, size, logo selection, user-supplemented text, session ID, etc.). The output is a "Generate Request" message body, submitted to the API gateway. The API gateway performs authentication (token / employee ID), parameter compliance checks (size / ratio / null values), rate limiting, and routing. Output content can be: "workNo":"A12345"; "sessionId":null; "generationType":"openScreen"; "payload":{...},; "traceId":"T-2025-12-01-001". Forward it to the conversation and AI middleware layer.

[0106] The second layer is the conversation and AI platform layer. The task semantic parsing module includes: Task semantic extraction: Based on rules and LLM, it performs semantic understanding of user input, extracting scene elements, text fields, and explicit user needs (such as "bedroom," "green plants," and "realistic"); Parameter standardization: It merges the size, style, and brand elements in the task configuration with the user's selections to generate structured parameters; Placeholder replacement: It fills in fixed task slots (APP name, activity name, title, subtitle, size, logo) from the prompt template, without style optimization or image expansion. The AI ​​orchestration and fusion module automatically selects the corresponding style template based on style tags in the structured parameters (such as "realistic photography," "3D," and "Chinese trend illustration"). It inputs the user's basic prompt, style system prompt, and task parameters into the LLM to generate a stable and controllable expanded version of the prompt. It injects the logo, color scheme, font, and template pointers returned by the resource library, and the template version and hyperparameters returned by the prompt parameter repository. It selects the image engine based on the task type and the image merging workflow based on the task to generate the final task ticket. The compliance and rules engine can automatically generate multiple versions of images in 1:1 / 4:5 / 9:16 / banner / custom sizes based on the task scenario (splash screen / external projection / waterfall layout). The canvas can be expanded using models or image merging workflows. It performs version consistency checks, font compliance checks, copyright checks, and portrait rights checks.

[0107] The third layer is the engine and execution layer. This includes task scheduling and queue 155, container orchestration controller, model adapters, font and typography services, masking and cutout services, brand injection services, typography and compositing engine, quality enhancement services, formatting and packaging services, and object storage. Task scheduling and queue 155 can take task tickets as input. The process includes queuing, scheduling, timeout and retry strategies, and callback address registration (for asynchronous notifications). Outputs include queue status and consumption allocation records, which are distributed to the model adapter pool via the container orchestration controller. The typography and compositing engine takes initial output images, brand resources, typography parameters, and masks (optional). The process includes: creating canvases and layers, placing logos (considering safety margins / alignment lines / priority levels), text placement (titles / subtitles / legal text), and cropping, alignment, gridding, and whitespace processing according to template specifications. Outputs a composite image and composition list, which are forwarded to the quality enhancement module for quality enhancement.

[0108] The fourth layer is the data and output layer. The data storage layer 160 takes sessions / tasks / products / compliance reports / composite manifests / counts as input. Processing includes: transactional writes, updating statistics and usage counts, and archiving compliance reports and parameter snapshots (auditable / reproducible). Output is a persistent primary key and retrieval index (for historical image library, bounce, and reproduction). This is then provided to the output and distribution layer 180 and the front end. The output and distribution layer 180 outputs a visible result box and a "fine-tuning entry" (replace logo / resize / regenerate) to the front end.

[0109] See Figure 9 , Figure 9 A flowchart of a moving image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0110] Step 902: Receive the activity image processing request submitted by the user for the target activity, and parse the activity image processing request to obtain activity image information; Step 904: Determine the activity image element that matches the activity image information based on the multiple visual attribute dimensions associated with the target activity, and generate target prompt information based on the activity image element and the initial prompt information; Step 906: Input the moving image information and the target prompt information into the moving image generation module, and generate an initial moving image containing the target moving information by querying the moving element database; Step 908: Perform an activity content compliance check on the initial activity image to obtain compliance check information.

[0111] This specification provides an embodiment of an active image processing method that receives an active image processing request submitted by a user for a target activity, parses the request, and obtains active image information. Active image elements matching the active image information are determined across multiple visual attribute dimensions associated with the target activity. Target prompt information is generated based on the active image elements and initial prompt information, used to prompt the active image generation module in the image's visual dimensions. The active image information and target prompt information are input into the active image generation module, which generates an initial active image containing the target activity information by querying an activity element database. By accessing the activity element database, active image elements that assist the active image generation module in generating the initial active image are obtained, thus generating a visually rich initial active image. The initial active image is then inspected to obtain compliance inspection information. This ensures the compliance of the initial image while improving the generation efficiency and usability of the initial image.

[0112] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 10 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 10 As shown, the device includes: The parsing unit 1002 is configured to receive an image processing request and parse the image processing request to obtain image information; The generation unit 1004 is configured to determine image elements that match the image information in multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information; The input unit 1006 is configured to input the image information and the target prompt information into the image generation module, and generate an initial image by querying a reference information database. The detection unit 1008 is configured to detect the initial image and obtain image detection information.

[0113] In an optional embodiment, the parsing unit 1002 is further configured to: The image processing request is parsed to obtain image type information and input information. The image type information corresponds to the image type selected by the user on the image generation page, and the input information is the information entered by the user through the image generation page. The image type information and the input information are used as the image information.

[0114] In an optional embodiment, the generation unit 1004 is further configured to: Identify the editable attribute dimensions and associated attribute dimensions included in multiple visual attribute dimensions; In the editable attribute dimension, an editable image element matching the image information is determined, and in the associated attribute dimension, an associated image element matching the editable image element is determined; The editable image element and the associated image element are used as the image element.

[0115] In an optional embodiment, the generation unit 1004 is further configured to: In the editable attribute dimension, a reference image element matching the image information is determined; In response to an element update request submitted for the reference image element, an element selection page containing at least one candidate image element is generated; In response to a confirmation request submitted for a target element contained in at least one candidate image element in the element selection page, the editable image element corresponding to the target element is determined.

[0116] In an optional embodiment, the input unit 1006 is further configured to: The image information and the target prompt information are input into the image generation module, and the text processing model and image processing model contained in the image generation module after input are determined. The text processing model and the image processing model are used to query the reference information database to obtain reference text and reference images; The reference text, the image information, and the target prompt information are processed using the text processing model to obtain image text, and the reference image, the image information, and the target prompt information are processed using the image processing model to obtain a layout image, and the initial image is generated based on the image text and the layout image.

[0117] In an optional embodiment, the input unit 1006 is further configured to: The text prompt information is determined from the target prompt information; The text processing model is used to process the reference text, the image information, and the text prompt information to obtain image text.

[0118] In an optional embodiment, the input unit 1006 is further configured to: Image prompt information is determined from the target prompt information; The image processing model is used to process the reference image, the image information, and the image prompt information to obtain a layout image.

[0119] In an optional embodiment, the detection unit 1008 is further configured to: Determine at least one detection dimension, and perform detection on the initial image according to the at least one detection dimension to obtain the image detection information; If it is determined that the initial image fails the detection based on the image detection information, the initial image is updated to an intermediate image based on the image detection information. The target image is generated based on the image template and the image information corresponding to the intermediate image.

[0120] In an optional embodiment, the detection unit 1008 is further configured to: Text information and image content information are determined from the image information corresponding to the intermediate image, and text regions and image regions are determined from the image template; Text content is generated in the text area based on the text information, and image content is generated in the image area based on the image information; The target image is generated based on the text content and the image content.

[0121] In an optional embodiment, the detection unit 1008 is further configured to: Determine at least one image processing type, and determine a target image processing type among the at least one image processing type; The target image is processed according to the target image processing type to obtain the image to be displayed.

[0122] An embodiment of this specification provides an image processing apparatus that receives an image processing request, parses the request, and obtains image information. It determines image elements matching the image information across multiple visual attribute dimensions and generates target prompt information based on these image elements and initial prompt information. This target prompt information is used to prompt the image generation module in the visual dimensions of the image. The image information and target prompt information are input into the image generation module, which generates an initial image by querying a reference information library. By accessing the reference information library, elements that assist the image generation module in generating the initial image are obtained, resulting in a richly detailed initial image. Detection of the initial image and obtaining image detection information ensures compliance while improving the efficiency and usability of initial image generation.

[0123] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0124] Corresponding to the above method embodiments, this specification also provides embodiments of a moving image processing apparatus. Figure 11 A schematic diagram of the structure of an active image processing apparatus according to one embodiment of this specification is shown. Figure 11 As shown, the device includes: The parsing unit 1102 is configured to receive an activity image processing request submitted by a user for a target activity, and to parse the activity image processing request to obtain activity image information; The generation unit 1104 is configured to determine an active image element that matches the active image information in multiple visual attribute dimensions associated with the target activity, and generate target prompt information based on the active image element and initial prompt information; The input unit 1106 is configured to input the activity image information and the target prompt information into the activity image generation module, and generate an initial activity image containing target activity information by querying the activity element database. The detection unit 1108 is configured to perform activity content compliance detection on the initial moving image and obtain compliance detection information.

[0125] This specification provides an embodiment of an active image processing apparatus that receives an active image processing request submitted by a user for a target activity, parses the request, and obtains active image information. It determines active image elements matching the active image information across multiple visual attribute dimensions associated with the target activity, and generates target prompt information based on the active image elements and initial prompt information to prompt the active image generation module in the image's visual dimensions. The active image information and target prompt information are input into the active image generation module, which generates an initial active image containing the target activity information by querying an activity element database. By accessing the activity element database, it obtains active image elements that assist the active image generation module in generating the initial active image, thereby generating a visually rich initial active image. The initial active image is then inspected to obtain compliance inspection information. This ensures the compliance of the initial image while improving its generation efficiency and usability.

[0126] The above is an illustrative scheme of a moving image processing apparatus according to this embodiment. It should be noted that the technical solution of this moving image processing apparatus and the technical solution of the moving image processing method described above belong to the same concept. For details not described in detail in the technical solution of the moving image processing apparatus, please refer to the description of the technical solution of the moving image processing method described above.

[0127] Figure 12A structural block diagram of a computing device 1200 according to an embodiment of this specification is shown. The components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.

[0128] The computing device 1200 also includes an access device 1240, which enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1240 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0129] In one embodiment of this specification, the aforementioned components of the computing device 1200 and Figure 12 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 12 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0130] The computing device 1200 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1200 can also be a mobile or stationary server.

[0131] The processor 1220 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above method.

[0132] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computing device can be referred to the description of the technical solution of the above method.

[0133] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described method.

[0134] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be referred to the description of the technical solution of the above method.

[0135] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0136] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to in the description of the technical solution of the above method.

[0137] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0138] The computer program / instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0139] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0140] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0141] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: Receive an image processing request and parse the image processing request to obtain image information; Image elements matching the image information are determined across multiple visual attribute dimensions, and target prompt information is generated based on the image elements and initial prompt information. The image information and the target prompt information are input into the image generation module, and the image generation module generates an initial image by querying the reference information database. The initial image is then detected to obtain image detection information.

2. The image processing method according to claim 1, wherein parsing the image processing request to obtain image information includes: The image processing request is parsed to obtain image type information and input information. The image type information corresponds to the image type selected by the user on the image generation page, and the input information is the information entered by the user through the image generation page. The image type information and the input information are used as the image information.

3. The image processing method according to claim 1, wherein determining the image elements matching the image information in multiple visual attribute dimensions includes: Identify the editable attribute dimensions and associated attribute dimensions included in multiple visual attribute dimensions; In the editable attribute dimension, an editable image element matching the image information is determined, and in the associated attribute dimension, an associated image element matching the editable image element is determined; The editable image element and the associated image element are used as the image element.

4. The image processing method according to claim 3, wherein determining the editable image element matching the image information in the editable attribute dimension includes: In the editable attribute dimension, a reference image element matching the image information is determined; In response to an element update request submitted for the reference image element, an element selection page containing at least one candidate image element is generated; In response to a confirmation request submitted for a target element contained in at least one candidate image element in the element selection page, the editable image element corresponding to the target element is determined.

5. The image processing method according to claim 1, wherein inputting the image information and the target prompt information into the image generation module, and generating an initial image by the image generation module by querying a reference information database, comprises: The image information and the target prompt information are input into the image generation module, and the text processing model and image processing model contained in the image generation module after input are determined. The text processing model and the image processing model are used to query the reference information database to obtain reference text and reference images; The reference text, the image information, and the target prompt information are processed using the text processing model to obtain image text, and the reference image, the image information, and the target prompt information are processed using the image processing model to obtain a layout image, and the initial image is generated based on the image text and the layout image.

6. The image processing method according to claim 5, wherein processing the reference text, the image information, and the target prompt information using the text processing model to obtain image text includes: The text prompt information is determined from the target prompt information; The text processing model is used to process the reference text, the image information, and the text prompt information to obtain image text.

7. The image processing method according to claim 5, wherein processing the reference image, the image information, and the target prompt information using the image processing model to obtain a layout image includes: Image prompt information is determined from the target prompt information; The image processing model is used to process the reference image, the image information, and the image prompt information to obtain a layout image.

8. The image processing method according to claim 1, wherein detecting the initial image to obtain image detection information includes: Determine at least one detection dimension, and perform detection on the initial image according to the at least one detection dimension to obtain the image detection information; If it is determined that the initial image fails the detection based on the image detection information, the initial image is updated to an intermediate image based on the image detection information. The target image is generated based on the image template and the image information corresponding to the intermediate image.

9. The image processing method according to claim 8, wherein generating the target image based on the image template and the image information corresponding to the intermediate image comprises: Text information and image content information are determined from the image information corresponding to the intermediate image, and text regions and image regions are determined from the image template; Text content is generated in the text area based on the text information, and image content is generated in the image area based on the image information; The target image is generated based on the text content and the image content.

10. The image processing method according to claim 8, further comprising, after detecting the initial image and obtaining image detection information: Determine at least one image processing type, and determine a target image processing type among the at least one image processing type; The target image is processed according to the target image processing type to obtain the image to be displayed.

11. A method for processing moving images, comprising: Receive an activity image processing request submitted by a user for a target activity, and parse the activity image processing request to obtain activity image information; The activity image element that matches the activity image information is determined by identifying multiple visual attribute dimensions associated with the target activity, and target prompt information is generated based on the activity image element and initial prompt information; The activity image information and the target prompt information are input into the activity image generation module, and an initial activity image containing the target activity information is generated by querying the activity element database. Perform a compliance check on the initial moving image to obtain compliance check information.

12. An image processing system, comprising a client and a server, including: The client is configured to generate an image processing request in response to a user's touch operation on the image processing page, and send the image processing request to the server. The server is used to parse the image processing request and obtain image information; Image elements matching the image information are determined across multiple visual attribute dimensions, and target prompt information is generated based on the image elements and initial prompt information. The image information and the target prompt information are input into the image generation module, and the image generation module generates an initial image by querying the reference information database. The initial image is detected to obtain image detection information; the image detection information and the initial image are then fed back to the client.

13. An image processing system, comprising a request processing module, an image generation module, an image detection module, and a reference information database; The request processing module is used to receive image processing requests, parse the image processing requests to obtain image information, determine image elements that match the image information in multiple visual attribute dimensions, and generate target prompt information based on the image elements and initial prompt information. The image information and the target prompt information are sent to the image generation module; The image generation module is used to generate an initial image based on the image information and the target prompt information by querying the reference information database. The initial image is sent to the image detection module; The image detection module is used to detect the initial image and obtain image detection information.

14. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

16. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.