Data processing method, moving image generation method, and data processing system
By receiving image generation requests, parsing and matching image templates, and generating target images, the problem of low image generation efficiency in existing technologies is solved, and efficient and accurate image generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG TMALL TECH CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115618A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of data processing technology, and in particular to data processing methods, moving image generation methods, and data processing systems. Background Technology
[0002] In fields such as intelligent design, ad generation, and content creation, the demand for automatically generating high-quality images based on user-provided content is growing. Current technologies typically use machine learning techniques such as diffusion models to generate images. While this method can synthesize images based on input descriptions, it still has significant shortcomings in terms of layout control and semantic understanding, resulting in low accuracy and usability of the generated images. Often, manual adjustments are required after image generation, leading to inefficiency. Therefore, a data processing method is urgently needed to address these issues. Summary of the Invention
[0003] In view of the above, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a moving image generation method, a data processing system, a data processing apparatus, a moving image generation apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0004] According to a first aspect of the embodiments of this specification, a data processing method is provided, comprising: Receive an image generation request and parse the image generation request to obtain image attribute data; A target image template matching the image attribute data is determined in the image template library. The target image template contains at least one image region and region description information corresponding to the at least one image region. The image attribute data and the target image template are input into the image generation model to obtain the target image, which contains the region images corresponding to the at least one image region.
[0005] According to a second aspect of the embodiments of this specification, a method for generating moving images is provided, comprising: Receive a user's request to generate an activity image for a target activity, and parse the request to obtain activity image attribute data; In the image template library associated with the target activity, an activity image template that matches the activity image attribute data is determined. The activity image template includes at least one activity content display area and area description information corresponding to each of the at least one activity content display areas. The activity image attribute data and the activity image template are input into the image generation model to obtain an activity content image containing target activity information. The activity content image includes regional content images corresponding to the at least one activity content display area, and the at least one regional content image contains at least one type of activity information.
[0006] According to a third aspect of the embodiments of this specification, a data processing system is provided, including a client and a server, comprising: The client is used to send an image generation request to the server; The server is configured to parse the image generation request to obtain image attribute data; determine a target image template that matches the image attribute data in the image template library, the target image template containing at least one image region and region description information corresponding to the at least one image region; input the image attribute data and the target image template into the image generation model to obtain a target image, and send the target image to the client, the target image containing region images corresponding to the at least one image region.
[0007] According to a fourth aspect of the embodiments of this specification, a data processing apparatus is provided, comprising: The receiving module is configured to receive an image generation request and parse the image generation request to obtain image attribute data; The determination module is configured to determine a target image template that matches the image attribute data in an image template library. The target image template contains at least one image region and region description information corresponding to the at least one image region. The input module is configured to input the image attribute data and the target image template into the image generation model to obtain the target image, wherein the target image includes the region images corresponding to the at least one image region.
[0008] According to a fifth aspect of the embodiments of this specification, a moving image generation apparatus is provided, comprising: The receiving module is configured to receive a user's request to generate an activity image for a target activity, and to parse the request to obtain activity image attribute data. The determination module is configured to determine an activity image template that matches the activity image attribute data in an image template library associated with the target activity. The activity image template includes at least one activity content display area and area description information corresponding to each of the at least one activity content display areas. The input module is configured to input the activity image attribute data and the activity image template into the image generation model to obtain an activity content image containing target activity information. The activity content image includes regional content images corresponding to the at least one activity content display area, and the at least one regional content image contains at least one type of activity information.
[0009] According to a sixth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above method.
[0010] According to a seventh aspect of an embodiment of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the method described above.
[0011] According to an eighth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0012] This specification provides a data processing method in one embodiment that obtains image attribute data by receiving and parsing an image generation request. A target image template matching the image attribute data is determined from an image template library. The target image template contains at least one image region and corresponding region description information for each of the at least one image region. The image attribute data and the target image template are input into an image generation model to obtain a target image. This enables automatic generation of the target image based on the image attribute data and the target image template, improving the generation efficiency of the target image. The target image contains region images corresponding to at least one image region. Using at least one image region, region description information, and image attribute data as generation constraints for the target image improves the controllability of the target image's layout and simultaneously increases the generation efficiency of each image region within the target image, avoiding mutual interference between image regions. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating a data processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of a target image template for a data processing method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the target image of a data processing method provided in one embodiment of this specification; Figure 4This is a schematic diagram of an image region illustrating a data processing method provided in one embodiment of this specification; Figure 5 This is a schematic diagram of a region image illustrating a data processing method provided in one embodiment of this specification; Figure 6 This is a flowchart illustrating an image generation process for a data processing method provided in one embodiment of this specification. Figure 7 This is a flowchart illustrating a method for generating moving images according to one embodiment of this specification; Figure 8 This is a schematic diagram of the structure of a data processing system provided in one embodiment of this specification; Figure 9 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this specification; Figure 10 This is a schematic diagram of the structure of a moving image generation apparatus provided in one embodiment of this specification; Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] The technical solutions provided in this application can employ deep learning models with relatively large parameter scales. However, this large model is merely an example; this application does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application can be artificial intelligence-based language models (LM) or multimodal models (MM).
[0019] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0020] Region Template: Refers to the layout specification of the screen divided by pixel size, used to define the generated areas and restricted areas (such as the header logo area and the bottom search box area).
[0021] Mask Image: A visual guide that controls the generated area during the generation process. AI generates content based on the masked area.
[0022] Workflow: refers to a customized generation process for different task types. Each workflow contains independent template, mask, and prompt word configurations.
[0023] To address the aforementioned technical problems, this specification provides a data processing method, and also relates to a moving image generation method, a data processing system, a data processing apparatus, a moving image generation apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0024] See Figure 1 , Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0025] Step 102: Receive an image generation request and parse the image generation request to obtain image attribute data.
[0026] Specifically, an image generation request can be a computer instruction submitted by a client to a server, requesting image processing. The client can be a mobile terminal capable of running applications or web pages with image generation capabilities, while the server can be a service provider offering data processing, logical operations, and image generation services to the client; the server can be a cloud server. The image generation request carries image attribute data, which is image description data determined based on user-provided interaction information. This user-provided interaction information includes, but is not limited to, image type information, the image itself, image description text, and requirement description text. This interaction information can be determined by the image generation page displayed on the client, such as the user entering text content, uploading an image, or selecting an image type. The image attribute data can represent the user's image generation requirements.
[0027] Based on this, the server receives the image generation request submitted by the client, parses the request, and obtains the image attribute data. The image generation request can be generated based on the content selected, input, and / or uploaded by the user on the image generation page provided by the client, representing the user's image generation needs.
[0028] Furthermore, considering that the image generation request submitted by the client is generated based on the user's image generation requirements, parsing the image generation request can yield data from multiple dimensions, such as image type data and image generation data. The specific implementation is as follows: The image generation request is parsed to obtain the image type data and image generation data determined by the user through the interactive page; the image type data and the image generation data are used as the image attribute data.
[0029] Specifically, the interactive page is an editable page provided by the client for users, offering an information input interface. The interactive page provides users with a variety of selectable image types, representing the user's image requirements. Image types include, but are not limited to, promotional posters, application splash screen images, marketing images, and product images. Users can select an image type from the multiple options by interacting with touch elements provided on the interactive page, or they can input the image type directly on the interactive page, including but not limited to text input and voice input. Image generation data can be data determined based on information interaction with the user, including but not limited to a description of the image generation requirement, and text or images uploaded by the user. Information interaction methods include, but are not limited to, text interaction, image interaction, and voice interaction. The description of the image generation requirement can be user input information determined through text interaction, image interaction, and / or voice interaction.
[0030] Based on this, the image generation request is parsed to obtain the image type data and interaction information determined by the user through an interactive page via selection and / or input. Image generation data containing text and / or images is then generated by parsing the interaction information. The image type data and image generation data are used as image attribute data to provide reference data for the subsequent generation of the target image.
[0031] For example, in image generation scenarios, image generation can be applied to the generation of various images in e-commerce, such as marketing images, promotional posters, product detail images, and application splash screen images. To meet users' customized needs, interactive pages can be provided. These interactive pages can be user-facing applications or web pages. Users can determine the image type data corresponding to their image generation requirements through input or selection, and can also interact with the client through the interactive web page, inputting text and uploading image materials. The client analyzes the text and / or image elements provided by the user to generate image generation data, and generates an image generation request based on this data and the image type data. The server responds to the image generation request submitted by the client, parses the request, and obtains image attribute data from multiple image generation dimensions, including the image type data and image generation data encapsulated by the client in the request.
[0032] In summary, using image type data and image generation data as image attribute data improves the comprehensiveness of image attribute data and facilitates the generation of more accurate target images.
[0033] Step 104: Determine a target image template that matches the image attribute data in the image template library. The target image template contains at least one image region and region description information corresponding to the at least one image region.
[0034] Specifically, after receiving and parsing the image generation request to obtain image attribute data, a target image template matching the image attribute data can be determined from the image template library. The target image template contains at least one image region and corresponding region description information for each of the at least one image region. The image template library stores multiple image templates, each containing at least one image region and corresponding region description information. Image regions within an image template are divided by boundary lines, and these regions are connected to each other by these boundary lines. The image templates in the library are used to generate images of various functional types, including but not limited to promotional posters, application splash screens, marketing images, and product images. The target image template is the one matching the image attribute data, found in the image template library based on the target image type contained in the image attribute data. The at least one image region within the target image template is adjacent to each other, collectively forming the target image template. At least one image region corresponds to a region type. The region function type includes, but is not limited to, restricted drawing region, title region, content region, and touch region. The region description information includes, but is not limited to, a description of the content type that can be generated corresponding to the image region, a description of the image region attributes, and a description of the image generation constraints.
[0035] Based on this, after receiving and parsing the image generation request to obtain image attribute data, template matching is performed in the image template library based on the image attribute data. According to the matching result, a target image template is extracted from the image template library. The target image template contains at least one image region and corresponding region description information for each of the at least one image region. The at least one image region in the target image template is divided from each other by region boundaries.
[0036] Furthermore, considering that the image template library contains multiple candidate image templates, in order to improve selection efficiency when selecting the target image template from among the candidate image templates, template matching can be performed in the image template library based on image attribute data. The specific implementation is as follows: Determine the target image type corresponding to the image attribute data, and match the target image type with the candidate image templates contained in the image template library; select the target image template that matches the target image type from the image template library according to the matching result.
[0037] Specifically, the target image type is the image type of the image to be generated that corresponds to the image generation request. The matching method for matching the target image type with the candidate image templates contained in the image template library can be either image-text matching or text matching. That is, matching the candidate image templates of the target image type in the image template library, or matching the candidate image types of the candidate image templates in the candidate template library in the text dimension based on the target image type, to determine the target image template.
[0038] Based on this, the target image type corresponding to the image attribute data is determined, and the target image type is matched with candidate image templates contained in the image template library. This matching is done using either image-text matching or text matching to find the target image template corresponding to the target image type in the image template library. Based on the matching results, a target image template matching the target image type can be selected from the image template library.
[0039] Continuing with the previous example, the target image type can be an application splash screen image, i.e., the image displayed after launching the application. There are various image types, such as resource image, marketing event image, and product detail image. When the target image type is a splash screen image, then the target image type is the splash screen image type. The image target library stores candidate image templates corresponding to various image types, such as resource image, marketing event image, and product detail image. By querying the image template library based on the splash screen image type, the splash screen image template corresponding to that type can be obtained; this splash screen image template is the target image template.
[0040] In summary, by matching the target image type with the candidate image templates in the image template library based on the matching results obtained by matching the target image type with the candidate image templates contained in the image template library, the efficiency and accuracy of target image determination can be improved.
[0041] Step 106: Input the image attribute data and the target image template into the image generation model to obtain the target image, wherein the target image contains the region images corresponding to the at least one image region.
[0042] Specifically, after determining the target image template matching the image attribute data in the image template library, and after the target image template contains at least one image region and corresponding region description information for each of the at least one image region, the image attribute data and the target image template can be input into the image generation model to obtain the target image. The target image contains region images corresponding to at least one image region. The image generation model can be a machine learning model with text-based image generation capabilities, or a machine learning model with text and image-based image generation capabilities. The image generation model can also be a large language model or a multimodal large language model used to process multimodal data composed of text, images, or images and text. The target image is the image corresponding to the image generation request, and is the image that matches the image attribute data.
[0043] Based on this, after determining the target image template matching the image attribute data in the image template library, and ensuring that the target image template contains at least one image region and corresponding region description information for each of the at least one image region, the image attribute data and the target image template are input into the image generation model. The image generation model then generates a target image containing the region images corresponding to each of the at least one image region. In the target image, the region images are arranged according to the positions of the image regions in the target image template. Generating the target image using AI can improve the generation efficiency of the target image.
[0044] Furthermore, considering that the target image template contains at least one image region, and each image region corresponds to a region mask template for image generation, the region description information of the region mask template describes the attributes of the region mask template. The region mask template and the corresponding region description information constitute the template information pair, which is input to the image generation model to facilitate the generation of the target image. The specific implementation is as follows: Generate region mask templates corresponding to at least one image region in the target image template; perform region function parsing on at least one region mask template to obtain at least one region description information; construct template information pairs based on at least one region mask template and at least one region description information, and input at least one template information pair and the image attribute data into the image generation model to obtain the target image.
[0045] Specifically, region masking templates are used to mask image regions within a target image template. The masked region represents the content generation region. Region masking templates allow for precise layout control of the generated target image. A region masking template is a pixelated representation of an image region. At least one image region corresponds to a different image function region via a region masking template. These image function regions include, but are not limited to, background areas, search box areas, object areas, person areas, and text areas. Region description information describes the image function of the region corresponding to the region masking template, including but not limited to function type, region color, region size, and region reference image. A template information pair consists of a region masking template and its corresponding region description information. When the image generation model processes multiple region masking templates to generate image content, the region description information corresponding to each region masking template can be clearly defined.
[0046] Based on this, region mask templates corresponding to at least one image region in the target image template are generated, with each region mask template covering one of the image regions in the target image template. Region function parsing is performed on the at least one region mask template to obtain at least one region description, which describes the functional type and image type of the image region. Template information pairs are constructed based on the at least one region mask template and the at least one region description, and these template information pairs, along with image attribute data, are input into the image generation model to obtain the target image. The image generation model processes each template information pair to generate the image corresponding to each image region, and the images corresponding to at least one image region constitute the target image.
[0047] Continuing the previous example, precise layout control of AI images can be achieved through masking regions (region masking templates). A region masking template is essentially a mask image, a pixelated representation of the target image template. Different functional areas in the target image template (such as the character area, title area, background area, and button area) are mapped to masking regions of different colors or numbers. During image generation, each masking region corresponds to a region description, which is a descriptive term (e.g., "Generate main visual character," "Generate main title text," "Generate background scene"). Each masking region, along with its corresponding local descriptive term, is input into an image generation model that supports region control. The image generation model can then generate image content within each masking region. For example... Figure 2 As shown, the blue areas in the target image template represent the areas where content can be generated, which may include background areas, people areas, and text areas. The background area corresponds to... Figure 2 The masking diagram in the image shows the background layer. The main title can be generated within the masked area corresponding to the background layer. The background area also corresponds to... Figure 2The masking diagram in Figure 2 shows the background layer. Background content can be generated within the masked area corresponding to the background layer. The character area corresponds to... Figure 2 The masking diagram in Figure 3 shows the main character layer. The masked area corresponding to the main character layer can generate the main character image. The text area corresponds to... Figure 2 The masking diagram in Figure 4 illustrates the subtitle text layer. The subtitle text layer's masked area can generate the promotional poster's subtitle. Each masked area corresponds to a local descriptive phrase, such as "scientific background scene," "model image," or "event title text." The content corresponding to each masked area is generated by AI. The image generation model generates a sci-fi scene in the blue background area, a model image in the blue character area, and title layout in the blue text bar area, while other gray areas retain the original template. The final output image content is strictly positioned within its respective functional area, achieving a result such as... Figure 3 The promotional poster shown is the target image.
[0048] In summary, by inputting at least one template information pair and image attribute data into the image generation model, a target image can be obtained. The image generation model can generate image content for each image region, resulting in a target image with stable layout and content that does not cross boundaries.
[0049] Furthermore, considering that there is at least one template information pair, and each template information pair corresponds to an image region, the target image can be obtained by generating image content for each image region one by one. The specific implementation is as follows: The image generation model is invoked, and the at least one template information pair is processed based on the image attribute data to obtain the region images corresponding to the at least one template information pair; the target image is constructed based on the at least one region image.
[0050] Based on this, an image generation model is invoked, and at least one template information pair is processed based on the image attribute data to obtain region images corresponding to each template information pair. Region images can be generated based on user-provided reference elements in the image attribute data; that is, images are generated in the regions corresponding to the region masking templates in the template information pairs according to the region description information in the image attribute data and template information pairs, thus obtaining at least one region image. A target image is then constructed based on at least one region image. For at least one template information pair, image generation can be performed in parallel or region images can be generated sequentially.
[0051] Continuing with the previous example, such as Figure 3As shown, at least one template information pair includes template information pairs corresponding to the background area, the person area, and the text area, respectively. The template information pairs corresponding to the background area, the person area, and the text area can be processed in parallel; that is, image content can be generated in parallel within the background area, the person area, and the text area. After obtaining the region images corresponding to the background area, the person area, and the text area, they can be integrated into the target image.
[0052] In summary, a target image is constructed based on at least one region image. For at least one template information pair, image generation can be performed in parallel or region images can be generated sequentially, improving both the efficiency and accuracy of target image generation.
[0053] Furthermore, considering that the generation processes of each region's image are independent of each other, when generating a region's image, to prevent the region's image from exceeding the boundaries of the image region, the non-image regions corresponding to the template information pair can be determined. This ensures that the non-image regions corresponding to the image regions in the template information pair do not generate image content. The specific implementation is as follows: The image generation model is invoked to determine the non-image regions corresponding to the at least one template information pair; the at least one template information pair is processed based on the image attribute data and the at least one non-image region to obtain the region images corresponding to the at least one template information pair.
[0054] Based on this, an image generation model is invoked to determine at least one non-image region corresponding to each template information pair. The non-image region of the template information pair is the area in the target image template other than the image region contained in the template information pair; non-image regions represent blank areas where image generation is prohibited. Each template information pair corresponds to at least one non-image region. Based on image attribute data and at least one non-image region, each template information pair is processed to obtain the region image corresponding to each template information pair. Processing the template information pair involves generating image content in the region masking template according to the region description information in the template information pair, thus obtaining the region image corresponding to the region masking template.
[0055] Continuing with the previous example, such as Figure 4 In the target image template shown, when generating the image corresponding to the core content area, it can be generated by AI. The core content area and its corresponding area description information form a template information pair. When generating the image in the core content area, other areas in the target image template (white space, logo display area, search box area) are considered non-image areas, and image generation is prohibited in these areas. The image attribute data provides user-specified parameters such as image style and tone. By processing at least one template information pair based on the image attribute data and at least one non-image area, images can be generated independently in each image area. Figure 5The image of each area in the middle is as follows: the core content area generates the core marketing content "Projector 2 Pro, subsidy up to 15% and a picture of the projector", the blank area can be filled with background color, the LOGO display area generates "Super New Product and background color", and the search box area generates the search box for "Super New Product".
[0056] In summary, by processing image attribute data and at least one non-image region for at least one template information pair respectively, regional images corresponding to at least one template information pair are obtained, so that at least one image region can generate image content independently of each other, avoiding interference between image regions.
[0057] Furthermore, considering that the generated target image is user-facing and may contain sensitive words, false advertising, or other non-compliant content, it is necessary to perform detection on the target image according to preset detection dimensions after generation. The specific implementation is as follows: Determine at least one preset detection dimension, and perform detection on the target image in each of the at least one preset detection dimensions to obtain detection data corresponding to each of the at least one preset detection dimensions; update the target image based on the at least one detection data to obtain an image to be displayed.
[0058] Specifically, preset detection dimensions can be determined based on actual image detection needs. Preset detection dimensions represent preset detection conditions and standards used to perform image detection on the target image from multiple aspects. At least one preset detection dimension includes, but is not limited to, image size, font, copyright, image version, compliance, and color dimensions. The detection data corresponding to each preset detection dimension is the detection result obtained by detecting the target image through each preset detection dimension. Detection data includes, but is not limited to, detection conclusions, image anomaly data, and image modification suggestions, among other data.
[0059] Based on this, at least one of the following dimensions—image size, font, copyright, image version, compliance, and color—is selected as at least one preset detection dimension. The target image is then detected using each of these preset dimensions, yielding detection data for each dimension. This detection data may include data from multiple dimensions, such as detection conclusions, image anomaly data, and image modification suggestions. The target image is then updated based on the at least one detection data point, achieving fine-tuning and resulting in the image to be displayed.
[0060] Following the previous example, after generating the base image (target image), based on preset scenario requirements, at least one detection dimension is selected from dimensions such as image size, font, copyright, image version, compliance, and color as preset detection dimensions, and the target image is then inspected based on at least one preset detection dimension. In the size dimension, the detection data can include size extension suggestions. The target image is then size-extended, and multiple aspect ratio versions are automatically generated, such as 1:1, 4:5, and 9:16. In the compliance dimension, the target image undergoes compliance detection to check for sensitive words or image elements. If any are found, modification suggestions are proposed, and corresponding compliance detection data is generated. Furthermore, the target image can also undergo font compliance detection (identifying text in the output image and comparing it with font authorization information to avoid using unauthorized fonts), copyright detection (identifying whether the image contains elements of well-known brands or protected content and comparing it with the authorization library), and image infringement detection (identifying recognizable faces and confirming whether they are authorized materials to avoid infringing on the portrait rights of public figures or third parties).
[0061] In summary, the target image is detected in at least one preset detection dimension, and updated based on at least one detection data point to obtain the image to be displayed. This achieves compliance detection of the target image followed by compliance processing, improving the usability, reliability, and compliance of the image to be displayed.
[0062] Furthermore, after obtaining the image to be displayed, it indicates that the image generation task corresponding to this image generation request has been completed. Various data generated during the execution of this image generation task, such as image attribute data, target image template, and target image, can be persistently stored. The specific implementation is as follows: An image generation record is constructed based on the image attribute data, the target image template, the target image, and / or the image to be displayed, and the image generation record is persistently stored.
[0063] Based on this, an image generation record is constructed using image attribute data, target image templates, target images, and / or images to be displayed. This record documents the execution process and results of the image generation task corresponding to the image generation request. The image generation record is persistently stored. This persistent storage facilitates the backtracking of the image generation task's execution process. The image attribute data, target image templates, and target images within the image generation record can also be used to construct training samples for the image generation model, further optimizing the model.
[0064] Continuing with the previous example, when generating the "Future Fashion Week" poster (the image to be displayed), key data related to the generation process (image generation record) can be recorded, including but not limited to: image template ID, mask image version, mask files corresponding to each area such as the background area, character area, and title area; the user's original input content (such as event theme, product information, title text, etc.); local descriptive words for the background, character descriptive words, and title descriptive words; the model version used; generation parameters such as inference steps and random seed; and the file ID and generation time of the final generated image. Persistently storing the image generation record facilitates backtracking of the image generation task's execution process.
[0065] In summary, persistent storage of image generation records facilitates the backtracking of the image generation task execution process and the updating of the image generation model.
[0066] This specification provides a data processing method in one embodiment that obtains image attribute data by receiving and parsing an image generation request. A target image template matching the image attribute data is determined from an image template library. The target image template contains at least one image region and corresponding region description information for each of the at least one image region. The image attribute data and the target image template are input into an image generation model to obtain a target image. This enables automatic generation of the target image based on the image attribute data and the target image template, improving the generation efficiency of the target image. The target image contains region images corresponding to at least one image region. Using at least one image region, region description information, and image attribute data as generation constraints for the target image improves the controllability of the target image's layout and simultaneously increases the generation efficiency of each image region within the target image, avoiding mutual interference between image regions.
[0067] The following is in conjunction with the appendix Figure 6 Taking the data processing method provided in this specification in the application of marketing image generation as an example, the data processing method will be further explained. Among other things, Figure 6 This specification illustrates an image generation flowchart of a data processing method according to an embodiment of the present invention, which specifically includes the following steps.
[0068] Step 602: Parse the image generation request to obtain the image type data and image generation data determined by the user through the interactive page.
[0069] Image generation requests can be image generation instructions provided by users (design teams) through the image generation page, used to express image generation needs. Image generation needs include, but are not limited to, image type requirements and image style requirements. Image type data refers to the image type corresponding to the image the user needs to generate, including but not limited to splash screen images, ad placements, marketing event images, and product detail images. Image generation data can be multimodal materials uploaded by users through the image generation page, or it can be dialogue between users on the image generation page, aimed at clarifying the user's image generation intent.
[0070] Step 604: Match the image type data with the candidate image templates contained in the image template library, and select the target image template that matches the target image type from the image template library based on the matching results. The target image template contains at least one image region and region description information corresponding to each of the at least one image region.
[0071] In practical applications, the image area is the masking area, and the area description information is the descriptive word. Specifically, it can be "the two sides of this image are clean gradient backgrounds, without text or people." Its specific function is to clarify the specific restrictions and controls on the content generated by each masking area.
[0072] Step 606: Input the image generation data and the target image template into the image generation model to obtain the target image, which contains at least one region image corresponding to each image region.
[0073] The specific steps for generating the target image include: ① loading the target image template; ② identifying the semantics of each masked region and generating descriptive words; ③ encoding the masked regions and descriptive words as model input; ④ the image generation model generating content in the corresponding masked regions; ⑤ outputting a target image with a regional layout consistent with the target image template. Precise layout control of the AI image is achieved through masked regions. The mask image corresponding to the masked region is a pixelated representation of the target image template. Different functional areas in the target image template (such as the character area, title area, background area, and button area) are mapped to areas of different colors or numbers. The image generation model generates a science fiction scene in the blue background area, a model image in the blue character area, and a title layout in the blue text bar area, while other gray areas retain the original template. The final output poster content is strictly placed within its respective functional area.
[0074] Step 608: Determine at least one preset detection dimension, and perform detection on the target image in each of the at least one preset detection dimensions to obtain detection data corresponding to each of the at least one preset detection dimensions.
[0075] Step 610: Update the target image based on at least one detection data point to obtain the image to be displayed.
[0076] The target image detection is performed by the output and verification module, which conducts multi-size processing, version consistency checks, and compliance checks on the generated target images. After generating the target image, this module extends the image size based on preset scenario requirements and automatically generates image versions in various aspect ratios, such as 1:1, 4:5, and 9:16. Before the image is added to the database, this module further performs compliance checks, including: font compliance checks (identifying text in the output image and comparing it with font authorization information to avoid using unauthorized fonts), IP copyright checks (identifying whether the image contains elements of well-known brands or protected content and comparing it with the authorization library), and portrait infringement checks (identifying recognizable faces and confirming whether they are authorized materials to avoid infringing on the portrait rights of public figures or third parties). If any of the multi-size images or compliance checks fail to meet the predetermined rules, the database entry operation is blocked or the content is regenerated. This significantly improves the usability, reliability, and compliance of the image generation method, ensuring that multiple versions of the output meet requirements at both the technical and legal levels.
[0077] Step 612: Construct an image generation record based on image attribute data, target image template, target image and / or image to be displayed, and persist the image generation record.
[0078] After generating the target image, key parameters related to the generation process can be recorded, including but not limited to: template ID, mask image version; mask files corresponding to each region such as the background area, character area, and title area; the user's original input content (such as event theme, product information, title text, etc.); local descriptive words for the background, character, and title; the model version used; generation parameters such as inference steps and random seed; and the file ID and generation time of the final generated image. The recorded data is then persistently stored.
[0079] The data processing method provided in one or more embodiments of this specification defines screen area templates (including logo area, content area, white space area, search box, etc.) in the generation workflow. Combined with masking images and descriptive word constraint mechanisms, it enables AI to generate content within the specified area and dynamically adapt to different task types (such as splash screen images, resource positions, marketing venues, product detail images, etc.). This significantly improves the layout controllability, compliance, and consistency of AI-generated images, reduces the degree of human intervention, ensures that the generated images meet the advertising specifications and brand visual standards, and possesses high reusability and engineering value.
[0080] See Figure 7 , Figure 7 A flowchart of a moving image generation method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0081] Step 702: Receive the activity image generation request submitted by the user for the target activity, and parse the activity image generation request to obtain activity image attribute data; Step 704: Determine an activity image template that matches the activity image attribute data in the image template library associated with the target activity. The activity image template includes at least one activity content display area and area description information corresponding to each of the at least one activity content display areas. Step 706: Input the activity image attribute data and the activity image template into the image generation model to obtain an activity content image containing target activity information. The activity content image contains area content images corresponding to the at least one activity content display area, and the at least one area content image contains at least one type of activity information.
[0082] This specification provides an embodiment of an activity image generation method. By receiving an activity image generation request submitted by a user for a target activity and parsing the request, activity image attribute data can be obtained. An activity image template matching the activity image attribute data is selected from an image template library associated with the target activity. The activity image template includes at least one activity content display area and corresponding area description information for each of the at least one activity content display area. The activity image attribute data and the activity image template are input into an image generation model to obtain an activity content image containing target activity information. This enables automatic generation of activity content images based on the activity image attribute data and the activity image template, improving the generation efficiency of activity content images. The activity content image includes at least one area content image corresponding to each activity content display area. Each area content image contains at least one type of activity information. Using at least one activity content display area, area description information, and activity image attribute data as generation constraints improves the controllability of the activity content image's layout and increases the generation efficiency of each activity content display area within the activity content image, while avoiding mutual interference between activity content display areas.
[0083] Corresponding to the above method embodiments, this specification also provides data processing system embodiments. Figure 8 A schematic diagram of the structure of a data processing system according to one embodiment of this specification is shown. Figure 8As shown, the data processing system 800 includes a client 810 and a server 820. The client 810 is used to send an image generation request to the server 820. The server 820 is used to parse the image generation request to obtain image attribute data; determine a target image template that matches the image attribute data in an image template library, the target image template containing at least one image region and region description information corresponding to the at least one image region; input the image attribute data and the target image template into an image generation model to obtain a target image, and send the target image to the client 810, the target image containing region images corresponding to the at least one image region.
[0084] In practical applications, when a user on the client side needs to generate an image, they can configure parameters or upload information on the client side. The client then generates an image generation request and sends it to the server. The server receives the image generation request, parses it, and obtains the image attribute data. A target image template matching the image attribute data is determined from the image template library. The target image template contains at least one image region and corresponding region description information for each of the at least one image region. The image attribute data and the target image template are input into the image generation model to obtain the target image. This enables automatic generation of the target image based on the image attribute data and the target image template, improving the generation efficiency. The target image is then sent to the client, which displays it to the user and provides a save interface for easy saving. The target image contains region images corresponding to at least one image region. Using at least one image region, region description information, and image attribute data as generation constraints improves the controllability of the target image's layout and increases the generation efficiency of each image region within the target image, preventing mutual interference between different image regions.
[0085] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 9 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 9 As shown, the device includes: The receiving module 902 is configured to receive an image generation request and parse the image generation request to obtain image attribute data; The determining module 904 is configured to determine a target image template that matches the image attribute data in the image template library. The target image template includes at least one image region and region description information corresponding to the at least one image region. The input module 906 is configured to input the image attribute data and the target image template into the image generation model to obtain the target image, wherein the target image includes the region images corresponding to the at least one image region.
[0086] In an optional embodiment, the receiving module 902 is further configured to: The image generation request is parsed to obtain the image type data and image generation data determined by the user through the interactive page; The image type data and the image generation data are used as the image attribute data.
[0087] In an optional embodiment, the determining module 904 is further configured to: Determine the target image type corresponding to the image attribute data, and match the target image type with the candidate image templates contained in the image template library; Based on the matching results, the target image template that matches the target image type is determined in the image template library.
[0088] In an optional embodiment, the input module 906 is further configured to: Generate region mask templates corresponding to at least one image region in the target image template; Perform region function parsing on at least one region masking template to obtain at least one region description information; Based on the at least one region masking template and the at least one region description information, a template information pair is constructed, and the at least one template information pair and the image attribute data are input into the image generation model to obtain the target image.
[0089] In an optional embodiment, the input module 906 is further configured to: The image generation model is invoked, and the at least one template information pair is processed based on the image attribute data to obtain the region images corresponding to the at least one template information pair respectively. The target image is constructed based on at least one region image.
[0090] In an optional embodiment, the input module 906 is further configured to: The image generation model is invoked to determine the non-image regions corresponding to the at least one template information pair. The at least one template information pair is processed based on image attribute data and at least one non-image region to obtain the region image corresponding to each of the at least one template information pair.
[0091] In an optional embodiment, the input module 906 is further configured to: Determine at least one preset detection dimension, and perform detection on the target image in the at least one preset detection dimension respectively to obtain detection data corresponding to the at least one preset detection dimension respectively; The target image is updated based on at least one detection data point to obtain the image to be displayed.
[0092] In an optional embodiment, the input module 906 is further configured to: An image generation record is constructed based on the image attribute data, the target image template, the target image, and / or the image to be displayed, and the image generation record is persistently stored.
[0093] The data processing apparatus provided in one embodiment of this specification obtains image attribute data by receiving and parsing an image generation request. A target image template matching the image attribute data is determined from an image template library. The target image template includes at least one image region and corresponding region description information for each of the at least one image region. The image attribute data and the target image template are input into an image generation model to obtain a target image. This enables automatic generation of the target image based on the image attribute data and the target image template, improving the generation efficiency of the target image. The target image includes region images corresponding to at least one image region. Using at least one image region, region description information, and image attribute data as generation constraints for the target image improves the controllability of the target image's layout and enhances the generation efficiency of each image region within the target image, while also preventing mutual interference between image regions.
[0094] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.
[0095] Corresponding to the above method embodiments, this specification also provides embodiments of a moving image generation apparatus. Figure 10 A schematic diagram of a moving image generation apparatus according to one embodiment of this specification is shown. Figure 10 As shown, the device includes: The receiving module 1002 is configured to receive a user's request to generate an activity image for a target activity, and to parse the request to obtain activity image attribute data. The determining module 1004 is configured to determine an activity image template that matches the activity image attribute data in an image template library associated with the target activity. The activity image template includes at least one activity content display area and area description information corresponding to each of the at least one activity content display areas. The input module 1006 is configured to input the activity image attribute data and the activity image template into the image generation model to obtain an activity content image containing target activity information. The activity content image includes regional content images corresponding to the at least one activity content display area, and the at least one regional content image contains at least one type of activity information.
[0096] This specification provides an embodiment of an activity image generation apparatus that receives and parses a user's activity image generation request for a target activity to obtain activity image attribute data. An activity image template matching the activity image attribute data is selected from an image template library associated with the target activity. The activity image template includes at least one activity content display area and corresponding area description information for each of the at least one activity content display area. The activity image attribute data and the activity image template are input into an image generation model to obtain an activity content image containing target activity information. This enables automatic generation of activity content images based on the activity image attribute data and the activity image template, improving the generation efficiency of activity content images. The activity content image includes at least one area content image corresponding to each activity content display area, and each area content image contains at least one type of activity information. Using at least one activity content display area, area description information, and activity image attribute data as generation constraints improves the controllability of the activity content image's layout and increases the generation efficiency of each activity content display area within the activity content image, while avoiding mutual interference between activity content display areas.
[0097] The above is a schematic scheme of a moving image generation apparatus according to this embodiment. It should be noted that the technical solution of this moving image generation apparatus and the technical solution of the moving image generation method described above belong to the same concept. For details not described in detail in the technical solution of the moving image generation apparatus, please refer to the description of the technical solution of the moving image generation method described above.
[0098] Figure 11 A structural block diagram of a computing device 1100 according to one embodiment of this specification is shown. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.
[0099] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0100] In one embodiment of this specification, the aforementioned components of the computing device 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0101] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1100 can also be a mobile or stationary server.
[0102] The processor 1120 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above method.
[0103] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computing device can be referred to the description of the technical solution of the above method.
[0104] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described method.
[0105] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be referred to the description of the technical solution of the above method.
[0106] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0107] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to in the description of the technical solution of the above method.
[0108] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0109] The computer program / instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0110] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0111] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0112] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Receive an image generation request and parse the image generation request to obtain image attribute data; A target image template matching the image attribute data is determined in the image template library. The target image template contains at least one image region and region description information corresponding to the at least one image region. The image attribute data and the target image template are input into the image generation model to obtain the target image, which contains the region images corresponding to the at least one image region.
2. The data processing method according to claim 1, wherein parsing the image generation request to obtain image attribute data includes: The image generation request is parsed to obtain the image type data and image generation data determined by the user through the interactive page; The image type data and the image generation data are used as the image attribute data.
3. The data processing method according to claim 1, wherein determining the target image template matching the image attribute data in the image template library comprises: Determine the target image type corresponding to the image attribute data, and match the target image type with the candidate image templates contained in the image template library; Based on the matching results, the target image template that matches the target image type is determined in the image template library.
4. The data processing method according to claim 1, wherein inputting the image attribute data and the target image template into the image generation model to obtain the target image includes: Generate region mask templates corresponding to at least one image region in the target image template; Perform region function parsing on at least one region masking template to obtain at least one region description information; Based on the at least one region masking template and the at least one region description information, a template information pair is constructed, and the at least one template information pair and the image attribute data are input into the image generation model to obtain the target image.
5. The data processing method according to claim 4, wherein inputting at least one template information pair and the image attribute data into the image generation model to obtain the target image comprises: The image generation model is invoked, and the at least one template information pair is processed based on the image attribute data to obtain the region images corresponding to the at least one template information pair respectively. The target image is constructed based on at least one region image.
6. The data processing method according to claim 5, wherein calling the image generation model and processing the at least one template information pair based on image attribute data to obtain the region images corresponding to the at least one template information pair respectively includes: The image generation model is invoked to determine the non-image regions corresponding to the at least one template information pair. The at least one template information pair is processed based on image attribute data and at least one non-image region to obtain the region image corresponding to each of the at least one template information pair.
7. The data processing method according to claim 1, after inputting the image attribute data and the target image template into the image generation model to obtain the target image, further comprising: Determine at least one preset detection dimension, and perform detection on the target image in the at least one preset detection dimension respectively to obtain detection data corresponding to the at least one preset detection dimension respectively; The target image is updated based on at least one detection data point to obtain the image to be displayed.
8. The data processing method according to claim 7, after updating the target image based on at least one detection data to obtain the image to be displayed, further comprising: An image generation record is constructed based on the image attribute data, the target image template, the target image, and / or the image to be displayed, and the image generation record is persistently stored.
9. A method for generating moving images, comprising: Receive a user's request to generate an activity image for a target activity, and parse the request to obtain activity image attribute data; In the image template library associated with the target activity, an activity image template that matches the activity image attribute data is determined. The activity image template includes at least one activity content display area and area description information corresponding to each of the at least one activity content display areas. The activity image attribute data and the activity image template are input into the image generation model to obtain an activity content image containing target activity information. The activity content image includes regional content images corresponding to the at least one activity content display area, and the at least one regional content image contains at least one type of activity information.
10. A data processing system, comprising a client and a server, including: The client is used to send an image generation request to the server; The server is used to parse the image generation request and obtain image attribute data; A target image template matching the image attribute data is determined in the image template library. The target image template contains at least one image region and region description information corresponding to the at least one image region. The image attribute data and the target image template are input into the image generation model to obtain the target image, and the target image is sent to the client. The target image contains the region images corresponding to the at least one image region.
11. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.