Data processing method and device
By using large-scale deep learning models to identify and match target images and texts, the problem of low efficiency in manually designed promotional images is solved, and efficient and high-quality promotional image generation is achieved, which is suitable for image design of Internet services.
Patent Information
- Application Number
- CN202510793051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the manual design of promotional images is inefficient and cannot meet the actual needs of Internet services. In particular, when designing product promotional images in online shopping scenarios, there are problems of low efficiency and high cost.
A large-scale deep learning model is used to identify and match target images and text to generate promotional images. This includes extracting target text and images from target publication content, using object recognition models to determine object locations, selecting appropriate text templates, and generating promotional images.
It improves the efficiency and quality of generating promotional images, reduces the cost of manual design, makes the generated images more attractive, meets the needs of advertising communication, and enhances the visual effect and communication effect of posters.
Smart Images

Figure CN120635255A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and more particularly to a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the continuous development of Internet technology, in the process of providing Internet services to users, it is necessary to design promotional images to promote Internet services. For example, in an online shopping scenario, it is necessary to design promotional images for the products being sold.
[0003] Currently, promotional images are designed and generated manually. However, this method is inefficient and cannot meet actual needs. Therefore, how to improve the efficiency of generating promotional images has become an urgent problem that needs to be solved. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0005] According to a first aspect of an embodiment of this specification, there is provided a data processing method, including:
[0006] Determining target text and target image from target published content, and selecting a plurality of candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include preset text filling positions;
[0007] Identifying a target object in the target image using an object recognition model to obtain an object position of the target object;
[0008] Based on the object position and the text filling position, selecting a target text template corresponding to the target image from the multiple candidate text templates;
[0009] A promotional image corresponding to the target published content is generated using the target text, the target image, and the target text template.
[0010] According to a second aspect of the embodiments of this specification, there is provided a data processing device, including:
[0011] A first template selection module is configured to determine a target text and a target image from the target published content, and select a plurality of candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include a preset text filling position;
[0012] a position recognition module, configured to recognize a target object in the target image using an object recognition model to obtain an object position of the target object;
[0013] A second template selection module is configured to select a target text template corresponding to the target image from the plurality of candidate text templates based on the object position and the text filling position;
[0014] The image generation module is configured to generate a promotional image corresponding to the target published content by using the target text, the target image and the target text template.
[0015] According to a third aspect of an embodiment of this specification, a computing device is provided, including:
[0016] memory and processor;
[0017] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.
[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the above-mentioned data processing method are implemented.
[0019] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0020] One or more embodiments of the present specification provide a data processing method. In the process of generating a promotional image, the target text and target image are first determined from the target published content, and multiple candidate text templates matching the target text are selected from the text template set; secondly, in order to generate a high-quality promotional image, the target object in the target image can be identified by using an object recognition model to obtain the object position of the target object, and based on the object position and the text filling position, a target text template matching the target image is selected from multiple candidate text templates; finally, based on the target text, target image and target text template, a promotional image corresponding to the target published content is quickly generated, thereby improving the generation efficiency of the promotional image and meeting actual needs, and through the target text template, target text and target image matching the target text and target image, a relatively high-quality promotional image can be generated, thereby improving the quality of the promotional image. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a schematic diagram of an application of a data processing method provided by an embodiment of this specification;
[0022] Figure 2 is a flow chart of a data processing method provided by one embodiment of this specification;
[0023] Figure 3 This is a flowchart of a data processing method provided by one embodiment of this specification;
[0024] Figure 4 This is a flow chart of a data processing method provided by an embodiment of this specification;
[0025] Figure 5 This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification;
[0026] Figure 6 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0027] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0028] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0029] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0030] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0031] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.
[0032] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0033] First, the terms involved in one or more embodiments of this specification are explained.
[0034] OCR (Optical Character Recognition): refers to optical character recognition; this OCR technology enables computers to recognize text content in images, thereby achieving automated processing.
[0035] Bounding box: refers to the detection bounding box.
[0036] CLIP (Contrastive Language-Image Pre-training): is a multimodal model.
[0037] With the continuous development of internet technology, the process of providing users with internet services requires the design of promotional images to promote these services. For example, in online shopping scenarios, promotional images for products need to be designed. Currently, promotional images are designed and generated manually. However, manual design methods are inefficient and cannot meet practical needs.
[0038] For example, in the current digital marketing landscape, small and medium-sized businesses in online shopping scenarios are increasingly relying on online advertising and social media to promote their products. However, high-quality poster design (i.e., promotional imagery) typically requires professionals to use specialized design skills and software to create it, increasing the difficulty and cost of self-operation for small and medium-sized businesses.
[0039] Based on this, this specification provides a text poster generation technology. However, this technology relies solely on simple templates and suffers from numerous deficiencies in processing complex visual elements. For example, issues such as text occlusion, lack of text clarity, and partial obscuration of the product body can lead to suboptimal click-through rates, impressions, and conversion rates. Additionally, this specification provides some automated tools for poster generation, but these tools often struggle to strike a good balance between aesthetics and efficiency.
[0040] Based on this, a data processing method is provided in this specification. One or more embodiments of this specification also involve a data processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0041] Considering the huge amount of model parameters and the limited computing resources of the mobile terminal, the data processing method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to this. Figure 1 In the illustrated application scenario, the large model is deployed on a server 10. Server 10 can be connected to one or more client devices 20 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Client devices 20 herein include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users via a graphical user interface (GUI) to access the large model and implement the methods provided in the embodiments of this specification.
[0042] In an embodiment of the present specification, the system composed of the client device 20 and the server 10 can perform the following steps: the client device 20 executes the step of sending a note to the server 10; the server 10 executes the steps of selecting a candidate template that matches the text, selecting a target template from multiple candidate templates, and generating a promotional poster based on the text, image, and target template. Selecting a candidate template that matches the text refers to comparing the number of characters in the text in the note with the maximum number of characters in the text slot, selecting multiple template materials whose maximum number of characters in the text slot is greater than or equal to the number of characters, and forming a candidate template list. Selecting a target template from multiple candidate templates refers to detecting the original note image using a visual detection model to obtain bounding box information, and selecting a target template from the candidate template list based on the bounding box information and the text slot. Generating a promotional poster based on the text, image, and target template refers to calling a template rendering service, generating a promotional poster using multiple note images, a target template, and the text, and sending the promotional poster to the client device 20.
[0043] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device.
[0044] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0045] Step 202: Determine target text and target image from target published content, and select multiple candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include preset text filling positions.
[0046] The data processing method provided in this embodiment can be applied to scenarios where promotional images are constructed based on user-posted content on any social platform. It can be used to perform secondary creation on user-posted content, increase the dissemination range of information in the posted content, reach more users, and increase users' enthusiasm for sharing information on social platforms.
[0047] The target published content can be understood as content published on the Internet, including text and images; the target published content can be content published on blogs, forums, websites, or content display programs. For example, the target published content can be a video or data containing text and images.
[0048] The target text can be understood as text extracted from the target published content, for example, the target text can be a copy, text displayed on a poster, etc. The target image can be understood as an image extracted from the target published content; the target image can be a background image displayed on a poster.
[0049] The text template set can be understood as a set composed of multiple text templates, wherein the various text templates contained in the text template set can be pre-constructed, and each template can define the text filling position, image filling position, promotional style, etc., so that after the template is selected, the text and image in the published content can be filled into the selected template to generate a promotional image for dissemination on the platform; wherein, different templates in the set correspond to different promotional styles, such as cartoon style, realistic style, movie style, oil painting style, comic style, etc., so that when the image and text are combined to generate a promotional image, it has the style corresponding to the template, so as to achieve the purpose of adapting to different user groups and attracting more users to browse the promotional image.
[0050] The multiple candidate text templates can be understood as multiple text templates that match the target text in the text template set. That is to say, after the target published content is determined, a template that matches the target text in the published content can be selected from the text template set for use. When selecting the candidate text template, it can be achieved by calculating the cosine similarity between each template and the text, or it can be selected by a selection instruction submitted by a secondary creative user corresponding to the target published content, or the system can automatically recommend a template with relatively high popularity as a candidate text template. In specific implementation, the method of selecting a template from the set can be set according to actual needs, and this embodiment does not impose any restrictions here.
[0051] The text filling position can be understood as the position where the target text needs to be filled or embedded, that is, the text filling position is the position where the target text is displayed or embedded in the promotional image; the text filling position can be a text filling position coordinate or a text filling position area.
[0052] Based on this, in the process of generating promotional images, the target text and target image will first be determined from the target publication content, and multiple candidate text templates matching the target text will be selected from the text template set; secondly, in order to generate high-quality promotional images, the object recognition model can be used to identify the target object in the target image, obtain the object position of the target object, and based on the object position and the text filling position, select the target text template matching the target image from multiple candidate text templates; finally, based on the target text, target image and target text template, the promotional image corresponding to the target publication content is quickly generated, thereby improving the generation efficiency of the promotional image and meeting actual needs, and through the target text template, target text and target image matching the target text and target image, relatively high-quality promotional images can be generated, thereby improving the quality of the promotional images.
[0053] In one or more embodiments provided herein, when extracting target text and target images from target published content, considering that the target published content may correspond to various formats, it is necessary to combine different types of published content to complete the text and image extraction operations. In this embodiment, the target published content is a target published video, and determining the target text and target image from the target published content includes:
[0054] Speech recognition is performed on the target published video to obtain video speech text, and a preset number of video frames are extracted from the target published video; a target text is determined based on the video speech text, and a target image is determined based on the video frames.
[0055] Determining the target text based on the video speech text refers to inputting the video speech text into a text rewriting model and rewriting the video speech text using the text rewriting model to obtain the target text. Determining the target image based on the video frame refers to performing image preprocessing on the video frame to obtain the target image.
[0056] In specific implementation, when the target release content is video release content, if a promotional image needs to be constructed for it, it is necessary to combine the lines involved in the video and the video screen to complete the construction of the promotional image. Before that, it is necessary to determine the target text and target image. Therefore, voice recognition can be performed on the target release video to obtain video voice text, and a preset number of video frames can be extracted from the target release video. On this basis, the target text can be determined based on the video voice text, such as inputting the video voice text into a large language model for summarization to obtain a target text with a set number of words. At the same time, the target image can be determined based on the video frame, such as randomly extracting any one video frame from a preset number of video frames as the target image, or fusing multiple video frames to generate a frame of image as the target image for subsequent use.
[0057] In one or more embodiments provided herein, if the target published content is a text-and-graphics published content, after the text and image are extracted, they may be processed to enable them to meet the requirements for subsequent promotional image generation. In this embodiment, determining the target text and target image from the target published content includes:
[0058] Extracting published text and published image from the target published content; rewriting the published text using a text rewriting model to obtain the target text; and performing image preprocessing on the published image to obtain the target image.
[0059] The published text can be understood as the text included in the target published content, for example, a title or description. The published image can be an image included in the target published content, for example, a product image or a scenic spot image. The text rewriting model can be understood as a neural network model capable of text rewriting. For example, the text rewriting model can be a large language model, a large model, a deep learning model, etc. When using a large language model as a text rewriting model, it is necessary to fine-tune the rewriting capabilities of the large language model to ensure that it has rewriting capabilities.
[0060] Based on this, when the target published content is text and picture content, the published text and published image can be extracted from the target published content; considering that the extracted images and text may not be suitable for subsequent use, such as the text is too long, or the wording of the text content is not suitable for publication, the image clarity is not enough, or the image size is not appropriate, etc., the published text can be rewritten using a text rewriting model to obtain a target text that meets the usage requirements; at the same time, the published image can be preprocessed to obtain a target image that meets the usage requirements.
[0061] In practice, if the rewritten text doesn't meet the requirements, the target text can be regenerated based on the semantics of the published text. Preprocessing of published images can also generate new images as target images, facilitating the subsequent generation of high-quality promotional images without deviating from the corresponding meaning of the published content.
[0062] The data processing method provided in this manual is described by taking its application in an advertising poster generation scenario as an example; the data processing method in this manual provides an intelligent template matching generation system based on multiple modal models; the user manually selects the target published content that he wishes to optimize for secondary creation in the client, and the client sends the target published content to the system; after the system receives the target published content, it automatically obtains the image (i.e., published image) and text (i.e., published text) contained in the target published content.
[0063] The system then inputs the text into a large language model (i.e., a text rewriting model) for rewriting, obtaining the promotional copy for the poster (i.e., the target text), and stores the promotional copy in the poster copy library. Furthermore, the system preprocesses the image to obtain the background image for the poster. Image preprocessing operations include, but are not limited to, cropping, filtering, formatting, and resizing the image.
[0064] Based on the above embodiments, it can be seen that the data processing method provided in this specification is applied to the advertising poster generation scenario. By introducing an intelligent template matching and content generation mechanism based on a multimodal model, the user's original published content (including images and text) is automatically analyzed and optimized for reconstruction, thereby significantly improving the efficiency and quality of advertising poster generation. By sending the published text input by the user into a large language model (i.e., a text rewriting model) for semantic understanding and intelligent rewriting, it is possible to automatically generate more attractive promotional copy that meets the needs of advertising communication, reducing the cost of manual writing, and to a certain extent ensuring the professionalism and consistency of the copy, thereby improving the efficiency and quality of advertising copy creation. In addition, by performing pre-processing operations such as cropping, filtering, format conversion, and size adjustment on the original published image, high-quality image materials suitable for poster backgrounds are obtained, which not only improves the visual expressiveness of the image, but also enhances its adaptability to the subsequently generated template, ensuring that the final output advertising poster has a good visual effect.
[0065] In one or more embodiments provided in this specification, when selecting multiple candidate text templates associated with a target text, considering that images have a better communication effect in the template, the candidate text templates can be screened by calculating the similarity between the image and the template. In this embodiment, the selection of multiple candidate text templates corresponding to the target text from the text template set includes steps 1 and 2:
[0066] Step 1: Based on the text attribute information of the target text, the text template set is selected from the template storage unit, wherein the text template set includes multiple text templates.
[0067] Among them, the text attribute information can be understood as information that characterizes the specific attributes of the target text. For example, the text attribute information can be the number of words in the target text or the text language type. The template storage unit can be understood as a unit for storing multiple text templates. The template storage unit can be a database, a server, or a local disk. For example, the template storage unit is a template material library. When selecting a text template set from the template storage unit, all text templates in the storage unit can be combined into a text template set, or a part of them can be selected from them as needed to form a text template set. For example, text templates of the same field or style can be selected according to the text attribute information to form a text template set.
[0068] Specifically, in the process of selecting a text template set, this method first analyzes the target text to obtain the text attribute information of the target text. The text attribute information can be used to represent the style, field, and other information of the target published content. Thereafter, the template text attribute information of the text template stored in the template storage unit can be determined, wherein the template text attribute information can be the text attribute information that represents the text adjusted by the text template; on this basis, based on the text attribute information and the template text attribute information, multiple text templates associated with the target text can be selected from the template storage unit, and the multiple text templates can be constructed into a text template set.
[0069] In one or more embodiments provided in this specification, when the text attribute information is the number of words in the text, selecting the text template set from the template storage unit based on the text attribute information of the target text includes:
[0070] Determine the text word count of the target text, and select multiple text templates whose template text word count is greater than or equal to the text word count from the template storage unit, wherein the template text word count is the text word count embedded in the text template; construct the text template set based on the multiple text templates.
[0071] Based on this, when selecting text templates to form a text template set according to the number of text words corresponding to the target text, multiple text templates whose template text word count is greater than or equal to the text word count can be selected from the template storage unit, wherein the template text word count is the number of text words embedded in the text template. On this basis, a text template set is constructed based on multiple text templates, so that candidate text templates can be screened and used subsequently.
[0072] Using the above example, in the actual production process, the copy and the template materials in the template library will be pre-matched to ensure that the number of characters in the copy meets the maximum number of characters required by the template slot, and then a candidate template list (i.e., text template set) is obtained. The specific method is:
[0073] First, identify the number of characters in the copy (i.e., the number of words in the text), and obtain the maximum number of words in the copy slot of the template material in the template material library (i.e., the number of words in the template text); compare the number of characters with the maximum number of words in the copy slot, and select multiple template materials whose maximum number of words in the copy slot is greater than or equal to the number of characters to form a candidate template list.
[0074] Based on the contents of the above embodiments, it can be seen that the data processing method provided in this specification effectively improves the matching efficiency and matching accuracy in the process of generating advertising posters by introducing a pre-matching mechanism between the copy and the template material. By making a preliminary comparison between the number of characters in the copy and the maximum number of characters in the copy slot in the template material, a list of candidate templates that meet the text capacity limit is screened out in advance, thereby avoiding invalid matching operations in the subsequent complex image and text fusion process, thereby significantly reducing the system computing load, improving the overall processing efficiency and template matching efficiency, and excluding mismatched items whose copy length exceeds the template carrying capacity during the matching stage, ensuring the adaptability of the copy and template in the final generated poster, reducing typesetting problems caused by excessively long copy or incorrect formatting, and improving the aesthetics and readability of the generated results.
[0075] In one or more embodiments provided in this specification, in addition, when the text attribute information is a text language type, selecting the text template set from a template storage unit based on the text attribute information of the target text includes:
[0076] Determine the text language type of the target text, and select multiple text templates whose template text language types are consistent with the text language type from the template storage unit, wherein the template text language type is the text language type of the text embedded in the text template; and construct the text template set based on the multiple text templates.
[0077] The text-to-speech type can be understood as the speech type corresponding to the target text, such as Chinese type, English type, Japanese type, etc. The text-to-speech type can be of multiple types.
[0078] Based on this, since different social platforms can support the publication of various types of content, after determining the target content, if a promotional image needs to be constructed for it, the language type needs to be ensured to be the same when selecting the template. Therefore, the text language type of the target text can be determined, and then multiple text templates whose template text language types are consistent with the text language type are selected from the template storage unit, and the text template set is constructed based on the multiple text templates, ensuring that the multiple text templates contained in the text template set are of the same language type as the target text, so that subsequent promotional images can also match the language type of the target text to reach more users.
[0079] In actual applications, in addition to the above-mentioned process of selecting text templates according to language type to form a text template set, all text templates stored in the template storage unit can also be translated, and the text template of the language type corresponding to the target text can be obtained according to the translation processing results. Constructing a text template set in this way can ensure that the text templates contained in the set are richer. During specific implementation, the method of constructing the text template set can be selected according to actual needs, and this embodiment does not make any restrictions here.
[0080] Using the above example, in the actual production process, the copy and the template materials in the template material library will be given priority for a round of pre-matching operations to ensure that the copy language type meets the copy language type requirements of the template, and then obtain a candidate template list (i.e., text template set). The specific method is:
[0081] First, identify the language type in the copy (i.e., the text language type), and obtain the template copy language type (i.e., the template text language type) corresponding to the template material in the template material library; compare the language type and the template copy language type, and select multiple template materials with the same language type and template copy language type to form a candidate template list.
[0082] Based on the above embodiments, it can be seen that the data processing method provided in this specification further introduces a pre-matching mechanism based on the text language type to ensure the language consistency between the copy and the template material, thereby achieving more efficient and accurate template matching in the advertising poster generation process. By identifying the language type of the copy (i.e., the text language type) and comparing it with the copy language type (i.e., the template text language type) corresponding to each template in the template material library, template materials with consistent language types are screened out to form a candidate template list. This process effectively avoids content incompatibility problems caused by language inconsistency and improves the accuracy and applicability of template matching.
[0083] Step 2: Calculate the similarity between each text template and the target image, and based on the similarity, select multiple candidate text templates corresponding to the target text from the text template set.
[0084] Specifically, after selecting the text template set from the template storage unit as mentioned above, considering that the text template set contains a large number of available templates, in order to be able to screen out text templates that are more suitable for use in the current scenario, the similarity between each text template and the target image can be calculated. The similarity can reflect the matching degree between the template and the current usage requirements, and then based on the similarity, multiple candidate text templates corresponding to the target text can be selected from the text template set.
[0085] In specific implementation, after calculating the similarity between the target image and the text template, the similarity can be sorted, and then the top N ranked ones are selected as candidate text templates, where the value of N can be set according to actual needs, and this embodiment does not impose any restrictions here.
[0086] In one or more embodiments provided herein, screening multiple candidate text templates can be accomplished by calculating similarities and then sorting them. In this embodiment, calculating the similarities between each text template and the target image, and selecting multiple candidate text templates corresponding to the target text from the text template set based on the similarities, includes:
[0087] A similarity analysis model is used to perform similarity analysis on each text template and the target image to obtain the similarity between each text template and the target image; based on the similarity, each text template in the text template set is sorted in descending order to obtain a text template sequence, and a preset number of text templates are selected from the text template sequence in a top-down manner as multiple candidate text templates corresponding to the target text.
[0088] The similarity analysis model can be understood as a model capable of performing similarity analysis to obtain similarity. For example, the similarity analysis model can be a CLIP model, a multimodal model, etc. Correspondingly, the text template sequence specifically refers to a sequence obtained by sorting the text templates included in the text template set in descending order. The sequence can be constructed using text template corresponding identifiers.
[0089] Based on this, when screening candidate text templates by calculating similarity, the similarity analysis model can be used to perform similarity analysis on the text templates and the target image, and then the similarity between the text templates and the target image can be obtained; on this basis, the text templates in the text template set can be sorted in descending order based on the similarity to obtain a text template sequence, and then a preset number of text templates can be selected from the text template sequence in a top-down manner as multiple candidate text templates corresponding to the target text.
[0090] In addition to descending sorting, an ascending sorting method can also be used to construct a text template sequence, and then a preset number of text templates are selected from the text template sequence in a bottom-up manner as multiple candidate text templates corresponding to the target text. In specific implementation, the selection can be based on actual needs, and this embodiment does not impose any restrictions here.
[0091] Continuing with the above example, after obtaining the candidate template list, this method will input the candidate templates and the original note image in the candidate template list into the CLIP multimodal model. The candidate template list and the original note image will be calculated through the CLIP multimodal model to calculate the image-text similarity. For example, the image-text similarity is any value in the range of [0,1]. The greater the image-text similarity, the more similar the candidate template and the original note image are; conversely, the less similar the candidate template and the original note image are.
[0092] After obtaining the image-text similarity, they will be sorted according to the similarity, and a fixed target number (i.e., a preset number) of image-text pair results will be selected from high to low, so as to obtain a fixed target number (e.g., 10) of image-text results (i.e., the optimal candidate results of image-template-text pair pairs), thereby taking into account the richness of the background image, template and text.
[0093] Based on the above embodiments, it can be seen that after obtaining a list of candidate templates, the data processing method provided in this specification further introduces a graphic-text similarity calculation mechanism based on the CLIP multimodal model to intelligently evaluate and sort the visual semantic consistency between the candidate templates and the original image, thereby achieving more accurate and efficient graphic-text matching and template recommendation. By inputting the candidate templates and the original note image into the CLIP multimodal model, the system can automatically extract the semantic features of the image and text and calculate the graphic-text similarity between the two (such as a value in the interval [0,1]), thereby achieving an assessment of the quality of the graphic-text combination. This mechanism significantly improves the intelligence of the graphic-text matching process and reduces the need for manual judgment.
[0094] Step 204: using an object recognition model to identify the target object in the target image, and obtaining the object position of the target object.
[0095] Specifically, when constructing a promotional image for the target published content, in order to be able to select a template that meets the current usage requirements from multiple candidate text templates, it is also necessary to use an object recognition model to identify the target object in the target image and obtain the object position of the target object, so as to facilitate the subsequent selection of the target text template in combination with the text filling position.
[0096] The object recognition model can be understood as a model that can recognize the target object. For example, the object recognition model can be a deep learning model or an object recognition model framework containing multiple subunits. The target object can be understood as a specific object of interest in the target image. For example, the target object can be text, a face, a human body, or a product in the target image. The object position can be understood as the location information of the target object in the target image. The object position can be the object position coordinates or the object position area.
[0097] In one or more embodiments provided in this specification, when performing object position recognition, in order to ensure recognition accuracy and comprehensiveness, recognition processing can be completed from multiple dimensions. In this embodiment, the target object includes text, human face, human body, and commodity, and the object position includes the position of text, human face, human body, and commodity; using the object recognition model to recognize the target object in the target image and obtain the object position of the target object includes:
[0098] The target image is input into the object recognition model, wherein the target image includes a text detection unit, a face detection unit, a human body detection unit and a commodity detection unit; the text detection unit is used to perform text detection on the target image to obtain the text position of the text; the face detection unit is used to perform face detection on the target image to obtain the face position of the face; the human body detection unit is used to perform human body detection on the target image to obtain the human body position of the human body; the commodity detection unit is used to perform commodity detection on the target image to obtain the commodity position of the commodity.
[0099] The text detection unit can be understood as a unit capable of detecting text in a target image. The text detection unit can be a submodel or network layer in an object recognition model. For example, the text detection unit can be a neural network model that implements OCR. The text position can be understood as the position information of the text in the target image. The text position can be the text coordinates or the text's bounding box information.
[0100] The face detection unit can be understood as a unit capable of performing face detection on a target image. This face detection unit can be a submodel or network layer within an object recognition model. For example, this face detection unit can be a face detection model or a convolutional neural network. The face position can be understood as the position of the face in the target image. This face position can be face coordinates or face bounding box information.
[0101] The human body detection unit can be understood as a unit capable of detecting human bodies in a target image. The human body detection unit can be a submodel or network layer in an object recognition model. For example, the human body detection unit can be a human body detection model or a convolutional neural network. The human body position can be understood as the position information of the human body in the target image. The human body position can be the human body coordinates or the human body's bounding box information.
[0102] The product detection unit can be understood as a unit capable of detecting products in a target image. This unit can be a submodel or network layer within an object recognition model. For example, this unit can be a product detection model or a convolutional neural network. The product location can be understood as the location information of the product in the target image. This can be the product coordinates or the product's bounding box information.
[0103] Based on this, considering that the target content published on the social platform may involve multiple scenes, such as shooting products, people, buildings, etc., it is necessary to cover multiple dimensions when performing object position recognition, so that the subsequent promotional image generation can be more accurate. In this process, the target image can be input into the object recognition model, wherein the target image includes a text detection unit, a face detection unit, a human body detection unit and a product detection unit; the text detection unit is used to perform text detection on the target image to obtain the text position of the text; the face detection unit is used to perform face detection on the target image to obtain the face position of the face; the human body detection unit is used to perform human body detection on the target image to obtain the human body position of the human body; the product detection unit is used to perform product detection on the target image to obtain the product position of the product.
[0104] In addition, in addition to the above-mentioned text detection, face detection, human body detection and commodity detection, building detection, animal detection, plant detection, etc. can also be performed to ensure that the object recognition model can handle multi-dimensional position recognition processing. It should be noted that in order to make the object recognition model have multi-dimensional recognition capabilities, an expert model architecture can be used to construct an object recognition model, so that it encapsulates multiple expert modules for identifying the positions of objects in different fields to meet actual usage needs. In addition, the recognition model corresponding to each field can also be trained separately. During specific implementation, the object recognition model can be selected according to actual needs, and this embodiment does not make any restrictions here.
[0105] Continuing with the above example, after obtaining the optimal candidate result for the image-template-text pair, this method can also use OCR, face and body detection, product detection and other models to further detect and process the original note image to obtain the specific bounding box information corresponding to the text, face, body and product. The specific method is:
[0106] The preprocessed note image is fed into the object recognition model, which includes a neural network model for OCR, a face detection model, a human body detection model, and a product detection model. The note image is then fed into the neural network model for OCR to detect text and obtain the text's bounding box information. The note image is then fed into the face detection model for face detection and obtain the face's bounding box information. The note image is then fed into the human body detection model for human body detection and obtain the human's bounding box information. The note image is then fed into the product detection model for product detection and obtain the product's bounding box information.
[0107] Based on the above embodiments, it can be seen that this method, by introducing a multi-task image analysis model, can perform fine-grained content recognition on the original note image, not only identifying the text information in the image (through OCR), but also detecting visual elements such as faces, human postures, and product objects, thereby achieving comprehensive perception and structured expression of image content, significantly enhancing the system's ability to understand input images; based on the detected boundingbox information of text, faces, bodies, and products, the system can effectively avoid the overlap between the copy template and existing text or important visual objects in the original image during the subsequent poster generation process, thereby improving the overall aesthetics and readability of the layout.
[0108] Step 206: Based on the object position and the text filling position, select a target text template corresponding to the target image from the multiple candidate text templates.
[0109] Specifically, after obtaining the object position and the text filling position in each candidate text template as mentioned above, in order to ensure that the text and image can be filled into the template to generate a promotional image with better display effect, the target text template corresponding to the target image can be selected from the multiple candidate text templates based on the object position and the text filling position, and then the promotional image is generated by combining the text, template and image.
[0110] The target text template may be understood as a text template that matches the target image.
[0111] In specific implementation, when selecting the target text template based on the object position and the text filling position, it can be understood as determining the text filling position in each candidate text template, as well as the position where the object can be placed without blocking the object, and then selecting the template that meets these two positions for filling and generating a promotional image as the target text template.
[0112] In addition, in addition to selecting the target text template in the above method, you can also select a universal text template as the target text template. Among them, the universal text template refers to a template that can generate a promotional image after filling in text and images. It can be applied to any scene, but the effect is inferior to other templates. When the selected target text template meets the user's needs, you can recommend a universal template for the user to generate promotional images, thereby meeting the different needs of users.
[0113] In one or more embodiments provided in this specification, considering that there are multiple candidate text templates, in order to ensure that the selected target text template better meets the current usage requirements, template selection can be completed by overlapping detection. In this embodiment, selecting the target text template corresponding to the target image from the multiple candidate text templates based on the object position and the text fill position includes:
[0114] Based on the object position and the text filling position, overlap detection is performed on each candidate text template and the target object to obtain overlap detection results between each candidate text template and the target object, wherein the overlap detection results are overlap results and non-overlap results; the candidate text template whose overlap detection result is a non-overlap result is determined as the target text template corresponding to the target image.
[0115] Specifically, overlap detection refers to the detection process of whether there is an overlapping display effect between each candidate text template and the target object, so as to avoid the generated promotional image having occlusion display attributes and affecting the promotional effect.
[0116] Based on this, in order to be able to select a template with better promotional effect, the target text template can be determined by overlapping detection. Specifically, based on the object position and the text filling position, each candidate text template can be overlapped with the target object to obtain the overlap detection results between each candidate text template and the target object, wherein the overlap detection results are overlapping results and non-overlapping results; on this basis, the candidate text template whose overlap detection result is a non-overlapping result can be determined as the target text template corresponding to the target image.
[0117] In addition, after the overlap detection process is performed, if multiple non-overlapping text templates are determined from multiple candidate text templates, if all of them are used for the subsequent promotional image construction, it may make the user's selection more difficult. Therefore, for multiple non-overlapping text templates, one can be randomly selected as the target text template, and multiple non-overlapping templates can be fused with the target text and the target image to obtain candidate promotional images for the user to choose from according to their needs. It is also possible to perform style matching calculations between multiple non-overlapping text templates and the target image to determine the target text template that meets the current usage requirements. In specific implementation, the selection can be made according to actual needs, and this embodiment does not make any restrictions here.
[0118] Continuing with the above example, after obtaining the Bounding box information, this method will further match the candidate results with the bounding box information after image processing. The specific method is: first, obtain the template copy slot (text filling position) from the optimal candidate template, and perform overlap detection on the bounding box of the image and the template copy slot to obtain the overlap detection result between the image and the template, wherein the overlap detection result is an overlapping result and a non-overlapping result; secondly, the candidate template whose overlap detection result is a non-overlapping result is determined as the template corresponding to the original note image (i.e., the target text template); and filter out the candidate template whose overlap detection result is an overlapping result.
[0119] Based on the above embodiments, it can be seen that after obtaining the bounding box information of key visual elements such as text, faces, human bodies and products in the image, the data processing method provided in this specification further performs overlap detection on it with the copy slot position in the candidate template to screen out the target text template that is highly adapted to the original image content layout; by effectively identifying and excluding candidate templates that may cause visual conflicts between the copy and the original important content of the image, it is ensured that the final selected template is clearer and more reasonable in the graphic layout, thereby significantly improving the reading experience and visual expression of the advertising poster.
[0120] Step 208: Generate a promotional image corresponding to the target published content using the target text, the target image, and the target text template.
[0121] Specifically, after determining the target text template, target text and target image, the target text, the target image and the target text template can be further fused to generate a promotional image corresponding to the target published content based on the fusion result.
[0122] The promotional image may be understood as an image used for promotion. For example, the promotional image may be a poster, an advertisement image, or the like.
[0123] In specific implementation, when generating a promotional image based on the target text, target image and target text template, considering that the target text template contains text filling position, and the target image can be used as a background image, the promotional image corresponding to the target published content can be obtained by filling the target text into the text filling position and embedding the target image into the template.
[0124] In addition, in order to improve the effect of the promotional image, after the target text, the target image and the target text template are fused, the fused promotional image can also be processed, such as image enhancement, image restoration, resolution adjustment, etc., to achieve a better display effect of the promotional image.
[0125] In one or more embodiments provided in this specification, when fusing the target text, target image, and target text template, in order to achieve a better fusion effect, the fusion process can be performed by calculating the proportion. In this embodiment, the target text, target image, and target text template are used to generate a promotional image corresponding to the target published content, including steps 1 to 3:
[0126] Step 1: Filter the target image using the object position to obtain a filtered image.
[0127] During specific implementation, considering that the target publication content contains multiple target images, and the display content of each target image may be different, when generating a promotional image, it is necessary to select one target image from them for use. Therefore, the selection of the filter image can be performed according to the object position, which can be understood as selecting the most appropriate image placed in the target text template as the filter image.
[0128] In actual applications, when screening images from multiple target images, images may be selected randomly or according to a set strategy (such as the degree of matching with a template), and this embodiment does not impose any limitation thereto.
[0129] In one or more embodiments provided in this specification, when selecting a screening image, it can be achieved by calculating a proportion. In this embodiment, the method of screening the target image using the object position to obtain the screening image includes:
[0130] Based on the object position, an object position image of the target object in the target image is determined, and an object proportion parameter between the object position image and the target image is calculated; and a target image whose object proportion parameter is less than or equal to a preset proportion threshold is determined as the screening image.
[0131] Specifically, the object position image refers to the image obtained by cropping the area corresponding to the target object in the target image, the object proportion parameter refers to the proportion of the object in the image to the overall area size of the image, and the preset proportion threshold refers to the threshold for screening target images with a smaller proportion, which can be set according to actual needs.
[0132] Based on this, when determining the screening image, considering that the smaller the proportion of the target object in the promotional image, the better the effect may be, the object position image of the target object in the target image can be determined based on the object position, that is, the image cropped from the corresponding area of the target object is determined from each target image, and then the object proportion parameter between the object position image and the target image can be calculated; the object proportion parameter can be used to represent the proportion of the object in the image, and then the target image with the object proportion parameter less than or equal to the preset proportion threshold can be determined as the screening image for subsequent use.
[0133] During specific implementation, in order to avoid a decrease in effect due to the object proportion parameter being too small, a second proportion threshold can also be set to use it as a screening image when the object proportion parameter is greater than the second proportion threshold and less than or equal to the preset proportion threshold, thereby making the promotional image more effective.
[0134] Continuing with the previous example, after obtaining the bounding box information, this method will filter the note images based on the bounding box information, thereby filtering out images with a large proportion of text areas through rules; the specific method is:
[0135] Based on the bounding box information of the text, the text area in the note image is identified, and the area ratio between the text area and the entire note image is calculated; note images with an area ratio less than or equal to a preset ratio threshold are determined to be images that can generate a poster; and note images with an area ratio greater than the preset ratio threshold are filtered out.
[0136] Based on the above embodiment, it can be seen that after obtaining the bounding box information of the text area in the image, the data processing method provided in this specification further filters the note image based on this information. By calculating the area ratio between the text area and the entire image, it selects original images suitable for generating high-quality advertising posters. By automatically excluding images that may affect the poster's aesthetics or the clarity of the information due to excessive text, this ensures that the images ultimately selected for poster generation are visually balanced, thereby improving the overall visual effect and quality of the generated poster.
[0137] Step 2: Using a text adjustment model, adjust the target text according to the screened image and the target text template to obtain an adjusted text.
[0138] Specifically, after obtaining the filtered image as described above, in order to further improve the generation effect of the promotional image, the text adjustment model can be combined to adjust the target text according to the filtered image and the target text template, so that the target text has a better display effect in the promotional image, and the adjusted text is obtained for subsequent use.
[0139] Among them, the text adjustment model can be understood as a model for adjusting the target text. For example, the text adjustment model can be a deep learning model, an adaptive adjustment model, etc.
[0140] In specific implementations, when adjusting the target text, multiple adjustments can be made to the target text, including font color, font style, content, and word count, so that the adjusted text better matches the overall presentation effect of the promotional image. Furthermore, text translation and other processing can also be performed. This embodiment is not limited in any way.
[0141] In one or more embodiments provided herein, the adjustment of the target text may be implemented in combination with the cleaning detection unit and the text adjustment unit in the model. In this embodiment, using the text adjustment model to adjust the target text according to the target image and the target text template to obtain the adjusted text includes:
[0142] Inputting the target text, the target image, and the target text template into a text processing model, wherein the text processing model includes a clarity detection unit and a text adjustment unit;
[0143] The clarity detection unit is used to determine the text-filled image area corresponding to the text-filled position from the target image, and clarity detection is performed based on the color information of the text-filled image area and the text color information of the target text to obtain a clarity detection result, wherein the clarity detection result is a clear result or an unclear result; when the clarity detection result is an unclear result, the text color information of the target text is adjusted by the text adjustment unit to obtain the adjusted text.
[0144] The clarity detection unit can be understood as a model or network layer capable of performing clarity detection, and the text adjustment unit can be understood as a model or network layer capable of adjusting the text color information of the target text.
[0145] Based on this, when performing text adjustment, the clarity detection unit can be used to determine the text-filled image area corresponding to the text-filled position from the target image, and perform clarity detection based on the color information of the text-filled image area and the text color information of the target text to obtain a clarity detection result, wherein the clarity detection result is a clear result or an unclear result; when the clarity detection result is an unclear result, the text adjustment unit also needs to be used to adjust the text color information of the target text to obtain the adjusted text.
[0146] Continuing with the above example, the system in this method will determine whether the color of the text matches the color of the image in the area and whether it is clear. If not, the system will adaptively adjust the color of the text to obtain the final candidate cover result for rendering. The specific method is as follows:
[0147] First, the image, text and template are input into the adaptive adjustment model, which includes a clarity detection model and a text adjustment model; the clarity detection model is used to identify the text area corresponding to the template text slot from the note image, and the text area color parameters of the text area are obtained; clarity detection is performed based on the text area color parameters and the color parameters pre-set for the text (such as white) to determine whether the color of the text is compatible with the color of the text area picture; if not, it means that the two are not clear, and the colors of the text and the text area picture need to be input into the text adjustment model; the text adjustment model is used to adaptively adjust the color of the text to obtain text that is compatible with the color of the text area picture, and then the final candidate cover result that can be rendered is obtained; if it is compatible, the above steps do not need to be performed, and the final candidate cover result that can be rendered is directly obtained.
[0148] Based on the above examples, the method provided in this specification uses an adaptive adjustment model to determine and optimize the compatibility of text color with the background image, ensuring the clarity and readability of the text on various images, thereby generating candidate cover results with excellent visual effects and efficient information transmission. The adaptive adjustment model can automatically assess the contrast and clarity between the text and the background image based on the color parameters of the text area, and adjust the text color when necessary. This mechanism effectively avoids the problem of information being difficult to read due to the background image and text color being too similar.
[0149] Step 3: Render the adjusted text, the filtered image, and the target text template to obtain a promotional image corresponding to the target published content.
[0150] Specifically, after obtaining the adjusted text and the filtered image, the adjusted text, the filtered image and the target text template may be rendered to obtain a promotional image corresponding to the target published content according to the rendering result.
[0151] In one or more embodiments provided herein, when there are multiple screening images, the construction of the promotional image can be performed by filtering by calculating a score. In this embodiment, the image rendering of the adjustment text, the screening image, and the target text template to obtain the promotional image corresponding to the target published content includes:
[0152] Image rendering is performed on multiple filtered images, the adjusted text, and the target text template to obtain multiple candidate promotional images corresponding to the target published content; aesthetic detection is performed on the multiple candidate promotional images using an aesthetic detection model to obtain aesthetic scores of the multiple candidate promotional images; and candidate promotional images having aesthetic scores greater than or equal to a preset aesthetic threshold are determined as promotional images corresponding to the target published content.
[0153] Specifically, a candidate promotional image refers to a promotional image generated by adjusting text, a target text template, and multiple screening images, when multiple screening images are present. Accordingly, an aesthetic detection model refers to a model that scores each promotional image. This model takes candidate promotional images as input and outputs an aesthetic score corresponding to the image. This model can be constructed using a classification model. Accordingly, a preset aesthetic threshold refers to a threshold for screening promotional images.
[0154] Based on this, when there are multiple screening images, the multiple screening images, the adjustment text, and the target text template can be rendered to obtain multiple candidate promotional images corresponding to the target published content; the multiple candidate promotional images are aesthetically detected using an aesthetic detection model to obtain the aesthetic scores of the multiple candidate promotional images; and the candidate promotional images whose aesthetic scores are greater than or equal to a preset aesthetic threshold are determined as the promotional images corresponding to the target published content.
[0155] Continuing with the above example, the system in this method will call the template rendering service in batches and in parallel, and use multiple note images and texts to generate a result image (i.e., a promotional poster); and after recording the result image accordingly, the aesthetic detection model is used to perform aesthetic detection on the promotional poster to obtain the aesthetic score corresponding to each promotional poster; the promotional poster with an aesthetic score greater than or equal to the preset score threshold is regarded as the optimal promotional poster, and the optimal promotional poster is fed back to the user.
[0156] Based on the above examples, the method provided in this specification ensures that users are provided with posters with the best visual effects and highest aesthetic value by invoking the template rendering service in batches and in parallel, and then using an aesthetic detection model to evaluate and screen the resulting images (i.e., promotional posters). The aesthetic detection model then assigns an aesthetic score to each resulting image. This automated evaluation mechanism objectively measures the poster's design aesthetics, ensuring that only works that meet the preset aesthetic standards are selected as optimal posters. This not only improves the quality of the final output, but also reduces the workload of manual review, thereby increasing work efficiency.
[0157] One or more embodiments of the present specification provide a data processing method. In the process of generating a promotional image, the target text and target image are first determined from the target published content, and multiple candidate text templates matching the target text are selected from the text template set; secondly, in order to generate a high-quality promotional image, the target object in the target image can be identified by using an object recognition model to obtain the object position of the target object, and based on the object position and the text filling position, a target text template matching the target image is selected from multiple candidate text templates; finally, based on the target text, target image and target text template, a promotional image corresponding to the target published content is quickly generated, thereby improving the generation efficiency of the promotional image and meeting actual needs, and through the target text template, target text and target image matching the target text and target image, a relatively high-quality promotional image can be generated, thereby improving the quality of the promotional image.
[0158] The following combined Figure 3 , taking the application of the data processing method provided in this specification in the promotional poster generation scenario as an example, the data processing method is further explained. Figure 3 A flowchart of a data processing method provided by an embodiment of this specification is shown.
[0159] based on Figure 3 It can be seen that the data processing method provided in this specification provides a multi-model driven advertising template intelligent matching recommendation system, and the specific method of generating promotional posters includes the following steps.
[0160] Step 302: User selects a note. Specifically, the user on the user side can manually select the note that he wishes to perform secondary optimization on the client, and the client will send the note to the system.
[0161] Step 304: Generate a copy library. Specifically, after receiving the note, the system automatically obtains the text contained in the note. Then, the text is input into the large language model for text rewriting to obtain the promotional copy for the promotional poster, and the promotional copy is stored in the poster copy material library.
[0162] Step 306: Extract the original note image. Specifically, after the system receives the note, it will automatically obtain the original note image contained in the note.
[0163] Step 308: Determine the template material library. Specifically, in the process of generating a poster, this method needs to determine a pre-built template material library; the template material library contains a pre-designed copy template. Moreover, the copy template contains a template copy slot. Moreover, in the actual production process, the copy and the template material in the template material library will be given priority for a round of pre-matching operations to ensure that the number of characters in the copy meets the maximum number of characters required for the copy slot of the template, and then obtain a candidate template list; the specific method is: first, identify the number of characters in the copy, and obtain the maximum number of characters in the copy slot of the template material in the template material library; second, compare the number of characters with the maximum number of characters in the copy slot, and select multiple template materials whose maximum number of characters in the copy slot is greater than or equal to the number of characters to form a candidate template list.
[0164] Step 310: Determine the candidate results using the text-image matching model. Specifically, after obtaining the candidate template list, this method will input the candidate templates and the original note image in the candidate template list into the CLIP multimodal model (i.e., text-image matching model). The candidate template list and the original note image will calculate the image-text similarity through the CLIP multimodal model. For example, the image-text similarity is any value in the range of [0,1]. The greater the image-text similarity, the more similar the candidate template and the original note image are; conversely, the less similar the candidate template and the original note image are. After obtaining the image-text similarity, they will be sorted according to the similarity, and a fixed target number (e.g., 10) of image-text pair results will be selected from high to low, thereby obtaining a fixed target number of image-text results (i.e., the optimal candidate results of the image-template-text pair).
[0165] Step 312: Use the visual detection model to detect the original note image and obtain the bounding box information. First, the system will perform image preprocessing on the image to obtain the background image of the promotional poster. Among them, the image preprocessing operations include but are not limited to: cropping, filtering, format adjustment, size adjustment and other preprocessing operations on the image. Secondly, use OCR, face, body detection, product detection and other models to further detect and process the original note image to obtain the specific bounding box information corresponding to the text, face, body and product. The specific method is:
[0166] 1. Input the preprocessed note image into the visual detection model, which includes a neural network model for implementing OCR, a face detection model, a human body detection model, and a product detection model.
[0167] 2. In the visual detection model, the original note image is input into the neural network model that implements OCR to detect text and obtain the bounding box information of the text;
[0168] 3. In the visual detection model, input the original note image into the face detection model for face detection to obtain the bounding box information of the face;
[0169] 4. In the visual detection model, input the original note image into the human body detection model for human body detection to obtain the human body's bounding box information;
[0170] 5. In the visual inspection model, input the original note image into the product inspection model for product inspection to obtain the product's bounding box information.
[0171] Step 314: Use the adaptive adjustment model to make adjustments. Specifically, after obtaining the bounding box information, this method will further match the candidate results with the bounding box information after image processing. The specific method is as follows: first, the template text slot (text filling position) is obtained from the optimal candidate template, and the image bounding box and the template text slot are overlapped to obtain the overlap detection results between the image and the template, wherein the overlap detection results are overlapping results and non-overlapping results; the candidate template with the overlap detection result of non-overlapping results is determined as the template corresponding to the original note image (i.e., the target text template); and the candidate template with the overlap detection result of overlapping results is filtered out. Secondly, after obtaining the bounding box information, this method will filter the note image based on the bounding box information, thereby filtering out images with a large text area ratio through rules; the specific method is: based on the bounding box information of the text, identify the text area of the text in the note image, and calculate the area ratio between the text area and the entire note image; note images with an area ratio less than or equal to a preset ratio threshold are determined to be images that can generate posters; and note images with an area ratio greater than the preset ratio threshold are filtered. Finally, the system in this method will determine whether the color of the text is compatible with the color of the image in the area and whether it is clear. If it is not clear, the color of the text will be adaptively adjusted to obtain the final candidate cover result that can be rendered. The specific method is:
[0172] 1. Input the image, text and template into the adaptive adjustment model, which includes the clarity detection model and the text adjustment model.
[0173] 2. Using the clarity detection model, identify the text area corresponding to the template text slot from the note image and obtain the text area color parameters of the text area. Perform clarity detection based on the text area color parameters and the color parameters pre-set for the text (e.g., white) to determine whether the color of the text is consistent with the color of the text area image.
[0174] 3. If they are not compatible, it means that the two are not clear. The colors of the text and the image in the text area need to be input into the text adjustment model. The text adjustment model is used to adaptively adjust the color of the text to obtain the text that is compatible with the color of the image in the text area, and then the final candidate cover result that can be rendered is obtained.
[0175] 4. If it is suitable, there is no need to perform the above steps, and the final candidate cover result that can be rendered is directly obtained.
[0176] Step 316: Execute template rendering. Specifically, the system in this method will call the template rendering service in batches and in parallel, using multiple note images, target templates and text to generate a result image (i.e., a promotional poster).
[0177] Step 318: Obtain the result graph. Specifically, after recording the result graph, the posters are aesthetically tested using the aesthetics detection model to obtain an aesthetic score for each poster. Posters with an aesthetic score greater than or equal to a preset threshold are considered optimal posters, and this optimal poster is fed back to the user.
[0178] Based on the above steps, it can be seen that this method provides an intelligent template matching generation system based on multiple modal models, which aims to solve the above problems. By integrating a variety of advanced artificial intelligence technologies, such as deep learning and computer vision technologies, this system can automatically identify and analyze the main elements in the image, and perform template selection, copy adaptation, aesthetic estimation and other optimized typesetting operations based on them. The system can not only intelligently avoid the overlap of text and important visual elements during the generation process, but also adjust the text style and color to ensure its readability in various backgrounds, thereby enhancing the visual appeal of the poster and the efficiency of information transmission. In addition, the system optimizes the time-consuming template matching and the aesthetics of the generated cover, which greatly reduces the time and design costs of small and medium-sized businesses in producing advertising covers, and ultimately helps small and medium-sized businesses produce exquisite poster covers with high click-through rates.
[0179] This approach established a complete SOP process for template library expansion, from template production -> template optimization -> template launch -> template feedback -> template update, thus establishing a complete poster cover production process involving the coordination of multiple models. The system uses parallel processing for both model processing and service calls, ensuring a service response speed of approximately 5 seconds to produce 20 poster cover results, thereby improving service throughput.
[0180] The following combined Figure 4 , taking the application of the data processing method provided in this specification in the promotion of goods as an example, the data processing method is further explained. Figure 4 A flow chart of a data processing method provided by an embodiment of this specification is shown.
[0181] The data processing method is applied to the content application platform, which includes a client and a server. The client is the terminal device held by the user who browses the content application platform to publish content and create content; the server specifically refers to the server that provides users with content publishing and editing functions. Figure 4 As shown:
[0182] The client displays the list of published content to the user and determines the target published content based on the user's selection operation.
[0183] The client sends a request to the server to process the target published content.
[0184] The server determines the target text and target image from the target published content in response to the processing request, and selects multiple candidate text templates corresponding to the target text from the text template set; wherein the candidate text templates include preset text filling positions.
[0185] The server uses the object recognition model to identify the target object in the target image and obtain the object position of the target object.
[0186] The server selects a target text template corresponding to the target image from multiple candidate text templates based on the object position and the text filling position.
[0187] The server uses the target text, target image, and target text template to generate a promotional image corresponding to the target published content and sends it to the client.
[0188] The client receives the promotional image and displays it to the user.
[0189] When extracting target text and target images from target published content, considering that the target published content may correspond to various forms, it is necessary to combine different types of published content to complete the text and image extraction operation. In this embodiment, the target published content is a target published video, and determining the target text and target image from the target published content includes:
[0190] Speech recognition is performed on the target published video to obtain video speech text, and a preset number of video frames are extracted from the target published video; a target text is determined based on the video speech text, and a target image is determined based on the video frames.
[0191] In addition, when the target published content is a graphic content, after the text and image are extracted, the text and image can be processed to make them suitable for subsequent promotional image generation. In this embodiment, the determination of the target text and target image from the target published content includes:
[0192] Extracting published text and published image from the target published content; rewriting the published text using a text rewriting model to obtain the target text; and performing image preprocessing on the published image to obtain the target image.
[0193] When selecting multiple candidate text templates associated with the target text, considering that images have a better dissemination effect in the template, the candidate text templates can be screened by calculating the similarity between the image and the template. In this embodiment, the selection of multiple candidate text templates corresponding to the target text in the text template set includes steps 1 and 2:
[0194] Step 1: Based on the text attribute information of the target text, the text template set is selected from the template storage unit, wherein the text template set includes multiple text templates.
[0195] Step 2: Calculate the similarity between each text template and the target image, and based on the similarity, select multiple candidate text templates corresponding to the target text from the text template set.
[0196] Wherein, when the text attribute information is the number of words in the text, the selecting the text template set from the template storage unit based on the text attribute information of the target text includes:
[0197] Determine the text word count of the target text, and select multiple text templates whose template text word count is greater than or equal to the text word count from the template storage unit, wherein the template text word count is the text word count embedded in the text template; construct the text template set based on the multiple text templates.
[0198] In addition, when the text attribute information is a text language type, the selecting the text template set from the template storage unit based on the text attribute information of the target text includes:
[0199] Determine the text language type of the target text, and select multiple text templates whose template text language types are consistent with the text language type from the template storage unit, wherein the template text language type is the text language type of the text embedded in the text template; and construct the text template set based on the multiple text templates.
[0200] When screening multiple candidate text templates, it can be completed by calculating similarities and then sorting them. In this embodiment, the similarities between each text template and the target image are calculated, and based on the similarities, multiple candidate text templates corresponding to the target text are selected from the text template set, including:
[0201] A similarity analysis model is used to perform similarity analysis on each text template and the target image to obtain the similarity between each text template and the target image; based on the similarity, each text template in the text template set is sorted in descending order to obtain a text template sequence, and a preset number of text templates are selected from the text template sequence in a top-down manner as multiple candidate text templates corresponding to the target text.
[0202] When performing object position recognition, in order to ensure recognition accuracy and comprehensiveness, recognition processing can be completed from multiple dimensions. In this embodiment, the target objects include text, human faces, human bodies, and commodities, and the object positions include text positions, human face positions, human body positions, and commodity positions; using the object recognition model to recognize the target object in the target image and obtain the object position of the target object includes:
[0203] The target image is input into the object recognition model, wherein the target image includes a text detection unit, a face detection unit, a human body detection unit and a commodity detection unit; the text detection unit is used to perform text detection on the target image to obtain the text position of the text; the face detection unit is used to perform face detection on the target image to obtain the face position of the face; the human body detection unit is used to perform human body detection on the target image to obtain the human body position of the human body; the commodity detection unit is used to perform commodity detection on the target image to obtain the commodity position of the commodity.
[0204] Considering that there are multiple candidate text templates, in order to ensure that the selected target text template better meets the current usage requirements, template selection can be completed by overlapping detection. In this embodiment, the target text template corresponding to the target image from the multiple candidate text templates based on the object position and the text filling position includes:
[0205] Based on the object position and the text filling position, overlap detection is performed on each candidate text template and the target object to obtain overlap detection results between each candidate text template and the target object, wherein the overlap detection results are overlap results and non-overlap results; the candidate text template whose overlap detection result is a non-overlap result is determined as the target text template corresponding to the target image.
[0206] When the target text, target image and target text template are fused, in order to achieve a better fusion effect, the fusion process can be performed by calculating the proportion. In this embodiment, the target text, target image and target text template are used to generate a promotional image corresponding to the target published content, including steps 1 to 3:
[0207] Step 1: Filter the target image using the object position to obtain a filtered image.
[0208] When selecting the screening image, it can be achieved by calculating the proportion. In this embodiment, the target image is screened by using the object position to obtain the screening image, including:
[0209] Based on the object position, an object position image of the target object in the target image is determined, and an object proportion parameter between the object position image and the target image is calculated; and a target image whose object proportion parameter is less than or equal to a preset proportion threshold is determined as the screening image.
[0210] Step 2: Using a text adjustment model, adjust the target text according to the screened image and the target text template to obtain an adjusted text.
[0211] Adjustment of the target text can be achieved by combining the cleaning detection unit and the text adjustment unit in the model. In this embodiment, the text adjustment model is used to adjust the target text according to the target image and the target text template to obtain the adjusted text, including:
[0212] The target text, the target image and the target text template are input into a text processing model, wherein the text processing model includes a clarity detection unit and a text adjustment unit; using the clarity detection unit, a text-filled image area corresponding to the text-filled position is determined from the target image, and clarity detection is performed based on the color information of the text-filled image area and the text color information of the target text to obtain a clarity detection result, wherein the clarity detection result is a clear result or an unclear result; when the clarity detection result is an unclear result, the text color information of the target text is adjusted by using the text adjustment unit to obtain the adjusted text.
[0213] Step 3: Render the adjusted text, the filtered image, and the target text template to obtain a promotional image corresponding to the target published content.
[0214] In the case where there are multiple screening images, the construction of the promotional image can be performed by screening by calculating scores. In this embodiment, the image rendering of the adjustment text, the screening image, and the target text template to obtain the promotional image corresponding to the target published content includes:
[0215] Image rendering is performed on multiple filtered images, the adjusted text, and the target text template to obtain multiple candidate promotional images corresponding to the target published content; aesthetic detection is performed on the multiple candidate promotional images using an aesthetic detection model to obtain aesthetic scores of the multiple candidate promotional images; and candidate promotional images having aesthetic scores greater than or equal to a preset aesthetic threshold are determined as promotional images corresponding to the target published content.
[0216] One or more embodiments of the present specification provide a data processing method. In the process of generating a promotional image, the target text and target image are first determined from the target published content, and multiple candidate text templates matching the target text are selected from the text template set; secondly, in order to generate a high-quality promotional image, the target object in the target image can be identified by using an object recognition model to obtain the object position of the target object, and based on the object position and the text filling position, a target text template matching the target image is selected from multiple candidate text templates; finally, based on the target text, target image and target text template, a promotional image corresponding to the target published content is quickly generated, thereby improving the generation efficiency of the promotional image and meeting actual needs, and through the target text template, target text and target image matching the target text and target image, a relatively high-quality promotional image can be generated, thereby improving the quality of the promotional image.
[0217] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 5 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 5 As shown, the device includes:
[0218] The first template selection module 502 is configured to determine a target text and a target image from the target published content, and select a plurality of candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include a preset text filling position;
[0219] A position recognition module 504 is configured to recognize the target object in the target image using an object recognition model to obtain the object position of the target object;
[0220] A second template selection module 506 is configured to select a target text template corresponding to the target image from the plurality of candidate text templates based on the object position and the text filling position;
[0221] The image generation module 508 is configured to generate a promotional image corresponding to the target published content using the target text, the target image, and the target text template.
[0222] Optionally, the image generation module 508 is further configured to: use the object position to filter the target image to obtain a filtered image; use a text adjustment model to adjust the target text according to the filtered image and the target text template to obtain an adjusted text; perform image rendering on the adjusted text, the filtered image and the target text template to obtain a promotional image corresponding to the target published content.
[0223] Optionally, the image generation module 508 is further configured to: input the target text, the target image and the target text template into a text processing model, wherein the text processing model includes a clarity detection unit and a text adjustment unit; using the clarity detection unit, determine the text-filled image area corresponding to the text-filled position from the target image, and perform clarity detection based on the color information of the text-filled image area and the text color information of the target text to obtain a clarity detection result, wherein the clarity detection result is a clear result or an unclear result; when the clarity detection result is an unclear result, use the text adjustment unit to adjust the text color information of the target text to obtain the adjusted text.
[0224] Optionally, there are multiple filtered images; the image generation module 508 is further configured to: perform image rendering on the multiple filtered images, the adjustment text, and the target text template to obtain multiple candidate promotional images corresponding to the target published content; use an aesthetic detection model to perform aesthetic detection on the multiple candidate promotional images to obtain aesthetic scores of the multiple candidate promotional images; and determine the candidate promotional images whose aesthetic scores are greater than or equal to a preset aesthetic threshold as the promotional images corresponding to the target published content.
[0225] Optionally, the image generation module 508 is further configured to: determine the object position image of the target object in the target image based on the object position, and calculate the object proportion parameter between the object position image and the target image; and determine the target image whose object proportion parameter is less than or equal to a preset proportion threshold as the screening image.
[0226] Optionally, the second template selection module 506 is further configured to: perform overlap detection on each candidate text template and the target object based on the object position and the text filling position, and obtain overlap detection results between each candidate text template and the target object, wherein the overlap detection results are overlapping results and non-overlapping results; and determine the candidate text template whose overlap detection result is a non-overlapping result as the target text template corresponding to the target image.
[0227] Optionally, the target object includes text, face, body and commodity, and the object position includes text position, face position, body position and commodity position; the position identification module 504 is further configured to: input the target image into the object recognition model, wherein the target image includes a text detection unit, a face detection unit, a body detection unit and a commodity detection unit; use the text detection unit to perform text detection on the target image to obtain the text position of the text; use the face detection unit to perform face detection on the target image to obtain the face position of the face; use the body detection unit to perform body detection on the target image to obtain the body position of the body; use the commodity detection unit to perform commodity detection on the target image to obtain the commodity position of the commodity.
[0228] Optionally, the first template selection module 502 is further configured to: extract published text and published images from the target published content; rewrite the published text using a text rewriting model to obtain the target text; and preprocess the published image to obtain the target image.
[0229] Optionally, the first template selection module 502 is further configured to: select the text template set from the template storage unit based on the text attribute information of the target text, wherein the text template set includes multiple text templates; calculate the similarity between each text template and the target image, and based on the similarity, select multiple candidate text templates corresponding to the target text from the text template set.
[0230] Optionally, the text attribute information is the number of text words; the first template selection module 402 is further configured to: determine the number of text words of the target text, and select multiple text templates whose template text word count is greater than or equal to the text word count from the template storage unit, wherein the template text word count is the number of text words embedded in the text template; and construct the text template set based on the multiple text templates.
[0231] Optionally, the first template selection module 502 is further configured to: perform similarity analysis on each text template and the target image using a similarity analysis model to obtain the similarity between each text template and the target image; sort the text templates in the text template set in descending order based on the similarity to obtain a text template sequence, and select a preset number of text templates from the text template sequence in a top-down manner as multiple candidate text templates corresponding to the target text.
[0232] One or more embodiments of the present specification provide a data processing device, which, in the process of generating a promotional image, first determines the target text and target image from the target published content, and selects multiple candidate text templates that match the target text from the text template set; secondly, in order to generate a high-quality promotional image, the target object in the target image can be identified by using an object recognition model to obtain the object position of the target object, and based on the object position and the text filling position, selects a target text template that matches the target image from multiple candidate text templates; finally, based on the target text, target image and target text template, a promotional image corresponding to the target published content is quickly generated, thereby improving the efficiency of generating the promotional image and meeting actual needs, and through the target text template, target text and target image that match the target text and target image, a relatively high-quality promotional image can be generated, thereby improving the quality of the promotional image.
[0233] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.
[0234] Figure 6 6 shows a block diagram of a computing device 600 according to one embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0235] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0236] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0237] Computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 may also be a mobile or stationary server.
[0238] The processor 620 is configured to execute the following computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by the processor.
[0239] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computing device embodiment is generally similar to the data processing method embodiment, so the description is relatively simple. For relevant parts, refer to the description of the data processing method embodiment.
[0240] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0241] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the data processing method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the data processing method embodiment.
[0242] An embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0243] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned data processing method.
[0244] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0245] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0246] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0247] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0248] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, characterized in that: include: Determining target text and target image from target published content, and selecting a plurality of candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include preset text filling positions; Identifying a target object in the target image using an object recognition model to obtain an object position of the target object; Based on the object position and the text filling position, selecting a target text template corresponding to the target image from the multiple candidate text templates; A promotional image corresponding to the target published content is generated using the target text, the target image, and the target text template.
2. The data processing method according to claim 1, wherein: The step of generating a promotional image corresponding to the target published content by using the target text, the target image, and the target text template includes: Filtering the target image using the object position to obtain a filtered image; Using a text adjustment model, adjusting the target text according to the screening image and the target text template to obtain an adjusted text; Image rendering is performed on the adjusted text, the filtered image, and the target text template to obtain a promotional image corresponding to the target published content.
3. The data processing method according to claim 2, characterized in that: The method of adjusting the target text according to the target image and the target text template using the text adjustment model to obtain the adjusted text includes: Inputting the target text, the target image, and the target text template into a text processing model, wherein the text processing model includes a clarity detection unit and a text adjustment unit; Determining, by means of the clarity detection unit, a text-filled image region corresponding to the text-filled position in the target image, and performing clarity detection based on color information of the text-filled image region and text color information of the target text to obtain a clarity detection result, wherein the clarity detection result is a clarity result or an unclarity result; In a case where the clarity detection result is an unclear result, the text color information of the target text is adjusted by using the text adjustment unit to obtain the adjusted text.
4. The data processing method according to claim 2, wherein: There are multiple screening images; The performing image rendering on the adjusted text, the filtered image, and the target text template to obtain a promotional image corresponding to the target published content includes: Performing image rendering on the plurality of screening images, the adjustment text, and the target text template to obtain a plurality of candidate promotional images corresponding to the target published content; Performing aesthetic detection on the plurality of candidate promotional images using an aesthetic detection model to obtain aesthetic scores of the plurality of candidate promotional images; The candidate promotional images having the aesthetic score greater than or equal to a preset aesthetic threshold are determined as the promotional images corresponding to the target published content.
5. The data processing method according to claim 2, wherein: The step of screening the target image by using the object position to obtain a screened image includes: determining an object position image of the target object in the target image based on the object position, and calculating an object ratio parameter between the object position image and the target image; The target image whose object proportion parameter is less than or equal to a preset proportion threshold is determined as the screening image.
6. The data processing method according to claim 1, wherein: The selecting, based on the object position and the text filling position, a target text template corresponding to the target image from the plurality of candidate text templates includes: Performing overlap detection on each candidate text template and the target object according to the object position and the text filling position, and obtaining overlap detection results between each candidate text template and the target object, wherein the overlap detection results include overlap results and non-overlap results; The candidate text template whose overlap detection result is a non-overlap result is determined as the target text template corresponding to the target image.
7. The data processing method according to claim 1, wherein: The target objects include text, human faces, human bodies and commodities, and the object positions include text positions, human face positions, human body positions and commodity positions; Identifying a target object in the target image using an object recognition model to obtain an object position of the target object includes: Inputting the target image into the object recognition model, wherein the target image includes a text detection unit, a face detection unit, a human body detection unit, and a commodity detection unit; Utilizing the text detection unit to perform text detection on the target image to obtain the text position of the text; Performing face detection on the target image using the face detection unit to obtain the face position of the face; Performing human body detection on the target image using the human body detection unit to obtain the human body position of the human body; The commodity detection unit is used to perform commodity detection on the target image to obtain the commodity position of the commodity.
8. The data processing method according to claim 1, wherein: Determining the target text and target image from the target published content includes: extracting published text and published images from the target published content; Rewriting the published text using a text rewriting model to obtain the target text; Perform image preprocessing on the published image to obtain the target image.
9. The data processing method according to claim 1, wherein: The selecting of a plurality of candidate text templates corresponding to the target text from the text template set includes: Based on the text attribute information of the target text, selecting the text template set from the template storage unit, wherein the text template set includes a plurality of text templates; The similarity between each text template and the target image is calculated, and based on the similarity, a plurality of candidate text templates corresponding to the target text are selected from the text template set.
10. The data processing method according to claim 9, characterized in that: The text attribute information is the number of words in the text; The selecting the text template set from the template storage unit based on the text attribute information of the target text includes: Determining the text word count of the target text, and selecting a plurality of text templates whose template text word count is greater than or equal to the text word count from the template storage unit, wherein the template text word count is the text word count embedded in the text template; The text template set is constructed based on the multiple text templates.
11. The data processing method according to claim 9, characterized in that: The calculating the similarity between each text template and the target image, and selecting a plurality of candidate text templates corresponding to the target text from the text template set based on the similarity, includes: Performing similarity analysis on each of the text templates and the target image using a similarity analysis model to obtain the similarity between each of the text templates and the target image; Based on the similarity, the text templates in the text template set are sorted in descending order to obtain a text template sequence, and a preset number of text templates are selected from the text template sequence in a top-down manner as multiple candidate text templates corresponding to the target text.
12. A data processing device, characterized in that: include: A first template selection module is configured to determine a target text and a target image from the target published content, and select a plurality of candidate text templates corresponding to the target text from a text template set; wherein the candidate text templates include a preset text filling position; a position recognition module, configured to recognize a target object in the target image using an object recognition model to obtain an object position of the target object; A second template selection module is configured to select a target text template corresponding to the target image from the plurality of candidate text templates based on the object position and the text filling position; The image generation module is configured to generate a promotional image corresponding to the target published content by using the target text, the target image and the target text template.
13. A computing device, characterized in that include: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium, characterized in that It stores a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product, characterized in that The method comprises a computer program / instruction which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.