Material optimization method and device

Through automated editing technology, the pre-trained target model is used to edit the modified materials, which solves the problem of low material modification efficiency in the existing technology, realizes efficient and automated material optimization, and improves the efficiency and quality of material delivery.

CN120198541APending Publication Date: 2025-06-24SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249229.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, there are a large number of non-compliant materials on social and video platforms, resulting in a long modification cycle, high cost and low material delivery efficiency.

Method used

By obtaining the material to be modified and the modification instructions information, based on the pre-trained target model and modification instructions information, the area to be modified is determined, a redrawn picture is generated, and the target material that meets the specification is finally generated.

Benefits of technology

It realizes automated processing of materials to be modified, reduces modification cycle and cost, improves material delivery efficiency, and improves the standardization and quality of materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198541A_ABST
    Figure CN120198541A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a material optimization method, and the method comprises the steps: obtaining a to-be-modified material and modification indication information, and the modification indication information comprises the problem description of the to-be-modified material; and based on the modification indication information, editing the to-be-modified material to generate a corresponding target material. According to the technical scheme provided by the embodiment of the invention, the to-be-modified material is automatically processed, so that the modification cost is reduced, and the material putting efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of Internet technologies, and in particular, to a method, device, computer device, computer-readable storage medium, and computer program product for optimizing materials. Background Art

[0002] There are a large number of non-compliant materials on social and video platforms, and these materials need to be modified and resubmitted. However, the existing manual material modification is inefficient, resulting in a long modification cycle, high cost, and low material placement efficiency.

[0003] It should be noted that the above information is not necessarily prior art and is not used to limit the patent protection scope of the present application.

[0004] Invention Information Embodiments of the present application provide a method, device, computer device, computer-readable storage medium, and computer program product for optimizing materials to solve or alleviate one or more of the above technical problems.

[0005] One aspect of the embodiments of the present application provides a method for optimizing materials, the method including: Obtaining a material to be modified and modification instruction information, where the modification instruction information includes a problem description of the material to be modified; and Based on the modification instruction information, editing the material to be modified to generate a corresponding target material.

[0006] Optionally, the material to be modified includes a picture to be modified; based on the modification instruction information, editing the material to be modified to generate a corresponding target material includes: Based on a pre-trained target model and the modification instruction information, determining a region to be modified of the picture to be modified; Generating a redrawn picture based on the region to be modified; and Generating a target picture based on the picture to be modified and the redrawn picture.

[0007] Optionally, the target model includes a target detection model for identifying the chest; The training operation of the target model includes: Performing model-assisted annotation on each training data in the training dataset to generate pseudo-labels for each training data; Training the target detection model through each training data and the corresponding label; where the label corresponding to each training data is obtained after calibrating the pseudo-label of each training data.

[0008] Optionally, generating a redrawn picture based on the region to be modified includes: Generate an initial mask region based on the region to be modified; When the image to be modified includes specific type of elements, determine a protected region that will not be redrawn according to the specific elements; Determine the mask region to be redrawn based on the protected region and the initial mask region.

[0009] Optionally, the image to be modified includes human limbs; Generate a redrawn image based on the region to be modified, including: Preprocess the image to be modified to obtain key point information in the image to be modified, where the key point information includes limb information; Generate the redrawn image based on the image to be modified, the mask region to be redrawn, and the key point information.

[0010] Optionally, the redrawing is implemented through an image redrawing model; The image redrawing model includes one or more of the following features: (1) The guiding coefficient value is lower than the default guiding coefficient value; (2) The intensity value of the image redrawing model is determined according to the characteristics of the base model in the image redrawing model.

[0011] Optionally, the redrawing is guided by a target prompt to instruct the image redrawing model to implement; The operation of obtaining the target prompt includes: Generate an initial prompt based on the image to be modified; Based on the initial prompt, add corresponding scene words and delete sensitive words to generate a target prompt.

[0012] Optionally, the material to be modified includes a title to be modified; edit the material to be modified based on the modification instruction information to generate a corresponding target material, including: Generate a target title through a target language model based on the modification instruction information and the title to be modified.

[0013] Optionally, the material to be modified includes an image to be modified; edit the material to be modified based on the modification instruction information to generate a corresponding target material, including: Generate a target text based on the modification instruction information and the original text in the image to be modified; Remove the original text of the image to be modified; Generate a target image with the target text embedded therein based on the target text and the image to be modified from which the original text has been removed.

[0014] Optionally, the removal is implemented by a text cleaning model; The training operation of the text cleaning model includes: Inputting the training data into a pre-trained model with fixed parameters to output text image features and corresponding textless image features; the training data includes multiple image groups, and each image group includes a text image and a corresponding textless image; Based on the text image features and the textless image features, obtaining a text cleaning loss value; Training the text cleaning model based on the loss value of the text cleaning and the text image.

[0015] Another aspect of the embodiments of the present application provides a material optimization device, and the device includes: An acquisition module, configured to acquire a material to be modified and modification instruction information, where the modification instruction information includes a problem description of the material to be modified; A generation module, configured to edit the material to be modified based on the modification instruction information to generate a corresponding target material.

[0016] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0017] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0018] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0019] The embodiments of the present application adopting the above technical solutions may include the following advantages: automatically editing the modified material based on the modification instruction information to generate a target material, realizing the automated processing of the material to be modified, thereby improving the modification efficiency and automation, reducing the modification cycle and cost, and improving the material delivery efficiency. Description of the Drawings

[0020] The accompanying drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0021] Figure 1 Schematically shows the operating environment diagram of the material optimization method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the material optimization method according to Embodiment 1 of the present application; Figure 3 Schematically shows according to Figure 2 The sub-step flowchart of step S202 in; Figure 4 Schematically shows the new flowchart of the material optimization method according to Embodiment 1 of the present application; Figure 5 Schematically shows according to Figure 3 The sub-step flowchart of step S302 in; Figure 6 Schematically shows according to Figure 3 The sub-step flowchart of step S302 in; Figure 7 Schematically shows the new flowchart of the material optimization method according to Embodiment 1 of the present application; Figure 8 Schematically shows according to Figure 2 The sub-step flowchart of step S202 in; Figure 9 Schematically shows the new flowchart of the material optimization method according to Embodiment 1 of the present application; Figure 10 Schematically shows the visualization effect diagram of the material optimization method according to Embodiment 1 of the present application; Figure 11 Schematically shows the quantitative index diagram of the material optimization method according to Embodiment 1 of the present application; Figure 12 Schematically shows the exemplary application diagram of the material optimization method according to Embodiment 1 of the present application; Figure 13 Schematically shows the block diagram of the material optimization device according to Embodiment 2 of the present application; and Figure 14 Schematically shows the schematic diagram of the hardware architecture of the computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0022] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0023] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0024] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0025] First of all, the following provides the term explanations involved in the present application: Mask: An image processing method used to control the visibility of an image or edit certain parts.

[0026] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies will be described below: The applicant has learned that among the materials with good view volume effects on the platform, many materials are non-compliant, and content producers often test the compliance bottom line when producing materials (promotional materials), which leads to frequent non-compliance situations. In addition, in order to prevent compliance detection from being targeted and bypassed, the platform's opinions on non-compliant materials are relatively ambiguous. For content producers, on the one hand, they need to understand the meaning of such opinions, and on the other hand, they need to modify the materials and re-conduct compliance detection. The entire process has a long cycle, which greatly affects the delivery efficiency.

[0027] Therefore, the embodiments of the present application provide a technical solution for material optimization. In this technical solution, (1) the automated processing of materials to be modified is realized, thereby reducing the modification cycle and modification cost and improving the material delivery efficiency; (2) the standardization and quality of materials are improved, the user experience is enhanced, and the brand image related to the materials is protected; (3) the materials can be delivered smoothly and efficiently and comply with the platform's specifications. See the following for details.

[0028] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0029] As Figure 1 shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where: The service platform 2 may be composed of a single or multiple computing devices. The multiple computing devices may include virtualized computing instances. The virtualized computing instances may include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices may load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device.

[0030] The service platform 2 may be configured to communicate with the client 6, etc. via the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 may include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0031] The service platform 2 may provide services such as material modification and material submission, such as providing a material modification service for the client.

[0032] The client 6 may be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle terminal, or a smart TV. Based on the above operating systems, various applications, such as video platforms, may be run.

[0033] The client 6 may provide / configure a material submission page for uploading materials, etc.

[0034] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices are adjustable.

[0035] Next, taking the service platform 2 as the execution entity, the technical solutions of this application will be introduced through multiple embodiments. It should be noted that these embodiments may be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0036] Embodiment 1 Figure 2The flowchart of the material optimization method according to Embodiment 1 of the present application is schematically shown.

[0037] As Figure 2 shown, the material optimization method may include steps S200 to S202, where: Step S200, obtaining the material to be modified and modification instruction information, where the modification instruction information includes a problem description of the material to be modified.

[0038] Step S202, editing the material to be modified based on the modification instruction information to generate a corresponding target material.

[0039] The material optimization method provided in this embodiment realizes the automatic processing of the material to be modified by automatically editing the modified material through the modification instruction information to generate a corresponding target material, thereby reducing the modification cost and improving the material delivery efficiency.

[0040] The following Figure 2 elaborates in detail each step in steps S200 to S202 and optional other steps.

[0041] Step S200 Obtain the material to be modified and modification instruction information, where the modification instruction information includes a problem description of the material to be modified.

[0042] The material to be modified may be a material detected by the platform as not conforming to the platform specifications, such as patterns, texts, or titles in pictures. The modification instruction information may be a description of the place to be modified identified during the material detection process. For example, the non-compliant position and non-compliant reason of a non-compliant picture, the non-compliant sentences of a non-compliant title and non-compliant picture copywriting, etc.

[0043] Step S202 Edit the material to be modified based on the modification instruction information to generate a corresponding target material.

[0044] When the modification instruction information is used to indicate modifying a pattern in a picture, after identifying the pattern area in the material to be modified specified by the modification instruction information, the difference in this part can be locally redrawn to generate a picture that meets the specifications. When the modification instruction information is used to indicate modifying some text in a picture, the text in this part of the picture can be recognized, erased, and rewritten. When the modification instruction information is used to indicate modifying the title of the material to be modified, some words in the title can also be removed, and the title can be rewritten and optimized. After generating the target material, it can be submitted for detection again to re-verify whether the modified material meets the requirements.

[0045] The following will exemplarily introduce the specific process of editing the material to be modified to generate a corresponding target material.

[0046] When the modification instruction information is used to indicate modifying the pattern in the picture, in an alternative embodiment, as Figure 3 shown, step S202 includes: Step S300, based on the pre-trained target model and the modification instruction information, determine the area to be modified of the picture to be modified.

[0047] Step S302, generate a redrawn picture based on the area to be modified.

[0048] Step S304, generate a target picture based on the picture to be modified and the redrawn picture.

[0049] A corresponding mask can be generated according to the area to be modified, and a corresponding redrawn picture can be generated through this mask, and then the redrawn picture can be pasted back to the corresponding position of the picture to be modified. The target model can include a target detection model, and the area to be modified can be confirmed through the target detection model. The target detection model is used to identify and locate specific target objects from images or videos. In some embodiments, in order to retain as much as possible the information and content conveyed by the original material, and at the same time to avoid hitting text when redrawing the area to be modified, resulting in the failure of the redrawing effect, the text in the picture can be recognized and segmented, and the recognized text mask area can be removed from the original mask as a protected area, so as to ensure that the original text on the picture will not be changed, so as to achieve a more accurate editing effect.

[0050] In some embodiments, the modification instruction information can indicate that there is a non-compliance problem in the chest. The picture to be modified is recognized by the target detection model to generate a mask for the corresponding area. For example, after the target detection model recognizes the modified picture, the area that does not need to be modified can be represented by the black part. The white part corresponds to the mask of the chest area, and then the mask is redrawn and pasted back into the picture to be modified to obtain a target picture that meets the requirements.

[0051] In some embodiments, a scheme of joint detection using the NudeNet model and the GroundingDINO model can be adopted to realize the automatic generation of the mask. The NudeNet model is a sensitive content detection model used to identify sensitive content in images and videos. The GroundingDINO model is a target detection model with the ability of target recognition and positioning.

[0052] In some embodiments, the generation of the mask can be achieved through the VTON masker. The VTON Masker is used to perform regional segmentation on human body parts and generate masks for specific clothing areas, such as tops, pants, etc.

[0053] In this embodiment, the area to be modified is identified through a model, and then a redrawn picture is generated for the area to be modified. Then, a target picture is generated based on the redrawn picture and the picture to be modified. Thereby, while ensuring the integrity and visual effect of the picture to be modified, the complexity and cost of modification are reduced, and the compliance of the picture is improved.

[0054] Regarding the target model, in an alternative embodiment, the target model includes a target detection model for identifying the chest; as Figure 4 shown, the training operation of the target model includes: Step S400, performing model-assisted annotation on each training data in the training dataset to generate pseudo-labels for each training data.

[0055] Step S402, training the target detection model through each training data and its corresponding label; wherein, the label corresponding to each training data is obtained after calibrating the pseudo-label of each training data.

[0056] The detection data of the materials within the platform can be used as the training dataset. The pseudo-labels can be calibrated by manual calibration or other calibration methods, such as multi-model verification, etc., which are not limited herein. Before training the target detection model, data cleaning can be performed, such as removing duplicate or useless data, ensuring that the format of the labels is consistent with the model requirements, etc. In some embodiments, pseudo-labels can be generated through GroundingDINO-assisted annotation and combined with manual calibration for iterative training to expand the dataset.

[0057] In this embodiment, the target detection model is trained through each training data and its corresponding label, so that the quality of the chest mask generated by the target detection model is higher, and at the same time, the chest can be completely covered, thereby improving the redrawing effect of the chest.

[0058] It should be noted that some areas in the area to be modified do not need to be modified, or the picture quality will deteriorate after modification. Therefore, in the embodiment of the present application, the area to be redrawn is further determined in the area to be modified to alleviate the above problems.

[0059] Regarding determining the area to be redrawn, in an alternative embodiment, as Figure 5 shown, step S302 includes: Step S500, generating an initial mask area based on the area to be modified.

[0060] Step S502, when the picture to be modified includes specific type elements, determining a protected area that is not redrawn according to the specific elements.

[0061] Step S504: Determine the mask area to be redrawn based on the protected area and the initial mask area.

[0062] Specific types of elements can include the face, hands, feet, shoes, etc. In the case where the image to be modified involves complex postures of the human body, problems such as redundant limbs and distorted postures may occur in the generated redrawn picture. Therefore, specific elements can be identified, the areas where specific elements exist can be determined as protected areas that are not redrawn, and then the mask area to be redrawn can be confirmed. In some embodiments, the protected area can be generated through Human Parsing and DensePose. Human Parsing is a human body part semantic segmentation technology used to divide a human body image into multiple semantically related parts. DensePose is a human pose estimation technology used to map the human body pixels in a two-dimensional image to each point on the three-dimensional human body surface model, providing the precise three-dimensional position of each pixel, thereby enabling more complex pose and shape analysis.

[0063] In this embodiment, based on the area to be modified, the mask area to be redrawn is determined through the protected area and the initial mask area, thereby narrowing the redrawing range, alleviating the influence of specific elements on the redrawing effect, and ensuring the quality of material modification.

[0064] Regarding generating the redrawn picture, in an alternative embodiment, the picture to be modified includes human limbs; as Figure 6 shown, step S302 includes: Step S600: Preprocess the picture to be modified to obtain the key point information in the picture to be modified, where the key point information includes limb information.

[0065] Step S602: Generate the redrawn picture based on the picture to be modified, the mask area to be redrawn, and the key point information.

[0066] Limb information can be the key parts of the human body, such as joints, skeletons, connection points of limbs, etc. The key point information can ensure the consistency of the human body's posture and shape after redrawing, while the mask area to be redrawn ensures that only the parts that need to be modified are processed. In some embodiments, Stable Diffusion can be used to redraw the mask area to be redrawn. Stable Diffusion is an image generation technology based on the diffusion model used to generate high-quality images from text descriptions. In the case of redrawing the limbs, Stable Diffusion is combined with ControlNet to redraw the mask area to be redrawn. ControlNet is a neural network structure used to control the diffusion model, which is used to precisely control the image generation process by adding additional conditions.

[0067] In this embodiment, the key point information is used to redraw the mask area naturally and stably to generate a redrawn picture, thereby avoiding problems such as distorted and unnatural human postures in the target picture obtained by redrawing the picture.

[0068] In an alternative embodiment, the redrawing is implemented by an image redrawing model; the image redrawing model includes one or more of the following features: (1) The guidance coefficient value is lower than the default guidance coefficient value; (2) The strength value of the image redrawing model is determined according to the characteristics of the base model inside the image redrawing model.

[0069] The image redrawing model is a deep learning model used to repair, fill, or redraw missing or non-compliant areas in an image. In some embodiments, the guidance coefficient value can be the CFG (Classifier-Free Guidance) parameter value, and CFG is used to control the dependence of the generated content on the input prompt. An overly high CFG value is likely to result in inconsistent clothing with the original image. Therefore, lowering the CFG value helps to generate images with a consistent and natural style. The strength value can be the Strength parameter value, and the Strength parameter is used to control the intensity of modification in an image editing or generation task. Different base models have different sensitivities to Strength in different scenarios. Some base models can still maintain good shape consistency at a relatively high Strength, while others are more suitable for working at a lower Strength value. For example, for the case of redrawing the legs, when the base model is RealisticVisionV6-inpainting, the CFG and Strength can be adjusted to a lower value to improve the leg redrawing effect.

[0070] In this embodiment, by setting a guidance coefficient value lower than the default level and an appropriate strength value, the image redrawing model can generate a redrawn picture that is consistent with the style of the picture to be modified and natural, while achieving the best balance between diversity and stability.

[0071] In an alternative embodiment, the redrawing is realized by guiding the image redrawing model based on a target prompt; as Figure 7 shown, the target prompt acquisition operation includes: Step S700, generating an initial prompt based on the picture to be modified.

[0072] Step S702, adding corresponding scene words and deleting sensitive words based on the initial prompt to generate a target prompt.

[0073] The initial prompt can be deduced from the image to be modified, and appropriate scene words can be added to the initial prompt while removing sensitive words. For example, when redrawing the legs, prompt words such as pants and long skirts can be added; when redrawing the chest, prompt words such as high-necked clothes and round-necked clothes can be added. In some embodiments, reverse prompt words fixed through testing can also be added, and the reverse prompt words can be used to avoid generating images related to the reverse prompt words. In some embodiments, for parts with relatively little impact on redrawing, such as the chest, fixed prompt words can be used to save the time required for reverse deduction.

[0074] In this embodiment, the target prompt obtained based on the image to be modified can cover more feature details, thereby improving the consistency of redrawing.

[0075] When the modification instruction information is used to indicate modifying the title of the material to be modified, in an alternative embodiment, step S202 includes: generating a target title based on the modification instruction information and the title to be modified through a target language model.

[0076] The target language model can be an LLM (Large Language Model), which is used to understand and generate natural language text by training large-scale text data. In some embodiments, the LLM can be the Qwen model. In some embodiments, Qwen can be continuously pre-trained based on the data within the platform, and a finely customized Prompt engineering can be built based on this as the foundation. At the same time, the title can be rewritten through the CoT method. Qwen is a large-scale language and multimodal model used to perform various tasks, including natural language understanding, text generation, visual understanding, and audio understanding, etc. A prompt is a text prompt structure used to guide the model to generate specific outputs. CoT is a technology used to enhance the reasoning ability of large language models. CoT generates intermediate reasoning steps by guiding the model to solve problems step by step, thereby improving the performance of the model in complex tasks.

[0077] In this embodiment, generating a target title through the target language model realizes the automatic rewriting of the title to be modified, thereby improving the efficiency of title rewriting.

[0078] When the modification instruction information is used to indicate modifying some text in the image, in an alternative embodiment, as Figure 8 shown, step S202 includes: Step S800, generating target text based on the modification instruction information and the original text in the image to be modified.

[0079] Step S802, removing the original text of the image to be modified.

[0080] Step S804: Generate a target image with the target text embedded therein based on the target text and the image to be modified from which the original text has been removed.

[0081] The image to be modified can be a cover image containing the text to be modified. In some embodiments, the original text and the text box position detected by OCR can be obtained, the original text can be rewritten by an LLM to obtain the target text, and the target text can be pasted back into the text box of the image to be modified from which the original text has been removed to generate the target image. OCR (Optical Character Recognition) is used to convert the text in an image into editable text by detecting the dark and bright patterns in the image to determine the character shapes and translating these shapes into text that can be recognized by a computer. Median color picking logic can be used when picking colors, and the median is not affected by extreme values (such as abnormal brightness or color values in the image).

[0082] In this embodiment, by generating the target text from the original text and the image to be modified from which the original text has been removed, and generating a target image with the target text embedded therein, it is possible to accurately modify the text on the image to be modified without affecting the background content of the image to be modified.

[0083] In an alternative embodiment, as Figure 9 shown, the removal is achieved through a text removal model; the training operation of the text removal model includes: Step S900: Input the training data into a pre-trained model with fixed parameters to output text image features and corresponding textless image features; the training data includes multiple image groups, and each image group includes a text image and a corresponding textless image.

[0084] Step S902: Obtain a text removal loss value based on the text image features and the textless image features.

[0085] Step S904: Train the text removal model based on the text removal loss value and the text image.

[0086] Positive and negative example contrast training can be used to induce the text removal model not to generate text during redrawing. The textless image is used as a positive example, and the text-containing image is used as a negative example. By comparing the feature differences between the positive and negative examples, the text removal model can understand how to distinguish text regions and not generate text when redrawing the image. In some embodiments, the text removal model serves as the student network, and the pre-trained model with fixed parameters serves as the teacher network. The image group is input into the teacher network model to respectively output the text image features and the corresponding textless image features, and the text removal loss value is calculated based on these two outputs. Then, the student network, i.e., the text removal model, is trained based on this loss value. The teacher-student structure provides guidance for the student network by fixing the output of the teacher network.

[0087] The following illustrates the effect differences between the text removal model trained based on this embodiment and the text removal model not trained based on this embodiment through specific embodiments. As Figure 11 shown, "Ours" represents the cleaning effect diagram of the text removal model trained based on this embodiment; the strong redrawing SDxl-inpainting model and the weak redrawing SD-ControlNet-tile model correspond to the effect diagrams of the text removal models not trained based on this embodiment. The strong redrawing SDxl-inpainting is an image repair technology based on Stable Diffusion XL, which is used to redraw or repair specific regions in the image. The weak redrawing SD-ControlNet-tile is an auxiliary model based on Stable Diffusion, which is used to provide additional structured control during image generation and is suitable for content generation of large-size images or complex scenes.

[0088] From Figure 10 the shown effect diagrams, it can be seen that the strong redrawing model will randomly generate some text or patterns due to its overly strong redrawing ability, and the weak redrawing model has a poor erasing effect due to its overly weak redrawing ability. However, the text removal model trained based on this embodiment can not only maintain a strong redrawing ability but also not randomly generate text or patterns. The visualization results ( Figure 10 ), and the quantitative indicators ( Figure 11 ) both demonstrate the excellent effect of the text removal model.

[0089] In this embodiment, the text removal model trained through this embodiment improves the text removal effect and at the same time avoids generating redundant incorrect text.

[0090] To make this application easier to understand, the following combines Figure 12 to provide an exemplary application.

[0091] 1. Upload material A to the service platform through the client for the service platform to detect.

[0092] 2. When the service platform detects that material A does not meet the requirements, it outputs modification instruction information for material A.

[0093] 3. The server platform automatically edits material A based on the modification instruction information: (1) When the modification instruction information is used to indicate modifying the pattern in the picture: S1. Determine the picture to be modified.

[0094] S2. Based on the pre-trained target model and the modification instruction information, determine the area to be modified in the picture to be modified.

[0095] The target model can be a target detection model, and the model can be trained with training data and its corresponding labels. The labels are generated as pseudo-labels by model-assisted annotation and then calibrated.

[0096] S3. Generate a redrawn picture based on the area to be modified.

[0097] Generate an initial mask area according to the area to be modified. When the picture to be modified includes specific types of elements (such as faces, hands, feet, etc.), determine the protected area that will not be redrawn according to the specific elements. Combine the initial mask area and the protected area to determine the mask area to be redrawn. Then input the mask area to be redrawn into the image redrawing model, and guide the image redrawing model to redraw through the target prompt word to generate the redrawn picture. The target prompt word is obtained by adding scene words and deleting sensitive words after generating the initial prompt word according to the picture to be modified.

[0098] S4. Generate the target picture based on the picture to be modified and the redrawn picture.

[0099] Paste the redrawn picture back to the corresponding position of the picture to be modified to generate the target picture.

[0100] (2) When the modification instruction information is used to indicate modifying some text in the picture: S1. Determine some text in the picture to be modified.

[0101] S2. Generate the target text based on the modification instruction information and the original text in the picture to be modified.

[0102] The original text obtained by OCR detection can be retrieved, and the target text can be obtained by rewriting the original text through the LLM.

[0103] S3. Remove the original text in the picture to be modified.

[0104] Remove the original text in the image to be modified through the trained text removal model. Train the text removal model (Student network) based on a pre-trained model with fixed parameters (Teacher network). After inputting an image group including a text image and the corresponding text-free image into the Teacher network, extract the text image features and the text-free image features respectively, and calculate the text removal loss value based on the two. Then, use this loss value to train the Student network (i.e., the text removal model).

[0105] S4. Generate a target image with the target text embedded based on the target text and the image to be modified from which the original text has been removed.

[0106] Paste the target text back into the text box of the image to be modified from which the original text has been removed to generate a target image with the target text embedded.

[0107] (3) When the modification instruction information is used to indicate modifying the title of the material to be modified: S1. Determine the title to be modified.

[0108] S2. Generate a target title through a target language model based on the modification instruction information and the title to be modified.

[0109] The target language model can be an LLM model, and the removal of non-compliant vocabulary and the rewriting and optimization of the title are realized based on the LLM.

[0110] IV. The server platform submits the modified material A for detection again.

[0111] In this exemplary application, the automated processing of the material to be modified is realized, thereby improving the modification efficiency and automation, reducing the modification cycle and cost, and improving the material delivery efficiency.

[0112] Embodiment 2 Figure 13 Schematically shows a block diagram of a material optimization device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 13 shown, the device 1300 may include: an acquisition module 1310, a generation module 1320, where: The acquisition module 1310 is used to acquire the material to be modified and modification instruction information, and the modification instruction information includes a problem description of the material to be modified; A generation module 1320, configured to edit the material to be modified based on the modification indication information to generate a corresponding target material.

[0113] As an optional embodiment, the material to be modified includes a picture to be modified; the generation module 1420 is further configured to: Determine a region to be modified of the picture to be modified based on a pre-trained target model and the modification indication information; Generate a redrawn picture based on the region to be modified; and Generate a target picture based on the picture to be modified and the redrawn picture.

[0114] As an optional embodiment, the target model includes a target detection model for identifying the chest; the training operation of the target model includes: Perform model-assisted annotation on each training data in the training dataset to generate pseudo-labels for each training data; Train the target detection model through each training data and its corresponding label; wherein, the label corresponding to each training data is obtained after calibrating the pseudo-label of each training data.

[0115] As an optional embodiment, the generation module 1320 is further configured to: Generate an initial mask region based on the region to be modified; In the case where the picture to be modified includes specific type elements, determine a protected region that is not redrawn according to the specific elements; Determine a mask region to be redrawn based on the protected region and the initial mask region.

[0116] As an optional embodiment, the picture to be modified includes a human body limb; the generation module 1420 is further configured to: Preprocess the picture to be modified to obtain key point information in the picture to be modified, where the key point information includes limb information; Generate the redrawn picture based on the picture to be modified, the mask region to be redrawn, and the key point information.

[0117] As an optional embodiment, the redrawing is implemented through an image redrawing model; the image redrawing model includes one or more of the following features: (1) The guiding coefficient value is lower than the default guiding coefficient value; (2) The intensity value of the image redrawing model is determined according to the characteristics of the inner model of the image redrawing model.

[0118] As an optional embodiment, the redrawing is guided by a target prompt word to implement the image redrawing model; the operation for obtaining the target prompt word includes: Generate an initial prompt based on the picture to be modified; Based on the initial prompt, add corresponding scene words and delete sensitive words to generate a target prompt.

[0119] As an optional embodiment, the material to be modified includes a title to be modified; the generation module 1420 is further configured to: Generate a target title based on the modification instruction information and the title to be modified through a target language model.

[0120] As an optional embodiment, the material to be modified includes a picture to be modified; the generation module 1420 is further configured to: Generate target text based on the modification instruction information and the original text in the picture to be modified; Remove the original text of the picture to be modified; Generate a target picture with the target text embedded therein based on the target text and the picture to be modified from which the original text has been removed.

[0121] As an optional embodiment, the removal is implemented through a text clearing model; the training operation of the text clearing model includes: Input training data into a pre-trained model with fixed parameters to output text image features and corresponding text-free image features; the training data includes multiple image groups, and each image group includes a text image and a corresponding text-free image; Obtain a text clearing loss value based on the text image features and the text-free image features; Train the text clearing model based on the loss value of the text clearing and the text image.

[0122] Embodiment III Figure 14 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the material optimization method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack-mounted server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers). As Figure 14 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the material optimization method. In addition, the memory 10010 may also be used to temporarily store various types of data that have been output or will be output.

[0123] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0124] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0125] It should be noted that Figure 14 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0126] In this embodiment, the material optimization method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of this application.

[0127] Embodiment 4 This application embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the material optimization method in the embodiment are implemented.

[0128] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the material optimization method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0129] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.

[0130] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0131] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A material optimization method, characterized in that: The method comprises: Acquire a material to be modified and modification instruction information, wherein the modification instruction information includes a description of a problem of the material to be modified; and Based on the modification indication information, the material to be modified is edited to generate a corresponding target material.

2. The method according to claim 1, characterized in that The material to be modified includes a picture to be modified; based on the modification indication information, the material to be modified is edited to generate a corresponding target material, including: Determine the area to be modified of the image to be modified based on the pre-trained target model and the modification indication information; Based on the area to be modified, generating a redrawn picture; and A target picture is generated based on the picture to be modified and the redrawn picture.

3. The method according to claim 2, characterized in that The target model includes a target detection model for identifying a chest; The training operation of the target model includes: Performing model-assisted labeling on each training data in the training data set to generate a pseudo label for each training data; The target detection model is trained using the training data and the corresponding labels; wherein the labels corresponding to the training data are obtained by calibrating the pseudo labels of the training data.

4. The method according to claim 2, characterized in that: Based on the area to be modified, a redrawn picture is generated, including: Based on the area to be modified, generating an initial mask area; In the case where the image to be modified includes a specific type of element, determining a protection area not to be redrawn according to the specific element; Based on the protection area and the initial mask area, a mask area to be redrawn is determined.

5. The method according to claim 4, characterized in that The image to be modified includes human limbs; Based on the area to be modified, a redrawn picture is generated, including: Preprocessing the image to be modified to obtain key point information in the image to be modified, wherein the key point information includes limb information; The redrawn picture is generated based on the picture to be modified, the mask area to be redrawn, and the key point information.

6. The method according to claim 4, characterized in that The redrawing is achieved through an image redrawing model; The image redrawing model includes one or more of the following features: (1) The bootstrap coefficient value is lower than the default bootstrap coefficient value; (2) The strength value of the image redrawing model is determined according to the characteristics of the bottom mold within the image redrawing model.

7. The method according to claim 6, characterized in that The redrawing is implemented by guiding the image redrawing model based on the target prompt word; The target prompt word acquisition operation includes: Based on the picture to be modified, generate an initial prompt word; Based on the initial prompt words, corresponding scene words are added and sensitive words are deleted to generate target prompt words.

8. The method according to claim 1, characterized in that The material to be modified includes a title to be modified; based on the modification indication information, the material to be modified is edited to generate a corresponding target material, including: Based on the modification indication information and the title to be modified, a target title is generated by a target language model.

9. The method according to claim 1, characterized in that: The material to be modified includes a picture to be modified; based on the modification indication information, the material to be modified is edited to generate a corresponding target material, including: Generate target text based on the modification indication information and the original text in the picture to be modified; Removing the original text of the image to be modified; Based on the target text and the image to be modified from which the original text has been removed, a target image in which the target text is embedded is generated.

10. The method according to claim 1, characterized in that The removal is achieved by a text cleaning model; The training operation of the text cleaning model includes: Inputting training data into a pre-trained model with fixed parameters to output text image features and corresponding non-text image features; the training data includes a plurality of image groups, each image group includes a text image and a corresponding non-text image; Based on the text image feature and the text-free image feature, obtaining a text removal loss value; The text cleaning model is trained based on the loss value of the text cleaning and the text image.

11. A material optimization device, characterized in that: The device comprises: An acquisition module, used to acquire a material to be modified and modification instruction information, wherein the modification instruction information includes a description of a problem of the material to be modified; A generating module is used to edit the material to be modified based on the modification indication information to generate a corresponding target material.

12. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 10 are implemented.