An image processing method, apparatus, electronic device, storage medium, and program product.
By receiving prompt text and images to generate descriptive text and encoding it, the problem of image processing relying on manual operation in existing technologies is solved, and automated high-quality image fusion and visual effect consistency are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, image processing relies on manual operation, resulting in low consistency and cumbersome visual effects. It also fails to automatically extract visual information from reference objects, thus reducing the effectiveness of image processing.
By receiving prompt text, a first image, and a second image, descriptive text is generated and encoded. Combined with image encoding, noise reduction is performed to generate the target image, thus fusing visual information and content objects.
It achieves automated image processing, improves the effect and consistency of image processing, and generates high-quality fused images that meet the requirements.
Smart Images

Figure CN122134870A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing, and more particularly to an image processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] In image processing techniques, visual effects heavily rely on manual intervention, requiring the identification and adjustment of materials within the reference object. This process is tedious and prone to errors. Furthermore, it cannot automatically extract visual information from the reference object, resulting in low consistency between the transferred material effects and the visual effects of the reference object, thus reducing the overall effectiveness of image processing. Summary of the Invention
[0003] This disclosure proposes an image processing method, apparatus, electronic device, storage medium, and program product, which at least partially solves the technical problems of poor image processing performance in related technologies.
[0004] In a first aspect, this disclosure provides an image processing method, comprising:
[0005] Receive prompt text, a first image, and a second image; the prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image; Based on the prompt text, the first image, and the second image, generate a first description text for the first image, a second description text for the second image, and a target description text; Text features are obtained by performing text encoding on the first description text, the second description text, the target description text, and the prompt text; and image features are obtained by performing image encoding on the first image, the second image, and the noise image. The target image is obtained by performing noise reduction processing based on the text features and the image features.
[0006] A second aspect of this disclosure provides an image processing apparatus, comprising: A receiving module is used to receive a prompt text, a first image, and a second image; the prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image; The description text module is used to generate a first description text for the first image, a second description text for the second image, and a target description text based on the prompt text, the first image, and the second image; The feature encoding module is used to perform text encoding on the first descriptive text, the second descriptive text, the target descriptive text, and the prompt text to obtain text features; and to perform image encoding on the first image, the second image, and the noise image to obtain image features. The target image module is used to perform noise reduction processing based on the text features and the image features to obtain the target image.
[0007] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in the first aspect.
[0008] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect.
[0009] A fifth aspect of this disclosure provides a computer program product including computer program instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect.
[0010] As can be seen from the above description, the image processing method, apparatus, electronic device, storage medium, and program product provided in this disclosure receive prompt text and a first image and a second image, generate their respective descriptive text and target descriptive text, then encode and extract text features from the descriptive text, and simultaneously encode and extract image features from the first image, the second image, and the noisy image. Based on the text features and image features, noise reduction processing is performed to generate the target image. This effectively fuses specific information from different images to generate a high-quality fused image that meets the requirements, thus improving the image processing effect. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the image processing architecture according to an embodiment of the present disclosure.
[0013] Figure 2 This is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
[0014] Figure 3 This is a schematic flowchart of an image processing method according to an embodiment of the present disclosure.
[0015] Figure 4 This is a schematic diagram illustrating the principle of image processing according to an embodiment of the present disclosure.
[0016] Figure 5 This is a schematic diagram of image processing for a reference text style in an embodiment of this disclosure.
[0017] Figure 6 This is a schematic diagram of image processing for a reference layout structure according to an embodiment of the present disclosure.
[0018] Figure 7 This is a schematic diagram of image processing for reference text styles and layout structures in embodiments of this disclosure.
[0019] Figure 8 This is a schematic diagram illustrating the construction of training data for an embodiment of this disclosure.
[0020] Figure 9 This is a schematic diagram illustrating the construction of training data for an embodiment of this disclosure.
[0021] Figure 10 This is a schematic diagram of an image processing apparatus according to an embodiment of the present disclosure. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0023] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0024] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained. For example, in response to receiving a user's active request, a prompt message may be sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operations of the technical solutions of this disclosure, based on the prompt message.
[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0026] Figure 1 A schematic diagram of an image processing architecture according to an embodiment of the present disclosure is shown. (Reference) Figure 1 The image processing architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected via a wired or wireless network 130. The server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, security services, and CDN.
[0027] Terminal 120 can be implemented in hardware or software. For example, when terminal 120 is implemented in hardware, it can be various electronic devices with a display screen and support page display, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module, without specific limitations.
[0028] It should be noted that the image processing method provided in this embodiment can be executed by the terminal 120, by the server 110, or by both the terminal 120 and the server 110. It should be understood that... Figure 1 The number of terminals, networks, and servers shown is for illustrative purposes only and is not intended to be a limitation. Any number of terminals, networks, and servers can be used depending on implementation needs.
[0029] Figure 2A schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of this disclosure is shown. Figure 2 As shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. The processor 202, memory 204, network module 206, and peripheral interface 208 are interconnected within the electronic device 200 via the bus 210.
[0030] Processor 202 may be a Central Processing Unit (CPU), a Neural Processing Unit (NPU), a Microcontroller (MCU), a programmable logic device, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits. Processor 202 can be used to perform functions related to the techniques described in this disclosure. In some embodiments, processor 202 may also include multiple processors integrated as a single logic component. For example, such as... Figure 2 As shown, processor 202 may include multiple processors 202a, 202b and 202c.
[0031] Memory 204 can be configured to store data (e.g., instructions, computer code, etc.). Figure 2 As shown, the data stored in memory 204 may include program instructions (e.g., program instructions for implementing the image processing method of the embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration documents of other modules, etc.). Processor 202 may also access the program instructions and data stored in memory 204 and execute the program instructions to operate on the data to be processed. Memory 204 may include volatile storage devices or non-volatile storage devices. In some embodiments, memory 204 may include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive (SSD), flash memory, memory stick, etc.
[0032] Network module 206 can be configured to provide communication with other external devices to electronic device 200 via a network. This network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It is understood that the type of network is not limited to the specific examples described above. In some embodiments, network module 206 may include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.
[0033] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to enable information input and output. For example, peripheral devices may include input devices such as keyboards, mice, touchpads, touch screens, microphones, and various sensors, as well as output devices such as displays, speakers, vibrators, and indicator lights.
[0034] Bus 210 can be configured to transmit information between various components of electronic device 200 (e.g., processor 202, memory 204, network module 206, and peripheral interface 208), such as internal buses (e.g., processor-memory bus), external buses (USB port, PCI-E bus), etc.
[0035] It should be noted that although the architecture of the above-described electronic device 200 only shows the processor 202, memory 204, network module 206, peripheral interface 208, and bus 210, in specific implementations, the architecture of the electronic device 200 may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the architecture of the above-described electronic device 200 may only include the components necessary for implementing the embodiments of this disclosure, and does not necessarily include all the components shown in the figures.
[0036] See Figure 3 , Figure 3 A schematic flowchart of an image processing method according to an embodiment of the present disclosure is shown. The image processing method according to an embodiment of the present disclosure can be deployed on a server or a terminal. Figure 3 In the image processing method 300, the following steps may be further included.
[0037] In step S310, a prompt text, a first image, and a second image are received; the prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image.
[0038] The prompt text can be a textual description guiding the image processing process, such as how to fuse specific visual information from the first image with content objects in the second image. The first image can be an input image, containing the visual information to be extracted or fused. The second image can be another input image, containing content objects that will be fused with the visual information from the first image. Visual information can refer to visually perceptible information in the image, such as color, texture, shape, edges, and layout. Content objects can refer to specific things or entities in the image, such as people, objects, or backgrounds.
[0039] By receiving prompt text containing fusion instructions and two source images, specific visual elements (such as style, text style, typography, etc.) in the first image can be merged with the actual content objects (such as people, objects, or scenes) in the second image. For example, the second image can be made to have the same visual effect as the first image. The second image can be processed with reference to the visual effect of the first image to generate a target image that retains the content of the second image while incorporating the visual effect of the first image, achieving innovation and enhancement of image content and improving the effect of image processing.
[0040] Specifically, see Figure 4 , Figure 4 A schematic diagram illustrating the principle of image processing according to an embodiment of the present disclosure is shown. Figure 4 In this context, the Visual Language Model (VLM) receives prompt text from the user (such as editing instructions to indicate changes to text style), input a first image (e.g., a reference image), and a second image (poster or stock image).
[0041] In step S320, based on the prompt text, the first image, and the second image, a first description text for the first image, a second description text for the second image, and a target description text are generated.
[0042] The first descriptive text refers to descriptive information about the first image, used to describe the key features and information of the first image. The second descriptive text, similar to the first, refers to descriptive information about the second image, used to describe the key features and information of the second image. The target descriptive text refers to descriptive information that combines the requirements of the prompt text with the features of the first and second images, used to describe relevant information about the desired target image, and can guide the fusion process of the first and second images.
[0043] Based on the received prompt text, the first image, and the second image, a first descriptive text describing the characteristics of the first image and a second descriptive text describing the characteristics of the second image are generated, along with a target descriptive text that integrates the prompt requirements and image information, pointing to the final generated image. This allows for a comprehensive and accurate extraction and summarization of key information from the images and prompts, providing a more targeted and guiding information foundation for subsequent image processing steps in text form, thus helping to improve the accuracy and quality of image fusion generation. Figure 4 As shown, the visual language model can generate descriptive text for the first and second images, as well as descriptive text for the target image (the desired result image), namely, the first descriptive text, the second descriptive text, and the target descriptive text.
[0044] In step S330, text encoding is performed on the first description text, the second description text, the target description text, and the prompt text to obtain text features; and image encoding is performed on the first image, the second image, and the noise image to obtain image features.
[0045] Text encoding refers to converting text information into feature vectors, and text features refer to the feature representation obtained after text encoding, which can contain key information and semantic content in the text. Image encoding refers to converting image information into feature vectors, and image features refer to the feature representation obtained after image encoding, which can contain key visual elements and structural information in the image. Noisy images can be images containing random noise. Text encoding is performed on the generated first description text, second description text, target description text, and prompt text to extract text features containing semantic information. Simultaneously, image encoding is performed on the first image, second image, and noisy image to obtain image features reflecting the visual content of the images. This allows for the full extraction of text and image information, effectively improving the effect and performance of image processing.
[0046] like Figure 4 As shown, the first description text, the second description text, the target description text, and the original prompt text (such as editing instructions) are fed into a text encoder (e.g., a text encoder). Simultaneously, a position encoder (e.g., a Pos Emb module) is used for positional encoding to obtain their respective text tokens. Randomly generated noise is then fed into an image encoder (e.g., a VAE Encoder) along with the input first and second images. This noise is also positionally encoded by the position encoder to obtain their respective image tokens.
[0047] In step S340, noise reduction processing is performed based on the text features and the image features to obtain the target image.
[0048] Noise reduction processing refers to removing or reducing noise in an image to improve its quality and clarity. The target image can refer to the image processing result that fuses the visual information of the first image with the content of the second image. By utilizing the semantic guidance information contained in text features and the visual content information presented by image features, noise interference in the image can be identified and removed. This allows the generated target image to effectively improve image clarity and quality while retaining the required visual information and semantic content, ensuring visual consistency between the target image and the first image and meeting the requirements of high-quality image processing.
[0049] As shown in the figure, the preceding text terms and image terms are concatenated, and then the concatenated result is fed into a diffusion module (e.g., DiT) for training. After training, the result of the diffusion module is fed into an image decoder (e.g., VAEDecoder) to output the final target image.
[0050] In some embodiments, the visual information includes a text style, and the target image includes the content object displayed based on the text style and / or specified text in the prompt text; The visual information includes a layout structure, and the target image includes the content object displayed based on the layout structure.
[0051] Text style refers to the presentation of text, such as font, font size, color, weight, italics, underline, letter spacing, and line spacing. Layout structure refers to the arrangement and organization of text, images, and other elements on a page or screen. This includes the placement of elements, their size ratio, alignment, and hierarchical relationships. For example, in newspaper layout, headlines are usually placed in a prominent position, and the body text is arranged in columns.
[0052] Visual information can include text styles, which the target image uses to present specified text from the content object or prompt text. Visual information can also include typographic structure, which the target image uses to display the content object. This allows for the precise application of specific text styles and typographic structures to the content presentation of the target image, making the generated target image more closely match the requirements in terms of text representation and layout. This effectively improves the accuracy and professionalism of image generation, meeting diverse and customized image creation needs.
[0053] See Figure 5 , Figure 5 A schematic diagram illustrating image processing of a reference text style according to an embodiment of the present disclosure is shown. Figure 5 In the first image 510, the text "RRRRRRRRRRRRR" with text style S1 is included. The second image 520 includes the text "TTTTTTTTTTTTT" with text style S2 and an object 522. A prompt text indicates that the style of the text 521 in the second image 520 should be changed to the text style S1 in the first image 510, thus obtaining the target image 530. In the target image 530, the object 522 remains unchanged, while the text "TTTTTTTTTTTTT" is changed to the text 531 displayed with text style S1.
[0054] In some embodiments, the target image includes the content object displayed based on the layout structure, including: The layout position of the content object in the second image is determined based on the layout structure and the prompt text; The content object is displayed at the layout position in the second image to obtain the target image; the color of the content object matches the second image.
[0055] The process involves determining the content object's position within the second image based on the layout structure and prompt text, and then displaying the content object at that position while ensuring its color matches the second image. This allows the content object to naturally integrate into the established layout framework of the second image, guaranteeing not only the image's layout rationality and aesthetics but also enhancing overall visual consistency through color matching, thereby generating a high-quality target image that meets the expected layout effect.
[0056] In some embodiments, displaying the content object at the layout position in the second image includes: Replace the matching object corresponding to the layout position in the second image with the content object, and adaptively update the color of the content object to match the second image.
[0057] Specifically, by replacing the original matching object corresponding to the layout position in the second image with the desired content object, and adaptively adjusting the color of the content object to match the overall tone of the second image, the content object can be naturally integrated into the second image. This ensures visual harmony and unity of the replaced image, effectively improving the generation quality and visual effect of the target image.
[0058] See Figure 6 , Figure 6 A schematic diagram of image processing of a reference layout structure according to an embodiment of the present disclosure is shown. Figure 6 In the first image 610, text 611 and object 612 with a layout structure P1 are included. The second image 620 includes object 620. A prompt text can instruct that the second image 620 be displayed using the layout structure P1 of the first image 610, thus obtaining the target image 630. In the target image 630, text 611 remains unchanged, while object 612 becomes object 620.
[0059] See Figure 7 , Figure 7 A schematic diagram illustrating image processing of reference text styles and typographic structures according to embodiments of the present disclosure is shown. Figure 7In the first image 710, there are objects 711 with a layout structure P2, text 712 "A", text 713 "B", and text 714 "C", all of which use text style S3. The second image 720 includes objects 721 with a layout structure P3, text 722 "1", text 723 "2", and text 724 "3", all of which use text style S4. A prompt text can instruct the second image 720 to be displayed using the layout structure P2 of the first image 710, and to use the text style of the first image 710 for all text portions of the second image 720, thus obtaining the target image 730. In target image 730, object 711 in first image 710 is replaced with object 721 in second image 720. The text styles S3 of text 712 "A", text 713 "B", and text 714 "C" remain unchanged, and their text contents are replaced with "1", "2", and "3" respectively, resulting in text 73a, 73b, and 73c. Object 721, text 73a, 73b, and 73c in target image 730 are typeset based on typesetting structure P2.
[0060] In some embodiments, the prompt text includes prompt training text, the first image includes a first training image, the second image includes a second training image, and the prompt training text is used to instruct the fusion of training visual information in the first training image with training content objects in the second training image; Based on the prompt training text, the first training image, the second training image, and the target training image, an image processing model for the image processing is trained; wherein the first training image and the second training image are obtained based on the target training image.
[0061] The image processing model is trained using prompt training text, a first training image, a second training image, and a target training image as a reference standard. This allows the model to fuse different image information based on the prompt text, thereby more effectively generating high-quality, expected target images according to user prompts in practical applications.
[0062] In some embodiments, the first training image and the second training image are obtained based on the target training image, including: The first text content in the target training image is updated to the second text content to obtain the first intermediate image; The second text content in the first intermediate image is cropped to obtain the first training image for the second text content; Furthermore, the first text style of the first text content in the target training image is updated to the second text style to obtain the second training image.
[0063] This process involves reversing the process of a target training image. First, the first text content is replaced with second text content to obtain a first intermediate image. Then, the second text content in the first intermediate image is cropped to form the first training image. Simultaneously, the first text style of the first text content in the target training image is changed to the second text style to obtain a second training image. This simulates the user-input first training image (e.g., a font reference image) and the second training image (e.g., an image with the text style to be updated). This allows for the targeted generation of training images with different text content and styles, providing diverse text style data for image processing model training. This enables the model to better learn the fusion of visual information related to text styles, improving its ability to handle text-image fusion tasks.
[0064] See Figure 8 , Figure 8 A schematic diagram illustrating the construction of training data according to an embodiment of this disclosure is shown. Figure 8 In this method, the largest text region in the target training image is located using Optical Character Recognition (OCR) technology to determine the text portion requiring modification. Two modification operations can be performed on the located text region: one is text style modification, adjusting the font, size, color, weight, and other style attributes to generate a style-modified training image, which serves as the second training image; the other is text content modification, changing the actual text content to generate a content-modified first intermediate image. The text region of this first intermediate image is cropped to form an independent font reference image, which can be used as the second training image. This allows for the construction of rich and diverse training data, which can be used to train image editing models, especially for tasks involving text style and content modification.
[0065] In some embodiments, the first training image and the second training image are obtained based on the target training image, including: The third text content in the target training image is updated to the fourth text content, and the first content object is updated to the second content object to obtain the first training image; And, crop the second content object in the target training image and merge it with the first background to obtain the second intermediate image that is for the second content object and has the first background; The second training image is obtained based on the second intermediate image; The method of obtaining the second training image based on the second intermediate image further includes the following: The second intermediate image is determined to be the second training image; The first background in the second intermediate image is updated to the second background to obtain the second training image with the second background. The first background in the second intermediate image is updated to a third background, and at least one visual element is added to obtain the second training image.
[0066] The process involves reverse engineering the resulting image to perform a dual update on the target training image: replacing text content and objects to obtain the first training image; then, cropping a new object from the target training image and merging it with the first background to create a second intermediate image, which is then used to generate the second training image. The second training image can be generated by directly using the original image, replacing the background, or replacing the background and adding visual materials. This approach, through different material construction methods, simulates diverse user input methods, flexibly creating diverse and targeted training images. This provides rich data for the image processing model, helping it learn image fusion rules in different scenarios, improving its ability to handle complex image fusion tasks and its generalization ability, and ensuring the generalization and reference accuracy of the final model's performance.
[0067] See Figure 9 , Figure 9 A schematic diagram illustrating the construction of training data according to an embodiment of this disclosure is shown. Figure 9In this model, objects and text in the target training image can be modified using an image editing model to obtain a layout reference image, which serves as the first training image. Objects in the target training image can be identified and segmented to obtain object images. Various input scenarios can be simulated based on these object images. The object image can be placed on a white background to simulate a clean-background source image, resulting in a second intermediate image. This can be used to train the model to handle images with prominent subjects and simple backgrounds, helping the model learn how to properly layout the subject against a clean background. The background of the second intermediate image can be changed, for example, by replacing it with a background that includes scene information (such as a desktop background), to simulate a source image with a regular background, resulting in a third intermediate image. This exposes the model to images closer to real-world application scenarios, allowing it to learn how to maintain a reasonable layout and visual effect of the subject even with background interference. Furthermore, the background of the second intermediate image can be changed while adding additional materials, such as text or images, resulting in a fourth intermediate image. This helps the model comprehensively learn how to handle layout processing for more complex scenes, meeting diverse user needs for image processing. One of the second, third, or fourth intermediate images can be used as the second training image. It is evident that the construction of training data simulates a variety of input formats through various operations, generating diverse training data. This provides comprehensive and targeted training materials for image editing models related to typesetting, which helps improve the model's typesetting performance in different scenarios.
[0068] In some embodiments, an image processing model for the image processing is trained based on the prompt training text, the first training image, the second training image, and the target training image, including: Based on the prompt training text, the first training image, and the second training image, a prediction training image is obtained; A loss function is determined based on the predicted training image and the target training image; The model parameters are adjusted to minimize the loss function to obtain the image processing model.
[0069] The process involves first generating a predicted training image based on the prompt training text and the first and second training images. Then, the loss function is determined by comparing the predicted training image with the target training image. Finally, the model parameters are continuously adjusted to minimize the loss function, thus training the image processing model. This training method accurately measures the difference between the model-generated image and the expected target image. By optimizing parameters and gradually narrowing the gap, the trained image processing model possesses stronger image processing capabilities and can more accurately and effectively complete image fusion and other processing tasks. The image processing model disclosed here features a unified model architecture supporting multimodal and multi-reference images. It not only supports generation from multiple input images but also supports different reference tasks, achieving joint optimization of font reference and typography reference tasks. This effectively reduces model redundancy, improves cross-task knowledge transfer capabilities, and lowers the cost of model deployment and subsequent maintenance.
[0070] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and the multiple devices will perform image processing together to complete the method described.
[0071] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0072] Based on the same technical concept, corresponding to any of the above embodiments, this disclosure also provides an image processing apparatus, see [link to relevant documentation]. Figure 10 The image processing apparatus includes: A receiving module is used to receive a prompt text, a first image, and a second image; the prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image; The description text module is used to generate a first description text for the first image, a second description text for the second image, and a target description text based on the prompt text, the first image, and the second image; The feature encoding module is used to perform text encoding on the first descriptive text, the second descriptive text, the target descriptive text, and the prompt text to obtain text features; and to perform image encoding on the first image, the second image, and the noise image to obtain image features. The target image module is used to perform noise reduction processing based on the text features and the image features to obtain the target image.
[0073] In some embodiments, the visual information includes a text style, and the target image includes the content object displayed based on the text style and / or specified text in the prompt text; The visual information includes a layout structure, and the target image includes the content object displayed based on the layout structure; The target image includes the content object displayed based on the layout structure, including: The layout position of the content object in the second image is determined based on the layout structure and the prompt text; The content object is displayed at the layout position in the second image to obtain the target image; the color of the content object matches the second image.
[0074] In some embodiments, displaying the content object at the layout position in the second image includes: Replace the matching object corresponding to the layout position in the second image with the content object, and adaptively update the color of the content object to match the second image.
[0075] In some embodiments, the prompt text includes prompt training text, the first image includes a first training image, the second image includes a second training image, and the prompt training text is used to instruct the fusion of training visual information in the first training image with training content objects in the second training image; Based on the prompt training text, the first training image, the second training image, and the target training image, an image processing model for the image processing is trained; wherein the first training image and the second training image are obtained based on the target training image.
[0076] In some embodiments, the first training image and the second training image are obtained based on the target training image, including: The first text content in the target training image is updated to the second text content to obtain the first intermediate image; The second text content in the first intermediate image is cropped to obtain the first training image for the second text content; Furthermore, the first text style of the first text content in the target training image is updated to the second text style to obtain the second training image.
[0077] In some embodiments, the first training image and the second training image are obtained based on the target training image, including: The third text content in the target training image is updated to the fourth text content, and the first content object is updated to the second content object to obtain the first training image; And, crop the second content object in the target training image and merge it with the first background to obtain the second intermediate image that is for the second content object and has the first background; The second training image is obtained based on the second intermediate image; The method of obtaining the second training image based on the second intermediate image further includes the following: The second intermediate image is determined to be the second training image; The first background in the second intermediate image is updated to the second background to obtain the second training image with the second background. The first background in the second intermediate image is updated to a third background, and at least one visual element is added to obtain the second training image.
[0078] In some embodiments, an image processing model for the image processing is trained based on the prompt training text, the first training image, the second training image, and the target training image, including: Based on the prompt training text, the first training image, and the second training image, a prediction training image is obtained; A loss function is determined based on the predicted training image and the target training image; The model parameters are adjusted to minimize the loss function to obtain the image processing model.
[0079] The apparatus of the above embodiments is used to implement the corresponding image processing method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0080] Based on the same technical concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the image processing method as described in any of the above embodiments.
[0081] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable multimedia, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0082] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the image processing method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0083] Based on the same inventive concept, corresponding to the image processing methods of any of the above embodiments, this disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, when the computer program instructions are run on a computer, the computer performs each step of each embodiment of the image processing method. Corresponding to the execution entity for each step in each embodiment of the image processing method, the processor performing the corresponding step may belong to the corresponding execution entity.
[0084] The computer program products of the above embodiments are used to cause the processor to execute the image processing method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0085] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0086] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0087] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0088] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An image processing method, comprising: Receive prompt text, first image, and second image; The prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image; Based on the prompt text, the first image, and the second image, generate a first description text for the first image, a second description text for the second image, and a target description text; Text features are obtained by performing text encoding on the first description text, the second description text, the target description text, and the prompt text; and image features are obtained by performing image encoding on the first image, the second image, and the noise image. The target image is obtained by performing noise reduction processing based on the text features and the image features.
2. The method according to claim 1, wherein, The visual information includes text styles, and the target image includes the content object displayed based on the text styles and / or the specified text in the prompt text; The visual information includes a layout structure, and the target image includes the content object displayed based on the layout structure; The target image includes the content object displayed based on the layout structure, including: The layout position of the content object in the second image is determined based on the layout structure and the prompt text; The content object is displayed at the layout position in the second image to obtain the target image; the color of the content object matches the second image.
3. The method according to claim 2, wherein displaying the content object at the layout position in the second image includes: Replace the matching object corresponding to the layout position in the second image with the content object, and adaptively update the color of the content object to match the second image.
4. The method according to claim 1, wherein, The prompt text includes prompt training text, the first image includes a first training image, the second image includes a second training image, and the prompt training text is used to instruct the fusion of training visual information in the first training image with training content objects in the second training image; Based on the prompt training text, the first training image, the second training image, and the target training image, an image processing model for the image processing is trained; wherein the first training image and the second training image are obtained based on the target training image.
5. The method according to claim 4, wherein, The first training image and the second training image are obtained based on the target training image, including: The first text content in the target training image is updated to the second text content to obtain the first intermediate image; The second text content in the first intermediate image is cropped to obtain the first training image for the second text content; Furthermore, the first text style of the first text content in the target training image is updated to the second text style to obtain the second training image.
6. The method according to claim 4, wherein, The first training image and the second training image are obtained based on the target training image, including: The third text content in the target training image is updated to the fourth text content, and the first content object is updated to the second content object to obtain the first training image; And, crop the second content object in the target training image and merge it with the first background to obtain the second intermediate image that is for the second content object and has the first background; The second training image is obtained based on the second intermediate image; The method of obtaining the second training image based on the second intermediate image further includes the following: The second intermediate image is determined to be the second training image; The first background in the second intermediate image is updated to the second background to obtain the second training image with the second background. The first background in the second intermediate image is updated to a third background, and at least one visual element is added to obtain the second training image.
7. The method according to claim 4, wherein an image processing model for the image processing is trained based on the prompt training text, the first training image, the second training image, and the target training image, comprising: Based on the prompt training text, the first training image, and the second training image, a prediction training image is obtained; A loss function is determined based on the predicted training image and the target training image; The model parameters are adjusted to minimize the loss function to obtain the image processing model.
8. An image processing apparatus, comprising: The receiving module is used to receive prompt text, the first image, and the second image; The prompt text is used to instruct the fusion of visual information in the first image with content objects in the second image; The description text module is used to generate a first description text for the first image, a second description text for the second image, and a target description text based on the prompt text, the first image, and the second image; The feature encoding module is used to perform text encoding on the first descriptive text, the second descriptive text, the target descriptive text, and the prompt text to obtain text features; and to perform image encoding on the first image, the second image, and the noise image to obtain image features. The target image module is used to perform noise reduction processing based on the text features and the image features to obtain the target image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method according to any one of claims 1 to 7.