Local image replacement method and device based on latent diffusion model and storage medium

By preprocessing and multimodal feature recognition of source and reference images, combined with iterative sampling and denoising techniques based on the latent diffusion model, the accuracy and consistency issues in local image replacement are resolved, enabling efficient and accurate automatic replacement of socket specifications, thus improving design efficiency and image quality.

CN120471758BActive Publication Date: 2025-11-07SHENZHEN POWEROAK NEWENER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510961810.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-07
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as seams, misalignment, occlusion, inconsistency with standards, inconsistent styles, and low replacement efficiency in local image replacement, especially when converting socket specifications for energy storage products, making it difficult to achieve efficient and accurate automated replacement.

Method used

By performing image preprocessing and multimodal feature recognition on the source and reference images, multimodal feature information and masking maps are generated. Combined with preset prompt word paradigms and sampling parameter chains, the target image is generated by iterative sampling and denoising using a latent diffusion model.

Benefits of technology

It enables precise and controllable replacement of socket areas, generates results that are consistent with product style and meet target standards, improves design efficiency and image quality, and reduces manpower and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471758B_ABST
    Figure CN120471758B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, in particular to a local image replacement method and device based on a latent diffusion model and a storage medium. The method comprises the following steps: performing image preprocessing and multi-modal feature identification on a source image and a reference image to obtain multi-modal feature information and a mask image; a text prompt word is obtained, the text prompt word is encoded, and a prompt word vector is obtained; the source image, the reference image, the mask image, the multi-modal feature information and the prompt word vector are encoded to generate an initial latent representation; the initial latent representation is iteratively sampled and denoised based on a sampling parameter chain and the prompt word vector to obtain a final latent representation; and the final latent representation is decoded to generate a target image. According to the method, the automatic and high-consistency replacement of a target region on a source image is realized based on a latent diffusion model, the design efficiency and the image quality are significantly improved, and the generated result is consistent with the style of the source image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a local image replacement method and device based on a latent diffusion model and a storage medium. BACKGROUND

[0002] With the continuous growth of global market localization demand, the rapid adaptation of regional customized components (such as electrical plug, mechanical interface, etc.) in industrial products has become a core challenge for manufacturing enterprises. The traditional local image replacement process relying on manual operation needs to go through complex procedures such as three-dimensional modeling reconstruction, texture mapping matching, physical environment rendering and multi-view verification, which has significant defects such as low efficiency, poor consistency and compliance risk. For example: different countries and regions have different requirements for the specifications of the sockets on energy storage products. In the design and production process of existing energy storage products, in order to meet the electrical safety standards and user habits of different countries, it is often necessary to develop socket modules separately for different markets. This conversion process traditionally relies mainly on manual 3D modeling, texture mapping, position matching, post-rendering and detail refining, which is tedious and requires high experience of designers. For energy storage products that need to adapt to multiple national socket specifications, manual operation not only has low efficiency, but also is prone to problems such as unnatural joints, non-compliance of socket specifications, and mismatched light and shadow, ultimately resulting in poor product consistency and inability to be mass-produced and quickly put into the market.

[0003] In recent years, Artificial Intelligence Generated Content (AIGC) technology has been gradually applied in product design, and Stable Diffusion, as a representative latent diffusion model, has been able to achieve local redrawing and style transfer. However, when a general generation model is used to perform the task of replacing local components of industrial products, there are the following technical limitations: 1) Mask precision and boundary fusion failure: existing models rely on coarse-grained mask control and cannot achieve pixel-level accurate component boundary alignment, resulting in edge seams, geometric misalignment, and occlusion residues between the generated area and the product body; 2) The generation process lacks strong constraint mechanisms for physical specifications, resulting in randomness in key attributes such as socket hole diameter, spacing, and arrangement topology; 3) Decoupling of local generation and global scene physical properties leads to material mismatch, lighting logic conflicts, and style drift; 4) The redrawing mechanism based on noise iteration requires multiple samplings to obtain a usable result for a single generation, and the output quality is unstable.

[0004] For example, directly using existing AI models such as Stable Diffusion for replacing the socket specifications of energy storage products has the following shortcomings: (1) the replacement area cannot be accurately defined, and the socket and the product body often have joints, misalignment, and occlusion; (2) the socket specifications after replacement are easily randomized and difficult to fully comply with the target national standards; (3) the newly generated socket material, color, light source, and other factors do not match the overall style of the product; (4) multiple redraws are time-consuming and the quality is unstable. Currently, there are some automatic conversion tools on the market, but they can only handle simple specification conversion and are still inadequate for complex and variable socket specification conversion. With the rapid development of artificial intelligence technology, AI (Artificial Intelligence) creation methods have become an important tool for enterprise marketing and design efficiency improvement. However, directly using the local replacement image content function of stable diffusion cannot accurately migrate to replace different product national socket specifications, and there are always problems such as misalignment of the newly generated socket specifications, unnatural connection between the socket and the product, color deviation, light source deviation, and texture deviation.

[0005] Therefore, there is an urgent need for an AI workflow for local image replacement that can fully utilize the generation capabilities of latent diffusion models while overcoming the limitations of existing technologies in terms of local replacement accuracy and consistency, enabling efficient, accurate, and batch automatic replacement of local images. SUMMARY

[0006] Embodiments of the present application aim to provide a local image replacement method and device based on a latent diffusion model and a storage medium to solve the problems of joints, misalignment, occlusion, inconsistency with standards, style inconsistency, and low replacement efficiency in local image replacement in the prior art.

[0007] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0008] According to a first aspect of the present application, a local image replacement method based on a latent diffusion model is provided, the method comprising:

[0009] performing image preprocessing and multi-modal feature recognition on a source image containing a local part to be replaced and a reference image for replacement to obtain multi-modal feature information and a mask image;

[0010] obtaining a text prompt word conforming to a preset prompt word paradigm, encoding the text prompt word to obtain a prompt word vector;

[0011] encoding based on the source image, the reference image, the mask image, the multi-modal feature information, and the prompt word vector to generate an initial latent representation;

[0012] iteratively sample and denoise the initial latent representation based on a preset sampling parameter chain and the prompt vector to obtain a final latent representation;

[0013] decode based on the final latent representation to generate a target image.

[0014] Optionally, the image preprocessing of the source image containing the local part to be replaced and the reference image for replacement comprises:

[0015] static preprocessing of the source image and the reference image, unifying the source image and the reference image to the same pixel baseline, the same pixel baseline comprising the same geometric scale, the same color space and the same pixel value range.

[0016] Optionally, the multi-modal feature recognition comprises image recognition, edge recognition, line drawing recognition, color recognition and depth recognition.

[0017] Optionally, the multi-modal feature recognition of the source image containing the local part to be replaced and the reference image comprises:

[0018] sequentially recognizing the source image containing the local part to be replaced and the reference image based on each modal feature, and in each modal feature recognition, recognizing the feature and detecting the error based on the source image after preprocessing of the previous step and the reference image after preprocessing of the previous step to obtain the current modal feature information and the current modal error information;

[0019] The image preprocessing of the source image containing the local part to be replaced and the reference image further comprises:

[0020] based on the current modal error information, reversely optimizing the preprocessing of the source image and the reference image.

[0021] Optionally, the mask image is obtained based on the following steps:

[0022] based on the multi-modal feature information, automatically segmenting the local part to be replaced in the source image to separate an initial mask image;

[0023] aligning and optimizing the initial mask image with the local part to be replaced in the source image to generate a mask image.

[0024] Optionally, the prompt word paradigm is "style + picture content description + detail description + parameter description".

[0025] Optionally, the sampling parameter chain comprises a neural network configuration file, a seed number, an iteration step number, a sampling method, a scheduler and noise.

[0026] Optionally, the neural network configuration file comprises parameters of the latent diffusion model, an inner complement encoder structure, a decoder structure, and multi-modal feature extraction parameters.

[0027] According to a second aspect of the present application, an electronic device is provided, comprising at least one processor and a memory connected in communication with the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0028] According to a third aspect of the present application, a computer storage medium is provided, the computer storage medium storing instructions or programs, when the instructions or programs are executed by at least one processor, enabling the at least one processor to perform the method described above.

[0029] The beneficial effects of the embodiments of the present application are: different from the prior art, in the embodiments of the present application, a local image replacement method based on a latent diffusion model is provided, image preprocessing and multi-modal feature recognition are performed on a source image containing a local part to be replaced and a reference image used for replacement to obtain multi-modal feature information and a mask image, a text prompt word conforming to a preset prompt word paradigm is obtained, the text prompt word is encoded to obtain a prompt word vector; then, the source image, the reference image, the mask image, the multi-modal feature information and the prompt word vector are encoded to generate an initial latent representation, and the initial latent representation is iteratively sampled and denoised based on a pre-set sampling parameter chain and the prompt word vector to obtain a final latent representation, and finally the final latent representation is decoded to generate a target image. The method of the present application provides a local image automatic replacement AI workflow combining precise mask editing, prompt word paradigm and sampling parameter chain optimization, realizes automatic and high consistency replacement of a target region on a source image based on a latent diffusion model, significantly improves design efficiency and image quality, and the generated result is consistent with the style of the source image. BRIEF DESCRIPTION OF DRAWINGS

[0030] One or more embodiments are illustrated by way of example in the drawings that are not intended to be limiting of the embodiments so as to illustrate exemplary principles of the embodiments. The same reference numerals in different drawings represent the same element or components in the drawings in which an illustrative embodiment is provided. The drawings are not necessarily to scale, the emphasis instead being placed upon illustrating exemplary principles of the embodiments.

[0031] Figure 1 is a flowchart of a local image replacement method based on a latent diffusion model provided by the embodiments of the present application;

[0032] Figure 2 is a source image provided by the embodiments of the present application;

[0033] Figure 3is a reference image provided by an embodiment of the present application;

[0034] Figure 4 is a mask provided by an embodiment of the present application;

[0035] Figure 5 is a process diagram of image preprocessing and multi-modal feature recognition provided by an embodiment of the present application;

[0036] Figure 6 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0038] In addition, the technical features involved in each of the embodiments of the present application described below can be combined with each other as long as there is no conflict.

[0039] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0040] The embodiments of the present application take energy storage products and their sockets as examples, and propose an energy storage product socket specification automatic replacement workflow based on a latent diffusion model. By combining multi-modal feature recognition, accurate mask editing, prompt word paradigm, neural network configuration file, and sampling parameter chain optimization, and other key technologies, the technical difficulties such as low accuracy, specification mismatch, inconsistent style, and long time consumption in the socket specification conversion of traditional manual replacement and existing AI partial redrawing can be effectively overcome.

[0041] Please refer to Figure 1 , Figure 1 is a flowchart of a partial image replacement method based on a latent diffusion model provided by an embodiment of the present application, and the method comprises the following steps:

[0042] In step S101, the source image containing the local part to be replaced and the reference image used for replacement are subjected to image preprocessing and multi-modal feature recognition to obtain multi-modal feature information and a mask.

[0043] The source image is a product image, the local to be replaced is a socket, and the reference image is a socket style image.

[0044] Please refer to Figure 2 and Figure 3 , Figure 2 The source image is provided by the embodiment of the application, which is a product image of energy storage, Figure 3 The reference image is provided by the embodiment of the application, which is a socket style image. Because the product image and the socket style image have complex sources and uneven quality, if multi-modal feature recognition is directly performed, feature extraction errors may be caused due to resolution, color gamut, noise and distortion differences, thereby affecting the mask precision and the subsequent diffusion generation. To eliminate the chain errors caused by the inconsistency of the front-end input, the product image containing the socket and the socket style image are first subjected to static preprocessing.

[0045] In an embodiment, the static preprocessing is used to unify the product image and the socket style image to the same pixel baseline. Specifically, by means of resolution alignment, color gamut conversion, perspective correction and tensor normalization, the product image and the socket style image are unified to the same geometric scale, the same color space and the same pixel value range, so as to provide consistent receptive field and numerical distribution for the subsequent multi-modal recognition network.

[0046] In other embodiments, the static preprocessing is also used to suppress high-frequency artifacts using a denoising algorithm, and to perform slight sharpening and brightness and contrast self-adaption, so as to not only retain the socket edges and textures, but also reduce the noise points being mistaken for effective features, thereby improving the reliability of the edge, depth and other recognition results from the source.

[0047] In an embodiment, the multi-modal feature recognition includes image recognition, edge recognition, realistic sketch recognition, color recognition and depth recognition. The image recognition is used to recognize semantic content (such as contained object categories) and global features (such as image style, lighting conditions, overall composition) in the image. The edge recognition is used to detect the contour edges of the target region such as the socket in the image, to generate clear boundary information, and to serve as the basis for subsequent mask generation and local replacement. The realistic sketch recognition is used to extract the detailed line information of the socket and the surrounding structure in the image, to assist the model in understanding the shape and texture, and to improve the realism and consistency of the local replacement. The color recognition is used to analyze the color distribution and hue features of the socket and the whole image, to provide color matching basis for subsequent generation of the replacement socket, and to ensure that the new socket is consistent with the original product color. The depth recognition is used to extract the spatial depth information of the socket and its relative position in the image, to guide the model to maintain the stereoscopic sense and the real spatial hierarchy when replacing the socket.

[0048] After image preprocessing and multi-modal feature recognition are performed on the product image containing the socket and the socket style image, multi-modal feature information is obtained. The multi-modal feature information combines edge, line drawing, color, depth and other image information, and can comprehensively describe the shape, texture and spatial relationship of the product and the socket, thereby providing rich context semantic support for accurate replacement.

[0049] Since the exposure, noise, perspective distortion and other conditions of the product image and the socket style image cannot be predicted, one-time static preprocessing cannot simultaneously meet the optimal input conditions of the multi-modal recognition network of edge, line drawing, color, depth and the like. To solve this problem, confidence or error detection (such as edge breakage, depth cavity, color drift and the like) is further performed in the multi-modal feature recognition stage, and the detected error information is fed back to the image preprocessing module for reverse optimization, so that the image preprocessing module can only make incremental correction to the affected area or parameter, thereby avoiding full-image recalculation and saving computing resources. In order to support this reverse optimization closed loop, the image preprocessing module saves the metadata such as size, Gamma value, distortion matrix obtained in the preprocessing link, so that after the error signal is written back in the multi-modal recognition stage, only incremental correction can be made to the affected area. If these metadata are not saved, the recognition error cannot be quickly eliminated, and the closed loop cannot converge.

[0050] In an embodiment, the multi-modal feature recognition on the product image containing the socket and the socket style image comprises: sequentially recognizing the product image containing the socket and the socket style image based on each modal feature, and in each modal feature recognition, performing feature recognition and error detection based on the product image after preprocessing of the previous link and the socket style image after preprocessing of the previous link, to obtain current modal feature information and current modal error information; the image preprocessing on the product image containing the socket and the socket style image further comprises: reverse optimization preprocessing (incremental correction to the affected area or parameter, avoiding full-image recalculation and saving computing resources) of the product image and the socket style image based on the current modal error information. Wherein, the preprocessing of the previous link comprises static preprocessing and reverse optimization preprocessing.

[0051] Please refer to Figure 5 , Figure 5 is a process schematic diagram of image preprocessing and multi-modal feature recognition provided by the embodiment of the present application. As Figure 5As shown, the product image and the socket style image are first preprocessed statically, and then the statically preprocessed product image and the statically preprocessed socket style image are input into the image recognition network, the image recognition network recognizes the image features, obtains image feature information and image error information, the image error information is input into the image preprocessing module to perform reverse optimization preprocessing on the product image and the socket style image, and then the optimized product image and the socket style image are input into the edge recognition network for edge recognition. The edge recognition process is similar to the image recognition process, and until all modal feature recognitions are completed, the multi-modal feature information and the multi-round preprocessed product image and socket style image are obtained. It should be noted that reverse optimization preprocessing needs to be performed once for each different feature recognition. Through the reverse optimization preprocessing of the loop iteration, the recognition error is eliminated at the preprocessing level, which can ensure that the generated mask image meets the accuracy requirements of "neat and consistent with the original socket contour".

[0052] In an embodiment, the generation process of the mask image specifically includes: first, based on the multi-modal feature information, the socket region in the product image is automatically segmented to separate out an initial mask image; and then the initial mask image is aligned with the socket position in the product image for optimization processing (for example, cleaning irregular edges and noise), to generate the mask image. The mask image is a binary image (or grayscale image) with the same size as the product image, which is used to indicate the filling position of the socket style image in the product image. As shown in the figure, Figure 4 As shown, it is a mask image provided by an embodiment of the present application.

[0053] In step S102, a text prompt word conforming to a preset prompt word paradigm is obtained, and the text prompt word is encoded to obtain a prompt word vector.

[0054] The text prompt word includes a positive prompt word and a negative prompt word. The positive prompt word and the negative prompt word are parsed by a visual large model, and a prompt word vector is generated by the model according to the semantic information of the prompt word. The prompt word vector is input into the sampler as a semantic guidance condition. Specifically, a visual large model with a latent diffusion model as the core technology is selected to match the target product style. The visual large model parses the input positive prompt word to clearly generate the target elements, specifications and details required for the socket, and at the same time filters out low-quality, structure error and other undesirable outputs by using the negative prompt word. Both of them jointly constrain the model generation direction to ensure that the socket replacement result not only meets the standard but also is consistent with the product style.

[0055] In an embodiment, the positive prompt word conforms to a preset prompt word paradigm. The prompt word paradigm is used to decompose the demand into roles, goals, contexts, constraints, and evaluation criteria, expressed by structured sentences, facilitating accurate parsing by the model and result verification. In an embodiment, the preset prompt word paradigm is "style + picture content description + detail description + parameter description". For example, a text prompt word conforming to the prompt word paradigm is: "real product photography, socket port, Chinese specification, square 3-hole, neat edge, convex character ladder structure, black gray, high quality". The style, picture content description, detail description, and parameter description of the text prompt word conforming to the prompt word paradigm can all be replaced with corresponding content, but the order cannot be changed.

[0056] The model judges the weight of the prompt word when processing the text prompt word through order processing, attention mechanism learning, vectorization processing, gradient propagation, etc. The present application emphasizes the weight of each prompt word by placing the core prompt word in front and the secondary prompt word in the back through the use of a specific prompt word paradigm, and accurately describes the socket style through multiple dimensions such as style, picture content description, detail description, and parameter description, which can speed up the generation speed of socket replacement.

[0057] Step S103, based on the prompt word vector, the source image (product image), the reference image (socket style image), the multi-modal feature information, and the mask image, encoding is performed to generate an initial latent representation.

[0058] In an embodiment, the method of the present application uses an internal supplement encoder to integrate and encode the preprocessed product image, the preprocessed socket style image, the mask image, the multi-modal feature information, and the prompt word vector, realizing local directional replacement in the latent space and generating an initial latent representation. The internal supplement encoder (Inpainting Encoder) is a core component cooperating with the latent diffusion model, and is specially used for local editing tasks (such as image repair and object replacement).

[0059] Step S104, based on the pre-set sampling parameter chain and the prompt word vector, the initial latent representation is iteratively sampled and denoised to obtain a final latent representation.

[0060] In an embodiment, the sampling parameter chain includes a neural network configuration file, a seed number, an iteration step number, a sampling method, a scheduler, and noise. Among them, the neural network configuration file includes parameters of a latent diffusion model, an inner padding encoder structure, a decoder structure, a feature extractor used during training, and multi-modal feature extraction parameters, etc. By encapsulating all core parameters and modules as a callable configuration file and a workflow template, the model is highly matched with the actual production scene, users can quickly replace different specifications of sockets in batches, which is convenient for production docking and significantly improves design efficiency. The seed number is used to fix randomness, ensuring that the socket replacement process can generate consistent socket patterns under the same input conditions, which is convenient for later comparison and version management. In an embodiment, the Heun sampling algorithm (Heun sampling algorithm is a second-order ordinary differential equation solving method in diffusion model, which belongs to high-order numerical integration strategy, and achieves a significant balance between stability and accuracy) is used to stabilize the detail generation process under the same iteration step number, improve the local consistency and structural integrity of the socket, and is particularly suitable for the generation of socket interfaces and female socket parts with rich details. The iteration step number is generally selected to be 20-30 steps, which can balance the detail integrity of the replaced socket and the calculation efficiency. Too few iteration steps may lead to fuzzy or incomplete replacement results, and too many iteration steps may increase the risk of overfitting and operation time. The scheduler uses an adaptive scheduler to dynamically allocate the noise removal strength of each step according to the sampling method and step number, flexibly controls the gradual restoration process from the latent space to the image space, and avoids the occurrence of light and shadow faults or material abruptness between the socket and the product body. Controllable noise is introduced into the initial latent representation, which on the one hand maintains randomness and increases pattern diversity, and on the other hand cooperates with the seed number to ensure the reproducibility of the results, and finally guides the model to generate new socket patterns that meet the requirements of the prompt word paradigm.

[0061] Through the sampling parameter chain composed of the neural network configuration file, the seed number, the iteration step number, the sampling method, the scheduler, and the noise injection, and combined with the adaptive scheduler dynamically allocating the noise removal strength of each step, the details of the socket replacement area can be controlled, the generation is stable, and the reproducibility is high.

[0062] Step S105, decoding based on the final latent representation to generate a target image (replacement socket product image).

[0063] Specifically, the final latent representation obtained after completing sampling is input to the decoder. The decoder reconstructs by multi-layer deconvolution and style transfer, and outputs a replacement socket product image that is consistent with the product body style and correct in rules.

[0064] The local image replacement method based on the latent diffusion model provided by the embodiments of the present application effectively overcomes the technical problems of low precision, specification mismatch, inconsistent style, and long time consumption in the conversion of socket specifications by combining multiple key technologies such as multi-modal feature recognition, accurate mask editing, prompt word paradigm, neural network configuration file, and sampler parameter optimization, and has the following advantages:

[0065] 1) Accurate control of replacement area: high-precision mask images are generated through multi-modal image recognition and segmentation, and the mask images are used as constraint inputs in the inner padding encoder to realize directional replacement of the local area of the socket, effectively avoiding the influence of the socket replacement on other parts of the product main body, and the edges of the replacement area are natural and seamless.

[0066] 2) Socket specification consistent with target standard: through the design of special positive and negative prompt word paradigms, combined with a visual large model and an optimized sampler, the latent diffusion process can be guided to strictly generate a socket specification that meets the target national standard, and the problems of randomization of socket shape, hole position, or structure do not occur.

[0067] 3) Consistent style and high quality of generated results: through the reverse optimization of multi-modal feature recognition and image preprocessing modules, the replaced socket and the original product image are naturally connected in terms of material, light source, color, etc., without the need for tedious post-correction, and the generated results have real visual effects and high clarity.

[0068] 4) High efficiency and reproducibility: through the parameterization control of the neural network configuration file, fixed seed value, and adaptive scheduler, the replacement results can be quickly and stably generated in batches in different batches, and consistent socket styles can be repeatedly generated under the same input conditions, facilitating version management and large-scale production application.

[0069] 5) Saving labor and time cost: compared with traditional 3D modeling and post-rendering, the automatic replacement process of the present application significantly reduces the dependence on professional designers, and the replacement process is fully automated, greatly shortening the cycle from product design to adaptation to multiple national markets.

[0070] In summary, the local image replacement method based on the latent diffusion model provided by the present application can ensure that the replaced socket specification meets the standards of various countries, and realize accurate control of the replacement area, high consistency of the generated results with the product body, and the method also has the advantages of high efficiency, stability, and batch production, which can greatly improve the design flexibility and production efficiency of products in the global market.

[0071] According to the embodiments of the present application, an electronic device is provided, such as Figure 6A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. The electronic device 100 can include a processor 10, a communication interface 30, a memory 20, and a communication bus. The processor 10, the communication interface 30, and the memory 20 can communicate with each other through the communication bus. The processor 10 can invoke a logical instruction in the memory 20 to execute the local image replacement method based on the latent diffusion model described above.

[0072] In addition, the logical instruction in the memory 20 described above can be implemented in the form of a software functional unit and sold or used as an independent product. In this case, the logical instruction can be stored in several computer-readable storage media. Based on this understanding, the technical solution of the present application or the part of the technical solution that essentially contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the local image replacement method based on the latent diffusion model described above. The storage medium described above includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0073] According to an embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium is of the type described above and stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the local image replacement method based on the latent diffusion model described above.

[0074] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the technical solution described above or the part of the technical solution that essentially contributes to the related art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium such as a ROM / RAM, a magnetic disk, an optical disk, etc. and includes instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment or some part of the embodiment.

[0075] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application. Therefore, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for local image replacement based on a latent diffusion model, characterized in that, The method comprises: image preprocessing and multi-modal feature recognition are performed on a source image containing a local part to be replaced and a reference image for replacement to obtain multi-modal feature information and a mask image; a text prompt word conforming to a preset prompt word paradigm is obtained, and the text prompt word is encoded to obtain a prompt word vector; an initial latent representation is generated based on the source image, the reference image, the mask image, the multi-modal feature information, and the prompt word vector; the initial latent representation is iteratively sampled and denoised based on a preset sampling parameter chain and the prompt word vector to obtain a final latent representation; a target image is generated based on the final latent representation; the multi-modal feature recognition comprises image recognition, edge recognition, line drawing recognition, color recognition, and depth recognition, and the mask image is obtained based on the following steps: an initial mask image is separated out by automatically segmenting a local area to be replaced in the source image based on the multi-modal feature information; the initial mask image is aligned with the local area to be replaced in the source image for optimization processing to generate a mask image.

2. The method of claim 1, wherein, The image preprocessing performed on the source image containing a local part to be replaced and the reference image for replacement comprises: static preprocessing is performed on the source image and the reference image, and the source image and the reference image are unified to a same pixel baseline, which comprises a same geometric scale, a same color space, and a same pixel value range.

3. The method of claim 2, wherein, The multi-modal feature recognition performed on the source image containing a local part to be replaced and the reference image comprises: each modal feature is recognized in sequence based on the source image containing a local part to be replaced and the reference image, and in each modal feature recognition, feature recognition and error detection are performed based on the source image after preprocessing of a previous link and the reference image after preprocessing of the previous link to obtain current modal feature information and current modal error information; The image preprocessing performed on the source image containing a local part to be replaced and the reference image further comprises: The source image and the reference image are subjected to reverse optimization preprocessing based on the current modal error information.

4. The method of claim 1, wherein, The prompt word paradigm is "style + picture content description + detail description + parameter description".

5. The method of claim 1, wherein, The sampling parameter chain comprises a neural network configuration file, a seed number, an iteration step number, a sampling method, a scheduler, and noise.

6. The method of claim 5, wherein, The neural network configuration file comprises parameters of the latent diffusion model, an inner complement encoder structure, a decoder structure, and multi-modal feature extraction parameters.

7. An electronic device, comprising: The computer storage medium stores instructions or programs that, when executed by at least one processor, enable the at least one processor to perform the method of any one of claims 1 to 6.

8. A computer storage medium, characterized in that The computer storage medium stores instructions or programs that, when executed by at least one processor, enable the at least one processor to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Semantic image editing method based on text condition diffusion model

    CN117541684A

  • Multi-modal guided progressive image generation method

    CN119888015A