Local image replacement method and device based on potential diffusion model, and storage medium
Through image preprocessing and multimodal feature recognition based on the potential diffusion model, combined with the prompt word paradigm and sampling parameter chain, the accuracy and consistency problems in local image replacement are solved, and the efficient and accurate automatic replacement of socket specifications are achieved, which improves design efficiency and image quality.
Patent Information
- Application Number
- CN202510961810.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-14
AI Technical Summary
The prior art has problems such as seams, misalignment, occlusion, inconsistency with standards, inconsistent styles and low efficiency in local image replacement, especially when switching socket specifications of energy storage products, it is difficult to achieve efficient and accurate automatic replacement.
Using a method based on the potential diffusion model, the source image and reference image are preprocessed and multimodal feature recognition are generated, and the target image is iteratively sampled and denoised by combining the preset prompt word paradigm and sampling parameter chain to generate the target image.
It realizes accurate replacement of socket areas, the generated results are consistent with the source image style, meet the target standards, improves design efficiency and image quality, and supports mass production.
Smart Images

Figure CN120471758A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a local image replacement method, device, and storage medium based on a potential diffusion model. Background Art
[0002] With the continued growth of global market demand for localization, rapidly adapting regionally customized components (such as electrical plugs and mechanical interfaces) in industrial products has become a core challenge for manufacturers. Traditional, manually-operated partial image replacement processes require complex steps, including 3D modeling and reconstruction, texture mapping and matching, physical environment rendering, and multi-view verification. These processes present significant drawbacks, including low efficiency, poor consistency, and compliance risks. For example, different countries and regions have varying requirements for socket specifications for energy storage products. During the design and production of existing energy storage products, to meet national electrical safety standards and user habits, separate socket modules are often developed for each market. This conversion process traditionally relies on manual 3D modeling, texture mapping, position matching, post-production rendering, and detailed refinement. This process is tedious and requires a high level of designer experience. For energy storage products that need to adapt to multiple national socket specifications simultaneously, manual labor is not only inefficient but also prone to unnatural seams, non-compliant socket specifications, and lighting mismatches. Ultimately, these issues lead to poor product consistency and prevent rapid, large-scale market launch.
[0003] In recent years, generative artificial intelligence (AIGC) technology has been gradually applied to product design. Stable diffusion, a representative latent diffusion model, has enabled local redrawing and style transfer. However, applying general generative models to the task of replacing local components in industrial products suffers from the following technical limitations: 1) Mask accuracy and boundary fusion fail. Existing models rely on coarse-grained mask control, failing to achieve pixel-level precise component boundary alignment. This results in edge seams, geometric misalignment, and residual occlusion between the generated area and the main product. 2) The generation process lacks strong constraints on physical specifications, resulting in randomness in key attributes (such as jack diameter, spacing, and arrangement topology). 3) Local generation is decoupled from the physical properties of the global scene, leading to material mismatches, lighting logic conflicts, and style drift. 4) The redrawing mechanism based on noise iteration requires multiple sampling cycles to obtain usable results, resulting in unstable output quality.
[0004] For example, when existing AI models such as Stable Diffusion are directly used to replace the socket specifications of energy storage products, the following shortcomings exist: (1) the replacement area cannot be accurately defined, and the socket and the product body often have seams, misalignments, and obstructions; (2) the socket specifications are easily randomized after replacement, making it difficult to be completely consistent with the target country standards; (3) the newly generated socket material, color, light source, etc. are inconsistent with the overall style of the product; (4) multiple redrawings are time-consuming and the quality is unstable. Currently, there are some automated conversion tools on the market, but they can often only handle simple specification conversions and are still unable to cope with complex and changing socket specification conversions. With the rapid development of artificial intelligence technology, AI (Artificial Intelligence) creative work methods have become an important tool for corporate marketing and design efficiency improvement. However, the function of directly using stable diffusion to locally replace image content cannot be accurately transferred to replacing socket specifications in different product countries. Problems such as the regenerated socket specifications not being in line with the version, the unnatural connection between the socket and the product, color deviation, light source deviation, and texture deviation often occur.
[0005] Therefore, there is an urgent need for an AI workflow for local image replacement that can fully utilize the generation capabilities of the potential diffusion model while overcoming the limitations of existing technologies in local replacement accuracy and consistency, and realize efficient, accurate, and batch automatic replacement of local images. Summary of the Invention
[0006] The embodiments of the present application aim to provide a method, device, and storage medium for local image replacement based on a potential diffusion model to solve the problems of seams, misalignment, occlusion, inconsistency with standards, inconsistent styles, and low replacement efficiency that are easily encountered in local image replacement in the prior art.
[0007] To solve the above technical problems, the embodiments of the present application provide the following technical solutions: According to a first aspect of the present application, a method for local image replacement based on a latent diffusion model is provided, the method comprising: Performing image preprocessing and multimodal feature recognition on a source image containing a part to be replaced and a reference image for replacement to obtain multimodal feature information and a mask image; Obtaining a text prompt word that conforms to a preset prompt word paradigm, encoding the text prompt word, and obtaining a prompt word vector; Encoding based on the source image, the reference image, the mask image, the multimodal feature information and the prompt word vector to generate an initial latent representation; Iteratively sampling and denoising the initial latent representation based on a preset sampling parameter chain and the prompt word vector to obtain a final latent representation; Decoding is performed based on the final latent representation to generate a target image.
[0008] Optionally, the performing image preprocessing on the source image containing the part to be replaced and the reference image for replacement includes: Static preprocessing is performed on the source image and the reference image to unify the source image and the reference image to a same pixel baseline, where the same pixel baseline includes a same geometric scale, a same color space, and a same pixel value range.
[0009] Optionally, the multimodal feature recognition includes image recognition, edge recognition, realistic line draft recognition, color recognition and depth recognition.
[0010] Optionally, performing multimodal feature recognition on the source image containing the part to be replaced and the reference image includes: The source image and the reference image containing the part to be replaced are sequentially identified based on each modal feature. In each modal feature identification, feature identification and error detection are performed based on the source image preprocessed in the previous link and the reference image preprocessed in the previous link to obtain the current modal feature information and the current modal error information. The image preprocessing of the source image and the reference image including the part to be replaced further includes: The source image and the reference image are subjected to reverse optimization preprocessing based on the current modal error information.
[0011] Optionally, the mask image is obtained based on the following steps: Automatically segmenting the local area to be replaced in the source image based on the multimodal feature information to separate an initial mask image; The initial mask image is aligned with the local area to be replaced in the source image and optimized to generate a mask image.
[0012] Optionally, the prompt word paradigm is "style+picture content description+detail description+parameter description".
[0013] Optionally, the sampling parameter chain includes a neural network configuration file, a seed number, an iteration number, a sampling method, a scheduler, and noise.
[0014] Optionally, the neural network configuration file includes parameters of the latent diffusion model, an inner complement encoder structure, a decoder structure and multimodal feature extraction parameters.
[0015] According to the second aspect of the present application, an electronic device is provided, comprising at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.
[0016] According to a third aspect of the present application, a computer storage medium is provided, wherein the computer storage medium stores instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor executes the method described above.
[0017] The beneficial effects of the embodiments of the present application are as follows: Different from the prior art, in the embodiments of the present application, a method for local image replacement based on a latent diffusion model is provided, which first performs image preprocessing and multimodal feature recognition on a source image containing a local part to be replaced and a reference image for replacement to obtain multimodal feature information and a mask map, then obtains a text prompt word that conforms to a preset prompt word paradigm, encodes the text prompt word, and obtains a prompt word vector; then, based on the source image, the reference image, the mask map, the multimodal feature information, and the prompt word vector, an initial latent representation is generated, and the initial latent representation is iteratively sampled and denoised based on a preset sampling parameter chain and the prompt word vector to obtain a final latent representation, and finally, decoding is performed based on the final latent representation to generate a target image. The method of the present application provides an AI workflow for automatic replacement of local images that combines precise mask editing, prompt word paradigm, and sampling parameter chain optimization. It realizes automated and highly consistent replacement of target areas on the source image based on the latent diffusion model, significantly improving design efficiency and image quality, and the generated result is consistent with the style of the source image. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0019] Figure 1 This is a flowchart of a local image replacement method based on a potential diffusion model provided in an embodiment of the present application; Figure 2 is a source image provided by an embodiment of the present application; Figure 3 is a reference image provided in the embodiment of the present application; Figure 4 is a mask diagram provided in an embodiment of the present application; Figure 5This is a schematic diagram of the process of image preprocessing and multimodal feature recognition provided by an embodiment of the present application; Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0022] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0023] Taking energy storage products and their sockets as examples, the embodiment of this application proposes an automatic replacement workflow for energy storage product socket specifications based on a potential diffusion model. By combining multiple key technologies such as multimodal feature recognition, precise mask editing, prompt word paradigm, neural network configuration file, and sampling parameter chain optimization, it can effectively overcome the technical difficulties of traditional manual replacement and existing AI partial redrawing in socket specification conversion, such as low precision, specification mismatch, style inconsistency, and long time consumption.
[0024] Please refer to Figure 1 , Figure 1 : is a flowchart of a local image replacement method based on a potential diffusion model provided in an embodiment of the present application, the method comprising: Step S101 : performing image preprocessing and multimodal feature recognition on a source image containing a part to be replaced and a reference image for replacement to obtain multimodal feature information and a mask image.
[0025] Among them, the source image is the product image, the part to be replaced is the socket, and the reference image is the socket style image.
[0026] Please refer to Figure 2 and Figure 3 , Figure 2 This is the source image provided by the embodiment of this application, which is an energy storage product image. Figure 3The reference image provided in the examples of this application is a diagram of a socket pattern. Due to the complex sources and varying quality of product images and socket pattern images, direct multimodal feature recognition may result in feature extraction errors due to differences in resolution, color gamut, noise, and distortion, which in turn may affect mask accuracy and back-end diffusion generation. To eliminate the chained errors caused by this "front-end input inconsistency," static preprocessing is first performed on the product image and socket pattern image containing the socket.
[0027] In one embodiment, static preprocessing is used to align the product image and the outlet pattern image to the same pixel baseline. Specifically, through resolution alignment, color gamut conversion, perspective correction, and tensor normalization, the product image and the outlet pattern image are aligned to the same geometric scale, color space, and pixel value range, providing a consistent receptive field and value distribution for the subsequent multimodal recognition network.
[0028] In other embodiments, static preprocessing is also used to suppress high-frequency artifacts using a denoising algorithm, as well as to perform mild sharpening and brightness contrast adaptation, thereby retaining the edges and textures of the sockets and reducing the chances of noise being misjudged as valid features, thereby improving the reliability of edge, depth, and other recognition results from the source.
[0029] In one embodiment, multimodal feature recognition includes image recognition, edge recognition, realistic line drawing recognition, color recognition, and depth recognition. Image recognition is used to identify semantic content (e.g., object categories) and global features (e.g., image style, lighting conditions, and overall composition) within an image. Edge recognition is used to detect the contour edges of target areas, such as sockets, within an image, generating clear boundary information that serves as the basis for subsequent mask generation and local replacement. Realistic line drawing recognition is used to extract detailed line information about the sockets and surrounding structures within the image, assisting the model in understanding shape and texture, and improving the realism and consistency of local replacement. Color recognition is used to analyze the color distribution and tonal characteristics of the sockets and the entire image, providing a basis for color matching when subsequently generating replacement sockets, ensuring that the new sockets match the color of the original product. Depth recognition is used to extract spatial depth information about the sockets and their relative positions within the image, guiding the model to maintain a sense of three-dimensionality and realistic spatial hierarchy when replacing the sockets.
[0030] After image preprocessing and multimodal feature recognition on product images and socket style images containing sockets, we obtain multimodal feature information. This multimodal feature information combines multiple image information such as edges, line drawings, color, and depth to comprehensively describe the shape, texture, and spatial relationship between the product and the socket, providing rich contextual semantic support for precise replacement.
[0031] Because product images and socket style images suffer from unpredictable exposure, noise, and perspective distortion, a single static preprocessing step often fails to simultaneously meet the optimal input requirements for edges, line art, color, and depth in the multimodal recognition network. To address this issue, confidence or error detection (e.g., edge breakage, depth holes, and color gamut shift) is performed during the multimodal feature recognition phase. This detected error information is fed back to the image preprocessing module for reverse optimization. This allows the module to incrementally correct only the affected areas or parameters, avoiding full image recalculation and conserving computing resources. To support this reverse optimization loop, the image preprocessing module stores metadata such as size, gamma value, and distortion matrix obtained during preprocessing. This allows the error signal to be written back during the multimodal recognition phase, allowing incremental corrections to be made only to the affected areas. Without this metadata, recognition errors cannot be quickly eliminated, and the closed loop cannot converge.
[0032] In one embodiment, performing multimodal feature recognition on a product image and a socket style image including a socket includes: sequentially identifying the product image and the socket style image including the socket based on each modal feature, and in each modal feature recognition, performing feature recognition and error detection based on the product image and the socket style image preprocessed in the previous step to obtain current modal feature information and current modal error information; performing image preprocessing on the product image and the socket style image including the socket also includes: performing reverse optimization preprocessing on the product image and the socket style image based on the current modal error information (incrementally correcting the affected areas or parameters to avoid recalculating the entire image and saving computing resources). The previous step preprocessing includes static preprocessing and reverse optimization preprocessing.
[0033] Please refer to Figure 5 , Figure 5 Schematic diagram of the process of image preprocessing and multimodal feature recognition provided by the embodiment of the present application. Figure 5 As shown, the product image and the socket style image are first statically preprocessed, and then the statically preprocessed product image and the statically preprocessed socket style image are input into the image recognition network. The image recognition network identifies the image features to obtain image feature information and image error information. The image error information is input into the image preprocessing module to perform reverse optimization preprocessing on the product image and the socket style image. The optimized product image and the socket style image are then input into the edge recognition network for edge recognition. The edge recognition process is similar to the image recognition process, and this process is continued until all modal feature recognitions are completed to obtain multimodal feature information and product images and socket style images after multiple rounds of preprocessing. It should be noted that each time a different feature is recognized, a reverse optimization preprocessing is required. Through cyclic iterative reverse optimization preprocessing, the recognition error is eliminated at the preprocessing level, which can ensure that the subsequently generated mask image meets the accuracy requirements of "neat and fitting the original socket outline."
[0034] In one embodiment, the mask image generation process specifically includes: first, automatically segmenting the socket area in the product image based on multimodal feature information to separate the initial mask image; then aligning the initial mask image with the socket position in the product image and performing optimization processing (for example, cleaning irregular edges and noise) to generate a mask image. The mask image is a binary image (or grayscale image) of the same size as the product image, used to indicate the filling position of the socket style image in the product image. Figure 4 , which is a mask diagram provided in an embodiment of the present application.
[0035] Step S102: obtaining a text prompt word that conforms to a preset prompt word paradigm, encoding the text prompt word, and obtaining a prompt word vector.
[0036] Text prompt words include positive prompt words and negative prompt words. The positive prompt words and negative prompt words are parsed by the visual macro model, and the model generates a prompt word vector based on the semantic information of the prompt words. The prompt word vector is input into the sampler as a semantic guidance condition. Specifically, a visual macro model that matches the target product style and is implemented with the latent diffusion model as the core technology is selected. The visual macro model parses the input positive prompt words to clearly define the target elements, specifications, and details required to generate the socket. At the same time, it uses negative prompt words to filter out undesirable outputs such as low quality and structural errors. The two together constrain the model generation direction to ensure that the socket replacement result is both in line with the standards and consistent with the product style.
[0037] In one embodiment, the positive prompt words conform to a preset prompt word paradigm. The prompt word paradigm is used to decompose the requirements into roles, goals, contexts, constraints, and evaluation criteria, and express them through structured statements to facilitate accurate model parsing and result verification. In one embodiment, the preset prompt word paradigm is "style + screen content description + detailed description + parameter description". For example, a text prompt word that conforms to this prompt word paradigm is: "real product photography, socket port, Chinese specifications, square 3 holes, neat edges, embossed ladder structure, black and gray, high quality". The style, screen content description, detailed description, and parameter description of the text prompt words that conform to this prompt word paradigm can all be replaced with corresponding content, but the order cannot be changed.
[0038] When processing text prompts, the model determines the weight of the prompt word through methods such as order processing, attention mechanism learning, vectorization processing, and gradient propagation. This application uses a specific prompt word paradigm to place the core prompt word first and the secondary prompt word later to emphasize the weight of each prompt word to the model. It also accurately describes the socket style through multiple dimensions such as style, image content description, detail description, and parameter description, which can accelerate the generation of socket replacements.
[0039] In step S103 , encoding is performed based on the prompt word vector, the source image (product image), the reference image (socket style image), the multimodal feature information, and the mask image to generate an initial latent representation.
[0040] In one embodiment, the method of this application uses an inpainting encoder to integrate and encode preprocessed product images, preprocessed socket style images, mask images, multimodal feature information, and hint word vectors, performing localized targeted replacement in the latent space and generating an initial latent representation. The inpainting encoder is a core component that collaborates with the latent diffusion model and is specifically designed for local editing tasks (such as image inpainting and object replacement).
[0041] Step S104 , iteratively sampling and denoising the initial latent representation based on the preset sampling parameter chain and the prompt word vector to obtain the final latent representation.
[0042] In one embodiment, the sampling parameter chain includes a neural network configuration file, a seed number, an iteration number, a sampling method, a scheduler, and noise. The neural network configuration file includes parameters for the latent diffusion model, the inner complement encoder structure, the decoder structure, the feature extractor used during training, and multimodal feature extraction parameters. By encapsulating all core parameters and modules into a callable configuration file and workflow template, the model is highly compatible with actual production scenarios. Users can quickly batch replace sockets of different specifications, facilitating production integration and significantly improving design efficiency. The seed number is used to maintain randomness, ensuring that consistent socket styles are reproducibly generated under the same input conditions during each socket replacement process, facilitating later comparison and version management. In one embodiment, the Heun sampling algorithm (a high-order numerical integration strategy for solving second-order ordinary differential equations in diffusion models) is employed to stabilize the detail generation process at a consistent number of iterations, improving the local consistency and structural integrity of the sockets. This algorithm is particularly suitable for generating socket interfaces and female jacks with rich details. The number of iterations is generally selected to be 20-30, balancing the detail completeness and computational efficiency of the replaced sockets. Too few iterations can easily lead to blurred results or incomplete replacement, while too many increases the risk of overfitting and computation time. The scheduler uses an adaptive scheduler, dynamically allocating denoising strength at each step based on the sampling method and number of steps. This flexibly controls the gradual restoration from latent space to image space, avoiding light and shadow discontinuities or material abruptness between the socket and the product itself. Controllable noise is introduced into the initial latent representation to maintain randomness and increase style diversity, while also ensuring reproducibility of results in conjunction with the number of seeds. Ultimately, this guides the model to generate new socket styles that meet the requirements of the prompt word paradigm.
[0043] By using a sampling parameter chain consisting of a neural network configuration file, number of seeds, number of iteration steps, sampling method, scheduler, and noise injection, combined with an adaptive scheduler to dynamically allocate the denoising strength of each step, the details of the socket replacement area can be controlled, the generation can be stable, and the performance can be highly reproducible.
[0044] Step S105 : Decoding is performed based on the final latent representation to generate a target image (replacing the socket product image).
[0045] Specifically, the final latent representation obtained after sampling is input into the decoder. The decoder reconstructs the image through multiple layers of deconvolution and style transfer, and outputs a replacement socket product image that is consistent with the main product style and has correct rules.
[0046] The local image replacement method based on the latent diffusion model provided in the embodiments of the present application effectively overcomes the technical difficulties of traditional manual replacement and existing AI local redrawing in socket specification conversion, such as low accuracy, specification mismatch, inconsistent style, and long time consumption, by combining multiple key technologies such as multimodal feature recognition, precise mask editing, prompt word paradigm, neural network configuration file, and sampler parameter optimization. The method has the following advantages: 1) Precise and controllable replacement area: Multimodal image recognition and segmentation are used to generate a high-precision mask image. This mask image is used as a constraint input in the internal complement encoder to achieve targeted replacement of the local area of the socket. This effectively avoids affecting other parts of the product when the socket is replaced, and the edges of the replacement area are natural and seamless. 2) Socket specifications are consistent with target standards: By designing dedicated positive and negative cue word paradigms, combined with a large visual model and an optimized sampler, we can guide the latent diffusion process to strictly generate socket specifications that meet the target country standards, without randomizing socket shape, hole position, or structure. 3) Consistent and high-quality generated results: Through reverse-optimized multimodal feature recognition and image preprocessing modules, the replaced socket is guaranteed to be seamlessly aligned with the original product image in terms of material, light source, and color, eliminating the need for tedious post-production corrections. The generated results are visually authentic and highly clear. 4) Efficient and Reproducible: Through parameterized control of neural network configuration files, fixed seed values, and an adaptive scheduler, replacement results can be generated quickly and stably across batches. Consistent socket patterns can be repeatedly generated under the same input conditions, facilitating version management and large-scale production applications. 5) Save manpower and time costs: Compared with traditional 3D modeling and post-rendering, the automatic replacement process of this invention significantly reduces the dependence on professional designers. The replacement process is fully automated, greatly shortening the cycle from product design to adaptation to multiple markets.
[0047] In summary, the local image replacement method based on the potential diffusion model provided by the present invention can ensure that the specifications of the replaced socket meet the standards of various countries, and achieve precise control of the replacement area, and the generated results are highly consistent with the product itself; this method also has the advantages of high efficiency, stability, and batch production, which can greatly improve the design flexibility and production efficiency of products in the global market.
[0048] According to an embodiment of the present application, an electronic device is provided, such as Figure 6 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 100 may include a processor 10, a communication interface 30, a memory 20, and a communication bus. The processor 10, the communication interface 30, and the memory 20 communicate with each other via the communication bus. The processor 10 may invoke logic instructions in the memory 20 to execute the aforementioned method for partial image replacement based on the latent diffusion model.
[0049] In addition, the logic instructions in the above-mentioned memory 20 can be implemented in the form of a software functional unit and can be stored in several computer-readable storage media when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the local image replacement method based on the potential diffusion model of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0050] According to an embodiment of the present application, a computer-readable storage medium is provided, the type of which is as described above, and the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the steps of the local image replacement method based on the potential diffusion model described above.
[0051] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or certain portions of the embodiments.
[0052] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions within the technical scope disclosed in the present application. Therefore, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A local image replacement method based on a latent diffusion model, characterized in that: The method comprises: Performing image preprocessing and multimodal feature recognition on a source image containing a part to be replaced and a reference image for replacement to obtain multimodal feature information and a mask image; Obtaining a text prompt word that conforms to a preset prompt word paradigm, encoding the text prompt word, and obtaining a prompt word vector; Encoding based on the source image, the reference image, the mask image, the multimodal feature information and the prompt word vector to generate an initial latent representation; Iteratively sampling and denoising the initial latent representation based on a preset sampling parameter chain and the prompt word vector to obtain a final latent representation; Decoding is performed based on the final latent representation to generate a target image.
2. The method according to claim 1, characterized in that The image preprocessing of the source image containing the part to be replaced and the reference image for replacement includes: Static preprocessing is performed on the source image and the reference image to unify the source image and the reference image to a same pixel baseline, where the same pixel baseline includes a same geometric scale, a same color space, and a same pixel value range.
3. The method according to claim 2, characterized in that The multimodal feature recognition includes image recognition, edge recognition, realistic line draft recognition, color recognition and depth recognition.
4. The method according to claim 3, characterized in that The performing multimodal feature recognition on the source image and the reference image containing the part to be replaced includes: The source image and the reference image containing the part to be replaced are sequentially identified based on each modal feature. In each modal feature identification, feature identification and error detection are performed based on the source image preprocessed in the previous link and the reference image preprocessed in the previous link to obtain the current modal feature information and the current modal error information. The image preprocessing of the source image and the reference image including the part to be replaced further includes: The source image and the reference image are subjected to reverse optimization preprocessing based on the current modal error information.
5. The method according to any one of claims 1 to 4, characterized in that The mask image is obtained based on the following steps: Automatically segmenting the local area to be replaced in the source image based on the multimodal feature information to separate an initial mask image; The initial mask image is aligned with the local area to be replaced in the source image and optimized to generate a mask image.
6. The method according to claim 1, characterized in that The prompt word paradigm is "style + picture content description + detail description + parameter description".
7. The method according to claim 1, characterized in that The sampling parameter chain includes a neural network configuration file, a seed number, an iteration number, a sampling method, a scheduler, and noise.
8. The method according to claim 7, characterized in that The neural network configuration file includes parameters of the latent diffusion model, an inner complement encoder structure, a decoder structure and multimodal feature extraction parameters.
9. An electronic device, characterized in that: The method comprises at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that The computer storage medium stores instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor is caused to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Semantic image editing method based on text condition diffusion model
CN117541684A
Multi-modal guided progressive image generation method
CN119888015A
Training and deployment of image generation models
US11995803B1
Cited By
Building effect picture generation method based on multi-condition weighted fusion and block control
CN120997368A
Prompt word sample generation method and device, storage medium and terminal
CN122219909A