Method for obtaining a synthetic image, client, device, apparatus and medium

By receiving initial photographic images input by the user and using a latent diffusion model for image transformation processing, the problem of low efficiency in synthesized image generation is solved, the operation steps are simplified, image generation efficiency and quality are improved, and the user experience is optimized.

CN118014864BActive Publication Date: 2026-02-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410303369.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2026-02-10
Estimated Expiration
2044-03-15

AI Technical Summary

Technical Problem

In existing technologies, the efficiency of synthesized image generation is low, users need to have certain technical skills and it is time-consuming, and the use of image generation tools is complicated.

Method used

By receiving the initial photographic image input by the user, obtaining the image conversion requirements, and using a latent diffusion model to perform image conversion processing, a synthetic image that meets the requirements is generated, simplifying the user's operation steps and reducing technical requirements.

Benefits of technology

It improves the efficiency of acquiring composite images, optimizes the user experience, reduces the technical requirements for users, simplifies the time required to acquire visual content materials, and enhances image quality and user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118014864B_ABST
    Figure CN118014864B_ABST
Patent Text Reader

Abstract

The disclosure provides a synthetic image acquisition method, a client, an apparatus, a device and a medium, relating to the fields of artificial intelligence such as natural language processing and computer vision, comprising: receiving an initial photographic image input by a user end, and acquiring image conversion requirements of the initial photographic image; acquiring conversion intervention information of the initial photographic image according to the image conversion requirements; performing image conversion processing on the initial photographic image through a latent diffusion model according to the conversion intervention information, to obtain a first candidate synthetic image after processing; and in response to the first candidate synthetic image meeting the image conversion requirements, obtaining a target synthetic image of the initial photographic image according to the first candidate image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and in particular to artificial intelligence fields such as natural language processing and computer vision. Background Technology

[0002] With the development of technology, intelligent image generation is playing an increasingly important role in people's lives and work. In related technologies, people can use image editing software to stitch together and integrate photographic subjects and visual content materials to generate the desired image.

[0003] Visual content materials can be obtained from open-source datasets or constructed using mapping tools in related technologies. This process is time-consuming, and the use of image editing tools to stitch together visual content materials and photographic subjects requires a certain level of technical expertise, resulting in poor image generation efficiency. Summary of the Invention

[0004] This disclosure presents a method, client, apparatus, device, and medium for acquiring synthetic images.

[0005] According to a first aspect of this disclosure, a method for obtaining a synthetic image is proposed. The method includes: receiving an initial photographic image input from a user terminal and obtaining an image conversion requirement for the initial photographic image; obtaining conversion intervention information for the initial photographic image based on the image conversion requirement; performing image conversion processing on the initial photographic image using a latent diffusion model based on the conversion intervention information to obtain a processed first candidate synthetic image; and obtaining a target synthetic image of the initial photographic image based on the first candidate image in response to the first candidate synthetic image satisfying the image conversion requirement.

[0006] According to a second aspect of this disclosure, a client for synthesizing images is proposed. The client includes: an image upload module, a reference style conversion image library module, a prompt word module, and an image synthesis icon. The image upload module is used to provide an upload port for an initial photographic image to a user terminal and transmit the initial photographic image uploaded by the user terminal to a server. The reference style conversion image library module is used to display various reference style conversion images in the reference style conversion image library to the user terminal, wherein, in response to the user terminal selecting a style conversion image of the initial photographic image from the reference style conversion image library, the style conversion image is sent to the server. The prompt word module is used to provide the user terminal with... A list of prompt word options is provided, wherein, in response to the user selecting a style conversion prompt word or a local redraw prompt word for the initial captured image from the prompt word option list, the style conversion prompt word or the local redraw prompt word is sent to the server; an image compositing icon is provided to provide a compositing instruction input port for the user, wherein, in response to the user clicking the image compositing icon, it is confirmed that the user has input a confirmation compositing instruction, and the confirmation compositing instruction is sent to the server, and the server starts to convert the initial captured image based on the received confirmation compositing instruction to obtain candidate compositing images, so as to obtain the target compositing image based on the candidate compositing images.

[0007] According to a third aspect of this disclosure, a synthetic image acquisition apparatus is proposed, comprising: a first acquisition module for receiving an initial photographic image input from a user terminal and acquiring an image conversion requirement of the initial photographic image; a second acquisition module for acquiring conversion intervention information of the initial photographic image according to the image conversion requirement; a first processing module for performing image conversion processing on the initial photographic image using a latent diffusion model according to the conversion intervention information to obtain a processed first candidate synthetic image; and a second processing module for obtaining a target synthetic image of the initial photographic image based on the first candidate image in response to the first candidate synthetic image satisfying the image conversion requirement.

[0008] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the synthetic image acquisition method proposed in the first aspect above.

[0009] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method for acquiring a synthetic image as described in the first aspect above.

[0010] According to a sixth aspect of this disclosure, a computer program product is proposed, comprising a computer program that, when executed by a processor, implements the method for acquiring a synthetic image as described in the first aspect.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is a schematic flowchart of a method for acquiring a synthesized image according to an embodiment of the present disclosure;

[0014] Figure 2 This is a schematic flowchart illustrating a method for acquiring a synthesized image according to another embodiment of this disclosure;

[0015] Figure 3 This is a schematic flowchart illustrating a method for acquiring a synthesized image according to another embodiment of this disclosure;

[0016] Figure 4 This is a schematic flowchart illustrating a method for acquiring a synthesized image according to another embodiment of this disclosure;

[0017] Figure 5 This is a schematic flowchart illustrating a method for acquiring a synthesized image according to another embodiment of this disclosure;

[0018] Figure 6 This is a schematic diagram of a synthetic image client according to an embodiment of the present disclosure;

[0019] Figure 7 This is a schematic diagram of a synthetic image client according to another embodiment of the present disclosure;

[0020] Figure 8 This is a schematic diagram of a synthetic image client according to another embodiment of the present disclosure;

[0021] Figure 9 This is a schematic diagram of a synthetic image client according to another embodiment of the present disclosure;

[0022] Figure 10 This is a schematic diagram of the structure of a device for acquiring a synthesized image according to an embodiment of the present disclosure;

[0023] Figure 11 This is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0025] Data processing is a fundamental aspect of systems engineering and automatic control. Data is a form of expression of facts, concepts, or instructions, which can be processed manually or by automated devices. After being interpreted and given meaning, data becomes information. Data processing involves the acquisition, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and difficult-to-understand data.

[0026] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Its main applications include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese OCR.

[0027] Computer vision (CV) refers to machine vision that uses cameras and computers to replace human eyes for target recognition, tracking, and measurement, and further performs image processing. It uses various imaging systems instead of visual organs as input sensors, with computers replacing the brain to complete the processing and interpretation. This allows computer-processed images to be more suitable for human observation or transmission to instruments for detection. Computer vision research studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting 'information' from images or multidimensional data. Because perception can be seen as extracting information from sensory signals, computer vision can also be seen as the science of studying how to enable artificial systems to 'perceive' from images or multidimensional data.

[0028] Artificial Intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. It attempts to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Since its inception, AI has matured in both theory and technology, and its applications have expanded continuously. It is conceivable that future AI-driven technological products will serve as "containers" of human wisdom. AI can simulate the information processes of human consciousness and thought.

[0029] Figure 1 This is a schematic flowchart illustrating a method for acquiring a synthesized image according to an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes:

[0030] S101 receives the initial photographed image input from the user terminal and obtains the image conversion requirements of the initial photographed image.

[0031] In this embodiment of the disclosure, after a user takes a photo of a person or object, the resulting photographic image may not meet the user's requirements. In this scenario, the user can perform conversion processing on the photographic image based on the image synthesis server to obtain a synthesized image that meets their requirements.

[0032] Among them, the photographic images that need style conversion processing can be marked as the initial photographic images.

[0033] In this embodiment of the disclosure, the initial photographic image can undergo global style conversion or local image redrawing. The user's corresponding requirements for converting the initial photographic image can be marked as the image conversion requirements of the initial photographic image.

[0034] S102, based on the image conversion requirements, obtain the conversion intervention information of the initial photographic image.

[0035] In this embodiment of the disclosure, when converting the initial photographic image, the image conversion process can be achieved by changing the shape, color, and outline of the elements included in the initial photographic image, deleting the elements included in the initial photographic image, and adding new elements to the initial photographic image.

[0036] In this scenario, information on how each element in the initial photographic image is processed during image conversion can be obtained, and this information can be marked as the conversion intervention information of the initial photographic image.

[0037] The transformation intervention information can determine the transformation processing method of each element in the initial photographic image. This can be understood as follows: for any element in the initial photographic image, the transformation intervention information can determine the color processing method and contour processing method of that element during image transformation processing.

[0038] In this embodiment of the disclosure, the user terminal may have various types of image conversion requirements for the initial photographic image. Among these, the conversion processing methods for each element of the initial photographic image differ for each type of image conversion requirement.

[0039] In this scenario, based on the user's image conversion requirements for the initial photographic image, the processing method information for each element in the initial photographic image during image conversion can be determined, thereby obtaining the conversion intervention information for the initial photographic image under the image conversion requirements.

[0040] S103, Based on the transformation intervention information, the initial photographic image is processed by a potential diffusion model to obtain the first candidate synthetic image after processing.

[0041] In this embodiment of the disclosure, the initial photographic image can be processed by a potential diffusion model in the related technology based on the conversion intervention information, and the processed image can be marked as the first candidate synthetic image.

[0042] Optionally, the image synthesis capability of the latent diffusion model can be invoked, and through this image synthesis capability, the elements in the initial photographic image can be transformed according to the transformation intervention information to obtain the first candidate synthesized image after processing.

[0043] Optionally, an image synthesis model based on a potential diffusion model can be constructed in related technologies, and the initial photographic image and transformation intervention information can be input into the image synthesis model. The image synthesis model can then perform transformation processing on each element in the initial photographic image based on the transformation intervention information, thereby outputting the processed first candidate synthesized image.

[0044] S104, in response to the first candidate synthetic image meeting the image conversion requirements, obtain the target synthetic image of the initial photographic image based on the first candidate image.

[0045] In this embodiment of the disclosure, the first candidate synthetic image obtained after performing image conversion processing on the initial photographic image based on the conversion intervention information may not meet the image conversion requirements of the user. In this scenario, it is necessary to identify whether the first candidate synthetic image meets the image conversion requirements of the user.

[0046] Optionally, reference information of each element carried in the image conversion requirement can be obtained, and the information of each element in the first candidate synthesized image can be compared with the corresponding reference information to identify whether the information of each element in the first candidate synthesized image matches the reference information.

[0047] As an example, for any element, if the reference information corresponding to that element is set to purple in the image conversion requirement, then the color information of that element in the first candidate composite image can be compared with the reference information. If the color information of that element in the first candidate composite image is also purple, then it can be determined that the information of that element matches its corresponding reference information.

[0048] Furthermore, when the information of each element in the first candidate synthesized image matches the reference information, it can be determined that the first candidate synthesized image meets the image conversion requirements of the user, and then the target synthesized image of the initial photographic image can be obtained based on the first candidate synthesized image.

[0049] The method for acquiring synthetic images proposed in this disclosure receives an initial photographic image input from a user terminal, obtains the image conversion requirements of the initial photographic image, and then obtains conversion intervention information of the initial photographic image based on the image conversion requirements. Based on the conversion intervention information, the initial photographic image is processed through a latent diffusion model to obtain a first candidate synthetic image. When the first candidate synthetic image is identified as meeting the user terminal's image conversion requirements, the target synthetic image of the initial photographic image is obtained based on the first candidate synthetic image. In this disclosure, the initial photographic image is converted through a latent diffusion model based on the conversion intervention information obtained from the image conversion requirements, thereby obtaining a target synthetic image that meets the user terminal's image conversion requirements. This eliminates the need for the user to obtain the visual content materials required for the image conversion process and eliminates the need for the user to perform image conversion processing on the initial photographic image. This reduces the technical requirements for the user in the process of acquiring synthetic images, simplifies the user's operation steps, saves the user's time in obtaining visual content materials, and improves the efficiency of acquiring synthetic images. The acquisition of synthetic images through the latent diffusion model improves the efficiency of acquiring synthetic images, optimizes the aesthetics of the synthetic images, reduces the user's operational difficulty, optimizes the user's operating experience, and increases user stickiness.

[0050] In the above embodiments, the acquisition of the target composite image under the requirement of global transformation of photographic images can be combined with... Figure 2 To understand further, Figure 2 This is a flowchart illustrating a method for acquiring a synthesized image according to another embodiment of the present disclosure, as shown below. Figure 2 As shown, the method includes:

[0051] S201, in response to the image conversion requirement for global conversion of photographic images, obtain the style conversion image selected by the user and / or the style conversion prompt word selected by the user and / or the style conversion description text input by the user to obtain global conversion intervention information for the initial photographic image.

[0052] In this embodiment of the disclosure, the user terminal can perform various operations based on the image conversion processing of the initial photographic image, and the server can obtain the conversion intervention information of the initial photographic image based on the user terminal's operations.

[0053] Among them, the conversion intervention information obtained when the image conversion requirement is a global conversion requirement can be marked as the global conversion intervention information of the initial photographic image.

[0054] In this embodiment of the disclosure, the server provides a client for the user. The user can upload and input the initial photographic image through the interactive interface provided by the client. The client server can display various reference style conversion images in its preset reference style conversion image library to the user. In this scenario, the user can select an image that meets its requirements from the reference style conversion image library as the style conversion image of the initial photographic image.

[0055] Furthermore, the style transfer images provided by the user to the server can be obtained from the reference style transfer image library provided by the server, or images uploaded by the user; no specific restrictions are imposed here.

[0056] Optionally, in response to the image transformation requirement of global transformation of photographic images, first style transformation information is extracted from the style transformation image.

[0057] In this scenario, the server can determine the style transfer image selected by the user through the operation information returned by the client, and obtain the style transfer intervention information of the style transfer image selected by the user from the reference style transfer intervention information of each reference style transfer image in the preset reference style transfer image library, and mark this information as the first style transfer information.

[0058] Correspondingly, users can also upload images through the client. In this scenario, users can upload style-transformed images themselves. The server can identify and analyze the image elements based on the style-transformed images uploaded by the users, thereby extracting the style-transformation intervention information in the style-transformed images and marking the information as the first style-transformation information.

[0059] Optionally, second style transfer information of the initial photographic image can be extracted from style transfer cue words.

[0060] In this embodiment of the disclosure, a preset list of prompt words is displayed on the client. The user can select any number of prompt words from the list as intervention information when performing image conversion processing on the initial photographic image.

[0061] In this scenario, the prompts selected by the user can be marked as style transfer prompts. These prompts can include terms like "high detail texture" and "fisheye lens." Based on these style transfer prompts, the server can determine the user's image conversion requirements for the initial photographic image and extract the intervention information to mark as second style transfer information.

[0062] Optionally, a pre-acquisition natural language parsing algorithm can be used to extract third style transfer information from the style transfer representation text.

[0063] In this embodiment of the disclosure, the user can input the text content of their requirements for image conversion processing of the initial photographic image through the client, and the required text content can be marked as style conversion description text.

[0064] In this scenario, the server can perform semantic recognition on the style transfer description text based on semantic recognition algorithms in related technologies, thereby extracting the style transfer intervention information carried in it and marking it as third style transfer information.

[0065] In this embodiment of the disclosure, the server can invoke a pre-acquired natural language parsing algorithm to process the style transfer description text input by the user, thereby extracting the third style transfer information carried therein.

[0066] Optionally, based on the first style transfer information and / or the second style transfer information and / or the third style transfer information, global transformation intervention information of the initial photographic image under the global transformation requirement of the photographic image is obtained.

[0067] In this embodiment of the present disclosure, the server may integrate the acquired first style conversion information and / or second style conversion information and / or third style conversion information based on a preset integration strategy, and mark the combined information as global conversion intervention information of the initial photographic image under the global conversion requirements of the photographic image.

[0068] Optionally, the first style transfer information and / or the second style transfer information and / or the third style transfer information have their own integration weights. The first style transfer information and / or the second style transfer information and / or the third style transfer information can be weighted and integrated based on these weights to obtain the global transformation intervention information of the initial photographic image.

[0069] Optionally, the first style transfer information and / or the second style transfer information and / or the third style transfer information have their own priorities. The first style transfer information and / or the second style transfer information and / or the third style transfer information can be sorted according to the order of priority from high to low. When there are different transformation intervention information for the same element in the first style transfer information and / or the second style transfer information and / or the third style transfer information, the global transformation intervention information of the initial photographic image is obtained based on the transformation intervention information with higher priority.

[0070] S202, Based on the transformation intervention information, the initial photographic image is processed by a potential diffusion model to obtain the first candidate synthetic image after processing.

[0071] Optionally, the subject of the photograph is extracted from the initial photographic image, and based on the global transformation intervention information, the initial photographic image is transformed using the potential diffusion model to obtain a first candidate synthetic image.

[0072] In this embodiment of the disclosure, the subject of the photograph in the initial photographic image can be a person. The initial photographic image can be segmented based on a preset subject segmentation method to obtain the subject.

[0073] Optionally, an image diffusion generation algorithm corresponding to the latent diffusion model can be used to perform global image transformation processing on the initial photographic image based on the photographic subject and the global transformation intervention information, and the processed image can be marked as the first candidate synthetic image.

[0074] This can be understood as follows: by using global transformation intervention information, the transformation processing method of each element in the initial photographic image can be determined. In this scenario, the image transformation processing of each element in the initial photographic image can be performed according to the transformation processing method to obtain the first candidate synthetic image after processing.

[0075] As an example, if we know from the global transformation intervention information that the background elements in the initial photographic image, excluding the subject, are covered by the background elements in the style-transformed image, then in this example, we can preserve the state of the subject in the initial photographic image based on the global transformation intervention information, and stitch the background elements in the style-transformed image with the subject to obtain the first candidate composite image in this example.

[0076] To better understand the above embodiments, it can be combined with Figure 3 To understand further, such as Figure 3 As shown, after the user inputs the initial photographic image, global transformation intervention information of the initial photographic image can be extracted from the style transformation image and / or style transformation prompt words and / or style transformation description text.

[0077] Furthermore, through Figure 3 The portrait image diffusion generation algorithm shown in AIGC performs image transformation processing on the initial photographic image to obtain... Figure 3 The first candidate composite image is shown.

[0078] like Figure 3 As shown, it is possible to identify whether the user is satisfied with the fusion effect of the people in the first candidate composite image. When it is identified that the user is satisfied with the fusion effect of the people in the first candidate composite image, the target composite image of the initial photograph can be generated based on the first candidate composite image.

[0079] Optionally, the first candidate synthetic image is subjected to high-definition restoration processing to obtain the processed target synthetic image.

[0080] In this embodiment of the disclosure, the first candidate synthetic image can be further processed to optimize the visual effect of the processed synthetic image, and the processed synthetic image can be marked as the target synthetic image of the initial photographic image.

[0081] As an example, such as Figure 3 As shown, it can be done through Figure 3 The illustrated high-definition restoration steps perform high-definition restoration processing on the first candidate composite image, wherein, it can be achieved through... Figure 3 The detail enhancement algorithm shown performs high-definition restoration processing on the first candidate synthetic image, and determines the high-definition merged image obtained after processing as the target synthetic image to be returned to the user.

[0082] In this embodiment of the disclosure, the first candidate synthesized image may not meet the image conversion requirements of the user. In response to the first candidate synthesized image not meeting the image conversion requirements, a first local image is cropped from the first candidate synthesized image using a pre-acquired image region segmentation algorithm.

[0083] Optionally, when the first candidate synthesized image does not meet the user's image conversion requirements, a local secondary conversion process can be performed on the unsatisfactory local images in the first candidate synthesized image. In this process, the first candidate synthesized image can be segmented based on an image region segmentation algorithm, and the unsatisfactory local images can be cropped from the first candidate synthesized image and marked as the first local image.

[0084] As an example, such as Figure 3 As shown, when the user is not satisfied with the fusion effect of the person in the first candidate synthesized image, it can be determined that the first candidate synthesized image does not meet the user's image conversion requirements. In this scenario, the first candidate synthesized image can be segmented using an image region segmentation algorithm to obtain... Figure 3 The first partial image shown.

[0085] Optionally, first local intervention information of the first local image can be obtained.

[0086] In this embodiment of the disclosure, a second local transformation process can be performed on the first local image based on the transformation intervention information corresponding to the first local image. The intervention information used when performing the second local transformation process on the first local image can be marked as the first local intervention information.

[0087] Optionally, a person fusion enhancement algorithm is obtained, and the first local image is fused and intervened based on the first local intervention information using the person fusion enhancement algorithm to obtain the processed second candidate synthetic image.

[0088] As an example, such as Figure 3 As shown, by Figure 3 As can be seen, the user is not satisfied with the fusion effect of the people in the first candidate synthesized image. In this example, a secondary fusion enhancement process is needed to enhance the fusion effect of the people in the first candidate synthesized image.

[0089] Specifically, the person fusion enhancement algorithm can be invoked to perform secondary person fusion enhancement processing on the first local image based on the first local intervention information, and the processed image can be marked as the second candidate composite image.

[0090] Optionally, a target composite image of the initial photographic image is obtained based on the second candidate composite image.

[0091] In this embodiment of the present disclosure, the second candidate synthetic image can be further processed based on a preset processing strategy, so as to mark the processed image as the target synthetic image of the initial photographic image.

[0092] In this process, the second candidate synthetic image undergoes high-definition restoration to obtain the processed target synthetic image.

[0093] In this embodiment of the disclosure, the second candidate synthetic image can be further processed to optimize the visual effect of the processed synthetic image, and the processed synthetic image can be marked as the target synthetic image of the initial photographic image.

[0094] As an example, such as Figure 3 As shown, it can be done through Figure 3 The illustrated high-definition restoration steps perform high-definition restoration processing on the second candidate composite image, wherein, it can be achieved through... Figure 3 The detail enhancement algorithm shown performs high-definition restoration processing on the second candidate synthetic image, and determines the high-definition merged image obtained after processing as the target synthetic image to be returned to the user.

[0095] It should be noted that, before obtaining the conversion intervention information of the initial photographic image according to the image conversion requirements, the initial photographic image also needs to be processed accordingly. This may include resolution compression of the initial photographic image and extraction of conversion reference information from the initial photographic image. The image reference information includes at least the subject pose information, image layer information, and contour information of the initial photographic image.

[0096] As an example, such as Figure 3 As shown, after the user inputs the initial photographic image, the resolution of the initial photographic image can be compressed according to the resolution compression method in related technologies.

[0097] Accordingly, when performing conversion processing on the initial photographic image, it is necessary to use some information from the initial photographic image as reference information, which can be marked as conversion reference information in the initial photographic image.

[0098] like Figure 3 As shown, the initial photographic image can be processed using image information extraction algorithms in related technologies to extract the transformation reference information of the initial photographic image.

[0099] It should be noted that the transformation reference information in the initial photographic image may include the subject's pose information, image information, and outline information in the initial photographic image, as well as other information that needs to be referenced when performing image transformation processing, without being specifically limited here.

[0100] like Figure 3 As shown, the target composite image can also be stored in the cloud, wherein, in response to receiving a cloud save instruction, the target composite image is uploaded to the cloud and stored.

[0101] In this embodiment of the disclosure, after receiving the target composite image, the user terminal can send a cloud save instruction for the target composite image to the server. When the server receives the cloud save instruction input by the user terminal, it can upload the acquired target composite image to the cloud and store it in a preset storage area through the cloud data storage function.

[0102] To better understand the above embodiments, it can be combined with Figure 3 ,like Figure 3 As shown:

[0103] S301, acquire the initial photographic image;

[0104] S302, Obtain the style transfer image and / or style transfer prompt words and / or style transfer description text;

[0105] S303, based on style transfer images and / or style transfer cue words and / or style transfer description text, obtain global intervention information;

[0106] S304. Based on global intervention information, the initial photographic image is transformed using a potential diffusion model to obtain the first candidate synthetic image.

[0107] S305, Identify whether the fusion effect of the people in the first candidate synthetic image is satisfactory;

[0108] S306, In response to the unsatisfactory character fusion effect, the first local image is cropped from the first candidate synthesized image through an image region segmentation algorithm;

[0109] S307, the first local image is enhanced by the person fusion enhancement algorithm to obtain the second candidate composite image;

[0110] S308, perform high-definition restoration processing on the first candidate synthetic image or the second candidate synthetic image using a detail enhancement algorithm;

[0111] S309, Obtain the target composite image after high-definition restoration processing of the first candidate composite image or the second candidate composite image;

[0112] S310 stores the target composite image in the cloud.

[0113] In this embodiment of the disclosure, the initial photographic image input by the user terminal can be obtained, and the initial photographic image can be compressed in resolution and the reference information extracted.

[0114] In this scenario, the user's image transformation requirements for the initial photographic image can be obtained. When the user's image transformation requirements for the initial photographic image are identified as global image transformation requirements, the user's style transformation image and / or the style transformation prompt words and / or the style transformation description text can be obtained. Then, the global intervention information of the initial photographic image can be extracted from the style transformation image and / or the style transformation prompt words and / or the style transformation description text.

[0115] Furthermore, based on the global intervention information, the initial photographic image is subjected to global image transformation processing through a latent model to obtain the first candidate synthetic image.

[0116] In this scenario, it is necessary to identify whether the user is satisfied with the fusion effect of the person in the first candidate synthesized image. When it is identified that the user is satisfied with the fusion effect of the person in the first candidate synthesized image, the first candidate synthesized image can be subjected to high-definition restoration processing through detail enhancement algorithm to obtain the corresponding target synthesized image, and then stored in the cloud through the cloud data storage function.

[0117] Accordingly, when it is detected that the user is not satisfied with the fusion effect of the person in the first candidate synthesized image, the first candidate synthesized image can be segmented by a pre-acquired image region segmentation algorithm to extract the part of the image that the user is not satisfied with as the first local image.

[0118] Furthermore, the first local image is subjected to person fusion enhancement processing through the pre-acquired person fusion enhancement algorithm to obtain the processed second candidate composite image. The second candidate composite image is then subjected to high-definition restoration processing through the detail enhancement algorithm to obtain the corresponding target composite image. The target composite image is then uploaded to the cloud for cloud storage through the cloud data storage function.

[0119] The proposed method for obtaining synthetic images, in response to the image conversion requirement being a global conversion requirement for photographic images, extracts global conversion intervention information of the initial photographic image from style conversion images and / or style conversion prompts and / or style conversion description text. Based on the global conversion intervention information, a latent diffusion model is used to perform global image conversion processing on the initial photographic image to obtain a first candidate synthetic image. When the first candidate synthetic image is identified as meeting the image conversion requirement, high-definition restoration processing is performed on the first candidate synthetic image to obtain a target synthetic image. When the first candidate synthetic image is identified as not meeting the image conversion requirement, a first local image and first local intervention information of the first candidate synthetic image are obtained, and a person fusion enhancement algorithm is used to perform fusion intervention processing on the first local image based on the first local intervention information to obtain a processed second candidate synthetic image. The second candidate synthetic image is then subjected to high-definition restoration to obtain the target synthetic image of the initial photographic image. In this disclosure, global conversion intervention information is obtained by selecting a style conversion image and / or style conversion prompts and / or style conversion description text on the user's end, thereby achieving global conversion of the initial photographic image and obtaining the target composite image of the initial photographic image. This eliminates the need for the user to obtain the visual content materials required for the image conversion process and eliminates the need for the user to perform image conversion processing on the initial photographic image. This reduces the technical requirements for the user in the process of obtaining the composite image, simplifies the user's operation steps, saves the user's time in obtaining visual content materials, and improves the efficiency of obtaining the composite image. Secondary human figure fusion enhancement processing is performed on the first candidate composite image to optimize the human figure fusion effect in the target composite image. High-definition processing is performed on the first candidate composite image or the second candidate composite image to improve the image quality of the target composite image.

[0120] In the above embodiments, the acquisition of the target composite image under the requirement of local redrawing of photographic images can be combined with... Figure 4 understand, Figure 4This is a flowchart illustrating a method for acquiring a synthesized image according to another embodiment of the present disclosure, as shown below. Figure 4 As shown, the method includes:

[0121] S401, in response to the image conversion requirement for local redrawing of the photographic image, obtains local conversion intervention information of the initial photographic image from the local redrawing text input by the user and / or the local redrawing prompt words selected by the user.

[0122] Optionally, in response to the image conversion requirement of local redrawing of a photographic image, a first local redrawing information is extracted from the local redrawing text using a pre-acquired natural language parsing algorithm. And / or, a second local redrawing information is extracted from local redrawing prompts.

[0123] In this embodiment of the disclosure, the user terminal may only have a need for partial redrawing of the initial photographic image. This need can be marked as a partial redrawing need of the photographic image. It can be understood that for any initial photographic image, the user terminal may only need to modify the color of the clouds in the image, and this need can be marked as a partial redrawing need of the photographic image.

[0124] In this scenario, the user can input the corresponding text content based on the text input function provided by the client, and this text can be marked as the locally redrawn text input by the user.

[0125] Optionally, the server can use a natural language parsing algorithm to process the local redrawing text input by the user, thereby extracting the local redrawing intervention information carried in it and marking it as the first local redrawing information.

[0126] Correspondingly, the user client can also select prompts for partial redrawing based on the prompt options provided by the client. The prompts selected by the user client can be marked as prompts for partial redrawing.

[0127] In this scenario, the server can also obtain the local redrawing intervention information carried in the local redrawing prompt word selected by the user and mark it as the second local redrawing information in the local redrawing prompt word.

[0128] Optionally, local transformation intervention information of the initial photographic image can be obtained based on the first local redrawing information and / or the second local redrawing information.

[0129] In this embodiment of the disclosure, when only one of the first local redrawing information and the second local redrawing information exists, the information can be determined as the local transformation intervention information of the initial photographic image.

[0130] Furthermore, when the first local redrawing information and the second local redrawing information exist simultaneously, the two information can be combined based on a preset combination strategy, and the combined information can be marked as the local transformation intervention information of the initial photographic image.

[0131] Furthermore, the first local redrawing information and the second local redrawing information can be integrated based on a preset information priority. This can be understood as follows: for the same element, when there is a difference between the redrawing information of the element in the first local redrawing information and the redrawing information in the second local redrawing information, the redrawing information with higher priority is retained as the final information used when the element is redrawn locally, thereby obtaining the local transformation intervention information of the initial photographic image.

[0132] S402, based on the transformation intervention information, the initial photographic image is processed by a potential diffusion model to obtain the first candidate synthetic image after processing.

[0133] Optionally, a second local image to be redrawn can be extracted from the initial photographic image.

[0134] Among them, the redrawing region segmentation of the initial photographic image can be performed by a pre-acquired image region segmentation algorithm to obtain the second local image to be redrawn in the initial photographic image.

[0135] In this embodiment of the disclosure, an image region segmentation and cropping can be performed on the initial photographic image using an image segmentation algorithm. Specifically, the user's local redrawing requirements for the photographic image can be analyzed to determine the local images in the initial photographic image that need to be redrawn by the user. The initial photographic image is then cropped and segmented using an image region segmentation algorithm to extract the local images in the initial photographic image that need to be redrawn, and these are marked as the second local images.

[0136] Optionally, based on the local transformation intervention information, the second local image is locally redrawn using a latent diffusion model to obtain a processed third candidate synthetic image, which is then used as the first candidate synthetic image.

[0137] As an example, such as Figure 5 As shown, the initial photographic image can be processed. Figure 5 After extracting the resolution compression and conversion parameter information shown, the redrawn region in the initial photographic image is extracted according to the local conversion requirements of the user's photographic image through an image region segmentation algorithm, thereby obtaining the second local image.

[0138] like Figure 5 As shown, the second local transformation intervention information of the second local image can be extracted from the local repainting text and local repainting prompts input by the user through natural language parsing algorithms, and then... Figure 5The portrait image diffusion generation algorithm shown, based on the latent diffusion model, performs local redrawing processing on the second local image using second local transformation intervention information, thereby obtaining... Figure 5 The third candidate composite image shown is the first candidate composite image after being transformed from the initial photographic image.

[0139] As another example, if the user inputs the text for local redrawing as "redraw a white wing on the left and right sides of the person in the initial photographic image," then the corresponding local transformation intervention information can be extracted based on this text.

[0140] In this example, the left and right regions of the person in the initial photographic image can be extracted as a second local image using an image region segmentation algorithm. Based on the local transformation intervention information, the second local image is then locally redrawn using a latent diffusion model. This redraws a white wing for the left and right regions of the person in the initial photographic image, thus achieving the local redrawing processing of the second local image.

[0141] Furthermore, the image processed by the second local image based on the local transformation intervention information is labeled as the third candidate synthetic image.

[0142] Optionally, a target composite image of the initial photographic image is obtained based on the third candidate composite image, wherein the third candidate composite image is subjected to high-definition restoration processing to obtain the target composite image of the initial photographic image.

[0143] As an example, such as Figure 5 As shown, it can be done through Figure 5 The detail enhancement algorithm shown performs high-definition restoration on the third candidate synthetic image, thereby obtaining the processed target synthetic image.

[0144] In this embodiment of the disclosure, after obtaining the target synthetic image, the user terminal may need to perform cloud storage processing on the target synthetic image, such as... Figure 5 As shown, the target composite image can be uploaded to the cloud, and the target composite image can be stored in a preset area through the cloud data storage function.

[0145] Optionally, based on the target synthesized image, an updated style transfer image for the user is constructed, and the updated style transfer image is stored in a reference style transfer image library.

[0146] In this embodiment of the disclosure, the image conversion processing of the target synthetic image obtained by the user terminal may use total intervention information composed of intervention information obtained from multiple channels. In this scenario, a new style conversion image can be constructed based on the target synthetic image, marked as the updated style conversion image of the user terminal, and stored in a reference style conversion image library on the server for storing all style conversion images.

[0147] This can be understood as follows: during the subsequent image synthesis process on the user's end, the updated style conversion image can be directly selected from the reference style conversion image library without the user's end needing to input any other information. The server can directly read the user's conversion requirements for the photographic image based on the updated style conversion image, thereby realizing the corresponding image conversion processing.

[0148] To better understand the above embodiments, it can be combined with Figure 5 ,like Figure 5 As shown:

[0149] S501, acquire the initial photographic image input by the user;

[0150] S502, Based on the initial photographic image, the local redrawn region in the initial photographic image is obtained through an image region segmentation algorithm and used as the second local image;

[0151] S503, obtain the local redrawing text input by the user terminal and / or the local redrawing prompt words selected by the user terminal, and extract the local transformation intervention information of the second local image from the local redrawing text and / or local redrawing prompt words through a natural language parsing algorithm;

[0152] S504, Based on the local transformation intervention information, the second local image is locally redrawn using a potential diffusion model to obtain a third candidate synthetic image;

[0153] S505, for the third candidate synthetic image, high-definition restoration is performed using a detail enhancement algorithm;

[0154] S506, Obtain the target composite image after high-definition restoration of the third candidate composite image;

[0155] S507 stores the target composite image in the cloud.

[0156] In this embodiment of the disclosure, the initial photographic image input by the user terminal can be obtained. The resolution of the initial photographic image can be compressed and the conversion reference information can be extracted. The image conversion requirements of the user terminal for the initial photographic image can be obtained. When the image conversion requirement of the user terminal for the initial photographic image is a local redrawing requirement of the photographic image, the initial photographic image can be divided into regions by a pre-acquired image region segmentation algorithm, the local redrawing region that needs to be locally redrawn in the initial photographic image can be extracted, and it can be determined as the second local image.

[0157] Optionally, the local redrawing text input by the user and / or the local redrawing prompt words selected by the user are obtained, and the second local redrawing information of the second local image is extracted from the local redrawing text and / or local redrawing prompt words through a pre-acquired natural language parsing algorithm.

[0158] Furthermore, based on the second local redrawing information, the second local image is locally redrawn using a latent diffusion model to obtain the processed third candidate synthetic image.

[0159] like Figure 5 As shown, it can be done through Figure 5 The detail enhancement algorithm shown performs high-definition restoration processing on the third candidate synthetic image to obtain the processed target synthetic image, and then uploads the target synthetic image to the cloud for cloud storage through the cloud data storage function.

[0160] The proposed method for obtaining synthetic images, in response to an image conversion requirement of local redrawing of a photographic image, extracts local conversion intervention information from local redrawing text and prompts input by the user. Based on this intervention information, a latent diffusion model is used to perform local redrawing processing on a second local image to obtain a processed third candidate synthetic image. This third candidate synthetic image is then subjected to high-resolution restoration to obtain the target synthetic image. In this method, the corresponding local conversion intervention information is extracted from the local redrawing text and / or prompts input by the user, thereby achieving local redrawing processing of the initial photographic image to obtain the corresponding target synthetic image. This eliminates the need for the user to acquire the visual content materials required for the image conversion process and to perform image conversion processing on the initial photographic image, reducing the technical requirements for obtaining the synthetic image and simplifying the user's operation.

[0161] This disclosure also proposes a synthetic image client that can be combined with Figure 6 understand, Figure 6 This is a schematic diagram of a synthetic image client according to an embodiment of the present disclosure, such as... Figure 6As shown, the image compositing client 600 includes: an image upload module 61, a reference style conversion image library module 62, a prompt word module 63, and an image compositing icon 64, wherein...

[0162] The image upload module 61 is used to provide an upload port for the initial photographed image to the user terminal and to transmit the initial photographed image uploaded by the user terminal to the server.

[0163] The reference style transfer image library module 62 is used to display various reference style transfer images in the reference style transfer image library to the user terminal. In response to the user terminal selecting the style transfer image of the initial photographic image from the reference style transfer image library, the style transfer image is sent to the server.

[0164] The prompt word module 63 is used to provide a prompt word option list to the user terminal, wherein, in response to the user terminal selecting a style transfer prompt word or a local redraw prompt word for the initial camera image from the prompt word option list, the style transfer prompt word or local redraw prompt word is sent to the server.

[0165] Image compositing icon 64 is used to provide a compositing instruction input port for the user terminal. In response to the user terminal clicking the image compositing icon, it is confirmed that the user terminal has entered a confirmation compositing instruction, and the confirmation compositing instruction is sent to the server. The server starts to convert the initial photographic image based on the received confirmation compositing instruction to obtain candidate compositing images, and then obtains the target compositing image based on the candidate compositing images.

[0166] In this embodiment of the disclosure, the server can... Figure 6 The synthesized image client 600 shown interacts with the user terminal. The synthesized image client 600 includes an image upload module 61. The user terminal can upload the initial photographic image through the image upload module 61. After receiving the initial photographic image uploaded by the user terminal, the synthesized image client 600 can transmit the received initial photographic image to the corresponding server.

[0167] As an example, such as Figure 7 As shown, Figure 7 The area marked 1 is the image upload module displayed by the composite image client to the user. The user can upload the initial photographed image through the image upload module displayed in this area.

[0168] like Figure 6 As shown, the composite image client 600 also includes a reference style conversion image library module 62, which stores multiple reference style conversion images. The user can filter from these images to determine the style conversion image used when performing image conversion processing on the initial photographic image.

[0169] The composite image client 600 can display various reference style-transformed images through the reference style-transformed image library module 62. In this scenario, the user can select the style-transformed image of the initial photographic image from the displayed reference style-transformed images.

[0170] As an example, such as Figure 7 As shown, Figure 7 The area corresponding to identifier 2 shown is the reference style-transformed images displayed by the synthetic image client to the user. The user can select the style-transformed image of the initial photographic image based on the information displayed in this area.

[0171] Furthermore, after receiving the style-transformed image selected by the user, the composite image client 600 can transmit the style-transformed image to the server.

[0172] like Figure 6 As shown, the image synthesis client 600 also includes a prompt word module 63, which allows the user to filter prompt words for image conversion processing based on the prompt word option list displayed by the module.

[0173] like Figure 7 As shown, Figure 7 The area marked 3 is the list of prompt options displayed by the image synthesis client to the user.

[0174] Furthermore, the user can filter the prompt words corresponding to the conversion process of the initial photographic image from the prompt word option list displayed in the prompt word module 63. When the image conversion requirement of the initial photographic image is a global image conversion requirement, the prompt word is a style conversion prompt word. When the image conversion requirement of the initial photographic image is a local image conversion requirement, the prompt word is a local redraw prompt word.

[0175] In this scenario, the composite image client 600 can send the style transfer prompt or local redraw prompt selected by the user to the server.

[0176] like Figure 6 As shown, the image compositing client 600 also includes an image compositing icon 64, which provides an input port for the user to confirm the compositing command. It can be understood that when the user clicks the image compositing icon 64, it means that the user has entered a confirmation command for compositing.

[0177] As an example, such as Figure 7 As shown, Figure 7 The area corresponding to the indicated symbol 4 is the image compositing icon displayed by the image compositing client. Users can click on this icon to input confirmation commands for compositing.

[0178] In this scenario, after receiving the confirmation synthesis instruction, the image synthesis client 600 can transmit the instruction to the server. Based on the received confirmation synthesis instruction, the server starts the image conversion processing flow of the initial photographic image and obtains the corresponding candidate synthesized image.

[0179] It should be noted that the candidate composite images displayed to the user by the composite image client 600 may not meet the user's requirements for the character fusion effect. In this scenario, it can be addressed by... Figure 8 The character fusion module shown enhances the character fusion effect by performing a secondary fusion on the candidate composite image.

[0180] like Figure 8 As shown, it can be done through Figure 8 The smart selection or manual selection shown is from Figure 8 The area that needs to be enhanced with the character fusion effect is selected from the candidate composite images shown and used as the first local image.

[0181] like Figure 8 As shown, Figure 8 The interface shown includes a character fusion icon. Users can click this icon to input a confirmation fusion command to the composite image client. After receiving the confirmation fusion command from the user, the composite image client can transmit the command to the server, thereby enabling the server to perform secondary enhancement processing on the character fusion effect of the first local image, and thus obtain the processed candidate composite image.

[0182] Optionally, the client also includes a high-definition processing component for performing high-definition restoration processing on the candidate synthetic image to obtain the processed target synthetic image.

[0183] In this embodiment of the present disclosure, the image synthesis client 600 further includes a high-definition processing component ( Figure 6 (not shown in the image), can be processed by a high-definition processing component ( Figure 6 (Not shown in the image) The candidate synthetic image is subjected to high-definition restoration processing to obtain the processed target synthetic image.

[0184] As an example, such as Figure 7 As shown, Figure 7 On the interface of the synthesized image client shown, the area corresponding to mark 5 is the display area of ​​the target synthesized image. The high-definition processing component in the synthesized image client can perform high-definition restoration processing on the candidate synthesized image, thereby obtaining the processed target synthesized image. Figure 7 The area corresponding to identifier 5 is displayed to the user.

[0185] Optionally, the composite image client also includes a cloud save icon. In response to the user clicking the cloud save icon, it is confirmed that the user has entered a cloud save command. Based on the cloud save command, the server uploads the target composite image to the cloud and stores it.

[0186] In this embodiment of the present disclosure, the image synthesis client 600 also includes a cloud save icon ( Figure 6 (Not shown in the image) After viewing the target synthesized image on the user's device, the user can save the image to the cloud by clicking "Save Image to Cloud". Figure 6 (Not shown in the image) Input the cloud save command to the composite image client 600.

[0187] In this scenario, the composite image client 600 can transmit the received cloud save instruction to the server. After receiving the instruction, the server can upload the target composite image to the cloud and store it.

[0188] Optionally, the composite image client 600 also includes a global conversion icon ( Figure 6 (not shown in the image) and local redrawing symbols ( Figure 6 (not shown in the image), where, in response to the user clicking the global conversion icon ( Figure 6 (not shown in the image), determine the global transformation requirements of the photographic image as the image transformation requirements of the initial photographic image, and send it to the server.

[0189] In this embodiment of the present disclosure, the image synthesis client 600 also includes a global conversion icon ( Figure 6 (not shown in the image) and local redrawing symbols ( Figure 6 (Not shown in the image).

[0190] Among them, when the user clicks the global conversion icon ( Figure 6 When (not shown in the image), the composite image client 600 can determine that the user's image conversion requirement for the initial photographic image is a global image conversion requirement and transmit it to the server.

[0191] Accordingly, in response to a user clicking on a local redraw icon ( Figure 6 (not shown in the image), determine the local redrawing requirement of the photographic image as the image transformation requirement of the initial photographic image, and send it to the server.

[0192] This can be understood as, when the user clicks on a local redraw icon ( Figure 6 In the image synthesis client 600 (not shown), the user's image conversion requirement for the initial photographic image is a local redraw conversion requirement for the photographic image, and sends it to the server.

[0193] As an example, when a user clicks the global conversion icon ( Figure 6When (not shown in the image), the client interface displayed by the composite image client 600 to the user terminal can be combined with... Figure 7 The interface shown illustrates this; correspondingly, when the user clicks the local redraw icon ( Figure 6 When (not shown in the image), the client interface displayed by the composite image client 600 to the user terminal can be combined with... Figure 9 understand.

[0194] like Figure 9 As shown, Figure 9 In the image, identifier 1 corresponds to the initial photographed image uploaded by the user, identifier 2 corresponds to the cropping and segmentation of the second local image that needs to be partially redrawn in the initial photographed image, identifier 3 corresponds to the local redraw prompt word selected by the user, and identifier 4 corresponds to the confirmation and compositing icon.

[0195] like Figure 9 As shown, the user can... Figure 9 The client's interface, which shows the partial redrawing of the corresponding composite image, demonstrates how the initial photographic image is partially redrawn to obtain the corresponding target composite image.

[0196] The image compositing client disclosed herein allows the user to upload an initial photographic image via an image upload module, select a style transfer image for image conversion processing via a reference style transfer image library module, select a corresponding style transfer prompt or local redraw prompt via a prompt word module, and then input a confirmation compositing command via an image compositing icon. The image compositing client transmits the information input by the user to the server, enabling the server to perform image conversion processing on the initial photographic image based on the information input by the user through the image compositing client, thereby obtaining the corresponding target composite image. This eliminates the need for the user to acquire the visual content materials required for the image conversion process, and also eliminates the need for the user to perform image conversion processing on the initial photographic image. This reduces the technical requirements for users in acquiring composite images, simplifies the user's operation steps, saves users time in acquiring visual content materials, reduces the difficulty of operation, and optimizes the user experience.

[0197] Corresponding to the synthetic image acquisition methods proposed in the above embodiments, an embodiment of this disclosure also proposes a synthetic image acquisition device. Since the synthetic image acquisition device proposed in this disclosure corresponds to the intention prediction + model training method proposed in the above embodiments, the implementation methods of the above synthetic image acquisition methods are also applicable to the synthetic image acquisition device proposed in this disclosure, and will not be described in detail in the following embodiments.

[0198] Figure 10 This is a schematic diagram of the structure of a synthetic image acquisition device according to an embodiment of the present disclosure, as shown below. Figure 10As shown, the image acquisition device 100 includes a first acquisition module 11, a second acquisition module 12, a first processing module 13, and a second processing module 14, wherein:

[0199] The first acquisition module 11 is used to receive the initial photographed image input by the user terminal and to acquire the image conversion requirements of the initial photographed image.

[0200] The second acquisition module 12 is used to acquire conversion intervention information of the initial photographic image according to the image conversion requirements.

[0201] The first processing module 13 is used to perform image transformation processing on the initial photographic image based on the transformation intervention information and through a potential diffusion model to obtain the processed first candidate synthetic image.

[0202] The second processing module 14 is used to obtain the target composite image of the initial photographic image based on the first candidate image in response to the first candidate image meeting the image conversion requirements.

[0203] In this embodiment of the disclosure, the second acquisition module 12 is further configured to: in response to an image conversion requirement that is a global conversion requirement for a photographic image, acquire a style conversion image selected by the user and / or a style conversion prompt word selected by the user and / or a style conversion description text input by the user, so as to obtain global conversion intervention information of the initial photographic image. In response to an image conversion requirement that is a local redrawing requirement for a photographic image, acquire local conversion intervention information of the initial photographic image from the local redrawing text input by the user and / or the local redrawing prompt word selected by the user.

[0204] In this embodiment of the disclosure, the second acquisition module 12 is further configured to: in response to an image conversion requirement that is a global conversion requirement for the photographic image, extract first style conversion information from the style conversion image; and / or extract second style conversion information of the initial photographic image from style conversion prompts; and / or extract third style conversion information from the style conversion description text using a pre-acquired natural language parsing algorithm. Based on the first style conversion information and / or the second style conversion information and / or the third style conversion information, acquire global conversion intervention information of the initial photographic image under the global conversion requirement for the photographic image.

[0205] In this embodiment of the present disclosure, the first processing module 13 is further configured to: extract the photographic subject from the initial photographic image; and, based on global transformation intervention information, perform image transformation processing on the initial photographic image using a latent diffusion model, thereby obtaining a first candidate synthetic image.

[0206] In this embodiment of the present disclosure, the second processing module 14 is further configured to: respond to the first candidate synthesized image not meeting the image conversion requirements, crop out a first local image to be fused from the first candidate synthesized image using a pre-acquired image region segmentation algorithm; acquire first local intervention information of the first local image; acquire a person fusion enhancement algorithm, and perform fusion intervention processing on the first local image based on the first local intervention information using the person fusion enhancement algorithm to obtain a processed second candidate synthesized image; and obtain a target synthesized image of the initial photographic image based on the second candidate synthesized image.

[0207] In this embodiment of the present disclosure, the second processing module 14 is further configured to: perform high-definition restoration processing on the first candidate composite image or the second candidate composite image to obtain the processed target composite image.

[0208] In this embodiment of the disclosure, the second acquisition module 12 is further configured to: in response to an image conversion requirement that is a local redrawing requirement for a photographic image, extract first local redrawing information from the local redrawing text using a pre-acquired natural language parsing algorithm; and / or extract second local redrawing information from local redrawing prompts; and obtain local conversion intervention information for the initial photographic image based on the first local redrawing information and / or the second local redrawing information.

[0209] In this embodiment of the present disclosure, the first processing module 13 is further configured to: extract a second local image to be redrawn from the initial photographic image; perform local redrawing processing on the second local image using a latent diffusion model based on local transformation intervention information to obtain a processed third candidate composite image, which serves as the first candidate composite image.

[0210] In this embodiment of the present disclosure, the first processing module 13 is further configured to: perform redrawing region segmentation on the initial photographic image using a pre-acquired image region segmentation algorithm to obtain a second local image to be redrawn in the initial photographic image.

[0211] In this embodiment of the present disclosure, the second processing module 14 is further configured to: perform high-definition restoration processing on the third candidate synthetic image to obtain the target synthetic image of the initial photographic image.

[0212] In this embodiment of the present disclosure, the first acquisition module 11 is further configured to: compress the resolution of the initial photographic image and extract the transformation reference information in the initial photographic image, wherein the image reference information includes at least the photographic subject pose information, image layer information and contour information of the initial photographic image.

[0213] In this embodiment of the present disclosure, the apparatus further includes an update module, configured to: construct an updated style transfer image for the user end based on the target synthesized image, and store the updated style transfer image in a reference style transfer image library.

[0214] In this embodiment of the present disclosure, the device further includes a storage module, configured to: upload the target composite image to the cloud and store it in response to receiving a cloud save instruction.

[0215] The synthetic image acquisition device proposed in this disclosure receives an initial photographic image input from a user terminal, obtains the image conversion requirements of the initial photographic image, and then obtains conversion intervention information of the initial photographic image based on the image conversion requirements. Based on the conversion intervention information, the initial photographic image is processed through a latent diffusion model to obtain a first candidate synthetic image. When the first candidate synthetic image is identified as meeting the user terminal's image conversion requirements, the target synthetic image of the initial photographic image is obtained based on the first candidate synthetic image. In this disclosure, the initial photographic image is converted through a latent diffusion model based on the conversion intervention information obtained from the image conversion requirements, thereby obtaining a target synthetic image that meets the user terminal's image conversion requirements. This eliminates the need for the user to obtain the visual content materials required for the image conversion process and eliminates the need for the user to perform image conversion processing on the initial photographic image. This reduces the technical requirements for the user in the synthetic image acquisition process, simplifies the user's operation steps, saves the user's time in obtaining visual content materials, and improves the efficiency of synthetic image acquisition. The acquisition of synthetic images through the latent diffusion model improves the efficiency of synthetic image acquisition, optimizes the aesthetics of the synthetic image, reduces the user's operational difficulty, optimizes the user's operating experience, and increases user stickiness.

[0216] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0217] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0218] like Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded into random access memory (RAM) 1103 from storage unit 1108. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0219] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0220] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the method for acquiring a composite image. For example, in some embodiments, the method for acquiring a composite image may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the method for acquiring a composite image described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform a method for acquiring a synthetic image by any other suitable means (e.g., by means of firmware).

[0221] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0222] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0223] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0224] To initiate interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user account can submit input to the computer. Other types of devices can also be used to initiate interaction with the user account; for example, feedback submitted to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including voice input, speech input, or tactile input).

[0225] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user account computer with a graphical user interface or web browser through which a user account can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0226] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0227] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0228] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for acquiring a synthesized image, characterized in that, The method includes: Receive the initial photographed image input from the user terminal, and obtain the image conversion requirements of the initial photographed image; Based on the image conversion requirements, obtain the conversion intervention information of the initial photographic image; Based on the transformation intervention information, the initial photographic image is transformed using a potential diffusion model to obtain the processed first candidate synthetic image. In response to the first candidate synthesized image satisfying the image conversion requirement, a target synthesized image of the initial photographic image is obtained based on the first candidate synthesized image, wherein: reference information of each element carried in the image conversion requirement is obtained, and the information of each element in the first candidate synthesized image is compared with the corresponding reference information to identify whether the information of each element in the first candidate synthesized image matches the reference information; The step of obtaining conversion intervention information for the initial photographic image based on the image conversion requirements includes: In response to the image conversion requirement being a global conversion requirement for photographic images, the style conversion image selected by the user terminal and / or the style conversion prompt word selected by the user terminal and / or the style conversion description text input by the user terminal are obtained to obtain the global conversion intervention information of the initial photographic image; In response to the image conversion request being a local redrawing request for a photographic image, local conversion intervention information for the initial photographic image is obtained from the local redrawing text input by the user and / or the local redrawing prompt word selected by the user.

2. The method according to claim 1, wherein, In response to the image conversion request being a global conversion request for the photographic image, the system acquires the style conversion image selected by the user and / or the style conversion prompt word selected by the user and / or the style conversion description text input by the user, to obtain the conversion intervention information for the initial photographic image, including: In response to the image conversion request being a global conversion request for the photographic image, first style conversion information is extracted from the style conversion image; and / or, Extract the second style transfer information of the initial photographic image from the style transfer prompts; and / or, A pre-acquired natural language parsing algorithm is used to extract third style transfer information from the style transfer description text; Based on the first style conversion information and / or the second style conversion information and / or the third style conversion information, obtain the global conversion intervention information of the initial photographic image under the global conversion requirement of the photographic image.

3. The method according to claim 2, wherein, The step of performing image transformation processing on the initial photographic image based on the transformation intervention information and using a latent diffusion model to obtain the processed first candidate synthetic image includes: Extract the main subject from the initial photographic image; Based on the global transformation intervention information, the initial photographic image is processed using the potential diffusion model based on the photographic subject to obtain the first candidate synthetic image.

4. The method according to claim 1, wherein, The method further includes: In response to the first candidate synthesized image not meeting the image conversion requirements, a first local image to be fused is cropped from the first candidate synthesized image using a pre-acquired image region segmentation algorithm; Obtain the first local intervention information of the first local image; A character fusion enhancement algorithm is obtained, and the first local image is fused and intervened based on the first local intervention information using the character fusion enhancement algorithm to obtain a processed second candidate composite image. The target composite image of the initial photographic image is obtained based on the second candidate composite image.

5. The method according to claim 4, wherein, The method further includes: The first candidate composite image or the second candidate composite image is subjected to high-definition restoration processing to obtain the processed target composite image.

6. The method according to claim 1, wherein, In response to the image conversion request being a local redrawing request for a photographic image, the local conversion intervention information of the initial photographic image is obtained from the local redrawing text input by the user and / or the local redrawing prompt words selected by the user, including: In response to the image conversion requirement being a local redrawing requirement for the photographic image, a first local redrawing information is extracted from the local redrawing text using a pre-acquired natural language parsing algorithm; and / or, Extract the second local redraw information from the local redraw prompts; The local transformation intervention information of the initial photographic image is obtained based on the first local redrawing information and / or the second local redrawing information.

7. The method according to claim 6, wherein, The step of performing image transformation processing on the initial photographic image based on the transformation intervention information and using a latent diffusion model to obtain the processed first candidate synthetic image includes: Extract the second local image to be redrawn from the initial photographic image; Based on the local transformation intervention information, the second local image is locally redrawn using the latent diffusion model to obtain a processed third candidate composite image, which is then used as the first candidate composite image.

8. The method according to claim 7, wherein, The step of extracting the second local image to be redrawn from the initial photographic image includes: The initial photographic image is segmented into redrawable regions using a pre-acquired image region segmentation algorithm to obtain the second local image to be redrawn in the initial photographic image.

9. The method according to claim 7, wherein, The method further includes: The third candidate composite image is subjected to high-definition restoration processing to obtain the target composite image of the initial photographic image.

10. The method according to claim 1, wherein, Before obtaining the conversion intervention information of the initial photographic image according to the image conversion requirements, the process includes: The initial photographic image is subjected to resolution compression, and the transformation reference information in the initial photographic image is extracted. The transformation reference information includes at least the subject pose information, image layer information, and contour information of the initial photographic image.

11. The method according to any one of claims 1-10, wherein, The method further includes: Based on the target synthesized image, an updated style transfer image for the user terminal is constructed, and the updated style transfer image is stored in a reference style transfer image library.

12. The method according to any one of claims 1-10, wherein, The method further includes: In response to receiving a cloud save command, the target composite image is uploaded to the cloud and stored.

13. A client for synthesizing images, characterized in that, The client includes: an image upload module, a reference style conversion image library module, a prompt word module, and an image synthesis icon, wherein, The image upload module is used to provide an upload port for the user terminal to capture the initial image, and to transmit the initial image uploaded by the user terminal to the server. The reference style transfer image library module is used to display each reference style transfer image in the reference style transfer image library to the user terminal, wherein, in response to the user terminal selecting the style transfer image of the initial photographic image from the reference style transfer image library, the style transfer image is sent to the server; The prompt word module is used to provide the user terminal with a prompt word option list, wherein, in response to the user terminal selecting a style conversion prompt word or a local redraw prompt word for the initial photographed image from the prompt word option list, the style conversion prompt word or the local redraw prompt word is sent to the server; The image compositing icon is used to provide a compositing instruction input port for the user terminal. In response to the user terminal clicking the image compositing icon, it is determined that the user terminal has input a confirmation compositing instruction, and the confirmation compositing instruction is sent to the server. Based on the received confirmation compositing instruction and conversion intervention information, the server begins conversion processing of the initial photographic image using a latent diffusion model to obtain candidate compositing images, and then obtains the target compositing image based on the candidate compositing images. This includes: acquiring reference information for each element carried in the image conversion requirement, and comparing the information of each element in the candidate compositing image with the corresponding reference information to identify whether the information of each element in the candidate compositing image matches the reference information. The client also includes a list of options for global transformation requirements and local redraw requirements, wherein... In response to the user clicking the global conversion request, the global conversion request for the photographic image is determined to be the image conversion request for the initial photographic image, and sent to the server. The server obtains the style conversion image selected by the user and / or the style conversion prompt word selected by the user and / or the style conversion description text input by the user to obtain the global conversion intervention information for the initial photographic image. In response to the user clicking the local redraw request, the local redraw request of the photographic image is determined to be the image transformation request of the initial photographic image, and sent to the server; the server obtains the local redraw text and / or the local redraw prompt word selected by the user to obtain the local transformation intervention information of the initial photographic image.

14. The client according to claim 13, wherein, The client also includes a high-definition processing component for performing high-definition restoration processing on the candidate synthetic image to obtain the processed target synthetic image.

15. The client according to claim 13, wherein, The client also includes a cloud save icon, wherein, in response to the user clicking the cloud save icon, it is determined that the user has entered a cloud save command; The server uploads and stores the target composite image to the cloud based on the cloud save instruction.

16. An apparatus for acquiring a composite image, characterized in that, The device includes: The first acquisition module is used to receive the initial photographed image input by the user terminal and acquire the image conversion requirements of the initial photographed image; The second acquisition module is used to acquire the conversion intervention information of the initial photographic image according to the image conversion requirements; The first processing module is used to perform image conversion processing on the initial photographic image based on the conversion intervention information and through a potential diffusion model to obtain the processed first candidate synthetic image. The second processing module is configured to, in response to the first candidate synthesized image satisfying the image conversion requirement, obtain a target synthesized image of the initial photographic image based on the first candidate synthesized image, wherein the processing module includes: obtaining reference information of each element carried in the image conversion requirement, and comparing the information of each element in the first candidate synthesized image with the corresponding reference information to identify whether the information of each element in the first candidate synthesized image matches the reference information. The second acquisition module is further configured to: In response to the image conversion requirement being a global conversion requirement for photographic images, the style conversion image selected by the user terminal and / or the style conversion prompt word selected by the user terminal and / or the style conversion description text input by the user terminal are obtained to obtain the global conversion intervention information of the initial photographic image; In response to the image conversion request being a local redrawing request for a photographic image, local conversion intervention information for the initial photographic image is obtained from the local redrawing text input by the user and / or the local redrawing prompt word selected by the user.

17. The apparatus according to claim 16, wherein, The second acquisition module is further configured to: In response to the image conversion request being a global conversion request for the photographic image, first style conversion information is extracted from the style conversion image; and / or, Extract the second style transfer information of the initial photographic image from the style transfer prompts; and / or, A natural language parsing algorithm is used to extract third style transfer information from the style transfer description text by generating content using artificial intelligence; Based on the first style conversion information and / or the second style conversion information and / or the third style conversion information, obtain the global conversion intervention information of the initial photographic image under the global conversion requirement of the photographic image.

18. The apparatus according to claim 17, wherein, The first processing module is further configured to: Extract the main subject from the initial photographic image; Based on the global transformation intervention information, the initial photographic image is processed using the potential diffusion model based on the photographic subject to obtain the first candidate synthetic image.

19. The apparatus according to claim 16, wherein, The second processing module is further configured to: In response to the first candidate synthesized image not meeting the image conversion requirements, a first local image to be fused is cropped from the first candidate synthesized image using a pre-acquired image region segmentation algorithm; Obtain the first local intervention information of the first local image; A character fusion enhancement algorithm is obtained, and the first local image is fused and intervened based on the first local intervention information using the character fusion enhancement algorithm to obtain a processed second candidate composite image. The target composite image of the initial photographic image is obtained based on the second candidate composite image.

20. The apparatus according to claim 19, wherein, The second processing module is further configured to: The first candidate composite image or the second candidate composite image is subjected to high-definition restoration processing to obtain the processed target composite image.

21. The apparatus according to claim 16, wherein, The second acquisition module is further configured to: In response to the image conversion requirement being a local redrawing requirement for the photographic image, a first local redrawing information is extracted from the local redrawing text using a pre-acquired natural language parsing algorithm; and / or, Extract the second local redraw information from the local redraw prompts; The local transformation intervention information of the initial photographic image is obtained based on the first local redrawing information and / or the second local redrawing information.

22. The apparatus according to claim 21, wherein, The first processing module is further configured to: Extract the second local image to be redrawn from the initial photographic image; Based on the local transformation intervention information, the second local image is locally redrawn using the latent diffusion model to obtain a processed third candidate composite image, which is then used as the first candidate composite image.

23. The apparatus according to claim 22, wherein, The first processing module is further configured to: The initial photographic image is segmented into redrawable regions using a pre-acquired image region segmentation algorithm to obtain the second local image to be redrawn in the initial photographic image.

24. The apparatus according to claim 22, wherein, The second processing module is further configured to: The third candidate composite image is subjected to high-definition restoration processing to obtain the target composite image of the initial photographic image.

25. The apparatus according to claim 16, wherein, The first acquisition module is further configured to: The initial photographic image is subjected to resolution compression, and the transformation reference information in the initial photographic image is extracted. The transformation reference information includes at least the subject pose information, image layer information, and contour information of the initial photographic image.

26. The apparatus according to any one of claims 16-25, wherein, The device further includes an update module for: Based on the target synthesized image, an updated style transfer image for the user terminal is constructed, and the updated style transfer image is stored in a reference style transfer image library.

27. The apparatus according to any one of claims 16-25, wherein, The device further includes a storage module for: In response to receiving a cloud save command, the target composite image is uploaded to the cloud and stored.

28. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

29. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.

30. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Image style conversion method and device, electronic equipment and storage medium

    CN117237185A

  • Figure image reloading method and device, storage medium and computer equipment

    CN117670656A

  • Apparatus and method for changing image style

    KR1020240033917A