Image processing method and terminal equipment

By acquiring and encoding images and text feature vectors on the terminal device and fusing these feature vectors to generate synthetic images, the problem of unnatural synthetic images in the prior art is solved, and more efficient user operations and personalized needs are achieved.

CN120070167APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311627926.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art requires a lot of manual interaction and adjustment when synthesising images on terminal devices, resulting in stiff characters in the composite image, mismatched styles and unnatural effects.

Method used

By obtaining the two images and descriptive text selected by the user, encoded images and text to obtain feature vectors, fuse feature vectors to generate synthetic images, simplify user operations and improve synthesis effect and authenticity.

Benefits of technology

It realizes the naturalness of synthetic images, simplifies user operations, meets user personalized needs, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070167A_ABST
    Figure CN120070167A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and terminal equipment, relates to the technical field of terminals, and aims to solve the problem that an image synthesis effect is not natural. The method comprises the steps that the terminal equipment acquires a first image, a second image and a first text; the terminal equipment encodes the first image, the second image and the first text to obtain a first image feature vector, a second image feature vector and a first text feature vector; the terminal equipment fuses the first image feature vector, the second image feature vector and the first text feature vector to obtain a first fusion feature vector; and the terminal device displays the first composite image according to the first fusion feature vector. The method is applied to an image processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of terminals, and in particular, to an image processing method and a terminal device. Background Art

[0002] With the rapid development of social networks and image processing technologies, users can synthesize images of multiple different people into one image in the form of a virtual group photo to meet the personalized needs of users for group photos. For example, users can synthesize their own photos with those of their family members who are in different places into a family portrait. Or, users can synthesize their own photos with those of celebrity idols to fulfill their star-chasing needs.

[0003] Currently, when users synthesize images on a terminal device, they need to perform a large amount of manual interaction and adjustment on the input images to obtain a relatively satisfactory synthesized image. Moreover, only simple splicing and fusion are performed, and the states of the people in the synthesized image (such as poses, expressions, etc.) are the same as those in the input images, resulting in problems such as stiffness of the people in the synthesized image, mismatch between the people and the overall style of the synthesized image, and unnatural synthesis effects.

[0004] Therefore, how to optimize image synthesis in a terminal device has become a technical problem that urgently needs to be solved. Summary of the Invention

[0005] The present application provides an image processing method and a terminal device, which can optimize the image processing strategy, make the synthesized image more natural, and effectively improve the image synthesis effect and authenticity.

[0006] To achieve the above technical objectives, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides an image processing method, which includes: the terminal device obtains a first image, a second image, and a first text; the terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector; the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fused feature vector; the terminal device displays a first synthesized image according to the first fused feature vector.

[0008] According to the first aspect, the first image and the second image are the images to be synthesized selected by the user, the first image includes a first target object, the second image includes a second target object, and the first text is the text information input by the user for describing the first target object and the second target object; the first image feature vector is the numericalized feature vector corresponding to the first image; the second image feature vector is the numericalized feature vector corresponding to the second image.

[0009] In this method, the terminal device only needs the user to provide a first image and a second image, and describe the relevant information of the first target object in the first image and the second target object in the second image in the first text, then the terminal device can intelligently generate the synthetic image expected by the user. That is to say, the text is used to define the synthesis effect of the synthetic image (such as expression, stance, environment, etc.), and the image is used to provide the key information in the synthetic image (such as portrait facial features information, limb information, etc.). During the process of generating the synthetic image, the image and the text are encoded into corresponding feature vectors so that various feature vectors can be effectively fused to improve the accuracy of the fusion result. Based on the fused feature vectors, image synthesis is performed to obtain the first synthetic image, making the synthesized image more natural, effectively improving the synthesis effect and authenticity, and optimizing the image synthesis strategy. Moreover, there is no need for the user to interact frequently with the terminal device, which simplifies the user operation, can also meet the user's personalized needs, and improves the user experience.

[0010] In some examples, the first image and the second image each include a target object, and the target object includes various entities, including living things (such as people, animals, plants, etc.), objects (such as buildings, vehicles, etc.) and / or natural landscapes, etc.

[0011] In some examples, if the target object in the image is a portrait, the portrait should be an image that includes the front face of the portrait or the face is deflected by no more than a preset angle, and the portrait can be a half-length portrait or a full-length portrait. Among them, the preset angle is the maximum threshold value of the face deflection determined in advance through experiments but still able to recognize the facial features of the face.

[0012] In this way, the facial features of the face can still be detected, accurate face information and portrait information can be obtained, the consistency of the portrait identity can be guaranteed, and distortion and distortion during the portrait synthesis process can be avoided.

[0013] In some examples, the terminal device provides an image synthesis function, and the user can start the image synthesis function to perform image synthesis.

[0014] In some examples, the terminal device obtains the first image, the second image and the first text, including: the terminal device starts the image synthesis function, the terminal device prompts the user to select the images to be synthesized and input the text, and the terminal device obtains the first image, the second image and the first text in response to the user operation.

[0015] In some examples, the terminal device obtains the first image, the second image and the first text, including: the terminal device obtains the first image and the second image in response to the user operation, the terminal device starts the image synthesis function, the terminal device prompts the user to input the text, and obtains the first text.

[0016] In some examples, when the terminal device acquires the first image and the second image, the terminal device determines the target object in each image in response to a user operation and labels each target object.

[0017] Exemplarily, if only one object is included in the first image, the terminal device determines this object as the target object in the first image, that is, the first target object. Correspondingly, if only one object is included in the second image, the terminal device determines this object as the target object in the second image, that is, the second target object.

[0018] Exemplarily, if multiple objects are included in the first image, the terminal device determines at least one object actively specified by the user as the target object in the first image, that is, the first target object. Correspondingly, if multiple objects are included in the second image, the terminal device determines at least one object actively specified by the user as the target object in the second image, that is, the second target object.

[0019] Exemplarily, if multiple objects are included in the first image, the terminal device prompts the user to select the target object, and the terminal device determines the target object in the first image in response to the user operation, that is, the first target object. Correspondingly, if multiple objects are included in the second image, the terminal device prompts the user to select the target object, and the terminal device determines the target object in the second image in response to the user operation, that is, the second target object.

[0020] Exemplarily, when the terminal device determines the target object in each image, the terminal device labels each target object in response to the user's operation of actively labeling each target object.

[0021] Exemplarily, when the terminal device determines the target object in each image, the terminal device actively labels each target object.

[0022] According to the first aspect, or any implementation manner of the above first aspect, the terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector, including: the terminal device inputs the first image and the second image into an image encoding network to obtain a first image feature vector and a second image feature vector; the terminal device inputs the first text into a text encoding network to obtain a first text feature vector.

[0023] Wherein, the image encoding network is a neural network trained with sample images for encoding image information; the text encoding network is a neural network trained with sample texts for encoding text information.

[0024] Exemplarily, the image encoding network is a trained convolutional neural network for encoding; the text encoding network is a trained long short-term memory network for encoding.

[0025] In this way, when encoding an image into a feature vector, the feature dimensions of images with any resolution are fixed and unified, and the encoded image features are more standardized; when encoding text into a feature vector, the feature dimensions of text with any length are fixed and unified, and the encoded text features are more standardized. Moreover, the fixed feature dimensions are also beneficial to feature fusion and improve the accuracy of the fused features.

[0026] According to the first aspect, or any implementation manner of the above first aspect, the terminal device inputs the first image and the second image into an image encoding network to obtain a first image feature vector and a second image feature vector, including: the terminal device preprocesses the first image and the second image; the terminal device inputs the preprocessed first image and the preprocessed second image into the image encoding network to obtain a first image feature vector and a second image feature vector.

[0027] According to the first aspect, or any implementation manner of the above first aspect, the terminal device preprocesses the first image and the second image, including: the terminal device segments a first target object in the first image according to a preset segmentation algorithm; the terminal device segments a second target object in the second image according to a preset segmentation algorithm.

[0028] In some examples, the terminal device segments a first target object in the first image according to a preset segmentation algorithm, extracts pixel information of the first target object, and removes pixel information irrelevant to the first target object; the terminal device segments a second target object in the second image according to a preset segmentation algorithm, extracts pixel information of the second target object, and removes pixel information irrelevant to the second target object.

[0029] Exemplarily, the preset segmentation algorithms include traditional threshold segmentation, edge detection-based segmentation, region-based segmentation, and deep learning-based semantic segmentation algorithms.

[0030] In this way, by segmenting the target object (key information) according to the preset segmentation algorithm and performing encoding processing on it, the encoding efficiency and accuracy of the main features of the target object can be improved.

[0031] According to the first aspect, or any implementation manner of the above first aspect, the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fused feature vector, including: the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector according to a preset fusion method to obtain a first fused feature vector.

[0032] In some examples, the terminal device combines the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fused feature vector.

[0033] In some examples, the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector by using weighted average to obtain the first fused feature vector.

[0034] In some examples, the terminal device couples the first image feature vector and the second image feature vector into the first text feature vector to obtain the first fused feature vector.

[0035] In this way, the fused feature vector includes text and image information, which is beneficial to accurately synthesize the desired synthesis effect and meet the personalized needs of users.

[0036] According to the first aspect, or any one of the above implementation manners of the first aspect, the terminal device displays the first synthesized image according to the first fused feature vector, including: the terminal device decodes and reconstructs the first fused feature vector through the StableDiffusion technology to generate the first synthesized image; the terminal device displays the first synthesized image.

[0037] According to the first aspect, or any one of the above implementation manners of the first aspect, the method further includes: the terminal device updates the first synthesized image in response to a user's update operation.

[0038] Among them, the update operation includes an operation in which the user triggers the terminal device to generate a synthesized image again, or an operation in which the user updates the first text.

[0039] According to the first aspect, or any one of the above implementation manners of the first aspect, the terminal device updates the first synthesized image in response to a user's update operation, including: the terminal device responds to an operation in which the user triggers the terminal device to generate a synthesized image again, synthesizes an image again according to the first image, the second image, and the first text, and updates the first synthesized image to the second synthesized image; the second synthesized image is different from the first synthesized image.

[0040] According to the first aspect, or any one of the above implementation manners of the first aspect, the terminal device updates the first synthesized image in response to a user's update operation, including: the terminal device updates the first text to the second text in response to a user's operation of updating the first text; the terminal device performs image synthesis according to the first image, the second image, and the second file to generate a third synthesized image; the terminal device updates the first synthesized image to the third synthesized image; the third synthesized image is different from the first synthesized image.

[0041] In some examples, based on the first synthesized image, the terminal device obtains a third text, and the terminal device generates a fourth synthesized image according to the third text and the first synthesized image; the fourth synthesized image is different from the first synthesized image.

[0042] According to the first aspect, or any implementation of the above first aspect, the first text includes one or more of the following: the interpersonal relationship, location, expression, body movement, and / or facial expression between the first target object and the second target object.

[0043] In some examples, the first text may further include the display effect that the user expects the synthesized image to present.

[0044] Exemplarily, the interpersonal relationship includes social relationships established in human interactions such as kinship, friendship, romantic relationship, classmate relationship, teacher-student relationship, employment relationship, comrade-in-arms relationship, or colleague relationship.

[0045] In this way, image synthesis based on more detailed text information can improve the synthesis effect and meet the personalized needs of users.

[0046] In a second aspect, the present application provides a terminal device, which includes: one or more processors; a memory; wherein, one or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the terminal device, the terminal device is caused to perform: the terminal device obtains a first image, a second image, and a first text; the terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector; the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fused feature vector; the terminal device displays a first synthesized image according to the first fused feature vector.

[0047] In a third aspect, the present application provides a chip system, which includes at least one processor and at least one interface circuit. The at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor. When the at least one processor executes the instructions, the at least one processor executes the method according to the above first aspect or any implementation of the first aspect.

[0048] In a fourth aspect, the present application provides a computer-readable storage medium, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the method according to the above first aspect or any implementation of the first aspect.

[0049] In a fifth aspect, the present application provides a computer program product, which includes: a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the method according to the above first aspect or any implementation of the first aspect.

[0050] It should be noted that for the technical effects brought by any of the designs in the second to fifth aspects above, reference can be made to the technical effects brought by the corresponding designs in the first aspect, and details will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 FIG.

[0052] Figure 2 is a schematic flowchart of image synthesis based on the correspondence between a person and height information provided by an embodiment of the present application;

[0053] Figure 3 is a schematic flowchart of image synthesis based on user control provided by an embodiment of the present application;

[0054] Figure 4 is a synthesized image with distortion provided by an embodiment of the present application;

[0055] Figure 5 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application Figure 1 ;

[0056] Figure 6 is a schematic diagram of the architecture of an image processing system provided by an embodiment of the present application;

[0057] Figure 7 is a schematic flowchart of an image processing method provided by an embodiment of the present application;

[0058] Figure 8 is a schematic diagram of a scenario example of an image processing method provided by an embodiment of the present application Figure 1 ;

[0059] Figure 9 is a schematic diagram of a scenario example of an image processing method provided by an embodiment of the present application Figure 2 ;

[0060] Figure 10 is a schematic diagram of a scenario example of an image processing method provided by an embodiment of the present application Figure 3 ;

[0061] Figure 11 is an interaction flowchart of each module in an image processing system provided by an embodiment of the present application;

[0062] Figure 12 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application Figure 2 ;

[0063] Figure 13 is a schematic diagram of the structure of a chip system provided by an embodiment of the present application. Detailed implementation manners

[0064] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.

[0065] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0066] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0067] To better understand the technical solution of the present application, the technical terms involved in the present application will be described first.

[0068] Image segmentation technology refers to the technology and process of dividing an image into several specific regions with unique properties and extracting the target of interest. Among them, the unique properties can be understood as edges, grayscales, colors, textures, etc.

[0069] Stable Diffusion (SD) is a steady-state diffusion model, a machine learning algorithm, applied in the field of image generation. By learning the inverse diffusion process to train a neural network, the "image generation" process is converted into a "diffusion" process of gradually removing noise. Specifically, the whole process starts from random Gaussian noise, gradually removes noise through training until there is no noise, and finally outputs an image closer to the text description. Or, denoise the image superimposed with Gaussian noise.

[0070] Feature fusion refers to combining the features extracted from an image into features with stronger discriminative ability than the extracted features. That is, different feature vectors are extracted from the same pattern and optimized and combined to achieve a more accurate and complete representation form.

[0071] In some examples, such as Figure 1As shown, the terminal device completes image synthesis based on various professional image synthesis software. In the face of users' group photo needs, various professional image synthesis software (which can also be described as applications (APPs)) have emerged like mushrooms after a spring rain. According to the characteristics of different terminal devices (such as mobile phones or computers, etc.), each type of terminal device has corresponding image synthesis software (such as mobile phone image synthesis software or computer image synthesis software, etc.). These software are diverse in types and have rich image synthesis functions, but the underlying technical principles are basically the same.

[0072] As Figure 1 Shown in (a) of [reference], it is a schematic flowchart of image synthesis by various image synthesis software in the terminal device. First, the user uses the image synthesis software to crop the target portrait from the original image. For example, the image synthesis software in the terminal device automatically crops the target portrait using the intelligent selection function based on the image segmentation algorithm, or the image synthesis software in the terminal device crops the target portrait through the cropping tool in response to the user's operation. Then the user drags the target portrait to the appropriate position in the target image and performs fusion. For example, the user drags the target portrait to the target position in the target image through tools such as moving and scaling in the image synthesis software, and fuses the target portrait to the target position in the target image based on the layer fusion technology. Subsequently, the user makes fine adjustments to the details of the fused image. For example, the user adjusts the person through techniques such as edge optimization and shadow adjustment to make it blend more naturally with the background. Finally, the user performs color consistency adjustment on the adjusted image to obtain the synthesized image. For example, the color balance, saturation, etc. are adjusted through the color harmony algorithm to make the tone of the synthesized image consistent.

[0073] Exemplarily, as Figure 1 Shown in (b) of [reference], it is a schematic diagram of the scene of image synthesis by the image synthesis software in the mobile phone. The user crops the original image through the cropping tool on the mobile phone to obtain the target portrait (the portrait of the boy) within the dashed box. Then the user drags the cropped target portrait to the right side of the girl in the target image and performs fusion to obtain a fused image with the girl on the left and the boy on the right. Subsequently, the user makes detail adjustments such as brightness and clarity to the fused image to make the edges of the portrait (such as the intersection of the shoulders of the boy and the girl) blend more naturally. Finally, color adjustments such as saturation, color temperature, and hue are performed, and filters ( Figure 1 not shown in (b) of [reference]) can also be used to make the synthesized image look like it was taken in the same environment.

[0074] However, in the above examples, in each step of image synthesis, the user needs to make corresponding adjustments. That is, a large amount of manual interaction is required to synthesize a relatively satisfactory synthesized image. Moreover, the setting process in each step of image synthesis is complex, the operation threshold is high, the synthesis time is long, the synthesis effect is not realistic, and it is not suitable for new users, which is not user-friendly.

[0075] In some other examples, in order to improve the authenticity of the synthesis effect, the corresponding relationship between the person and the height information is pre-stored, and the height ratio between the target persons is determined according to the pre-stored corresponding relationship when synthesizing the image, and the synthesized image is generated according to the height ratio. Specifically, as Figure 2 shown in (a) of [reference], it is a schematic flowchart of image synthesis based on the corresponding relationship between the person and the height information. Determine the target persons according to the group photo request instruction; determine the height ratio between the target persons according to the pre-stored corresponding relationship between the person and the height information. For example, determine the height ratio through human key point matching, and discard the body parts that affect the beauty of the group photo through the human segmentation algorithm; adjust the target persons according to the height ratio, and generate the synthesized image based on the image stitching algorithm.

[0076] Exemplarily, as Figure 2 shown in (b) of [reference], the user uploads the images of the target persons to be synthesized, and the system confirms that the target persons are user A and user B respectively according to the uploaded images. The system then determines the height ratio between user A and user B according to the uploaded images of the user and the pre-stored corresponding relationship between the person and the height information, adjusts the target persons according to the height ratio, such as enlarging user B and discarding the body parts within the dotted line frame of user B. Finally, synthesize the adjusted target persons to obtain the synthesized image, so as to ensure the authenticity and beauty of the group photo.

[0077] However, in the above examples, it is necessary to pre-acquire and store the height information of the person (the height information in the standing posture), and the posture of the person in the image uploaded by the user should be the standing posture. Only in this way can the height ratio between the target persons be calculated according to the pre-set height information. Moreover, using a simple image stitching algorithm to synthesize the image cannot adaptively adjust the expressions, actions and expressions of the persons according to the group photo background and the intimacy relationship of the persons. That is, only the height ratio can be guaranteed to be real, but the authenticity of the actions and expressions of the persons in the synthesized image and the overall style of the synthesized image cannot be guaranteed, the synthesis effect is not natural, and the personalized needs of the users cannot be met.

[0078] In some other examples, such as Figure 3As shown in the figure, it is a schematic flowchart of image synthesis based on user control. The user selects a target area in the input image. For example, if the input image only includes a person, the target area is the area within the dashed box. The target object (such as a cloud) is obtained and synthesized into the target area to obtain the synthesized image 1. The user can also update the target object according to the matching degree between the target object and the synthesized image in the synthesized image (such as updating the cloud to the sun) until the synthesized image 2 expected by the user is obtained.

[0079] However, in the above example, the posture of the target object in the synthesized image is fixed, and the user cannot customize and adjust the actions and expressions of the target object, which cannot meet the personalized needs of the user.

[0080] In some other examples, image synthesis is performed through SD technology. The user inputs pictures and / or text descriptions, and the system performs image synthesis according to the content input by the user. However, the text descriptions input by the user are usually used to define the overall environment of the synthesized image. For example, the user inputs two images, and the text description is to synthesize the two images into a family portrait. The system obtains the synthesized image as shown in Figure 4 However, in the above example, SD only performs image synthesis by simple application, which easily causes distortion of the human face in the synthesized image. That is, the expressions of the people in the synthesized image are strange, the facial features are distorted, or the appearance of the people in the synthesized image is quite different from that of the people in the input image, and it is impossible to recognize that they are the same person, resulting in the loss of the identity attributes of the people. As shown in Figure 4 In the figure, the human face in the dashed box is distorted and the nose is crooked and the mouth is askew.

[0081] In order to improve the above-mentioned technical problems, the embodiments of the present application provide an image processing method. The terminal device obtains a first image, a second image, and a first text. The terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector. The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector. The terminal device displays a first synthesized image according to the first fusion feature vector. The technical solution of the present application performs image synthesis based on the encoded feature vectors, optimizes the image synthesis strategy, the synthesized image is more natural, effectively improves the synthesis effect and authenticity, simplifies the user operation, meets the personalized needs of the user, and improves the user experience.

[0082] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0083] In some embodiments of the present application, the terminal device may include, but is not limited to, a smart phone, a netbook, a tablet computer, a smart drawing board, a graphics tablet, a smart watch, a smart bracelet, a phone watch, smart glasses, a smart camera, a handheld computer, an in-vehicle computer, a personal computer (PC), a personal digital assistant (PDA), a portable multimedia player (PMP), an augmented reality (AR) / virtual reality (VR) device, a smart TV, a projection device, or a body-sensing game console in a human-computer interaction scenario, etc. Alternatively, the terminal device may also be a terminal device of other types or structures, which is not limited in the present application.

[0084] As an example, please refer to Figure 5 , taking the terminal device as a mobile phone as an example, Figure 5 which shows a schematic diagram of the hardware structure of a terminal device provided by an embodiment of the present application.

[0085] The terminal device may include a processor 510, an external memory interface 520, an internal memory 521, a universal serial bus (USB) interface 530, a charging management module 540, a power management module 541, a battery 542, an antenna 1, an antenna 2, a mobile communication module 550, a wireless communication module 560, an audio module 570, a sensor module 580, a button 590, a motor 591, an indicator 592, a camera 593, a display screen 594, and a subscriber identification module (SIM) card interface 595, etc.

[0086] The processor 510 may include one or more processing units. For example, the processor 510 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0087] A memory may also be provided in the processor 510 for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can save the instructions or data that the processor 510 has just used or recycled. If the processor 510 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.

[0088] The external memory interface 520 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the terminal device.

[0089] The internal memory 521 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 521 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the terminal device (such as audio data, phone book, etc.). The processor 510 executes various functional applications and data processing of the terminal device by running the instructions stored in the internal memory 521 and / or the instructions stored in the memory provided in the processor.

[0090] In some embodiments, the processor 510 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, and / or a universal serial bus (USB) interface, etc.

[0091] The charging management module 540 is configured to receive a charging input from a charger. The power management module 541 is used to connect the battery 542, the charging management module 540, and the processor 510. The power management module 541 receives inputs from the battery 542 and / or the charging management module 540 to power the processor 510, the internal memory 521, the display screen 594, the camera 593, the wireless communication module 560, etc.

[0092] The wireless communication function of the terminal device can be implemented by antenna 1, antenna 2, the mobile communication module 550, the wireless communication module 560, the modulation and demodulation processor, and the baseband processor, etc.

[0093] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0094] The mobile communication module 550 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device. The mobile communication module 550 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.

[0095] The wireless communication module 560 can provide solutions for wireless communications applied to the terminal device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc.

[0096] In some embodiments, antenna 1 of the terminal device is coupled to the mobile communication module 550, and antenna 2 is coupled to the wireless communication module 560, enabling the terminal device to communicate with the network and other devices through wireless communication technologies.

[0097] The audio module 570 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 570 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 570 can be disposed in the processor 510, or some functional modules of the audio module 570 can be disposed in the processor 510.

[0098] The sensor module 580 can include one or more of a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0099] The keys 590 include a power-on key, volume keys, etc. The keys 590 can be mechanical keys or touch keys. The terminal device can receive key inputs and generate key signal inputs related to the user settings and function controls of the terminal device.

[0100] The motor 591 can generate a vibration prompt.

[0101] The indicator 592 can be an indicator light, which can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0102] The camera 593 is used to capture still images or videos. In some embodiments, the terminal device can include one or N cameras 593, where N is a positive integer greater than 1.

[0103] The display screen 594 is used to display images, videos, etc. The display screen 594 includes a display panel. In some embodiments, the terminal device may include one or N display screens 594, where N is a positive integer greater than 1. In some embodiments of the present application, the display screen 594 can be used to display the generated composite image.

[0104] The terminal device realizes the display function through the GPU, the display screen 594, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 594 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 510 may include one or more GPUs, which execute program instructions to generate or change display information.

[0105] The SIM card interface 595 is used to connect the SIM card.

[0106] It can be understood that the above Figure 5 merely takes a mobile phone as an example to illustrate the structure of the terminal device in the embodiments of the present application, but does not constitute a specific limitation on the structure and form of the terminal device. In other embodiments of the present application, the terminal device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0107] The image processing method provided by the embodiments of the present application can be applied to any image synthesis scenario. Optionally, the image synthesis scenario includes image synthesis (such as synthesizing multiple images into one image), image repair and reconstruction (such as fusing multiple photos to repair damaged old photos), virtual reality enhancement (such as fusing virtual elements with the real scene into one image), film and television special effects, and other scenarios. The embodiments of the present application do not impose special restrictions on the specific form of the image synthesis scenario.

[0108] For ease of understanding, Figure 6 a schematic diagram of the architecture of an image processing system applicable to the image processing method provided by the embodiments of the present application is shown. The image processing system includes a terminal device 601.

[0109] In the embodiments of the present application, the terminal device 601 includes an image encoding module, a text encoding module, a feature fusion module, and a decoding and synthesis module.

[0110] Specifically, the image encoding module can be used to obtain the images uploaded by the user, such as the first image and the second image described below. The image encoding module can also be used to encode the obtained images and convert them into image feature vectors, such as the first image feature vector and the second image feature vector described below. The image encoding module can also be used to send the converted image feature vectors to the feature fusion module.

[0111] The text encoding module can be used to obtain the text input by the user. The text encoding module can also be used to encode the input text and convert it into a text feature vector. The text encoding module can also be used to send the converted text feature vector to the feature fusion module. Among them, the text input by the user includes description information such as the position, expression, interpersonal relationship, action, and / or expression of the person in the synthetic image.

[0112] The feature fusion module can be used to perform feature fusion on the encoded and converted image feature vector and text feature vector to obtain a new fused feature vector. Exemplarily, the portrait feature vector in the image feature vector is coupled into the text feature vector to obtain the fused feature vector. The feature fusion module can also be used to send the fused feature vector to the decoding and synthesis module.

[0113] The decoding and synthesis module can be used to decode and synthesize the fused feature vector according to the SD algorithm to reconstruct the synthetic image. The decoding and synthesis module is also used to output the synthetic image so that the user can obtain the synthetic image.

[0114] It can be understood that the above Figure 6 The shown system architecture diagram is only an example. In practical applications, the terminal device 601 may include more or fewer modules. The embodiments of the present application do not limit the module division in the terminal device. For example, the terminal device 601 may also include an input module and a display module. Among them, the input module can be used to input images and texts. The display module can be used to display synthetic images, etc.

[0115] It can be understood that in the embodiments of the present application, the terminal device can execute some or all of the steps in the embodiments of the present application. These steps or operations are only examples. The embodiments of the present application can also execute other operations or various deformations of the operations. In addition, each step can be executed in a different order presented in the embodiments of the present application, and it is possible not to execute all the operations in the embodiments of the present application.

[0116] Exemplarily, the technical solutions involved in the following embodiments can all be implemented in the above terminal device. The following describes the image processing method provided by the embodiments of the present application in detail with reference to the drawings and application scenarios.

[0117] The following takes the application of the embodiments of the present application in Figure 6 the shown image processing system as an example for illustration.

[0118] In the embodiments of the present application, as Figure 7 shown, it is a schematic flowchart of an image processing method provided by the embodiments of the present application. The method includes:

[0119] S701. The terminal device obtains a first image, a second image, and a first text.

[0120] In the embodiments of the present application, the first image and the second image are the to-be-synthesized images selected by the user. The first image and the second image respectively include target objects, and the target objects include various entities, including living things (such as people, animals, plants, etc.), objects (such as buildings, vehicles, etc.) and / or natural landscapes, etc.

[0121] In some examples, the first image and / or the second image may be images already stored in the terminal device. For example, the first image and the second image are images in the album. Or, the first image and the second image are images downloaded by the user on various application programs (such as browsers, social application programs, and / or communication application programs, etc.).

[0122] In other examples, the first image and / or the second image may also be images uploaded by the user to the terminal device. For example, the user transmits the first image and the second image to the terminal device, and the terminal device saves the first image and the second image first for subsequent synthesis.

[0123] In other examples, the first image and / or the second image may also be images taken by the user on the terminal device. For example, the user takes the first image and the second image through the camera, and the terminal device stores the first image and the second image for subsequent synthesis.

[0124] It can be understood that the embodiments of the present application do not limit the specific implementation manner for the terminal device to obtain the to-be-synthesized images.

[0125] It can be understood that if the target object in the image is a portrait, the portrait should be an image including the front face of the portrait or the face deflection not exceeding a preset angle, and the portrait can be a half-body portrait or a full-body portrait.

[0126] Wherein, the preset angle is the maximum threshold value of the face deflection determined in advance through experiments but still able to recognize the facial features. For example, 10 degrees. The face in the image deflects to the left or right, but the deflection angle does not exceed ten degrees.

[0127] In this way, the facial features can still be detected, accurate facial information and portrait information can be obtained, the consistency of the portrait identity can be ensured, and distortion and distortion during the portrait synthesis process can be avoided.

[0128] In the embodiments of the present application, the first text is text information input by the user describing the interpersonal relationship, position, expression, body movement, and / or expression, etc. between the target object in the first image (which can also be described as the first target object) and the target object in the second image (which can also be described as the second target object).

[0129] Among them, interpersonal relationships include social relationships established in human interactions such as kinship, friendship, romantic relationships, classmate relationships, teacher-student relationships, employment relationships, comrade-in-arms relationships, or colleague relationships.

[0130] In the embodiments of the present application, the first text may further include the display effects that the user expects the synthesized image to present. Among them, the display effects may include the scene of the synthesized image, the background of the synthesized image, the atmosphere of the synthesized image, etc. For example, the background of the synthesized image is a classroom or a certain natural scenery.

[0131] It can be understood that compared with the input text of SD in the prior art, the first text in the embodiments of the present application is more detailed. The input text in the prior art can only limit the overall environment of the synthesized image, while the first text in the present application can limit information such as the relationship, position, expression, body movements, and / or facial expressions between target objects in the synthesized image. Image synthesis based on more detailed text information can improve the synthesis effect and meet the personalized needs of users.

[0132] In some embodiments of the present application, the terminal device provides an image synthesis function, and the user can perform image synthesis by enabling the image synthesis function.

[0133] In a possible implementation manner, the image synthesis function can be enabled by the user in the photo album or camera. For example, the user clicks on the synthesis control in the photo album of the terminal device to enable the image synthesis function.

[0134] In another possible implementation manner, the image synthesis function can also be enabled by the user by opening an application program in the terminal device that includes the image synthesis function. Thus, the image synthesis function of the terminal device is enabled.

[0135] Moreover, according to different types of terminal devices, the operation for the user to enable the image synthesis function can be achieved in different ways. For example, when the device type is a touch-operable terminal device such as a mobile phone or a tablet, the operation for the user to enable the image synthesis function may include touch operations such as clicking and swiping by the user's finger on the photo album, camera, or application program that includes the image synthesis function of the terminal device. When the device type is a large-screen device, such as a smart screen, the operation for the user to enable the image synthesis function may include the user using a remote control to send instructions to the large-screen device to perform operations such as selection and confirmation on the photo album, camera, or application program that includes the image synthesis function of the terminal device. When the device type is a computer, the operation for the user to enable the image synthesis function may include the user using input devices such as a keyboard or a mouse to perform operations such as selection and confirmation on the photo album, camera, or application program that includes the image synthesis function of the terminal device.

[0136] In some examples, the terminal device obtains a first image, a second image, and a first text, including: the terminal device enables the image synthesis function, the terminal device prompts the user to select the images to be synthesized and input the text, and the terminal device obtains the first image, the second image, and the first text in response to the user operation.

[0137] For example, the user first enables the image synthesis function in the photo album, and the terminal device pops up a dynamic pop-up window in response to the user operation to prompt the user to select an image and input text (such as the description information of the expected display effect of the synthesized image), and the terminal device determines the first image, the second image, and the first text in response to the user operation.

[0138] For another example, the user first enables the image synthesis function in the photo album, and the terminal device pops up a dynamic pop-up window in response to the user operation to prompt the user to select an image. Then, the terminal device determines the first image and the second image in response to the user selection operation. Finally, the terminal device pops up a dynamic pop-up window to prompt the user to input text, and the terminal device determines the first text in response to the user input operation.

[0139] In other examples, the terminal device obtains a first image, a second image, and a first text, including: the terminal device obtains the first image and the second image in response to the user operation, the terminal device enables the image synthesis function, and the terminal device prompts the user to input text to obtain the first text.

[0140] For example, the user first selects the first image and the second image in the photo album, and then enables the image synthesis function. The terminal device pops up a dynamic pop-up window to prompt the user to input text, and the terminal device determines the first text in response to the user input operation.

[0141] It can be understood that the embodiments of the present application do not limit the specific implementation manners for the terminal device to obtain the first image, the second image, and the first text.

[0142] In some embodiments of the present application, the terminal device obtains the first image and the second image in response to the user's operation of selecting the first image and the second image. The terminal device obtains the first text in response to the user's operation of inputting the first text.

[0143] In a possible implementation manner, when the terminal device obtains the first image and the second image, the terminal device determines the target object in each image in response to the user operation and labels each target object.

[0144] It can be understood that the embodiments of the present application do not limit the sequence of the terminal device obtaining the image, determining the target object, and labeling the target object.

[0145] Specifically, when the terminal device acquires the first image and the second image, the terminal device determines the first target object in the first image in response to a user operation, and the terminal device labels the first target object in response to the user operation of labeling the first target object. The terminal device determines the second target object in the second image in response to a user operation, and the terminal device labels the second target object in response to the user operation of labeling the second target object.

[0146] In some examples, if only one object is included in the first image, the terminal device determines the object as the target object in the first image, that is, the first target object. Correspondingly, if only one object is included in the second image, the terminal device determines the object as the target object in the second image, that is, the second target object.

[0147] In other examples, if multiple objects are included in the first image, the terminal device determines at least one object actively specified by the user as the target object in the first image, that is, the first target object. Correspondingly, if multiple objects are included in the second image, the terminal device determines at least one object actively specified by the user as the target object in the second image, that is, the second target object.

[0148] In other examples, if multiple objects are included in the first image, the terminal device prompts the user to select a target object, and the terminal device determines the target object in the first image in response to the user operation, that is, the first target object. Correspondingly, if multiple objects are included in the second image, the terminal device prompts the user to select a target object, and the terminal device determines the target object in the second image in response to the user operation, that is, the second target object.

[0149] It can be understood that the number of target objects in each image is at least one, and the embodiments of the present application do not limit the number of target objects in each image. The embodiments of the present application also do not limit the specific implementation manner for the terminal device to determine the target object in each image.

[0150] In some examples, when the terminal device determines the target object in each image, the terminal device labels each target object in response to the user operation of actively labeling each target object.

[0151] In other examples, when the terminal device determines the target object in each image, the terminal device actively labels each target object.

[0152] It can be understood that labeling each target object is for accurately obtaining the information of the target object during subsequent synthesis to improve the synthesis effect. The embodiments of the present application do not limit the specific implementation manner for the terminal device to label the target object.

[0153] Exemplarily, the first image is a solo photo of user A (such as a half-body or full-body photo including the front face of user A), and the target object in the first image (which can also be described as the first target object) is the portrait of user A. The user labels the identity document (ID) of the first target object as V1. The second image is a solo photo of user B (such as a half-body or full-body photo including the front face of user B), and the target object in the second image (which can also be described as the second target object) is the portrait of user B. The user labels the identity document (ID) of the second target object as V2.

[0154] Exemplarily, as Figure 8 shown in (a) of, the terminal device displays the main interface 801, and the main interface 801 includes an album 802. When the user wants to synthesize the images in the album, the user clicks on the album 802. In response to the user's operation of clicking on the album 802, the terminal device displays an album interface 803 as shown in Figure 8 (b) of. The user can arbitrarily select multiple images in the album interface 803, such as selecting a solo photo of a boy (i.e., the first image) and a solo photo of a girl (i.e., the second image). In response to the user's operation of selecting the first image (solo photo of a boy) and the second image (solo photo of a girl), the terminal device displays function controls at the bottom of the album interface 803, such as favorite, add to, delete, and more 804. At the same time, since both the first image and the second image are solo photos, the first image only contains a portrait of a boy (the first target object), and the second image only contains a portrait of a girl (the second target object). The terminal device can actively label each target object. For example, the first target object (portrait of a boy) is labeled as boy, and the second target object (portrait of a girl) is labeled as girl. As Figure 8 shown in (b) of, the function space currently displayed in the album interface 803 cannot implement image synthesis. Then the user clicks on the more 804 control to view more functions. In response to the user's operation of touching the more 804 function control, the terminal device displays a function interface 805 as shown in Figure 8 (c) of. This function interface 805 includes function controls such as photo movie, cutout, special effects, artistic photo, synthesis 806, and filter. The user clicks on the synthesis 806 function control. In response to the user's operation of touching the synthesis 806, the terminal device displays an interface 807 as shown in Figure 8 (d) of. This album interface includes a dynamic pop-up window 808, which is used for the user to input the first text, such as "The boy and the girl stand together. The boy is on the right. The boy is taller than the girl. Both of them are laughing heartily."

[0155] It can be understood that the above description takes the synthesis of the first image and the second image as an example. In practical applications, the number of images to be synthesized in the terminal device is at least two, that is, the number of images to be synthesized is at least two. For example, the terminal device obtains the first image, the second image, and the third image, a total of three images. Correspondingly, the first image includes a first target object, the second image includes a second target object, the third image includes a third target object, and the first text is the text information input by the user for describing the first target object, the second target object, and the third target object. The embodiments of the present application do not limit the number of images to be synthesized.

[0156] S702. The terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector.

[0157] In the embodiments of the present application, the first image feature vector refers to abstracting the information in the first image into a numerical feature vector for representing the features of the first image.

[0158] In the embodiments of the present application, the second image feature vector refers to abstracting the information in the second image into a numerical feature vector for representing the features of the second image.

[0159] Among them, the image feature vector is a set of numerical features for representing the attributes and features of the image.

[0160] Optionally, the image feature vector includes color features, texture features, shape features, structural features, etc. Among them, the color feature can be the color distribution of the image, including hue, saturation, brightness, etc. The texture feature can be the texture information of the image, including the roughness, direction, density, etc. of the texture. The shape feature includes the shape information of the image, including contour, curvature, angle, etc. The structural feature can be the overall structural information of the image, including proportion, symmetry, position, etc.

[0161] In the embodiments of the present application, the first text feature vector refers to converting the information of the first text into a vector form representation for representing the features of the first text.

[0162] In the embodiments of the present application, the terminal device inputs the first image and the second image into an image encoding network to obtain a first image feature vector and a second image feature vector; the terminal device inputs the first text into a text encoding network to obtain a first text feature vector.

[0163] In some examples, the terminal device inputs the first image and the second image into an image encoding network. The image encoding network encodes the first image to obtain a first image feature vector, and encodes the second image to obtain a second image feature vector. The terminal device inputs the first text into a text encoding network, and the text encoding network encodes the first text to obtain a first text feature vector.

[0164] In some other examples, the image encoding network of the terminal device includes a first preset dimension. The terminal device inputs the first image and the second image into the image encoding network, and the image encoding network encodes the first image into a first image feature vector and the second image into a second image feature vector according to the first preset dimension. Wherein, the dimensions of the first image feature vector and the second image feature vector are both the first preset dimension.

[0165] Wherein, the first preset dimension can be the dimension of the image feature vector determined in advance through experiments by the image encoding network, and the first preset dimension can be a fixed dimension.

[0166] Correspondingly, in some other examples, the text encoding network of the terminal device includes a second preset dimension. The terminal device inputs the first text into the text encoding network, and the text encoding network encodes the first text into a first text feature vector according to the second preset dimension. Wherein, the dimension of the first text feature vector is the second preset dimension.

[0167] Wherein, the second preset dimension can be the dimension of the text feature vector determined in advance through experiments by the text encoding network, and the second preset dimension can be a fixed dimension.

[0168] It can be understood that the embodiments of the present application do not limit the specific implementation manners of the terminal device encoding the first image, the second image, and the first text to obtain the first image feature vector, the second image feature vector, and the first text feature vector.

[0169] In this way, when encoding an image into a feature vector, regardless of the resolution of the input image, the dimension of the finally encoded and converted image feature vector is the same. That is to say, images with any resolution are encoded into image feature vectors with a fixed dimension, and the feature dimensions of multiple images with any resolution after encoding are fixed and unified, and the encoded image features are more standardized. Correspondingly, when encoding a text into a feature vector, regardless of the length of the input text, the dimension of the finally encoded and converted text feature vector is the same. That is to say, texts with any length are encoded into text feature vectors with a fixed dimension, the feature dimensions of texts with any length are fixed and unified, and the encoded text features are more standardized. Moreover, the fixed feature dimension is also beneficial to feature fusion, realizing the effective fusion of different modality data (image or text), and improving the accuracy of the fusion result.

[0170] Among them, the image encoding network can be a neural network trained with a large number of sample images for encoding image information. The text encoding network can be a neural network trained with a large number of sample texts for encoding text information.

[0171] Exemplarily, the image encoding network is a convolutional neural network (CNN) trained for encoding.

[0172] Specifically, the initial convolutional neural network includes an initial encoding convolutional neural network and an initial decoding convolutional neural network. The initial encoding convolutional neural network is used to encode an image into an image vector, and the initial decoding convolutional neural network is used to reconstruct the image vector into an image. The initial encoding convolutional neural network extracts and encodes features of the training image, and encodes the training image into an image vector. For example, convolutional and pooling operations are performed on the training image to perform high-level abstraction processing on the image, extract highly recognizable image features, and further combine these images into a feature vector with a fixed dimension. Then, the encoded image vector is input into the initial decoding convolutional neural network, and the initial decoding convolutional neural network reconstructs the image vector to obtain a reconstructed image. The reconstruction residual is obtained by comparing the training image and the reconstructed image, and the training and iteration of the initial encoding convolutional neural network and the initial decoding convolutional neural network are inversely constrained according to the reconstruction residual until the reconstruction residual meets the preset convergence condition, thereby ending the training and obtaining the trained convolutional neural network (including the encoding convolutional neural network and the decoding convolutional neural network).

[0173] It can be understood that during the training process, a convolutional neural network for encoding and a convolutional neural network for decoding are required, but only the trained convolutional neural network for encoding is required during the image synthesis process.

[0174] Exemplarily, the text encoding network is a long short-term memory (LSTM) network trained for encoding.

[0175] Specifically, the initial long short-term memory network includes an initial encoding long short-term memory network and an initial decoding long short-term memory network. The initial encoding long short-term memory network is used to encode text into a text vector, and the initial decoding long short-term memory network is used to reconstruct the text vector into text. The initial encoding long short-term memory network extracts and encodes features from the training text, encoding the training text into a text vector. For example, each word or character in the training text is processed step by step to convert the entire text sequence into a fixed-length text feature vector, which contains the semantic and syntactic information of the training text. Convolution and pooling operations are performed on the training image to perform high-level abstraction processing on the image, extract highly recognizable image features, and further combine these images into a feature vector of a fixed dimension. Then, the encoded text vector is input into the initial decoding long short-term memory network, and the initial decoding long short-term memory network reconstructs the text vector to obtain the reconstructed text. The reconstruction residual is obtained by comparing the training text and the reconstructed text, and the training and iteration of the initial encoding long short-term memory network and the initial decoding long short-term memory network are inversely constrained according to the reconstruction residual until the reconstruction residual meets the preset convergence condition, thereby ending the training and obtaining the trained long short-term memory network (including the encoding long short-term memory network and the decoding long short-term memory network).

[0176] It can be understood that during the training process, the long short-term memory network for encoding and the long short-term memory network for decoding are required, but only the trained long short-term memory network for encoding is used during the image synthesis process.

[0177] Exemplarily, the text editing network can also be a trained recurrent neural network (RNN) for encoding. The training method of the recurrent neural network is as described above and will not be elaborated here.

[0178] In some embodiments of the present application, after the terminal device acquires the first image, the second image, and the first text, the terminal device preprocesses the first image and the second image, and the terminal device inputs the preprocessed first image and the second image into the convolutional neural network. The preprocessed first image is encoded by the convolutional neural network to output a first image feature vector, and the preprocessed second image is encoded by the convolutional neural network to output a second image feature vector. The terminal device inputs the first text into the long short-term memory network, and the first text is encoded by the long short-term memory network to output a first text feature vector.

[0179] In some examples, preprocessing the first image and the second image includes: the terminal device segments the first target object in the first image according to a preset segmentation algorithm, extracts the pixel information of the first target object, and removes the pixel information irrelevant to the first target object (for example, replaces the pixel points irrelevant to the first target object with white pixels); the terminal device segments the second target object in the second image according to a preset segmentation algorithm, extracts the pixel information of the second target object, and removes the pixel information irrelevant to the second target object (for example, replaces the pixel points irrelevant to the second target object with white pixels).

[0180] Among them, the preset segmentation algorithms include traditional threshold segmentation, edge detection-based segmentation, region-based segmentation, and deep learning-based semantic segmentation algorithms, etc. The embodiments of the present application do not limit the specific implementation manners of preprocessing the first image and the second image.

[0181] Exemplarily, based on the example in the above S701 Figure 8 the terminal device preprocesses the first image to obtain a male portrait, preprocesses the second image to obtain a female portrait, and inputs the male portrait and the female portrait into the image encoding network to obtain a first image feature vector (such as vector 1) after encoding the male portrait and a second image feature vector (such as vector 2) after encoding the female portrait respectively. The terminal device inputs the first text into the text encoding network to obtain a first text feature vector after encoding the first text, such as "A boy and a girl stand together. The boy is on the right. The boy is taller than the girl. The two are laughing heartily."

[0182] S703. The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector.

[0183] In the embodiments of the present application, the first fusion feature vector contains image features and text features, and the dimension of the first fusion feature vector is fixed.

[0184] In the embodiments of the present application, the terminal device performs feature fusion on the first image feature vector, the second image feature vector, and the first text feature vector according to a preset fusion method to obtain a first fusion feature vector.

[0185] In some examples, the terminal device combines the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector.

[0186] In other examples, the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector in a weighted average manner to obtain a first fusion feature vector.

[0187] In some other examples, the terminal device couples the first image feature vector and the second image feature vector into the first text feature vector to obtain a first fused feature vector.

[0188] Specifically, the terminal device normalizes the first image feature vector, the second image feature vector, and the first text feature vector, and the terminal device fuses the normalized first image feature vector, the second image feature vector, and the first text feature vector according to a preset fusion method to obtain a first fused feature vector.

[0189] It can be understood that the embodiments of the present application do not limit the specific implementation manner in which the terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain the first fused feature vector.

[0190] Exemplarily, based on the example in S702 above, the terminal device couples the first image feature vector (such as vector 1) and the second image feature vector (such as vector 2) into the first text feature vector "A boy and a girl are standing together. The boy is on the right, the boy is taller than the girl, and both of them are laughing heartily" to obtain the first fused feature vector "Vector 1 and vector 2 are standing together. Vector 1 is on the right, vector 1 is taller than vector 2, and both of them are laughing heartily".

[0191] S704. The terminal device displays a first synthesized image according to the first fused feature vector.

[0192] In the embodiments of the present application, the terminal device decodes and reconstructs the first fused feature vector into a first synthesized image through the SD technology, and the terminal device displays the first synthesized image.

[0193] It can be understood that for the specific implementation manner of decoding and reconstructing the first fused feature vector into the first synthesized image through the SD technology, refer to the prior art and will not be elaborated here.

[0194] Exemplarily, based on the example in S703 above, the terminal device decodes and reconstructs the first fused feature vector into a first synthesized image through the SD technology, such as Figure 9 shown, the terminal device displays the first synthesized image 901 and function controls, such as edit, synthesize, and save. In the first synthesized image 901, the boy is on the right of the girl, the boy is taller than the girl, and the expressions of both of them are laughing heartily.

[0195] It can be understood that in the above example, the synthesis is based on two single-person photos. In practical applications, the terminal device performs image synthesis based on an image and text. The text is used to define the synthesis effect of the synthesized image (such as expression, stance, environment, etc.), and the image is used to provide key information in the synthesized image (such as portrait information (such as facial features information, limb information), etc.). For example, if the terminal device obtains three single-person ID photos uploaded by the user (with a slightly serious expression), and the portrait in each image is a half-body image including the front face of the person, and the first text is "Generate a graduation photo including three people's full-body images, the shooting environment is a classroom, and the expressions of the three people are all very happy", the finally generated synthesized image may be three people standing in the classroom wearing graduation gowns with bright smiles. In this example, there are only half-body images in the images, but the text defines that the people in the synthesized image should be full-body images. Therefore, the terminal device generates body parts matching the people in the image based on the text and the image. The terminal device sets the background of the synthesized image as a classroom based on the synthesis environment "classroom, graduation photo" defined by the text, and the costumes of the people in the synthesized image are school uniforms or graduation gowns, and completes the synthesis based on the text and the image.

[0196] Through the above technical solution, in the embodiment of the present application, the terminal device obtains a first image, a second image, and a first text. The terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector. The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector; the terminal device displays a first synthesized image according to the first fusion feature vector. The technical solution of the present application performs image synthesis based on the encoded feature vectors, optimizes the image synthesis strategy, the synthesized image is more natural, effectively improves the synthesis effect and authenticity, simplifies the user operation, meets the user's personalized needs, and enhances the user experience.

[0197] It can be understood that after the terminal device generates the first synthesized image according to the first fusion feature vector, the terminal device displays the first synthesized image. If the user is not satisfied with the synthesis effect of the first synthesized image, the user can choose to generate a synthesized image again or update the text description to re-synthesize.

[0198] In the embodiment of the present application, after step S704, when the user is not satisfied with the synthesis effect of the first synthesized image, the terminal device updates the first synthesized image in response to the user's update operation.

[0199] Among them, the update operation includes a preset operation in which the user triggers the terminal device to update the first composite image. Optionally, the preset operation may be some quick operations of the user, such as quick operations by means of gestures or key combinations. The preset operation may also be a voice command of the user, such as issuing a command to update the first composite image to the mobile phone through a voice assistant, etc. The embodiments of the present application do not limit the specific implementation form of the preset operation.

[0200] In some examples, the terminal device updates the first composite image to obtain a second composite image in response to the user triggering the terminal device to generate a composite image again.

[0201] It can be understood that generating a composite image again is also to update the first composite image, and the first image, the second image, and the first text remain unchanged. Exemplarily, based on the above Figure 9 example, the user touches and synthesizes this control, and the terminal device re-executes the above steps to re-perform image synthesis based on the first image, the second image, and the first text.

[0202] In a possible implementation manner, the terminal device re-encodes the first image, the second image, and the first text to obtain a third image feature vector, a fourth image feature vector, and a second text feature vector respectively; the terminal device fuses the third image feature vector, the fourth image feature vector, and the second text feature vector to obtain a second fusion feature vector; the terminal device generates a second composite image according to the second fusion feature vector and displays the second composite image.

[0203] Among them, the third image feature vector is the image feature vector obtained after the first image is re-encoded, which is different from the first image feature vector. The fourth image feature vector is the image feature vector obtained after the second image is re-encoded, which is different from the second image feature vector; the second text feature vector is the text feature vector obtained after the first text is re-encoded, which is different from the first text feature vector; the second fusion feature vector is different from the first fusion feature vector; the second composite image is the updated composite image, which is different from the first composite image.

[0204] In this way, the first image, the second image, and the first text are re-encoded and feature-fused to generate a new composite image.

[0205] In another possible implementation manner, the terminal device re-performs feature fusion when performing feature fusion. That is, the terminal device obtains a second fusion feature vector according to the first image feature vector, the second image feature vector, and the first text feature vector; the terminal device generates a second composite image according to the second fusion feature.

[0206] Among them, the second fusion feature vector is different from the first fusion feature vector; the second composite image is the updated composite image, which is different from the first composite image.

[0207] Specifically, when performing an update, there is no need to re-encode the image and text again. Only when performing feature fusion, other preset fusion methods are used for feature fusion. In this way, the update speed can be accelerated.

[0208] Exemplarily, as Figure 10 shown, the terminal device displays the updated composite image 1001 and function controls. In the updated composite image 1001, the distance between the boy and the girl is closer.

[0209] In some other examples, in response to the user's operation of updating the first text, the terminal device updates the first text to the second text; the terminal device encodes the second text to obtain a third text feature vector; the terminal device fuses the first image feature vector, the second image feature vector, and the third text feature vector to obtain a third fusion feature vector; the terminal device displays a third composite image according to the third fusion feature vector.

[0210] Among them, the third text feature vector is the text feature vector obtained after encoding the second text, and the third composite image is the updated composite image, which is different from the first composite image.

[0211] Among them, updating the first text includes modifying the first text or adding new text on the basis of the first text. In this way, the second text should be the text information modified by the user or the text information including the first text and the added text.

[0212] Exemplarily, based on the above Figure 9 shown example, when the user touches and edits this control, a dynamic pop-up window appears on the terminal device. The user can modify the first text "The boy and the girl stand together. The boy is on the right. The boy is taller than the girl. Both of them are laughing" to the second text "The boy and the girl stand together. The boy and the girl are a couple. The boy is on the right. The boy is taller than the girl. Both of them are laughing" on the dynamic pop-up window. The terminal device encodes, fuses, and synthesizes an image according to the modified second text, and obtains the updated composite image 1001 as Figure 10 shown. In the updated composite image 1001, the relationship between the boy and the girl is closer, and the distance between them is closer.

[0213] It can be understood that the specific implementation manners of the terminal device encoding the second text and obtaining the second fusion feature vector and displaying the second composite image can be referred to the above, and will not be elaborated here.

[0214] In some other examples, on the basis of the first composite image, text information can also be added to update the display effect of the first composite image. That is, the terminal device obtains a third text (that is, the added text), and the terminal device generates a fourth composite image according to the third text and the first composite image.

[0215] Among them, the fourth synthesized image is the updated synthesized image, which is different from the first synthesized image.

[0216] It can be understood that for the specific implementation of the terminal device to generate the fourth synthesized image, refer to the above, and details are not described here. The embodiments of the present application do not limit the specific implementation of the terminal device to update the first synthesized image.

[0217] In this way, when the user is not satisfied with the display effect of the first synthesized image, the terminal device responds to the user's update operation and updates the first synthesized image until a synthesized image satisfactory to the user is obtained, which can meet the personalized needs of the user and improve the user experience.

[0218] Exemplarily, as Figure 11 shown, it is an interaction flowchart of each module in another image processing system provided by the embodiments of the present application.

[0219] As Figure 11 shown, in the embodiments of the present application, the image encoding module in the terminal device acquires a first image (a solo photo of a boy) and a second image (a solo photo of a girl). The image encoding module preprocesses (such as portrait segmentation) the first image and the second image to obtain a boy portrait and a girl portrait respectively. The image encoding module encodes the boy portrait and the girl portrait to obtain corresponding first image feature vectors (such as vector 1) and second image feature vectors (vector 2). The image encoding module sends the first image feature vector and the second image feature vector to the feature fusion module.

[0220] The text encoding module in the terminal device acquires a first text (such as the boy and the girl standing together), and encodes the first text to obtain a first text feature vector (such as the boy and the girl standing together). The text encoding module sends the first text feature vector to the feature fusion module.

[0221] The feature fusion module fuses the first image feature vector, the second image feature vector and the first text feature vector, couples the first image feature vector and the second image feature vector with only the first text feature vector to obtain a first fusion feature vector (such as vector 1 and vector 2 standing together). The feature fusion module sends the first fusion feature vector to the decoding and synthesizing module.

[0222] The decoding and synthesizing module decodes and reconstructs the first fusion feature vector into a first synthesized image through SD technology, and the decoding and synthesizing module sends the first synthesized image to the terminal device so that the terminal device can display the first synthesized image.

[0223] The system architecture and business scenarios described in this application are to more clearly illustrate the technical solutions of this application, and do not constitute the only limitation to the technical solutions provided by this application. As those skilled in the art know, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by this application are equally applicable to similar technical problems.

[0224] The above mainly introduces the solutions provided by the embodiments of this application from the perspective of methods. It can be understood that in order for an electronic device to implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of each example described in the embodiments disclosed in this application, the embodiments of this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solutions of the embodiments of this application.

[0225] The embodiments of this application can divide the functional modules of an electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of units in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0226] As Figure 12 shown, it is a schematic structural diagram of another terminal device provided by the embodiments of this application. The terminal device 1200 can be used to implement the methods described in each of the above method embodiments. Exemplarily, the terminal device 1200 can specifically include: a processing module 1201, an acquisition module 1202, and a display module 1203.

[0227] Among them, the processing module 1201 is used to execute the processing functions that support the terminal device 1200 to execute Figure 7 any one of

[0228] The acquisition module 1202 is used to execute the acquisition functions that support the terminal device 1200 to execute Figure 7 any one of

[0229] The display module 1203 can be used to display a picture according to the display drive data, etc. And / or, the display module 1203 is further used to support other display operations executed by the terminal device 1200 in the embodiments of this application. In the embodiments of this application, the display module 1203 is used to display a composite image.

[0230] Optionally, Figure 12 the terminal device 1200 shown may further include a communication module ( Figure 12 not shown in the figure), and the communication module is used to support the terminal device 1200 to execute the steps of communication between the terminal device and other devices in the embodiments of the present application.

[0231] Optionally, Figure 12 the terminal device 1200 shown may further include a storage module ( Figure 12 not shown in the figure), and the storage module stores programs or instructions. When the processing module 1201 executes the programs or instructions, Figure 12 the terminal device 1200 shown can execute the methods shown in the above method embodiments.

[0232] Figure 12 For the technical effects of the terminal device 1200 shown, reference may be made to the technical effects of the methods described in the above method embodiments, and details are not described herein again. Figure 12 The processing module 1201 involved in the terminal device 1200 shown may be implemented by a processor or processor-related circuit components, and may be a processor or a processing module. The communication module may be implemented by a transceiver or transceiver-related circuit components, and may be a transceiver or a transceiver module. The display module 1203 may be implemented by components related to a display screen.

[0233] The embodiments of the present application further provide a chip system, as Figure 13 shown, the chip system 1300 includes at least one processor 1301 and at least one interface circuit 1302. As an example, when the chip system 1300 includes one processor and one interface circuit, then one processor may be Figure 13 the processor 1301 shown in the solid line box (or the processor 1301 shown in the dashed line box), and one interface circuit may be Figure 13 the interface circuit 1302 shown in the solid line box (or the interface circuit 1302 shown in the dashed line box). When the chip system 1300 includes two processors and two interface circuits, then the two processors include Figure 13 the processor 1301 shown in the solid line box and the processor 1301 shown in the dashed line box, and the two interface circuits include Figure 13 the interface circuit 1302 shown in the solid line box and the interface circuit 1302 shown in the dashed line box. There is no limitation on this.

[0234] The processor 1301 and the interface circuit 1302 can be interconnected by a line. For example, the interface circuit 1302 can be used to receive signals. For another example, the interface circuit 1302 can be used to send signals to other devices (such as the processor 1301). Exemplarily, the interface circuit 1302 can read the instructions stored in the memory and send the instructions to the processor 1301. When the instructions are executed by the processor 1301, each step in the above embodiments can be executed. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations on this.

[0235] Optionally, the processor in the chip system can be one or more. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that implements by reading the software code stored in the memory.

[0236] Optionally, the chip system can also include a memory ( Figure 13 not shown in the figure), and the memory can also be one or more. The memory can be integrated with the processor or can be separately arranged from the processor, and the present application does not limit this. Exemplarily, the memory can be a non-transitory processor, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be separately arranged on different chips. The present application does not make specific limitations on the type of the memory and the setting manner of the memory and the processor.

[0237] Exemplarily, the chip system can be a field programmable gate array (FPGA), can be an application specific integrated circuit (ASIC), can also be a system on chip (SoC), can also be a central processing unit (CPU), can also be a network processor (NP), can also be a digital signal processing circuit (DSP), can also be a micro controller unit (MCU), can also be a programmable logic device (PLD) or other integrated chips.

[0238] It should be understood that each step in the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The method steps disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor.

[0239] The embodiments of the present application further provide a computer storage medium. A computer instruction is stored in the computer storage medium. When the computer instruction runs on a terminal device, the terminal device is enabled to execute the method described in the above method embodiments.

[0240] The computer-readable storage medium includes, but is not limited to, any one of the following: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc., various media that can store program codes.

[0241] The embodiments of the present application provide a computer program product. The computer program product includes: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is enabled to execute the method described in the above method embodiments.

[0242] In addition, the embodiments of the present application further provide a device. This device may specifically be a chip, component, or module. The device may include a processor and a memory connected to each other. Among them, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory so that the device executes the methods in the above method embodiments.

[0243] Among them, the terminal device, computer storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved by them can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0244] The steps of the methods or algorithms described in connection with the disclosed embodiments of the present application may be implemented in hardware or by a processor executing software instructions. The software instructions may consist of corresponding software modules, which may be stored in a random access memory (RAM), flash memory, read-only memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable hard disk, compact disc read-only memory (CD-ROM), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a part of the processor. The processor and the storage medium may be located in an application specific integrated circuit (ASIC).

[0245] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions may be allocated to different functional modules according to needs; that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For the specific working processes of the systems, devices, and units described above, reference may be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein again.

[0246] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods may be implemented in other ways. The embodiments may be combined or referred to each other without conflict. The device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.

[0247] The unit described as a separation component may or may not be physically separated. The component shown as a unit may be a physical unit or multiple physical units, that is, it may be located in one place or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0248] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0249] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0250] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image processing method, characterized in that, the method includes: The terminal device acquires a first image, a second image, and a first text; The terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector; The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector; The terminal device displays a first composite image according to the first fusion feature vector.

2. The method according to claim 1, characterized in that, The first image and the second image are the to-be-combined images selected by the user. The first image includes a first target object, the second image includes a second target object, and the first text is the text information input by the user for describing the first target object and the second target object; the first image feature vector is the numericalized feature vector corresponding to the first image; the second image feature vector is the numericalized feature vector corresponding to the second image.

3. The method according to claim 1 or 2, characterized in that, The terminal device encodes the first image, the second image, and the first text to obtain a first image feature vector, a second image feature vector, and a first text feature vector, including: The terminal device inputs the first image and the second image into an image encoding network to obtain the first image feature vector and the second image feature vector; wherein, the image encoding network is a neural network trained by sample images for encoding image information; The terminal device inputs the first text into a text encoding network to obtain the first text feature vector; wherein, the text encoding network is a neural network trained by sample texts for encoding text information.

4. The method according to claim 3, characterized in that, The terminal device inputs the first image and the second image into an image encoding network to obtain the first image feature vector and the second image feature vector, including: The terminal device preprocesses the first image and the second image; The terminal device inputs the preprocessed first image and the preprocessed second image into the image encoding network to obtain the first image feature vector and the second image feature vector.

5. The method according to claim 4, characterized in that, The terminal device preprocesses the first image and the second image, including: The terminal device segments the first target object in the first image according to a preset segmentation algorithm; The terminal device segments the second target object in the second image according to a preset segmentation algorithm.

6. The method according to any one of claims 1-5, characterized in that, The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector to obtain a first fusion feature vector, including: The terminal device fuses the first image feature vector, the second image feature vector, and the first text feature vector according to a preset fusion method to obtain the first fusion feature vector.

7. The method according to any one of claims 1-6, wherein, the terminal device displays a first synthesized image according to the first fusion feature vector, including: the terminal device decodes and reconstructs the first fusion feature vector through the Stable Diffusion technique to generate the first synthesized image; the terminal device displays the first synthesized image.

8. The method according to any one of claims 1-7, wherein, the method further includes: the terminal device updates the first synthesized image in response to a user's update operation; wherein, the update operation includes an operation by which the user triggers the terminal device to generate a synthesized image again, or an operation by which the user updates the first text.

9. The method according to claim 8, wherein, the terminal device updates the first synthesized image in response to a user's update operation, including: the terminal device responds to an operation by which the user triggers the terminal device to generate a synthesized image again, synthesizes an image again according to the first image, the second image, and the first text, and updates the first synthesized image to a second synthesized image; the second synthesized image is different from the first synthesized image; or, the terminal device responds to an operation by which the user updates the first text, and updates the first text to a second text; the terminal device performs image synthesis according to the first image, the second image, and the second file to generate a third synthesized image; the terminal device updates the first synthesized image to the third synthesized image; the third synthesized image is different from the first synthesized image.

10. The method according to any one of claims 1-9, wherein, the first text includes one or more of the following: the interpersonal relationship, location, expression, body movement, and / or facial expression between the first target object and the second target object.

11. A terminal device, wherein, it includes: one or more processors; a memory; wherein, one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the terminal device, cause the terminal device to execute the method according to any one of claims 1-10.

12. A chip system, wherein, it includes at least one processor and at least one interface circuit, the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor, and the at least one processor executes the instructions, and the at least one processor executes the method according to any one of claims 1-10.

13. A computer-readable storage medium, wherein, the computer-readable storage medium includes a computer program or instructions that, when the computer program or instructions are run on a computer, cause the computer to execute the method according to any one of claims 1-10.

14. A computer program product, It is characterized in that the computer program product includes: a computer program or instruction, which, when running on a computer, causes the computer to execute the method according to any one of claims 1-10.