Method, apparatus, device and program product for generating image
By obtaining the feature information of the target image, the clone image is adaptively integrated, which solves the problem of users taking photos with friends during shooting, and achieves a convenient and low-threshold image fusion effect.
Patent Information
- Application Number
- CN202510300415.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is difficult to make it possible for users to take photos with friends during shooting, especially when friends are not around, and the post-image processing requirements are cumbersome and the threshold is high.
By acquiring the target image feature information of the first character, the clone image of the second character is adaptively integrated into the target image based on these feature information, and a fused image is generated.
It enables users to take photos with friends during shooting, and the fusion effect is adaptive, with convenience and low threshold, avoiding cumbersome image processing in the later stage.
Smart Images

Figure CN120147480A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and more particularly, to a method, apparatus, computing device, computer-readable storage medium, and computer program product for generating images. Background Art
[0002] Artificial Intelligence (AI) is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. It is a branch of computer science, aiming to produce intelligent machines that can respond in a way similar to human intelligence. The research fields of AI include language and image recognition, robotics, natural language processing, and expert systems, etc.
[0003] The AI image generation technology uses artificial intelligence algorithms. Through learning and analyzing a large amount of image data, it can quickly generate creative and specific-style images according to instructions such as text descriptions and reference images input by users. It covers various functions such as text-to-image, image editing, and image restoration, bringing new ideas and efficient solutions to image creation. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, computing device, computer-readable storage medium, and computer program product for generating images. According to an embodiment of the present disclosure, when taking a photo, a user can call the avatar image of a friend in real time to complete a group photo, and has an adaptive fusion effect.
[0005] According to a first aspect of the present disclosure, there is provided a method for generating an image. The method includes obtaining feature information of a target image including a first person, where the feature information includes multi-dimensional features of the first person. The method further includes adaptively integrating an avatar image of a second person into the target image based on the feature information to generate a fused image.
[0006] According to a second aspect of the present disclosure, there is provided an apparatus for generating an image. The apparatus includes: a feature information obtaining unit configured to obtain feature information of a target image including a first person, where the feature information includes multi-dimensional features of the first person; and a fused image generating unit configured to adaptively integrate an avatar image of a second person into the target image based on the feature information to generate a fused image.
[0007] According to a third aspect of the present disclosure, there is provided a computing device, including: at least one processing unit; at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the computing device to execute the method as described in the first aspect of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer storage medium including machine-executable instructions that, when executed by a device, cause the device to execute the method as described in the first aspect of the present disclosure.
[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product including machine-executable instructions that, when executed by a device, cause the device to execute the method as described in the first aspect of the present disclosure.
[0010] The Summary of the Invention is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the embodiments of the present disclosure will become more readily understood. In the drawings, multiple embodiments of the present disclosure will be illustrated by way of example and not limitation, where:
[0012] Figure 1A A schematic diagram of a display interface according to an embodiment of the present disclosure is shown;
[0013] Figure 1B A schematic diagram of a display interface in which a clone image is integrated according to an embodiment of the present disclosure is shown;
[0014] Figure 1C A schematic diagram of a display interface including a prompt input area according to an embodiment of the present disclosure;
[0015] Figure 2A A schematic diagram of a display interface according to an embodiment of the present disclosure is shown;
[0016] Figure 2B A schematic diagram of a display interface after uploading a portrait picture according to an embodiment of the present disclosure is shown;
[0017] Figure 3 A schematic flowchart of a method for generating an image according to an embodiment of the present disclosure is shown;
[0018] Figure 4 A block diagram of an apparatus for generating an image according to an embodiment of the present disclosure is shown; and
[0019] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0020] In all the drawings, the same or similar reference numerals denote the same or similar elements. Detailed implementation manners
[0021] The present disclosure will now be described with reference to several exemplary implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, rather than implying any limitation on the scope of the present disclosure.
[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0026] The embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0027] In the description of the embodiments of the present disclosure, the term "including" and its like terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.
[0028] In this document, a smart phone is taken as an example to show an exemplary user interface. However, those skilled in the art can understand that the embodiments of the present disclosure are equally applicable to other devices with a display screen that may have different screen aspect ratios, such as tablet computers, notebook computers, desktop computers, wearable devices with a screen, and devices with a foldable screen, etc. In addition, the interfaces provided herein are for illustrative purposes only. Some of the elements therein may be omitted or have different numbers, and there may also be more elements not shown. Furthermore, the interfaces according to the embodiments of the present disclosure may have a layout different from that shown in the drawings, and the positions of the respective elements may also be different. The present disclosure makes no limitations in these aspects.
[0029] In the field of image recognition, AI avatar is a new application technology that combines computer vision, natural language processing, and deep learning, etc. The avatar image generated based on AI avatar technology can simulate the appearance and image of a specific subject, and simulate the specific subject to perform activities and interactions in various scenarios. For example, it can act on behalf of users to interact in the social field, serve as a virtual anchor in content creation, act as a tutoring assistant in the education industry, and become a virtual customer service in customer service, etc. More application scenarios of AI avatars need to be further explored to enhance the functions of related application software.
[0030] The inventors noticed some problems. A user may hope to take a group photo with a friend, but it cannot be achieved because the friend is not around. Although it is also possible to add the friend to the already taken photo through post-image processing (for example, using a dedicated image processing tool), this requires cumbersome post-production and requires the user to have relevant processing techniques, bringing a relatively high learning cost. Therefore, it would be beneficial to provide a convenient and low-threshold image fusion method to achieve group photos.
[0031] In view of this, the present disclosure provides a method for generating an image based on an avatar image. The avatar image can be generated based on a user photo and can be used only after obtaining the authorization of the user. According to the method proposed by the present disclosure, an avatar image representing another person can be obtained, and the avatar image can be adaptively incorporated into the photo when taking or importing a photo of a person to obtain a group photo. Thus, even if a friend is not around, it is possible to take a group photo with the friend anytime and anywhere, and it has an adaptive fusion effect. The following refers to Figures 1A to 5Describe embodiments of the present disclosure in detail.
[0032] Figure 1A A schematic diagram of a display interface 100A according to an embodiment of the present disclosure is shown. The interface 100A is an example of a split-screen camera page according to an embodiment of the present disclosure. When a user uses an electronic device and clicks on a control with a split-screen camera function on an application, the electronic device can display Figure 1A the interface 100A shown. The electronic device can be a terminal with a camera and a display screen, such as a mobile phone, a tablet computer, etc. Regarding the creation of split-screen images, it will be described in detail in Figures 2A - 2B the corresponding part.
[0033] As shown in the figure, the interface 100A includes an image display area 102, an upload control 104, and a camera control 106. The image display area 102 can display the scene captured in real time by the camera of the electronic device that the user is using. In some implementations, the interface 100A may further include a camera switching control 108 to facilitate the user to freely switch between the front camera, the rear camera, etc. when using a multi-camera device. When the user clicks on the camera control 106, the image display area 102 can display the image captured by the user, including the first person 105 captured in real time. Alternatively, the user can also click on the upload control 104 to select any picture from the local photo album, so that the image display area 102 displays the uploaded image. It can be understood that the first interface 100A in FIG. 1 is only for illustrative purposes, and the split-screen camera page can have different layouts or designs, and the present disclosure does not limit this.
[0034] According to an embodiment of the present disclosure, a split-screen image can be integrated into a real-time captured or uploaded image (also referred to as a target image or original photo in this article). The split-screen image can be a split-screen image of the user himself or herself created, or a split-screen image shared and authorized by other users for use. In some implementations, the user can create his or her own split-screen through an image generation application on the electronic device, share it and authorize other users to use it. Similarly, the user can receive split-screens shared and authorized by other users for use. The user's own split-screen or the split-screens shared by other users can be integrated into the real-time captured or uploaded image to automatically generate a fused image. In some embodiments, the user can take a picture of a scene without people, and then integrate his or her own and / or other people's split-screens into the scene picture. In some embodiments, the user can take a picture including people, and then integrate one or more split-screens into the picture to achieve a group photo.
[0035] As shown in the figure, the interface 100A further includes a control 103 related to the avatar image. The user can click on the control 103 to integrate the avatar image into the image being displayed in the image display area 102. For example, in response to the user clicking on the control 103, a list of avatar images shared and authorized for use by other users may be popped up, with each avatar image corresponding to a different person from the current user. The user can select one or more of the avatar images and adaptively integrate them into the currently displayed image.
[0036] If the user has not obtained authorization to use a certain avatar image, the user can send a request to the owner of the avatar image to take a group photo using their avatar image. In response to this request, the owner of the avatar image can send their own avatar image to the current user, or confirm the authorization to allow the current user to use their avatar image. Alternatively, the original photo and the request can be sent together to the owner of the avatar image, who can then generate the group photo and send the generated group photo to the current user.
[0037] Figure 1B A schematic diagram of a display interface 100B in which an avatar image is integrated according to an embodiment of the present disclosure is shown. As Figure 1B shown, after generating the composite image, the composite image can be displayed in the image display area 102, which includes the existing person (which can be referred to as the first person) 105 and the second person 107 in the original photo. The user can preview the effect of the composite image in the camera in real time.
[0038] The shot type (close - up or long - shot), position, pose, expression, size, or clothing of the avatar image 107 in the composite image can be determined based on the feature information of the first person 105 and the scene, so that the generated composite image is more harmonious and achieves the effect of a real shot. The feature information can include multi - dimensional features describing, for example, the shot type, position, pose, expression, size, and clothing of the first person 105. The feature information can also include the category of the scene (such as natural scenery or buildings), season, time of day (day or night), geographical location, style, or objects in the scene (such as furniture or outdoor facilities).
[0039] The dof selection of the avatar image 107 and the first person 105 is the same. For example, both are full body or both are half body. The first person 105 and the avatar image 107 can be adjacent in position to achieve the effect of a group photo. In some embodiments, the poses, expressions, and clothing of the first person 105 and the avatar image 107 can have matching semantic information. For example, the clothing, clothing color, hairstyle, etc. of the avatar image 107 can be determined based on the first person 105 to achieve a visual match between the person and the scene. For another example, the pose of the avatar image 107 can also be determined based on the pose of the first person 105, making the two make the same or similar movements to avoid the sense of disharmony of a rigid fusion. In addition, other contents in the original image (such as other objects in the background) can also be recognized, and based on this, the image, pose, position, etc. of the avatar image 107 can be determined, so that the avatar image 107 is adapted to the overall image. The following lists several specific examples to illustrate how to determine the adaptive features of the avatar image 107 based on the multi-dimensional features of the first person 105 and the multi-dimensional features of the scene.
[0040] For example, in a seaside beach scene, the first person 105 is a full body image, wearing summer clothes, standing on the beach, facing the sea, with arms outstretched, head up enjoying the sea breeze, and hair fluttering in the wind. Then the avatar image 107 can also wear summer clothes, stand behind the first person 105, slightly closer, put both hands on the shoulders of the former, and the body also leans back slightly, feeling the sea breeze together with the former, looking out at the distant sea, and the poses of the two together create a comfortable and relaxing atmosphere by the sea. For another example, in a library reading area scene, the first person 105 is a half body image, wearing school uniforms, sitting at a desk, with a straight body, head down reading a book, one hand on the page, and the other hand holding a pen to take notes. Then the avatar image 107 can sit opposite, wear school uniforms, also hold a book, look at the other person, with a smile on the face, as if waiting for the other person to share the content of the book, creating a quiet and harmonious reading atmosphere.
[0041] As Figure 1B shown, the interface 100B can display a download control 110 and an edit control 112. The user can click on the download control 110 to download the generated composite image to the local for storage. The edit control 112 can be used to adjust the composite effect of the generated composite image.
[0042] In some embodiments, the user can select a preset style template to quickly switch the style of the composite image, such as a cartoon style, a comic style, etc. For example, the user can switch the preset style template by swiping the screen left or right to select the composite effect they want. Additionally or alternatively, the user can click on the edit control 112 to adjust the composite effect of the composite image. Specifically, the user can enter a prompt to change the composite effect.
[0043] Figure 1C Schematic diagram of a display interface 100C including a prompt input area according to an embodiment of the present disclosure. As Figure 1C shown, in response to the user clicking on the edit control 112, a prompt input area 114 pops up on the current page. The user can input the desired effect in the prompt input area 114 and click send. For example, changing the expression, clothing, position, or pose of the avatar image in the figure, or switching the entire composite image to another style. Then, the image display area 102 will display the composite image adjusted according to the prompt.
[0044] The following will be combined with Figure 2A and Figure 2B to detail the process of creating an avatar image. Figure 2A Schematic diagram of an interface 200A according to an embodiment of the present disclosure is shown. Herein, the interface 200A is an example of an avatar creation page according to an embodiment of the present disclosure. When the user clicks on a control with an avatar image creation function on other pages (for example, the dialogue interface of an AI large language framework, the personal information interface of an image generation application), the electronic device can switch the display to Figure 2A the interface 200A shown.
[0045] As Figure 2A shown, the interface 200A may include an example area 202, a portrait upload area 204, and a generation control 206. The example area 202 can be used to guide the user to upload correct and high-quality portrait pictures. The example area 202 may include at least one correct example and at least one incorrect example. As Figure 2A shown, the example area 202 may include a correct example 202-1 for guiding the user to upload a clear frontal portrait picture, a correct example 202-2 for guiding the user to upload portrait pictures with multiple backgrounds and angles, and incorrect examples 202-3 for guiding the user to avoid uploading portrait pictures with the face blocked and 202-4 for guiding the user to avoid uploading portrait pictures with too small a face. The user can click on the upload control 204-2 in the portrait upload area 204 according to the guidance to upload at least one real portrait picture of the same person for the AI model to extract facial features to generate an avatar image. If the user does not upload a portrait picture or the number of uploaded portrait pictures is insufficient, the generation control 206 can be set to an unselectable state.
[0046] Figure 2B Schematic diagram of an interface 200B after uploading a portrait picture according to an embodiment of the present disclosure is shown. As Figure 2BAs shown, after the user uploads a sufficient number of portrait images, the display effect of the upload control 204-4 can be changed, and the generate control 206 can change from an unselectable state to a selectable state. The user can click the generate control 206 to complete the creation of the avatar image and return. In some embodiments, the user can share and authorize other users to use the created avatar image.
[0047] Figure 3 FIG. 4 shows a flowchart of a method 300 for generating an image according to some embodiments of the present disclosure. The method 300 can be implemented by any electronic device having computing and display capabilities and including a camera and a display screen.
[0048] In block 310, the method 300 can obtain feature information of a target image including a first person, and the feature information includes multi-dimensional features of the first person. Wherein, the target image can be obtained by the user taking a photo with a camera or selecting a stored picture. In some embodiments, the multi-dimensional features of the first person can describe at least one of the following: scene selection, position, pose, expression, size, and clothing.
[0049] In block 320, the method 300 can adaptively integrate the avatar image of the second person into the target image based on the feature information to generate a fused image.
[0050] Optionally, the avatar image can be directly fused into the target image. In some embodiments, it can be determined which adaptive features the avatar image has based on the feature information, and then the fused image can be directly generated based on these adaptive features, the avatar image, and the scene image. Optionally, the user can also provide text content to indicate the desired effects or features.
[0051] Alternatively, an avatar image with adaptive features can be generated first, and then the image of the second person and the target image can be fused between images to obtain a fused image. For example, multiple avatar images with different adaptive features can be generated, and the user can select a satisfactory avatar image or indicate the desired effects or features to regenerate, and then the avatar image confirmed by the user can be fused into the target image.
[0052] Determining the adaptive features of the avatar image can include determining the following features of the avatar image, including but not limited to: scene selection, position, pose, expression, size, and clothing, etc. The determined adaptive features match the multi-dimensional features of the first person and have matching semantic information. The following lists several specific examples to further illustrate how to determine the adaptive features of the avatar image of the second person based on the multi-dimensional features of the first person.
[0053] In some embodiments, the shot of the avatar image of the second person may depend on the shot of the first person. For example, if the first person is a bust shot, the avatar image will also be a bust shot; if the first person is a full-body shot, the avatar image will also be a full-body shot.
[0054] In some embodiments, the position of the avatar image may depend on the position of the first person. For example, the position of the avatar image can be close to the first person to create a sense of interaction with the first person.
[0055] In some embodiments, the pose of the avatar image may depend on the pose of the first person. The first person and the second person may have the same or similar actions. For example, if the first person is standing, the second person will also stand; if the first person is running, the second person will also be running. Optionally, the first person and the second person may have complementary actions to form a certain semantic information. For example, if the first person makes a heart-shaped gesture, the second person will also make a similar heart-shaped gesture to create a sense of interaction.
[0056] In some embodiments, the expression of the avatar image may depend on the expression of the first person. For example, the second person and the first person may have the same or similar expressions.
[0057] In some embodiments, the size of the avatar image may depend on the size of the first person in the target image. For example, the second person and the first person may have a similar size. In some embodiments, the size of the avatar image can also be adjusted based on the depth-of-field relationship.
[0058] In some embodiments, the clothing of the avatar image may depend on the clothing of the first person. For example, considering the seasonal characteristics of the first person's clothing, they can wear clothing of the same season. For another example, considering the style of the first person's clothing, such as ancient or modern, they can wear clothing of the same style.
[0059] To further enhance the scene coordination of the fused image to be closer to a real group photo, in some embodiments, the multi-dimensional features of the scene in the target image can be used to determine the adaptive features of the avatar image, so that the adaptive features of the avatar image can match the feature information of the scene, for example, having matching semantic information. The multi-dimensional features of the scene can describe at least one of the following: the category of the scene (such as indoor or outdoor), the season of the scene (such as spring, winter, etc.); the time of day of the scene (such as morning, night, etc.), the geographical location of the scene (such as mountainous area, seaside, etc.), the style of the scene (such as retro or modern), and the target objects in the scene (such as furniture, outdoor facilities, etc.).
[0060] It should be noted that the adaptability characteristics of the avatar image can be determined based on various combinations of multi-dimensional characteristics of the first person and / or scene, rather than relying on a single characteristic.
[0061] Since the target image is taken by the user as a single person, the proportions of the characters in the fused image may be out of balance after adding the second person's avatar image to the group photo. Therefore, in some embodiments, method 300 may also adjust the target image by first changing the proportion of the first person compared to the target image or expanding the background of the target image, and then integrate the avatar image into the adjusted target image. In this way, the quality and coordination of the generated fused image can be guaranteed, so as to be closer to the real group photo effect.
[0062] In some cases, the target image may not have enough space to accommodate the newly added avatar image of the second person, or directly adding the avatar image may cause the layout of the characters in the target image to be unreasonable or unsightly. At this time, the target image can be adjusted. In some embodiments, the target image can be adjusted by changing the ratio of the first person compared to the target image or expanding the background of the target image, and then the avatar image can be integrated into the adjusted target image. Changing the ratio of the first person compared to the target image can provide more space. For example, the size of the first person can be reduced and moved to a more distant position (this can ensure the coordination of the person and the background), thereby allowing more people to be added to the original photo. By expanding the background of the target image without changing the first person, the effect of generating an image with a wider view can be achieved.
[0063] In some embodiments, the avatar image may include multiple characters, and the ratio of the existing characters to the original photo may be changed or the background thereof may be expanded based on the number of the multiple characters. For example, the more characters there are in the avatar image, the smaller the existing characters may be adjusted and the larger the expanded background may be.
[0064] Since the avatar image needs the authorization of the person to use, the user can send a request to the second person to use the second person's avatar image to take a photo together. If the second person agrees to the request, he can send the user his authorization for all of his avatar images, or select several avatar images from multiple avatar images for authorization and reply to the user. The user can continue with the image fusion only after obtaining the authorized avatar image or the second person's authorization for the avatar image. In this way, the privacy and security of personal information can be guaranteed.
[0065] To further meet the user's customization needs and improve playability, in some embodiments, method 300 may further include: displaying a composite image and editing controls for the composite image; in response to the editing controls being triggered, displaying an input area; and adjusting the composite image based on the content input in the input area. When the generated composite image does not meet the user's expectations, or the user wishes to add some special effects to the composite image, the editing controls can be triggered in the form of a click or the like after the composite image is generated, and the user's own needs can be input to achieve a custom editing effect.
[0066] For example, if the user feels that the proportion of the doppelganger image in the composite image is not good, the pose is not natural enough, is not satisfied with the clothing, or the interactivity between the characters is not strong enough, the user can adjust the proportion, pose, etc. of the doppelganger image in the input area. For another example, if the user hopes to hide the real portrait information of himself or the second person before reposting and sharing, an instruction to cartoonize the doppelganger image of the second person can be input after the composite image is generated to meet the requirement.
[0067] Figure 4 FIG. shows a schematic block diagram of an apparatus 400 for generating an image according to an embodiment of the present disclosure. As Figure 4 shown, the apparatus 400 includes a feature information acquisition unit 402 and a composite image generation unit 404. The feature information acquisition unit 402 is configured to acquire the feature information of a target image including a first person, and the feature information includes multi-dimensional features of the first person. The composite image generation unit 404 is configured to adaptively integrate the doppelganger image of a second person into the target image based on the feature information to generate a composite image.
[0068] It should be noted that more elements as Figures 1A to 3 shown can be implemented by Figure 4 the apparatus 400 shown. For example, the apparatus 400 may include more modules or units to implement the elements described above, or Figure 4 some of the units or modules shown can be further configured to implement the elements described above. Details are not repeated here.
[0069] The above has described the interface design and interaction process for generating a composite image and creating a doppelganger image according to an embodiment of the present disclosure with reference to FIGS. 1 to Figure 4 According to an embodiment of the present disclosure, when the user takes a photo, the doppelganger image of a friend can be called in real time to complete a group photo. The position and pose of the doppelganger image will adapt in real time according to the content in the camera view, the user's position and pose. In this way, it is not necessary for the friend to be around, and the user can take a group photo with the friend anytime and anywhere, thereby improving the user experience.
[0070] Figure 5FIG. shows a schematic block diagram of an exemplary device 500 that may be used to implement embodiments of the present disclosure. As shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 506 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0071] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a touch screen, a keyboard, a mouse, etc.; an output unit 507, such as various types of displays (e.g., an interactive display, such as a touch screen), a speaker, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0072] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as method 300. For example, in some embodiments, method 300 can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of method 300 described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute method 300 in any other suitable manner (e.g., by means of firmware).
[0073] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.
[0074] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0075] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0076] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages and conventional procedural programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0077] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause a computer, a programmable data - processing apparatus, and / or other devices to work in a particular manner. Thus, the computer - readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0078] The computer - readable program instructions can also be loaded onto a computer, other programmable data - processing apparatus, or other devices so that a series of operation steps are executed on the computer, other programmable data - processing apparatus, or other devices to produce a computer - implemented process, such that the instructions executed on the computer, other programmable data - processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0080] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for generating an image, comprising: Acquire feature information of a target image including a first person, wherein the feature information includes multi-dimensional features of the first person; as well as Based on the feature information, the clone image of the second person is adaptively integrated into the target image to generate a fused image.
2. The method according to claim 1, wherein the multi-dimensional feature of the first person describes at least one of the following: Shot selection, position, pose, expression, size and clothing. 3 . The method according to claim 1 , wherein the feature information further includes multi-dimensional features of the scene in the target image.
4. The method according to claim 3, wherein the multi-dimensional features of the scene describe at least one of the following: the category of the scenario; the season of said scene; The time of day of the described scene; The geographical location of the scene; the style of the scene; and The target object in the scene.
5. The method according to claim 1, wherein adaptively integrating the avatar image of the second person into the target image comprises: Determining an adaptive feature of the avatar image based on the feature information; as well as A fused image is generated based on the target image, the clone image and the adaptive feature.
6. The method of claim 5, wherein generating the fused image comprises: generating an image of the second person based on the avatar image and the adaptability feature; as well as The image of the second person and the target image are fused to obtain the fused image.
7. The method according to claim 1, wherein adaptively integrating the avatar image of the second person into the target image comprises: adjusting the target image by changing the proportion of the first person compared to the target image or expanding the background of the target image; as well as The clone image is integrated into the adjusted target image.
8. The method of claim 7, wherein the second person comprises a plurality of persons, and adjusting the target image further comprises: Based on the number of the plurality of characters, the ratio of the first character to the target image is changed or the background of the target image is expanded.
9. The method according to claim 1, further comprising: Sending a request for taking a photo with the avatar image of the second person; as well as Obtain the authorized avatar image or the authorization of the avatar image by the second person.
10. The method according to claim 1, further comprising: The target image is acquired by shooting with a camera or selecting a stored image.
11. The method according to claim 1, further comprising: displaying the fused image and an editing control for the fused image; In response to the editing control being triggered, displaying an input area; as well as The fused image is adjusted based on the content input in the input area.
12. A system for generating an image, comprising: A feature information acquisition unit, configured to acquire feature information of a target image including a first person, wherein the feature information includes a multi-dimensional feature of the first person; as well as The fused image generating unit is configured to adaptively integrate the clone image of the second person into the target image based on the feature information to generate a fused image.
13. A computing device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 11.
14. A computer storage medium comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.
15. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.