Method, apparatus, device and program product for generating image
By acquiring the multi-dimensional features of the scene image and adaptively integrating the clone image of the target person into the scene image, the problems of high image fusion cost and threshold in the prior art are solved, and a convenient and efficient image fusion effect is achieved.
Patent Information
- Application Number
- CN202510300462.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to integrate the clone image of the target person into the scene image with a low threshold and efficient manner, resulting in a high production cost and threshold for image fusion.
By obtaining multi-dimensional features in the scene image and using these features to adaptively integrate the clone image of the target character into the scene image, a fusion image is generated.
The production cost and threshold of fusion images are reduced, making the image fusion process more convenient and efficient.
Smart Images

Figure CN120147481A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method, an apparatus, a computing device, a computer-readable storage medium, and a computer program product for generating images. Background Art
[0002] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. It is a branch of computer science and aims to produce intelligent machines that can respond in a manner similar to human intelligence. The research fields of AI include language and image recognition, robotics, natural language processing, and expert systems, etc.
[0003] The AI image generation technology uses artificial intelligence algorithms. Through the learning and analysis of a large amount of image data, it can quickly generate images with creativity and specific styles according to instructions such as text descriptions and reference images input by users. It covers various functions such as text-to-image, image editing, and image restoration, bringing new ideas and efficient solutions to image creation. Summary of the Invention
[0004] The present disclosure provides a method, an apparatus, a computing device, a computer-readable storage medium, and a computer program product for generating images. By acquiring multi-dimensional features of a scene in a scene image and using these features to adaptively integrate a clone image of a target person into the scene image, it is possible to reduce the production cost and threshold of the integrated image.
[0005] According to a first aspect of the present disclosure, there is provided a method for generating an image. The method includes acquiring feature information of a scene image, where the feature information includes multi-dimensional features of a scene in the scene image. The method further includes adaptively integrating a clone image of a target person into the scene image based on the feature information to generate an integrated image.
[0006] According to a second aspect of the present disclosure, there is provided an apparatus for generating an image. The apparatus includes a feature information acquisition unit configured to acquire feature information of a scene image, where the feature information includes multi-dimensional features of a scene in the scene image. The apparatus further includes an integrated image generation unit configured to adaptively integrate a clone image of a target person into the scene image based on the feature information to generate an integrated image.
[0007] According to a third aspect of the present disclosure, there is provided a computing device, including: at least one processing unit; at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the computing device to execute the method as described in the first aspect of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer storage medium, including machine-executable instructions, the machine-executable instructions when executed by a device causing the device to execute the method as described in the first aspect of the present disclosure.
[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product, including machine-executable instructions, the machine-executable instructions when executed by a device causing the device to execute the method as described in the first aspect of the present disclosure.
[0010] The Summary of the Invention is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the embodiments of the present disclosure will become more readily understood. In the drawings, multiple embodiments of the present disclosure will be illustrated by way of example and not limitation, where:
[0012] Figure 1A A schematic diagram of a display interface according to an embodiment of the present disclosure is shown;
[0013] Figure 1B A schematic diagram of a display interface in which a split-image is integrated according to an embodiment of the present disclosure is shown;
[0014] Figure 1C A schematic diagram of a stylized display interface according to an embodiment of the present disclosure is shown;
[0015] Figure 1D A schematic diagram of a display interface including a prompt word input area according to an embodiment of the present disclosure is shown;
[0016] Figure 2A A schematic diagram of a display interface according to an embodiment of the present disclosure is shown;
[0017] Figure 2B A schematic diagram of a display interface after uploading a portrait picture according to an embodiment of the present disclosure is shown;
[0018] Figure 3 FIG. 1 shows a schematic flowchart of a method for generating an image according to an embodiment of the present disclosure;
[0019] Figure 4 FIG. 2 shows a block diagram of an apparatus for generating an image according to an embodiment of the present disclosure; and
[0020] Figure 5 FIG. 3 shows a block diagram of an electronic device according to an embodiment of the present disclosure.
[0021] In all the figures, the same or similar reference numerals denote the same or similar elements. DETAILED DESCRIPTION
[0022] The present disclosure will now be described with reference to several exemplary implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, rather than to imply any limitation on the scope of the present disclosure.
[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0024] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0025] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0028] In the description of the embodiments of the present disclosure, the term "comprising" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.
[0029] In this article, a smart phone is taken as an example to show an exemplary user interface. However, those skilled in the art can understand that the embodiments of the present disclosure are equally applicable to other devices with a display screen that may have different screen aspect ratios, such as tablet computers, laptop computers, desktop computers, wearable devices with a screen, and devices with a foldable screen, etc. In addition, the interfaces provided herein are for illustrative purposes only. Some of the elements may be omitted or have different numbers, and there may also be more elements not shown. Furthermore, the interfaces according to the embodiments of the present disclosure may have a layout different from the interfaces shown in the drawings, and the positions of the respective elements may also be different. The present disclosure makes no limitations in these aspects.
[0030] The AI avatar is a new application technology that integrates computer vision, natural language processing, and deep learning, etc. The avatar images generated based on the AI avatar technology can simulate the appearance and image of a specific subject and simulate the activities and interactions of the specific subject in various scenarios. For example, in the social field, it can interact on behalf of users; in content creation, it can be a virtual anchor; in the education industry, it can be a tutoring assistant; in customer service, it can be a virtual customer service, etc. More application scenarios of the AI avatar need to be further explored to meet the in-depth needs of users using related application software.
[0031] The inventors noticed some problems. Some users hope to take pictures of people in specific scenarios, but this cannot be achieved because the scenario is far from the user and there is no time to go there, or the photo taken has a poor effect due to limited personal selfie angles. Although it is also possible to add people to the already taken scene photos through post-processing of images (for example, using specialized image processing tools), this requires cumbersome post-production and requires the user to have relevant processing techniques, resulting in a relatively high learning cost. Therefore, it would be beneficial to provide a convenient and low-threshold image fusion method based on AI avatars.
[0032] In view of this, the present disclosure provides a method for generating an image, which adaptively integrates the avatar image of a target person into the scene image by obtaining multi-dimensional features of the scene in the scene image and using these features, thereby being able to reduce the production cost and threshold of the fused image. The following will refer to Figures 1A to 5 the embodiments of the present disclosure in detail.
[0033] Figure 1A shows a schematic diagram of a display interface 100A according to some embodiments of the present disclosure. Herein, the display interface 100A is an example of a page of an avatar camera according to the embodiments of the present disclosure. When the user clicks on a control with the function of an avatar camera on the application, the electronic device can display Figure 1A the display interface 100A shown. Among them, the electronic device can be a terminal with a camera and a display screen, such as a mobile phone, a tablet computer, etc. For example, the user can, through an image generation application of the electronic device, after creating an avatar image, generate a fused image automatically by taking or uploading a scene image. Regarding the creation of the avatar image, it will be described in detail in the following Figures 2A - 2B section.
[0034] As Figure 1A shown, the display interface 100A may include an image display area 102, an upload control 104, and a camera control 106. The image display area 102 can display the scene captured in real time by the camera of the electronic device that the user is using. In some implementations, the display interface 100A may further include a camera switching control 108 to facilitate the user to freely switch between the front camera, the rear camera, etc. when using a multi-camera device. When the user clicks on the camera control 106, the image display area 102 can display the scene image captured by the user, and then automatically start to adaptively integrate the avatar image created by the user into the scene to generate a fused image. The user can also click on the upload control 104 to select any scene image from the local album, and then similarly automatically start generating a fused image.
[0035] Some users may create multiple avatar images with different styles, and the avatar image that best matches the scene is selected by default during automatic fusion. In some implementations, the display interface 100A may also include an avatar image selection control 103, and the user can click the avatar image selection control 103 to manually select the desired avatar image and merge it into the scene image. The avatar image can be the user's own, or it can be an avatar image shared and authorized for use by other users. It can be understood that the display interface 100A of Figure 1 is for illustrative purposes only, and the avatar camera page can have different layouts or designs, and the present disclosure does not limit this.
[0036] Figure 1B FIG. 1 is a schematic diagram of a display interface 100B in which a clone image is integrated according to an embodiment of the present disclosure. Figure 1B As shown, after the fused image is generated, the fused image can be displayed in the image display area 102, wherein the features of the avatar image can be adapted to the scene image. In some embodiments, the adaptive features of the avatar image can be determined based on the feature information of the scene image uploaded or photographed by the user. The feature information of the scene image may include multi-dimensional features of the scene, such as the category of the scene (e.g., natural scenery or buildings), the season of the scene, the time of the scene (day or night), the geographical location of the scene, the style of the scene, and the target object in the scene. The adaptive features of the avatar image may include scene selection (close-up or long-range view), position, posture, expression, size, and clothing.
[0037] like Figure 1B As shown, the scene in the image display area 102 includes rich elements such as lakes, fountains, and high-rise buildings in the distance. The avatar image stands in front of the railing, and the railing as a linear element will guide the audience's sight. If the avatar image is in the center of the picture, the lines of the railing can be used to attract the audience's attention to the avatar image, which helps to balance the picture. At the same time, the high-rise buildings with bright lights in the distance serve as the background, and a relatively stable visual area can be formed behind the picture. The avatar image has his hands behind the railing and smiles. This relaxed and pleasant state is also consistent with the cozy atmosphere of the lake at night. The black hair of the avatar image naturally hangs on the left, and this hairstyle itself is also relatively natural. The avatar image wears a black low-necked sling and holds a white coat in his right hand. This kind of dress is more suitable for summer nights, which is both cool and convenient for coping with possible temperature differences at night. Black echoes the night environment, and the white coat is used as an embellishment to add highlights and coordinate with the overall scene tone.
[0038] In addition to the fused image, the display interface 100B also displays a download control 110 and an edit control 112. The user can click on the download control 110 to download the generated fused image to the local for storage. The edit control 112 can be used to adjust the fusion effect of the generated fused image. In addition, the camera control 106 can be switched to display a static icon indicating that the fusion is complete.
[0039] In Figure 1C FIG. shows a schematic diagram of the stylized display interface 100C according to an embodiment of the present disclosure. In some embodiments, the user can select a preset style template to quickly switch the style of the fused image, such as a cartoon style, a comic style, etc. As Figure 1C shown, the user can switch the preset style template, for example, by swiping the screen to the left, to select the desired fusion effect.
[0040] Figure 1D FIG. shows a schematic diagram of the display interface 100D including a prompt input area according to an embodiment of the present disclosure. In some embodiments, the user can further adjust the effect on the basis of the preset style template. As Figure 1D shown, the user can click on the edit control 112 to pop up a prompt input area 114 on the current page. The user can enter the desired effect in the prompt input area 114 and click send. For example, the user input can indicate changing the image, action, or expression of the avatar image in the figure, or changing the style of the fused image, etc. Then, the image display area 102 will display the fused image adjusted according to the prompt.
[0041] The following will be combined with Figures 2A - 2B to detail the process of creating an avatar image. Figure 2A FIG. shows a schematic diagram of the display interface 200A according to an embodiment of the present disclosure. Herein, the interface 200A is an example of an avatar creation page according to an embodiment of the present disclosure. When the user clicks on a control with an avatar image creation function on other pages (such as the dialogue interface of the AI large language framework, the personal information interface of the image generation application, etc.), the electronic device can switch to display Figure 2A the interface 200A as shown.
[0042] As Figure 2A shown, the interface 200A may include a sample area 202, a portrait upload area 204, and a generation control 206. The sample area 202 can be used to guide the user to upload a correct and high-quality portrait picture. The sample area 202 can include at least one correct sample and at least one incorrect sample. As Figure 2AAs shown, the example area 202 may include a correct example 202-1 that guides the user to upload a clear frontal portrait picture, a correct example 202-2 that guides the user to upload portrait pictures with multiple backgrounds and from multiple angles, a wrong example 202-3 that guides the user to avoid uploading portrait pictures with the face blocked, and a wrong example 202-4 that guides the user to avoid uploading portrait pictures with a too small face. The user can click the upload control 204-2 in the portrait upload area 204 according to the guidance to upload at least one real portrait picture including the same person, so that the AI model can extract facial features to generate a clone image. If the user does not upload a portrait picture or the number of uploaded portrait pictures is insufficient, the generate control 206 can be set to an unselectable state.
[0043] Figure 2B FIG. shows a schematic diagram of a display interface 200B after uploading a portrait picture according to an embodiment of the present disclosure. As Figure 2B shown, after the user uploads enough portrait pictures, the display effect of the upload control 204-4 can be changed, and the generate control 206 can be changed from an unselectable state to a selectable state. The user can complete the creation of the clone image and return by clicking the generate control 206.
[0044] Figure 3 FIG. shows a flowchart of a method 300 for image generation according to some embodiments of the present disclosure. The method 300 can be implemented by any electronic device having computing and display capabilities and including a camera and a display screen.
[0045] In block 310, the method 300 includes obtaining feature information of a scene image, where the feature information includes multi-dimensional features of the scene in the scene image. The scene image can be obtained by the user taking a real-time photo with a camera or selecting a stored picture. In some embodiments, the multi-dimensional features of the scene can describe the category of the scene (indoor or outdoor), the season of the scene (e.g., spring, winter, etc.), the time of day of the scene (e.g., morning, night, etc.), the geographical location of the scene (e.g., mountainous area, seaside, etc.), the style of the scene (e.g., vintage or modern, etc.), and the objects in the scene (e.g., desk, bench, etc.). A trained AI model can be used to extract this information from the scene image.
[0046] In block 320, the method 300 includes adaptively integrating a clone image of a target person into the scene image based on the feature information to generate a fused image.
[0047] In some embodiments, the clone image can be directly fused into the scene image. In some embodiments, based on the feature information, it is determined what adaptive features the clone image has, and then based on these adaptive features, the clone image and the scene image, a fused image is directly generated. Optionally, the user can also provide text content to indicate the desired effects or features.
[0048] In some embodiments, a clone image with adaptive features can be generated first, and then the generated clone image can be fused into a scene image to obtain a fused image. For example, multiple clone images with different adaptive features can be generated, and the user can select a clone image that they are satisfied with, or indicate the desired effect or features to regenerate it. Then, the clone image confirmed by the user is fused into the scene image.
[0049] Determining the adaptive features of the clone image may include determining the following features of the clone image, including but not limited to: shot selection, position, pose, expression, size, and clothing, etc. The determined adaptive features match the content described by the feature information of the scene image. The following lists several specific examples to further illustrate how to determine the adaptive features of the clone image based on the multi-dimensional features of the scene.
[0050] The adaptive features of the clone image can depend on the type of the scene. For example, in an indoor scene, the clone image character can be a sitting half-body image or a more relaxed standing full-body image, while in an outdoor sports scene, a clone image with actions such as running and jumping is selected.
[0051] The adaptive features of the clone image can depend on the season of the scene. For example, for a spring scene, the clone image character can be adjusted to wear a light long-sleeved shirt, paired with a light-colored skirt or pants, with bright and lively colors such as pink and light blue, to adapt to the vibrant atmosphere of spring. In terms of expression, with a smiling face, showing the love and joy for the beautiful scenery of spring. Such character features can better blend with the seasonal characteristics, reflecting the vitality and beauty of spring. For a winter scene, the clone image character is suitable to wear a down jacket, with a hat, scarf, etc., mainly in dark colors such as black and dark blue, to play a role in keeping warm and coordinating with the winter environment. The pose can be huddling the body to resist the cold wind, or walking slowly in the snow, showing the cold of winter. Such character features conform to the actual situation of winter and people's general perception of winter.
[0052] The adaptive features of the clone image can depend on the time of day. If the scene reflects the morning time, with soft sunlight and the sky showing a faint blue or orange color, the clone image character can wear comfortable sportswear, such as jogging suits and sports shoes, in a pose of morning running or morning exercise. While in the evening scene, as the sun sets, the sky is dyed orange-red and the light is relatively soft, the character is suitable to wear relatively casual clothes, such as a loose shirt paired with jeans, sitting on a park bench enjoying the sunset, or walking slowly on a path, with a relaxed and natural pose.
[0053] The adaptive features of the avatar image can depend on the geographical location of the scene. For example, for a seaside scene, the avatar image character can wear a swimsuit, shorts or beach skirt, a sun hat, and slippers, and can be running or playing on the beach, or lying on a beach chair in the sun with a relaxed and happy expression. For a mountain scene, the character can wear outdoor equipment, such as mountaineering clothes and hiking shoes, carry a backpack, and hold a trekking pole, showing the posture of climbing or walking on a mountain road.
[0054] The adaptive features of the avatar image can depend on the style of the scene. For example, in a retro-style scene, retro furniture and decorations are often arranged with rich and calm colors. Accordingly, the avatar image characters wear retro-style clothing, such as retro dresses for women and retro suits for men, and the hairstyles and accessories also follow the retro style. For modern minimalist style scenes, the characters can wear simple modern clothing.
[0055] The adaptive features of the avatar image can depend on the target object in the scene. For example, an amusement park scene may have amusement facilities such as a Ferris wheel and a roller coaster, so the avatar image character can wear colorful casual clothes, or have expressions and postures of screaming and laughing on the amusement facilities. In a library scene with bookshelves, desks and a large number of books, the character can wear neat clothes, such as plain shirts and long skirts, and look for books among the bookshelves, or sit quietly at the desk and read with a focused and serious expression.
[0056] In some embodiments, the target person may have multiple avatar images, each of which has its own characteristics, such as different appearances (such as hairstyles, clothes), different expressions (happy, sad, etc.), different personalities (extroverted or introverted), etc. Method 300 may also include selecting an avatar image from the multiple avatar images of the target person based on feature information or user selection. For example, the user takes or selects a photo such as Figure 1A The scene image shown in the figure is a lakeside in the summer night. When automatically generating a fusion image, a casual style avatar image that best matches the quiet scene can be selected by default. If the user who has multiple styles of avatar images is not satisfied with the style of the default selected avatar image, he or she can also manually select another style of avatar image.
[0057] In some embodiments, the avatar image is generated based on at least one photo of the target person and is authorized for use by the target person. Each avatar image is generated by a group of photos uploaded by the user at one time. Thus, the user can upload groups of photos of different styles (e.g., smiling, happy, sad, etc.) to generate a variety of avatar images.
[0058] In some embodiments, method 300 may also generate a clone image through the following steps: display a clone creation control for triggering the creation of a clone image; in response to the clone creation control being triggered, display an upload area for uploading photos and a clone image generation control; and in response to one or more photos of the target person being uploaded and the clone image generation control being triggered, generate a clone image.
[0059] The clone creation control can be displayed in various interfaces, including but not limited to the dialogue interface of the AI large language framework and the personal information interface of the image generation application. When the user triggers the clone creation control in the form of a mouse / finger click, etc., the user can select a certain number of photos by themselves for uploading to generate a clone image.
[0060] To further meet the user's customization needs, in some embodiments, method 300 may further include displaying a composite image and an editing control for the composite image; in response to the editing control being triggered, display an input area; and based on the content input in the input area, adjust the composite image. When the generated composite image does not meet the user's expectations, or the user hopes to add some special effects to the composite image, the editing control can be triggered in the form of a click, etc. after the composite image is generated and the user's own requirements can be input to achieve a customized editing effect. For example, if the user feels that the proportion of the clone image in the composite image is not good, the pose is not natural enough, or is not satisfied with the clothing, the user can adjust the proportion, pose, etc. of the clone image in the input area. Another example is that if the user hopes to hide their true portrait information before republishing and sharing, the user can input an instruction to cartoonize the clone image after the composite image is generated to meet the requirement.
[0061] Figure 4 FIG. shows a schematic block diagram of an apparatus 400 for generating an image according to an embodiment of the present disclosure. As Figure 4 shown, the apparatus 400 includes a feature information acquisition unit 402 configured to acquire feature information of a scene image, where the feature information includes multi-dimensional features of the scene in the scene image. The apparatus 400 further includes a composite image generation unit 404 configured to adaptively integrate a clone image of a target person into the scene image based on the feature information to generate a composite image.
[0062] It should be noted that more elements shown in FIGS. 1 to Figure 3 shown can be implemented by the Figure 4 shown apparatus 400. For example, the apparatus 400 may include more modules or units to implement the elements described above, or Figure 4 shown some units or modules may be further configured to implement the elements described above. This will not be elaborated here again.
[0063] The above reference Figures 1A to 4Describes the interface design and interaction process for generating a fused image and creating an avatar image according to an embodiment of the present disclosure. Embodiments of the present disclosure can adaptively integrate an avatar image of a target person into a scene image by obtaining multi-dimensional features of the scene in the scene image, thereby reducing the production cost and threshold of the fused image.
[0064] Figure 5 Shows a schematic block diagram of an example device 500 that can be used to implement embodiments of the present disclosure. As shown, device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 506 into a random access memory (RAM) 503. In RAM 503, various programs and data required for the operation of device 500 can also be stored. The computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0065] Multiple components in device 500 are connected to the I / O interface 505, including: an input unit 506, such as a touch screen, keyboard, mouse, etc.; an output unit 507, such as various types of displays (e.g., an interactive display, such as a touch screen), speakers, etc.; a storage unit 508, such as a disk, optical disc, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0066] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as method 300. For example, in some embodiments, method 300 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of method 300 described above may be executed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute method 300 by any other suitable means (e.g., by means of firmware).
[0067] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0068] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0069] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0070] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0071] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0072] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0073] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0074] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of technologies in the market, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating an image, comprising: Acquire feature information of a scene image, wherein the feature information includes multi-dimensional features of a scene in the scene image; as well as Based on the feature information, the target person's avatar image is adaptively integrated into the scene image to generate a fused image.
2. The method according to claim 1, wherein the multi-dimensional feature describes at least one of the following: the category of the scenario; the season of said scene; The time of day of the described scene; The geographical location of the scene; the style of the scene; and The target object in the scene.
3. The method according to claim 1 or 2, wherein adaptively integrating the target person's avatar image into the scene image comprises: Determining an adaptive feature of the avatar image based on the feature information; as well as The fused image is generated based on the scene image, the clone image and the adaptive feature.
4. The method according to claim 3, wherein generating the fused image comprises: Based on the avatar image and the adaptability feature, generating an image of the target person; as well as The image of the target person and the scene image are fused to obtain the fused image.
5. The method according to claim 3, wherein the adaptive features of the avatar image include at least one of the following: Shot selection, position, pose, expression, size and clothing.
6. The method according to claim 1, further comprising: Based on the feature information or user selection, the avatar image is selected from a plurality of avatar images of the target person.
7. The method according to claim 1, further comprising: The scene image is acquired by shooting with a camera or selecting a stored picture. 8 . The method according to claim 1 , wherein the avatar image is generated based on at least one photo of the target person and is authorized for use by the target person.
9. The method according to claim 1, further comprising generating the avatar image by the following steps: Displays an avatar creation control for triggering the creation of an avatar image; In response to the avatar creation control being triggered, displaying an upload area for uploading photos and an avatar image generation control; and In response to one or more photos of the target person being uploaded and the avatar image generation control being triggered, the avatar image is generated.
10. The method according to claim 1, further comprising: displaying the fused image and an editing control for the fused image; as well as In response to the editing control being triggered, displaying an input area; as well as The fused image is adjusted based on the content input in the input area.
11. A system for generating an image, comprising: A feature information acquisition unit is configured to acquire feature information of a scene image, wherein the feature information includes multi-dimensional features of a scene in the scene image; as well as The fused image generating unit is configured to adaptively integrate the target person's avatar image into the scene image based on the feature information to generate a fused image.
12. A computing device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 10.
13. A computer storage medium comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 10.
14. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 10.