An image generation method, apparatus, electronic device, and storage medium

By adjusting the key points of the target person's face and adding specified stickers, the problem of template GIF images failing to accurately express emotions was solved, and the generated images can better express emotions, thus improving the user experience.

CN117274309BActive Publication Date: 2026-04-03JOYME PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, template GIF images can only display template characters with changing facial expressions, which cannot accurately express emotions. As a result, the generated GIF images also cannot accurately express emotions, leading to a poor user experience.

Method used

By acquiring template information and the image to be processed, the positions of key facial features of the target person are adjusted, specified stickers are added to indicate emotions, and a dynamic image format is generated to ensure that the facial expressions of the target person are consistent with those of the template person.

Benefits of technology

It improves the ability of generated images to express emotions and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274309B_ABST
    Figure CN117274309B_ABST
Patent Text Reader

Abstract

This invention provides an image generation method, apparatus, electronic device, and storage medium, relating to the field of image processing technology. The method includes: acquiring template information and an image to be processed; adjusting the position of facial key points of a target person in the image to be processed according to a first motion trajectory to obtain a first video to be processed; adjusting the size of a designated sticker to which a designated sticker identifier belongs to a target size, and adding a designated sticker of the target size at a target position according to a target angle in each first video frame of the first video to be processed to obtain a second video frame corresponding to the first video frame; combining the second video frames corresponding to each first video frame in the first video to be processed to obtain a second video to be processed; and converting the second video to be processed into a dynamic image format to obtain a target image. The target image can better express emotions and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of image processing technology, image processing platforms are providing users with an increasing number of image processing functions. For example, image processing platforms can be apps (applications) on the Android system. Users can modify the colors of images through image processing platforms, or they can use image processing platforms to convert static images into GIF (Graphics Interchange Format) images.

[0003] In existing technology, image processing platforms can display template GIF images to users, where the facial expressions of the template characters can change. When the image processing platform receives a request including an identifier for a static image displaying a character and an identifier for a template GIF image, it can process the static image displaying the character according to the way the facial expressions of the template character in the template GIF image change, thus obtaining a GIF image corresponding to the static image. The way the facial expressions of the character displayed in the GIF image corresponding to the static image change is the same as the way the facial expressions of the template character in the template GIF image change. For example, if the template character in the template GIF image is blinking, then the character in the static image in the processed GIF image will also be blinking.

[0004] However, template GIFs only display template characters with changing expressions, which cannot effectively convey emotions. For example, when a template character blinks in a template GIF, it's impossible to determine whether the character is happy or sad. Consequently, GIFs generated from template GIFs that cannot adequately express emotions also fail to convey emotions effectively, resulting in a poor user experience. Summary of the Invention

[0005] The purpose of this invention is to provide an image generation method, apparatus, electronic device, and storage medium to obtain a target image that can better express the emotions that the target user wants the target character to express, thereby improving the user experience. The specific technical solution is as follows:

[0006] In a first aspect of the present invention, an image generation method is provided, the method comprising: acquiring template information and an image to be processed; wherein the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image containing facial key points of a template person changing according to a first motion trajectory; a specified sticker to which the specified sticker identifier belongs is used to represent: the emotion expressed by the template person in the specified template image; the image to be processed is a static image displaying a target person; adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain a first video to be processed; adjusting the size of the specified sticker to which the specified sticker identifier belongs to a target size, and adding the specified sticker of the target size at a target position and according to a target angle at each first video frame in the first video to be processed to obtain a second video frame corresponding to the first video frame; combining the second video frames corresponding to each first video frame in the first video to be processed to obtain a second video to be processed; and converting the second video to be processed into a dynamic image format to obtain a target image.

[0007] Optionally, obtaining template information and the image to be processed includes: obtaining the image identifier of a specified template image and the image to be processed; and obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between the image identifier and the sticker identifier, thereby obtaining the specified sticker identifier.

[0008] Optionally, before obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, the method further includes: obtaining a template video and multiple reference motion trajectories; wherein, the template video displays the template character; the reference motion trajectories are used to determine the emotion expressed by the template character displayed in the template video; the reference motion trajectories correspond to stickers; a sticker corresponding to a reference motion trajectory is used to represent: the emotion expressed by the template character when facial key points change according to the reference motion trajectory; obtaining the change trajectory of the facial key points of the template character in the template video to obtain a first motion trajectory; calculating the similarity between the first motion trajectory and each reference motion trajectory, and determining the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video; determining the template sticker corresponding to the template video from multiple stickers corresponding to the reference motion trajectory corresponding to the template video; converting the template video into a dynamic image format to obtain a template image; recording the correspondence between the image identifier of the template image and the sticker identifier of the template sticker to obtain the correspondence between image identifiers and sticker identifiers.

[0009] Optionally, obtaining the change trajectory of the facial key points of the template person in the template video to obtain the first motion trajectory includes: for each template video frame of the template video, determining the position of the facial key points of the template person in that template video frame; and obtaining the change trajectory of the position of the facial key points of the template person in the template video as the first motion trajectory according to the change of the position of the facial key points of the template person in each template video frame.

[0010] Optionally, the step of recording the correspondence between the image identifier of the template image and the sticker identifier of the template sticker to obtain the correspondence between the image identifier and the sticker identifier includes: adding the template sticker corresponding to the template video to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

[0011] Optionally, the step of obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, and obtaining the specified sticker identifier, includes: obtaining the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and obtaining the sticker identifier corresponding to the image identifier of the specified template image as the specified sticker identifier.

[0012] Optionally, obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, and obtaining the specified sticker identifier, includes: obtaining the sticker identifier of the target type corresponding to the image identifier of the specified template image according to the pre-recorded correspondence between image identifiers and sticker identifiers, and displaying the obtained sticker identifier of the target type; from the displayed sticker identifier of the target type, determining, according to the selection instruction for the specified sticker identifier in the sticker identifier of the target type, that the sticker identifier indicated by the selection instruction input by the user is the specified sticker identifier corresponding to the image identifier of the specified template image, and obtaining the specified sticker identifier.

[0013] Optionally, adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed includes: copying the image to be processed according to a first number of positions of facial key points in the first motion trajectory to obtain the first number of images to be processed; wherein, one image to be processed corresponds to one position of a facial key point in the first motion trajectory; for each position of a facial key point in the first motion trajectory, adjusting the position of the facial key point of the target person in the image to be processed corresponding to the position of the facial key point in the first motion trajectory to obtain the first video frame to be processed; sorting each first video frame to be processed according to the order of the positions of each facial key point in the first motion trajectory, and combining the sorted first video frames to be processed into the first video to be processed.

[0014] Optionally, before adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed, the method further includes: detecting whether the facial features of the target person are displayed in the image to be processed; if the facial features of the target person are not displayed in the image to be processed, outputting a prompt message indicating that the image to be processed should be re-uploaded; adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed includes: if the facial features of the target person are displayed in the image to be processed, adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0015] In a second aspect of the invention, an image generation apparatus is also provided, the apparatus comprising:

[0016] The first acquisition module is used to acquire template information and an image to be processed; wherein, the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the facial key points of the template character change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent the emotion expressed by the template character in the specified template image; the image to be processed is a static image displaying a target character;

[0017] The first adjustment module is used to adjust the position of the facial key points of the target person in the image to be processed according to the first motion trajectory, so as to obtain the first video to be processed;

[0018] The second adjustment module is used to adjust the size of the specified sticker to which the specified sticker identifier belongs to the target size, and add the specified sticker of the target size at the target position in each first video frame in the first video to be processed according to the target angle to obtain the second video frame corresponding to the first video frame;

[0019] The combination module is used to combine the second video frames corresponding to each first video frame in the first video to be processed to obtain the second video to be processed.

[0020] The format conversion module is used to convert the second video to be processed into a dynamic image format to obtain the target image.

[0021] Optionally, the first acquisition module is specifically used to: acquire the image identifier of the specified template image and the image to be processed;

[0022] Based on the pre-recorded correspondence between image identifiers and sticker identifiers, the sticker identifier corresponding to the image identifier of the specified template image is obtained, thus obtaining the specified sticker identifier.

[0023] Optionally, the device further includes:

[0024] The second acquisition module is used to acquire a template video and multiple reference motion trajectories before the first acquisition module executes the pre-recorded correspondence between image identifiers and sticker identifiers to acquire the sticker identifier corresponding to the image identifier of the specified template image and obtains the specified sticker identifier. The template video displays the template character; the reference motion trajectories are used to determine the emotions expressed by the template character displayed in the template video; the reference motion trajectories correspond to template stickers; and the template stickers are used to indicate the emotions corresponding to the corresponding reference motion trajectories.

[0025] The third acquisition module is used to acquire the change trajectory of the facial key points of the template character in the template video to obtain the first motion trajectory.

[0026] The similarity calculation module is used to calculate the similarity between the first motion trajectory and each reference motion trajectory, and to determine the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video.

[0027] The template sticker determination module is used to determine the template sticker corresponding to the template video from the template stickers corresponding to the reference motion trajectory corresponding to the template video;

[0028] The recording module is used to convert the template video into a dynamic image format to obtain a template image, and to record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker corresponding to the template video.

[0029] Optionally, the third acquisition module is specifically used to: for each template video frame of the template video, determine the position of the facial key points of the template person in the template video frame; and obtain the trajectory of the change in the position of the facial key points of the template person in the template video as the first motion trajectory according to the change in the position of the facial key points of the template person in each template video frame.

[0030] Optionally, the recording module is specifically used to: add the template sticker corresponding to the template video to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

[0031] Optionally, the first acquisition module is specifically used to: acquire the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and obtain the sticker identifier corresponding to the image identifier of the specified template image as the specified sticker identifier.

[0032] Optionally, the first acquisition module is specifically used to: acquire the sticker identifier of the target type corresponding to the image identifier of the specified template image according to the pre-recorded correspondence between image identifiers and sticker identifiers, and display the acquired sticker identifier of the target type; from the displayed sticker identifier of the target type, determine, according to the selection instruction for the specified sticker identifier in the sticker identifier of the target type, that the sticker identifier indicated by the selection instruction input by the user is the specified sticker identifier corresponding to the image identifier of the specified template image, and obtain the specified sticker identifier.

[0033] Optionally, the first adjustment module is specifically configured to: copy the image to be processed according to a first number of facial key points in the first motion trajectory to obtain the first number of images to be processed; wherein, one image to be processed corresponds to one position of a facial key point in the first motion trajectory; for each position of a facial key point in the first motion trajectory, adjust the position of the facial key point of the target person in the image to be processed corresponding to the position of the facial key point in the first motion trajectory to obtain a first video frame to be processed; sort each first video frame to be processed according to the order of the positions of each facial key point in the first motion trajectory, and combine the sorted first video frames to be processed into a first video to be processed.

[0034] Optionally, the device further includes:

[0035] The detection module is used to detect whether the facial features of the target person are displayed in the image to be processed before the first adjustment module performs the adjustment of the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0036] The prompt message output module is used to output a prompt message indicating that the image to be re-uploaded if the facial features of the target person are not displayed in the image to be processed.

[0037] The first adjustment module is specifically used to: if the facial features of the target person are displayed in the image to be processed, adjust the position of the key facial points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0038] In a third aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0039] A memory is used to store computer programs; a processor is used to execute the program stored in the memory to implement the image generation method steps described in any of the first aspects above.

[0040] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the image generation method described in any of the first aspects above.

[0041] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the image generation methods described in the first aspect above.

[0042] This invention provides an image generation method, comprising: acquiring template information and an image to be processed; wherein the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the facial key points of a template person change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent: the emotion expressed by the template person in the specified template image; the image to be processed is a static image displaying a target person; adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain a first video to be processed; adjusting the size of the specified sticker to which the specified sticker identifier belongs to a target size, and adding a specified sticker of the target size at a target position according to a target angle in each first video frame in the first video to be processed to obtain a second video frame corresponding to the first video frame; combining the second video frames corresponding to each first video frame in the first video to be processed to obtain a second video to be processed; and converting the second video to be processed into a dynamic image format to obtain a target image.

[0043] Based on the above processing, since the template information includes a designated sticker identifier, and the designated sticker to which the identifier belongs can indicate the emotion expressed by the template character, the electronic device adjusts the position of the facial key points of the target character in the image to be processed according to the position of the facial key points of the template character displayed in the template image. In the resulting first video to be processed, the expression changes of the target character are the same as those of the template character, thus allowing the target character to express the same emotion as the template character. Furthermore, by adding a designated sticker that indicates the emotion expressed by the template character to the first video to be processed—which is equivalent to adding a designated sticker that indicates the emotion expressed by the target character—the target image obtained from the first video to be processed with the added designated sticker can better express the emotion, improving the user experience.

[0044] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0046] Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention;

[0047] Figure 2 This is a second flowchart of the image generation method provided in an embodiment of the present invention;

[0048] Figure 3 This is a third flowchart of the image generation method provided in the embodiments of the present invention;

[0049] Figure 4 This is a fourth flowchart of the image generation method provided in the embodiments of the present invention;

[0050] Figure 5 This is a fifth flowchart of the image generation method provided in an embodiment of the present invention;

[0051] Figure 6 This is a sixth flowchart of the image generation method provided in an embodiment of the present invention;

[0052] Figure 7 This is a seventh flowchart of the image generation method provided in the embodiments of the present invention;

[0053] Figure 8 This is an eighth flowchart of the image generation method provided in the embodiments of the present invention;

[0054] Figure 9 This is a ninth flowchart of the image generation method provided in the embodiments of the present invention;

[0055] Figure 10 A schematic diagram of the image generation result provided in an embodiment of the present invention;

[0056] Figure 11 A structural diagram of an image generation apparatus provided in an embodiment of the present invention;

[0057] Figure 12 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on the present invention are within the scope of protection of the present invention.

[0059] Template GIF images only display a template character with changing facial expressions, which cannot effectively convey emotions. For example, when a template character in a template GIF image blinks, it's impossible to determine whether the character is happy or sad. Consequently, GIFs generated from template GIF images that cannot adequately express emotions also fail to convey emotions effectively, resulting in a poor user experience.

[0060] To address the aforementioned problems, this invention provides an image generation method applied to an electronic device. The electronic device acquires template information and an image to be processed. The template information includes an image identifier for a specified template image and a corresponding specified sticker identifier. The specified template image is a dynamic image in which the facial key points of a template character change according to a first motion trajectory. The specified sticker to which the specified sticker identifier belongs represents the emotion expressed by the template character in the specified template image. The image to be processed is a static image displaying a target character. The positions of the facial key points of the target character in the image to be processed are adjusted according to the first motion trajectory to obtain a first video to be processed. The size of the specified sticker to which the specified sticker identifier belongs is adjusted to a target size, and a specified sticker of the target size is added at the target position and angle in each first video frame of the first video to be processed to obtain a second video frame corresponding to that first video frame. The second video frames corresponding to each first video frame in the first video to be processed are combined to obtain a second video to be processed. The second video to be processed is converted into a dynamic image format to obtain the target image, which can improve the user experience.

[0061] See Figure 1 , Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention. The method may include the following steps:

[0062] S101: Obtain template information and image to be processed.

[0063] The template information includes an image identifier for a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the key facial features of the template character change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent the emotion expressed by the template character in the specified template image; and the image to be processed is a static image displaying the target character.

[0064] S102: According to the first motion trajectory, adjust the position of the facial key points of the target person in the image to be processed to obtain the first video to be processed.

[0065] S103: Adjust the size of the specified sticker to which the specified sticker identifier belongs to the target size, and add the specified sticker of the target size at the target position in each first video frame in the first video to be processed, according to the target angle, to obtain the second video frame corresponding to the first video frame.

[0066] S104: Combine the second video frames corresponding to each first video frame in the first video to be processed to obtain the second video to be processed.

[0067] S105: Convert the second video to be processed into a dynamic image format to obtain the target image.

[0068] Based on the image generation method provided in this embodiment of the invention, since the template information includes a designated sticker identifier, and the designated sticker to which the designated sticker identifier belongs can indicate the emotion expressed by the template character, the electronic device adjusts the position of the facial key points of the target character in the image to be processed according to the position of the facial key points of the template character displayed in the designated template image. In the resulting first video to be processed, the expression change pattern of the target character is the same as that of the template character, so the target character can express the same emotion as the template character. Furthermore, by adding a designated sticker that can indicate the emotion expressed by the template character to the first video to be processed, that is, by adding a designated sticker that can indicate the emotion expressed by the target character, the target image obtained from the first video to be processed with the designated sticker can better express the emotion, thereby improving the user experience.

[0069] For step S101, the image processing platform can be an app installed on the user terminal. Multiple template images can be displayed to the user on the image processing platform's display interface. Furthermore, the user can input an image identifier for a specified template image into the image processing platform by clicking on the template images displayed on the platform's display interface.

[0070] For example, the multiple template images can be directly displayed in the display interface of the image processing platform. Alternatively, the image identifiers of the multiple template images can be displayed in the display interface of the image processing platform; for example, the image identifier of a template image can be a thumbnail of that template image.

[0071] In one implementation, when the image processing platform displays the multiple template images to the user, it can also display a sticker corresponding to each template image. For example, the image processing platform can display template images with corresponding stickers added. Correspondingly, the user can input the image identifier of a specified template image and the corresponding sticker identifier into the image processing platform through a click operation.

[0072] In another implementation, after receiving the image identifier of the specified template image input by the user, the image processing platform can then display the sticker identifier corresponding to the image identifier of the specified template image to the user, so that the user can determine the specified sticker identifier from the sticker identifiers corresponding to the image identifier of the specified template image. The method by which the image processing platform obtains the specified sticker identifier can be found in the detailed description of the following embodiments.

[0073] After the user inputs the image identifier of a specified template image and the corresponding specified sticker identifier into the image processing platform, the platform obtains the template information input by the user. Furthermore, the user can also upload images to be processed to the image processing platform; these images are static images displaying the target person.

[0074] If a specified template image contains a template person whose facial key points change according to a first motion trajectory, then based on the specified template image's image identifier, the electronic device can determine the necessary motion trajectory to adjust the position of the target person's facial key points. Similarly, based on a specified sticker identifier, the electronic device can determine what kind of sticker needs to be added to the target image and what emotion it represents. In other words, based on the template information, the electronic device can determine what kind of processing is required for the image to be processed.

[0075] The electronic device can be a user terminal equipped with an image processing platform. The image processing platform acquires template information and the image to be processed; that is, the electronic device acquires the template information and the image to be processed. Examples of electronic devices include mobile phones and computers.

[0076] Alternatively, the electronic device can act as a server providing services to the image processing platform. After the image processing platform obtains the template information and the image to be processed, it sends the obtained template information and the image to be processed to the server, thus enabling the electronic device to obtain the template information and the image to be processed.

[0077] In some embodiments, Figure 1 Based on this, see Figure 2 Step S101 may include the following steps:

[0078] S1011: Obtain the image identifier of the specified template image and the image to be processed.

[0079] S1012: Based on the pre-recorded correspondence between image identifiers and sticker identifiers, obtain the sticker identifier corresponding to the image identifier of the specified template image, and obtain the specified sticker identifier.

[0080] After browsing through multiple template image identifiers on the client's display interface, users can select the image identifier of a specific template image that expresses the emotion they wish to convey in the target image. Correspondingly, the client can retrieve the image identifier of the specified template image based on the user's selection.

[0081] Furthermore, the client can determine the sticker identifier corresponding to the image identifier of the specified template image according to the pre-recorded correspondence between image identifiers and sticker identifiers, and then determine the specified sticker identifier from the determined sticker identifiers.

[0082] In some embodiments, Figure 2 Based on this, see Figure 3 Before step S1012, the method may further include the following steps:

[0083] S106: Obtain template video and multiple reference motion trajectories.

[0084] The template video displays a template character; a reference motion trajectory is used to determine the emotion expressed by the template character in the template video; the reference motion trajectory corresponds to a sticker; a sticker corresponding to a reference motion trajectory is used to indicate the emotion expressed by the template character when the facial key points change according to the reference motion trajectory.

[0085] S107: Obtain the change trajectory of the facial key points of the template character in the template video to obtain the first motion trajectory.

[0086] S108: Calculate the similarity between the first motion trajectory and each reference motion trajectory, and determine the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video.

[0087] S109: Determine the template sticker corresponding to the template video from among the multiple stickers corresponding to the reference motion trajectory of the template video.

[0088] S110: Convert the template video into a dynamic image format to obtain the template image.

[0089] S111: Record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker to obtain the correspondence between the image identifier and the sticker identifier.

[0090] Electronic devices can pre-acquire multiple reference motion trajectories, which can be uploaded in advance by technicians. Furthermore, for each reference motion trajectory, technicians can pre-set the corresponding emotion. This emotion can correspond to a type of sticker; that is, a reference motion trajectory and a sticker are associated. A sticker corresponding to a reference motion trajectory represents the emotion expressed by a template character whose facial key points change according to that reference motion trajectory. For example, if a reference motion trajectory corresponds to stickers including the text "happy," or stickers including a smiley face pattern representing the emotion "happy," then the template character whose facial key points change according to that reference motion trajectory expresses the emotion "happy."

[0091] Users who need to obtain the target image via the client, as well as technicians, can upload template videos. After the electronic device obtains the template video, it can capture the change trajectory of the key facial features of the template person in the video, thus obtaining the first motion trajectory.

[0092] In one implementation, after determining the facial key points in the first template video frame, the electronic device uses these key points as a matching metric. Then, it searches for this matching metric in the second template video frame and takes the position where the matching metric reaches its extreme value as the optimal matching point. The line connecting the coordinates of the facial key points to the coordinates of the optimal matching point represents the motion trajectory of the facial key points from the first to the second template video frame. This process continues until the optimal matching point is determined in the last video frame of the template video, thus establishing the first motion trajectory of the facial key points of the template character from the first to the last video frame.

[0093] In another implementation, Figure 3 Based on this, see Figure 4 Step S107 may include the following steps:

[0094] S1071: For each template video frame of the template video, determine the position of the facial key points of the template character in that template video frame.

[0095] S1072: Based on the changes in the position of the facial key points of the template person in each template video frame, obtain the trajectory of the changes in the position of the facial key points of the template person in the template video, and use it as the first motion trajectory.

[0096] For each template video frame, the electronic device can first determine the key facial features of the template person in that template video frame based on a facial feature point determination method. For example, the facial feature point determination method can be based on ASM (Active Shape Model) or LBF (Local Binary Fitting), etc.

[0097] Furthermore, for each facial key point of the template person, the electronic device can determine the coordinates of that key point in each template video frame. It then determines the sequence of coordinate changes for each key point according to the order of the template video frames within the template video. Subsequently, based on this sequence of coordinate changes, the electronic device obtains the trajectory of each key point's movement. Based on the trajectory of the changing positions of each facial key point of the template person in the template video, the electronic device obtains the first motion trajectory.

[0098] For example, according to the order in the template video, the template video includes: Template Video Frame 1, Template Video Frame 2, and Template Video Frame 3. The electronic device identifies facial keypoint 1 at coordinates (1,1); facial keypoint 2 at coordinates (1,2); and facial keypoint 3 at coordinates (1,3) in Template Video Frame 1. It identifies facial keypoint 1 at coordinates (2,1); facial keypoint 2 at coordinates (2,2); and facial keypoint 3 at coordinates (2,3) in Template Video Frame 2. It identifies facial keypoint 1 at coordinates (3,1); facial keypoint 2 at coordinates (3,2); and facial keypoint 3 at coordinates (3,3) in Template Video Frame 3.

[0099] The sequence of coordinate changes for facial key point 1 is (1, 1), (2, 1), (3, 1), meaning the trajectory 1 of facial key point 1 is: from coordinate (1, 1) to coordinate (2, 1), and then from coordinate (2, 1) to coordinate (3, 1). Correspondingly, the electronic device can determine the trajectory 2 of facial key point 2 and the trajectory 3 of facial key point 3. Trajectory 1, trajectory 2, and trajectory 3 together constitute the first motion trajectory of the facial key points of the template character.

[0100] The first motion trajectory includes multiple trajectory points, each trajectory point corresponding to a template video frame. This trajectory point can represent the position of each facial key point of the template character in the corresponding template video frame. As in the aforementioned embodiment, the trajectory point corresponding to template video 1 can be called trajectory point 1, and trajectory point 1 includes the coordinates of each facial key point in template video frame 1.

[0101] Based on the above processing, the electronic device can determine the position of the facial key points of the template character from each template video frame of the template video. Then, based on the position of each facial key point, a first motion trajectory representing the changes of the facial key points of the template character in the template video is obtained. That is, the first motion trajectory is determined directly based on the specific facial key points, rather than being located by facial key points determined from other template video frames. Therefore, the accuracy of the first motion trajectory is higher. Subsequently, based on the more accurate first motion trajectory, a target image expressing a more accurate emotion can be obtained, which can improve the user experience.

[0102] After acquiring a first motion trajectory, the electronic device, in order to determine the emotion corresponding to the first motion trajectory, calculates the similarity between the first motion trajectory and each reference motion trajectory, since each reference motion trajectory corresponds to a certain emotion. If the similarity between the first motion trajectory and a reference motion trajectory is high, then the probability that the emotion corresponding to the first motion trajectory is the same as the emotion corresponding to the reference motion trajectory is also high; if the similarity between the first motion trajectory and a reference motion trajectory is low, then the probability that the emotion corresponding to the first motion trajectory is the same as the emotion corresponding to the reference motion trajectory is also low.

[0103] Furthermore, the electronic device can determine the reference motion trajectory with the highest similarity to the first motion trajectory from among the various reference motion trajectories, and use this as the reference motion trajectory corresponding to the template video. Therefore, the probability that the first motion trajectory and the reference motion trajectory corresponding to the template video correspond to the same emotion is the highest; that is, the probability that the template character in the template video expresses the same emotion as the reference motion trajectory corresponding to the template video is the highest. Thus, the electronic device can determine that the template character in the template video is expressing the emotion corresponding to the reference motion trajectory corresponding to the template video.

[0104] Furthermore, the electronic device can determine the template sticker corresponding to the template video from among multiple stickers corresponding to the reference motion trajectory of the template video.

[0105] In one implementation, the electronic device can identify multiple stickers corresponding to the reference motion trajectory of the template video as template stickers corresponding to the template video. In this case, the template video can correspond to multiple template stickers.

[0106] In another implementation, the electronic device can display sticker identifiers for multiple stickers corresponding to the reference motion trajectory of the template video on a display interface, allowing the user or technician uploading the template video to select one. Then, based on the received selection instruction, the electronic device can determine the sticker corresponding to the sticker identifier carried in the selection instruction from among the multiple stickers, and use it as the template sticker for the template video. In this case, the template video can correspond to only one template sticker or multiple template stickers.

[0107] Correspondingly, the electronic device can convert the template video into a dynamic image format to obtain the template image, and then the electronic device can obtain the image identifier of the template image. Furthermore, the electronic device can record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker, thus obtaining the correspondence between the image identifier and the sticker identifier.

[0108] In some embodiments, Figure 3 Based on this, see Figure 5 Step S111 may include the following steps:

[0109] S1111: Add the template sticker corresponding to the template video to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

[0110] When the image identifier of the template image corresponds one-to-one with the sticker identifier of the template sticker, the electronic device, after obtaining the template image, can directly add the template sticker corresponding to the template video to the template image, thus obtaining a template image with the sticker added. This template image with the sticker added serves as a correspondence between image identifiers and sticker identifiers, indicating that the image identifier of the template image corresponds one-to-one with the sticker identifier of the template sticker added to the template image. Accordingly, the image processing platform can then display the template image with the sticker added to the user.

[0111] In some embodiments, step S1012 may include the following steps:

[0112] Step 1: Obtain the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and use it as the specified sticker identifier.

[0113] When a user clicks on a template image with added stickers displayed on the image processing platform, they can input the image identifier of the specified template image into the platform. Correspondingly, the image processing platform can determine the sticker identifier of the sticker added to the specified template image as the specified sticker identifier based on the image identifier of the specified template image and the correspondence between image identifiers and sticker identifiers.

[0114] Based on the above processing, when a target image needs to be generated, the user can input the image identifier of the specified template image and the specified sticker identifier into the image processing platform by clicking on the template image with stickers added. This simplifies the user operation and improves the user experience.

[0115] In some embodiments, Figure 2 Based on this, see Figure 6 Step S1012 may include the following steps:

[0116] S10121: According to the pre-recorded correspondence between image identifiers and sticker identifiers, obtain the sticker identifier of the target type corresponding to the image identifier of the specified template image, and display the obtained sticker identifier of the target type.

[0117] S10122: From the displayed sticker identifiers of the target type, based on the selection instruction for the specified sticker identifier in the sticker identifiers of the target type, determine that the sticker identifier indicated by the selection instruction entered by the user is the specified sticker identifier corresponding to the image identifier of the specified template image, and obtain the specified sticker identifier.

[0118] When the image identifier of the specified template image corresponds to the sticker identifier of the target type, that is, when the image identifier of the specified template image corresponds to multiple sticker identifiers, since the number of specified stickers that can be added to the target image to be generated is limited, the electronic device cannot obtain the target image to which all stickers to which the image identifier corresponding to the image identifier of the specified template image belong. Therefore, the electronic device needs to determine the specified sticker identifier from multiple sticker identifiers.

[0119] Therefore, after obtaining the image identifier of a specified template image, the electronic device can first obtain the target type sticker identifier corresponding to the image identifier of the specified template image according to the pre-recorded correspondence between image identifiers and sticker identifiers. The stickers to which these multiple sticker identifiers belong are all stickers that can express the emotions the user wishes to express.

[0120] Then, the electronic device can display the acquired target type sticker icons on the image processing platform's display interface. Correspondingly, users who need to obtain the target image can browse the target type sticker icons through the image processing platform's display interface and select sticker icons that meet their needs. For example, they can select sticker icons with their preferred shapes, sticker icons with their preferred fonts, or sticker icons with their preferred artistic text effects.

[0121] The sticker identifier selected by the user that matches their needs is the specific sticker identifier indicated by the user's input selection command. The user selects a sticker identifier that matches their needs from the sticker identifiers of the target type, which means inputting a selection command for a specific sticker identifier within the target type of sticker identifier into the image processing platform. Correspondingly, the electronic device can determine the specific sticker identifier indicated by the user's input selection command and obtain the specific sticker identifier corresponding to the image identifier of the specified template image.

[0122] Based on the above processing, users can select a specific sticker icon from multiple sticker icons according to their needs. Subsequently, a target image with the specified sticker icon belonging to the specified sticker icon that meets the user's needs is generated, which can meet the user's personalized needs and improve the user experience.

[0123] Regarding step S102, in one implementation, the electronic device can acquire a template video of the specified template image to which the image identifier used to generate the specified template image belongs, and obtain the first motion trajectory of the facial key points of the template person from the acquired template video. The method by which the electronic device obtains the first motion trajectory of the facial key points of the template person from the template video is similar to the method of obtaining the change trajectory of the facial key points of the template person in the template video in the foregoing embodiments, and can be referred to the relevant description in the foregoing embodiments.

[0124] In another implementation, when generating a specified template image, the electronic device can obtain the first motion trajectory of the facial key points of the template person in the template video used to generate the specified template image, and record the correspondence between the image identifier of the specified template image and the first motion trajectory. Subsequently, after obtaining the image identifier of the specified template image, the electronic device can directly obtain the first motion trajectory corresponding to the image identifier of the specified template image according to the correspondence between the image identifier of the specified template image and the first motion trajectory.

[0125] The first motion trajectory includes the positions of multiple facial key points. The position of each facial key point corresponds to a template video frame in the template video. The position of the facial key point includes the coordinates of each facial key point in the corresponding template video frame.

[0126] In some embodiments, Figure 1 Based on this, see Figure 7 Step S102 may include the following steps:

[0127] S1021: Copy the image to be processed according to the first number of facial key points in the first motion trajectory to obtain the first number of images to be processed.

[0128] In this process, one image to be processed corresponds to a position of a facial key point in the first motion trajectory.

[0129] S1022: For the position of each facial key point in the first motion trajectory, adjust the position of the facial key point of the target person in the image to be processed to the position of the facial key point in the first motion trajectory to obtain the first video frame to be processed.

[0130] S1023: Sort each first video frame to be processed according to the position of each facial key point in the first motion trajectory and the order in the first motion trajectory, and combine the sorted first video frames to be processed into a first video to be processed.

[0131] The first number of facial key points in the first motion trajectory is also the first number of template video frames used to generate the specified template image. In order to make the generated first video to be processed correspond to the template video, the electronic device needs to copy the image to be processed into a first number of images to be processed. These first number of images to be processed correspond one-to-one with the first number of template video frames in the template video, that is, one image to be processed corresponds to one position of a facial key point in the first motion trajectory.

[0132] For each image to be processed, the electronic device can determine the key facial features of the target person in the image based on a facial feature point determination method. For example, the facial feature point determination method can be: ASM-based facial feature point determination method, LBF-based facial feature point determination method, etc.

[0133] The location of a facial key point is a trajectory point in the first motion trajectory described in the aforementioned embodiment. Furthermore, for each facial key point in the first motion trajectory, that is, for each trajectory point in the first motion trajectory, this trajectory point includes the coordinates of each facial key point in the corresponding template video frame. Therefore, for each facial key point in the video frame to be processed corresponding to this trajectory point, the electronic device can adjust the coordinates of this facial key point to the coordinates of the facial key point included in this trajectory point. Furthermore, after the electronic device adjusts the coordinates of each facial key point in the video frame to be processed corresponding to this trajectory point, it achieves the adjustment of the position of the facial key point of the target person in the corresponding image to be processed according to the position of the facial key point in the first motion trajectory, thus obtaining the first video frame to be processed corresponding to the position of the facial key point.

[0134] Furthermore, according to the order of the positions of each facial key point in the first motion trajectory, the electronic device can sort the first video frames to be processed corresponding to the positions of each facial key point, thus obtaining sorted first video frames to be processed. In the sorted first video frames to be processed, the first first video frame to be processed corresponds to the first template video frame in the template video; the second first video frame to be processed corresponds to the second template video frame in the template video, and so on; the last first video frame to be processed corresponds to the last template video frame in the template video.

[0135] Therefore, the electronic device can combine the sorted first video frames into a first video to be processed. In this first video, the trajectory of the facial key points of the target person is the same as the trajectory of the facial key points of the template person in the template video. That is, the facial expression changes of the target person are the same as those of the template person in the template video.

[0136] Based on the above processing, the electronic device processes the images frame by frame. The resulting first video is more accurate in conveying the user's intended emotion, meaning it has higher accuracy. Subsequent target images obtained from this more accurate first video are also more accurate, thus improving the user experience.

[0137] In some embodiments, in order to improve the accuracy of the obtained first video to be processed, the electronic device determines the facial feature points of the target person in the image to be processed in the same way as it determines the facial feature points of the template person in the template video frame.

[0138] In steps S103 and S104, after obtaining the first video to be processed, in order to obtain a target image that can better express emotions, the electronic device can add a designated sticker that can indicate the emotions expressed by the target person to the first video to be processed.

[0139] The electronic device first retrieves the designated sticker corresponding to the specified sticker identifier, thus obtaining the designated sticker that indicates the emotion expressed by the target person. Then, the electronic device adjusts the size of the designated sticker to the target size and adds the designated sticker of the target size at the target position and angle in each first video frame of the first video to be processed, thus obtaining the second video frame corresponding to that first video frame. Subsequently, the electronic device combines the second video frames corresponding to each first video frame in the order of the first video frames in the first video to be processed, thus obtaining the second video to be processed with the designated sticker added. Since the designated sticker can indicate the emotion expressed by the target person, the second video to be processed with the designated sticker can better express the emotion that the user wants to express.

[0140] In one implementation, to simplify user operation, the target size, target angle, and target position can all be preset by the technician. For example, if the designated sticker to which the designated sticker belongs is a sticker with a lot of text, then to avoid the designated sticker obscuring the face of the target person, the target size can be set to a smaller size, such as 200*50 pixels; the target angle can be 0°; and the target position can be the bottom of the target image to be generated.

[0141] In another implementation, to enhance the user's personalized experience, the target size, target angle, and target position can be set by the user according to their needs during the acquisition of the target image through the image processing platform. For example, when the user inputs a selection command for a specific sticker identifier among the sticker identifiers of the target type, they can scale the selected sticker to adjust its size, thus inputting the target size to the electronic device; rotate the sticker to modify its angle, thus inputting the target angle to the electronic device; or adjust the sticker's position to the location selected by the user, thus inputting the target position to the electronic device. Correspondingly, the electronic device can determine the target size, target angle, and target position based on the user's operations on the specified sticker.

[0142] In another implementation, before generating the target image, the user can directly set the target size, target angle, and target position. For example, the template image can be obtained from a user-uploaded template video. After determining the template sticker corresponding to the template video, the user can directly set the target size, target angle, and target position of the template sticker. Correspondingly, the electronic device can record the target size, target angle, target position, sticker identifier, and image identifier of the template image. Subsequently, when the electronic device obtains the image identifier of the specified template image and the specified sticker identifier, it can directly obtain the target size, target angle, and target position corresponding to the specified sticker identifier. This eliminates the need for the user to re-enter the target size, target angle, and target position of the specified sticker, simplifying user operations while meeting personalized needs and improving the user's personalized experience.

[0143] Regarding step S105, after the electronic device obtains a second video to be processed that can better express the user's desired emotion, in order to obtain a target image that can better express the user's desired emotion, the electronic device can convert the second video to be processed into a dynamic image format according to a format conversion algorithm to obtain the target image. For example, the format conversion algorithm can be a compression algorithm based on RLE (Run-Length Encoding); the dynamic image format can be GIF format.

[0144] In some embodiments, Figure 1 Based on this, see Figure 8 Before step S102, the method may further include the following steps:

[0145] S112: Detect whether the facial features of the target person are displayed in the image to be processed.

[0146] S113: If the facial features of the target person are not displayed in the image to be processed, output a prompt message indicating that the image to be processed should be re-uploaded.

[0147] Accordingly, step S102 may include the following steps:

[0148] S1024: If the facial features of the target person are displayed in the image to be processed, adjust the position of the key facial points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0149] Since electronic devices need to adjust the positions of facial key points of the target person displayed in the image to be processed in order to obtain the target image, and these facial key points are generally the key points of the target person's facial features, after the electronic device acquires the image to be processed, it can first use a detection algorithm to detect whether the target person's facial features are displayed in the image. For example, the detection algorithm could be the Haar cascade detection algorithm.

[0150] When the facial features of the target person are not displayed in the image to be processed, the facial key points of the target person in the image acquired by the electronic device do not correspond to the positions of the facial key points included in the first motion trajectory. Therefore, the electronic device cannot adjust the image to be processed according to the first motion trajectory to obtain a first video in which the facial expressions of the target person are identical to those of the template person in the template video. Thus, the electronic device can output a prompt message indicating that the image to be processed should be re-uploaded, reminding the user to re-upload an image that displays the facial features of the target person.

[0151] When the facial features of the target person are displayed in the image to be processed, the facial key points of the target person in the image to be processed acquired by the electronic device can correspond to the positions of the facial key points included in the first motion trajectory. The electronic device can then adjust the image to be processed according to the first motion trajectory to obtain a first video to be processed where the facial expressions of the target person are identical to those of the template person in the template video. Therefore, the electronic device can adjust the positions of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0152] Based on the above processing, the electronic device can first detect whether the facial features of the target person are displayed in the image to be processed, and then process the image to be processed according to the first motion trajectory when it is determined that the facial features of the target person are displayed in the image to be processed, which can improve the accuracy of the generated target image.

[0153] See Figure 9 , Figure 9 A ninth flowchart of an image generation method provided in an embodiment of the present invention. The method may include the following steps:

[0154] S901: Users can see GIF (Graphics Interchange Format) animated materials containing videos and stickers within the software.

[0155] The GIF animation material containing video and stickers is the thumbnail of the template image with the specified stickers added in the aforementioned embodiment.

[0156] In this step, the image processing platform's display interface can directly show thumbnails of template images with specified stickers. After browsing the thumbnails of template images with stickers in the image processing platform's display interface, users can directly select the thumbnail of the template image with stickers, and thus directly input the image identifier of the specified template image and the identifier of the specified sticker into the image processing platform.

[0157] S902: The user uploads a clear portrait photo containing facial features; the client checks the photo to ensure that the user's photo contains facial features, and if the user's photo does not contain facial features, it guides the user to re-upload.

[0158] A clear portrait photograph containing facial features is the image to be processed that displays the facial features of the target person in the aforementioned embodiments.

[0159] In this step, after acquiring the image to be processed, the electronic device detects whether the facial features of the target person are displayed in the image. If the facial features of the target person are not displayed in the image, a prompt message indicating that the image to be processed should be re-uploaded is output. If the facial features of the target person are displayed in the image, the positions of the key facial points of the target person in the image are adjusted according to the first motion trajectory to obtain the first video to be processed.

[0160] S903: Uses preset video footage to drive photo-based animation, bringing characters to life.

[0161] The pre-set video material is the template video in the aforementioned embodiment; the process of driving the photo to make the person move is the process of obtaining the first video to be processed based on the image to be processed in the aforementioned embodiment.

[0162] In this step, the electronic device copies the image to be processed according to a first number of facial key points in the first motion trajectory, resulting in a first number of images to be processed; each image to be processed corresponds to a position of a facial key point in the first motion trajectory; for each position of a facial key point in the first motion trajectory, the position of the facial key point of the target person in the image to be processed corresponding to that position is adjusted to the position of that facial key point in the first motion trajectory, resulting in a first video frame to be processed; the first video frames to be processed are sorted according to the order of the positions of the facial key points in the first motion trajectory, and the sorted first video frames to be processed are combined into a first video to be processed.

[0163] S904: After receiving the animated GIF result, request the sticker service again to merge the matching stickers and video.

[0164] The animated result is the first video to be processed in the aforementioned embodiment, and requesting the sticker service means adding a specified sticker to the first video to be processed.

[0165] In this step, the electronic device first determines the designated sticker according to the designated sticker identifier, then adjusts the size of the designated sticker to the target size, and adds the designated sticker of the target size at the target position and the target angle in each first video frame of the first video to be processed to obtain the second video frame corresponding to the first video frame. Then, the second video frames corresponding to each first video frame in the first video to be processed are combined to obtain the second video to be processed.

[0166] S905: Convert the received dynamic video into GIF format and return a GIF image containing preset videos and stickers to the user.

[0167] The dynamic video is the second video to be processed in the aforementioned embodiment, and the GIF image obtained by converting it into GIF format is the target image in the aforementioned embodiment.

[0168] In this step, after the electronic device acquires a second video that effectively expresses the user's intended emotion, it converts the second video into a dynamic image format to obtain the target image. Subsequently, the electronic device can send the target image to the user's client application.

[0169] See Figure 10 , Figure 10 This is a schematic diagram of the image generation result provided in an embodiment of the present invention. Figure 10 The image on the left is a thumbnail of a template image with the specified sticker added. The template image displays the face of the template character, and the specified sticker is displayed below the template image with the text "OMG" and corresponding artistic text effects. Figure 10 The intermediate image in the image is the image to be processed, which displays the facial features of the target person. Figure 10 The image on the right is a schematic diagram of the target image obtained from the image to be processed. In the target image, the specified sticker is also displayed below the template image, and the displayed text content is also "OMG", and the displayed text content also has the same artistic text effect.

[0170] Based on the image generation method provided in this embodiment of the invention, since the template information includes a designated sticker identifier, and the designated sticker to which the designated sticker identifier belongs can indicate the emotion expressed by the template character, the electronic device adjusts the position of the facial key points of the target character in the image to be processed according to the position of the facial key points of the template character displayed in the designated template image. In the resulting first video to be processed, the expression change pattern of the target character is the same as that of the template character, so the target character can express the same emotion as the template character. Furthermore, by adding a designated sticker that can indicate the emotion expressed by the template character to the first video to be processed, that is, by adding a designated sticker that can indicate the emotion expressed by the target character, the target image obtained from the first video to be processed with the designated sticker can better express the emotion, thereby improving the user experience.

[0171] Compared to existing facial expression-driven solutions that are limited to entertainment or pranks, this solution provides a highly efficient image processing platform. This image generation platform can serve as an AI (Artificial Intelligence) emoji creation tool, offering a variety of stickers with different meanings that can be paired with video emotions. Users can upload a photo and, based on the template information they input, generate multiple emoji results. These results can include various facial expressions and stickers, quickly conveying the user's emotions. Users can then rapidly express their emotions in social media and chat scenarios, expanding the scope of AI emoji usage.

[0172] Based on the same inventive concept as the image generation method described above, embodiments of the present invention also provide an image generation apparatus. See [link to previous document]. Figure 11 , Figure 11 This is a structural diagram of an image generation apparatus provided in an embodiment of the present invention, the apparatus comprising:

[0173] The first acquisition module 1101 is used to acquire template information and an image to be processed; wherein, the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the facial key points of the template character change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent the emotion expressed by the template character in the specified template image; the image to be processed is a static image displaying a target character;

[0174] The first adjustment module 1102 is used to adjust the position of the facial key points of the target person in the image to be processed according to the first motion trajectory, so as to obtain the first video to be processed;

[0175] The second adjustment module 1103 is used to adjust the size of the specified sticker to which the specified sticker identifier belongs to the target size, and add the specified sticker of the target size at the target position in each first video frame in the first video to be processed according to the target angle to obtain the second video frame corresponding to the first video frame.

[0176] The combination module 1104 is used to combine the second video frames corresponding to each first video frame in the first video to be processed to obtain the second video to be processed.

[0177] The format conversion module 1105 is used to convert the second video to be processed into a dynamic image format to obtain the target image.

[0178] Optionally, the first acquisition module 1101 is specifically used for:

[0179] Retrieve the image identifier of the specified template image and the image to be processed;

[0180] Based on the pre-recorded correspondence between image identifiers and sticker identifiers, the sticker identifier corresponding to the image identifier of the specified template image is obtained, thus obtaining the specified sticker identifier.

[0181] Optionally, the device further includes:

[0182] The second acquisition module is used to acquire a template video and multiple reference motion trajectories before the first acquisition module 1101 executes the pre-recorded correspondence between image identifiers and sticker identifiers to acquire the sticker identifier corresponding to the image identifier of the specified template image and obtains the specified sticker identifier; wherein, the template video displays the template character; the reference motion trajectories are used to determine the emotions expressed by the template character displayed in the template video; the reference motion trajectories correspond to template stickers; and the template stickers are used to indicate the emotions corresponding to the corresponding reference motion trajectories.

[0183] The third acquisition module is used to acquire the change trajectory of the facial key points of the template character in the template video to obtain the first motion trajectory.

[0184] The similarity calculation module is used to calculate the similarity between the first motion trajectory and each reference motion trajectory, and to determine the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video.

[0185] The template sticker determination module is used to determine the template sticker corresponding to the template video from the template stickers corresponding to the reference motion trajectory corresponding to the template video;

[0186] The recording module is used to convert the template video into a dynamic image format to obtain a template image, and to record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker corresponding to the template video.

[0187] Optionally, the third acquisition module is specifically used for:

[0188] For each template video frame of the template video, determine the position of the facial key points of the template character in that template video frame;

[0189] Based on the changes in the position of the facial key points of the template character in each template video frame, the trajectory of the change in the position of the facial key points of the template character in the template video is obtained, which is used as the first motion trajectory.

[0190] Optionally, the recording module is specifically used for:

[0191] The template sticker corresponding to the template video is added to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

[0192] Optionally, the first acquisition module 1101 is specifically used for:

[0193] Obtain the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and use the sticker identifier corresponding to the image identifier of the specified template image as the specified sticker identifier.

[0194] Optionally, the first acquisition module 1101 is specifically used for:

[0195] According to the pre-recorded correspondence between image identifiers and sticker identifiers, obtain the sticker identifier of the target type corresponding to the image identifier of the specified template image, and display the obtained sticker identifier of the target type;

[0196] From the displayed sticker identifiers of the target type, based on the selection instruction for the specified sticker identifier in the sticker identifiers of the target type, the sticker identifier indicated by the selection instruction input by the user is determined to be the specified sticker identifier corresponding to the image identifier of the specified template image, and the specified sticker identifier is obtained.

[0197] Optionally, the first adjustment module 1102 is specifically used for:

[0198] According to a first number of facial key points in the first motion trajectory, the image to be processed is copied to obtain the first number of images to be processed; wherein, one image to be processed corresponds to one position of a facial key point in the first motion trajectory;

[0199] For each facial key point in the first motion trajectory, the position of the facial key point of the target person in the image to be processed corresponding to the position of the facial key point in the first motion trajectory is adjusted to the position of the facial key point in the first motion trajectory to obtain the first video frame to be processed.

[0200] According to the order of the positions of each facial key point in the first motion trajectory, each first video frame to be processed is sorted, and the sorted first video frames to be processed are combined into a first video to be processed.

[0201] Optionally, the device further includes:

[0202] The detection module is used to detect whether the facial features of the target person are displayed in the image to be processed before the first adjustment module 1102 performs the adjustment of the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed.

[0203] The prompt message output module is used to output a prompt message indicating that the image to be re-uploaded if the facial features of the target person are not displayed in the image to be processed.

[0204] The first adjustment module 1102 is specifically used for:

[0205] If the facial features of the target person are displayed in the image to be processed, the positions of the key facial points of the target person in the image to be processed are adjusted according to the first motion trajectory to obtain the first video to be processed.

[0206] Based on the image generation apparatus provided in this embodiment of the invention, since the template information includes a designated sticker identifier, and the designated sticker to which the designated sticker identifier belongs can indicate the emotion expressed by the template character, the electronic device adjusts the position of the facial key points of the target character in the image to be processed according to the position of the facial key points of the template character displayed in the designated template image. In the resulting first video to be processed, the expression change pattern of the target character is the same as that of the template character, so the target character can express the same emotion as the template character. Furthermore, by adding a designated sticker that can indicate the emotion expressed by the template character to the first video to be processed, that is, by adding a designated sticker that can indicate the emotion expressed by the target character, the target image obtained from the first video to be processed with the added designated sticker can better express the emotion, thereby improving the user experience.

[0207] This invention also provides an electronic device, such as... Figure 12 As shown, it includes a processor 1201, a communication interface 1202, a memory 1203, and a communication bus 1204. The processor 1201, communication interface 1202, and memory 1203 communicate with each other via the communication bus 1204.

[0208] Memory 1203 is used to store computer programs;

[0209] When the processor 1201 executes the program stored in the memory 1203, it implements the steps of the image generation method described in any of the above embodiments.

[0210] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0211] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0212] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0213] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0214] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described image generation methods.

[0215] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the image generation methods described above.

[0216] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0217] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0218] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0219] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An image generation method, characterized in that, The method includes: Acquire template information and an image to be processed; wherein, the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the facial key points of the template character change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent: the emotion expressed by the template character in the specified template image; the image to be processed is a static image displaying the target character; According to the first motion trajectory, adjust the position of the facial key points of the target person in the image to be processed to obtain the first video to be processed; Adjust the size of the designated sticker to which the designated sticker identifier belongs to the target size, and add the designated sticker of the target size at the target position in each first video frame of the first video to be processed according to the target angle to obtain the second video frame corresponding to the first video frame; The second video frames corresponding to each first video frame in the first video to be processed are combined to obtain the second video to be processed. The second video to be processed is converted into a dynamic image format to obtain the target image; The acquisition of template information and the image to be processed includes: Retrieve the image identifier of the specified template image and the image to be processed; Based on the pre-recorded correspondence between image identifiers and sticker identifiers, the sticker identifier corresponding to the image identifier of the specified template image is obtained, thus obtaining the specified sticker identifier; Before obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, the method further includes: A template video and multiple reference motion trajectories are obtained; wherein, the template video displays the template character; the reference motion trajectories are used to determine the emotions expressed by the template character displayed in the template video; the reference motion trajectories correspond to stickers; a sticker corresponding to a reference motion trajectory is used to represent: the emotions expressed by the template character whose facial key points change according to the reference motion trajectory; The change trajectory of the facial key points of the template character in the template video is obtained to obtain the first motion trajectory; Calculate the similarity between the first motion trajectory and each reference motion trajectory, and determine the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video; From the multiple stickers corresponding to the reference motion trajectory of the template video, determine the template sticker corresponding to the template video; The template video is converted into a dynamic image format to obtain a template image; Record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker to obtain the correspondence between the image identifier and the sticker identifier; The step of obtaining the change trajectory of the facial key points of the template person in the template video to obtain the first motion trajectory includes: For each template video frame of the template video, the facial key points of the template person in the template video frame are determined using a facial feature point determination method. For each facial key point of the template character, determine the coordinates of that facial key point in each template video frame; Determine the order in which the coordinates of the facial key points change according to the arrangement order of each template video frame in the template video; Based on the sequence of changes in the coordinates of the facial key points, the trajectory of the changes in the facial key points can be obtained; The trajectory of the position change of each facial key point of the template character is determined as the first motion trajectory.

2. The method according to claim 1, characterized in that, The process of recording the correspondence between the image identifier of the template image and the sticker identifier of the template sticker to obtain the correspondence between the image identifier and the sticker identifier includes: The template sticker corresponding to the template video is added to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

3. The method according to claim 2, characterized in that, The process of obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, to obtain the specified sticker identifier, includes: Obtain the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and use the sticker identifier corresponding to the image identifier of the specified template image as the specified sticker identifier.

4. The method according to claim 1, characterized in that, The process of obtaining the sticker identifier corresponding to the image identifier of the specified template image based on the pre-recorded correspondence between image identifiers and sticker identifiers, to obtain the specified sticker identifier, includes: According to the pre-recorded correspondence between image identifiers and sticker identifiers, obtain the sticker identifier of the target type corresponding to the image identifier of the specified template image, and display the obtained sticker identifier of the target type; From the displayed sticker identifiers of the target type, based on the selection instruction for the specified sticker identifier in the sticker identifiers of the target type, the sticker identifier indicated by the selection instruction input by the user is determined to be the specified sticker identifier corresponding to the image identifier of the specified template image, and the specified sticker identifier is obtained.

5. The method according to claim 1, characterized in that, The step of adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed includes: According to a first number of facial key points in the first motion trajectory, the image to be processed is copied to obtain the first number of images to be processed; wherein, one image to be processed corresponds to one position of a facial key point in the first motion trajectory; For each facial key point in the first motion trajectory, the position of the facial key point of the target person in the image to be processed corresponding to the position of the facial key point in the first motion trajectory is adjusted to the position of the facial key point in the first motion trajectory to obtain the first video frame to be processed. According to the order of the positions of each facial key point in the first motion trajectory, each first video frame to be processed is sorted, and the sorted first video frames to be processed are combined into a first video to be processed.

6. The method according to claim 1, characterized in that, Before adjusting the positions of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed, the method further includes: Detect whether the facial features of the target person are displayed in the image to be processed; If the facial features of the target person are not displayed in the image to be processed, a prompt message indicating that the image to be processed should be re-uploaded will be output. The step of adjusting the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed includes: If the facial features of the target person are displayed in the image to be processed, the positions of the key facial points of the target person in the image to be processed are adjusted according to the first motion trajectory to obtain the first video to be processed.

7. An image generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire template information and an image to be processed; wherein, the template information includes an image identifier of a specified template image and a corresponding specified sticker identifier; the specified template image is a dynamic image in which the facial key points of the template character change according to a first motion trajectory; the specified sticker to which the specified sticker identifier belongs is used to represent the emotion expressed by the template character in the specified template image; the image to be processed is a static image displaying a target character; The first adjustment module is used to adjust the position of the facial key points of the target person in the image to be processed according to the first motion trajectory, so as to obtain the first video to be processed; The second adjustment module is used to adjust the size of the specified sticker to which the specified sticker identifier belongs to the target size, and add the specified sticker of the target size at the target position in each first video frame in the first video to be processed according to the target angle to obtain the second video frame corresponding to the first video frame; The combination module is used to combine the second video frames corresponding to each first video frame in the first video to be processed to obtain the second video to be processed. The format conversion module is used to convert the second video to be processed into a dynamic image format to obtain the target image; The first acquisition module is specifically used for: Retrieve the image identifier of the specified template image and the image to be processed; Based on the pre-recorded correspondence between image identifiers and sticker identifiers, the sticker identifier corresponding to the image identifier of the specified template image is obtained, thus obtaining the specified sticker identifier; The device further includes: The second acquisition module is used to acquire a template video and multiple reference motion trajectories before the first acquisition module executes the pre-recorded correspondence between image identifiers and sticker identifiers to acquire the sticker identifier corresponding to the image identifier of the specified template image and obtains the specified sticker identifier. The template video displays the template character; the reference motion trajectories are used to determine the emotions expressed by the template character displayed in the template video; the reference motion trajectories correspond to template stickers; and the template stickers are used to indicate the emotions corresponding to the corresponding reference motion trajectories. The third acquisition module is used to acquire the change trajectory of the facial key points of the template character in the template video to obtain the first motion trajectory. The similarity calculation module is used to calculate the similarity between the first motion trajectory and each reference motion trajectory, and to determine the reference motion trajectory with the highest similarity to the first motion trajectory as the reference motion trajectory corresponding to the template video. The template sticker determination module is used to determine the template sticker corresponding to the template video from the template stickers corresponding to the reference motion trajectory corresponding to the template video; The recording module is used to convert the template video into a dynamic image format to obtain a template image, and to record the correspondence between the image identifier of the template image and the sticker identifier of the template sticker corresponding to the template video; The third acquisition module is specifically used for: For each template video frame of the template video, the facial key points of the template person in the template video frame are determined using a facial feature point determination method. For each facial key point of the template character, determine the coordinates of that facial key point in each template video frame; Determine the order in which the coordinates of the facial key points change according to the arrangement order of each template video frame in the template video; Based on the sequence of changes in the coordinates of the facial key points, the trajectory of the changes in the facial key points can be obtained; The trajectory of the position change of each facial key point of the template character is determined as the first motion trajectory.

8. The apparatus according to claim 7, characterized in that, The recording module is specifically used for: The template sticker corresponding to the template video is added to the template image to obtain the template image with the sticker added, which serves as the correspondence between the image identifier and the sticker identifier.

9. The apparatus according to claim 8, characterized in that, The first acquisition module is specifically used for: Obtain the sticker identifier of the sticker added to the specified template image to which the image identifier of the specified template image belongs, and use the sticker identifier corresponding to the image identifier of the specified template image as the specified sticker identifier.

10. The apparatus according to claim 7, characterized in that, The first acquisition module is specifically used for: According to the pre-recorded correspondence between image identifiers and sticker identifiers, obtain the sticker identifier of the target type corresponding to the image identifier of the specified template image, and display the obtained sticker identifier of the target type; From the displayed sticker identifiers of the target type, based on the selection instruction for the specified sticker identifier in the sticker identifiers of the target type, the sticker identifier indicated by the selection instruction input by the user is determined to be the specified sticker identifier corresponding to the image identifier of the specified template image, and the specified sticker identifier is obtained.

11. The apparatus according to claim 7, characterized in that, The first adjustment module is specifically used for: According to a first number of facial key points in the first motion trajectory, the image to be processed is copied to obtain the first number of images to be processed; wherein, one image to be processed corresponds to one position of a facial key point in the first motion trajectory; For each facial key point in the first motion trajectory, the position of the facial key point of the target person in the image to be processed corresponding to the position of the facial key point in the first motion trajectory is adjusted to the position of the facial key point in the first motion trajectory to obtain the first video frame to be processed. According to the order of the positions of each facial key point in the first motion trajectory, each first video frame to be processed is sorted, and the sorted first video frames to be processed are combined into a first video to be processed.

12. The apparatus according to claim 7, characterized in that, The device further includes: The detection module is used to detect whether the facial features of the target person are displayed in the image to be processed before the first adjustment module performs the adjustment of the position of the facial key points of the target person in the image to be processed according to the first motion trajectory to obtain the first video to be processed. The prompt message output module is used to output a prompt message indicating that the image to be re-uploaded if the facial features of the target person are not displayed in the image to be processed. The first adjustment module is specifically used for: If the facial features of the target person are displayed in the image to be processed, the positions of the key facial points of the target person in the image to be processed are adjusted according to the first motion trajectory to obtain the first video to be processed.

13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-6.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Expression recognition method and related device

    CN110852224A

  • Video display method and device and electronic equipment

    CN111372029A

  • Method and device for generating dynamic expression image sequence by utilizing static face images

    CN111696185A