Image synthesis method, electronic device and readable storage medium

By selecting frame images from multiple video streams, performing facial superposition and fusion and analysis of expression changes, the problem of poor image synthesis effect in the prior art is solved, and high-quality dynamic expression generation is achieved.

CN114170124BActive Publication Date: 2025-05-02MIGU CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111612863.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-05-02
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The existing image synthesis methods have poor results, especially under the influence of image quality and application scenarios, it is difficult to generate high-quality dynamic expressions.

Method used

By selecting frame images from multiple video streams, forming an image set, and performing facial superposition and fusion and analysis of expression changes, the synthetic image is processed by using expression changes to generate continuous dynamically changing expressions.

Benefits of technology

The expression generation effect is improved, and universal fusion based on face features is achieved. The generated images have better dynamic changes and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170124B_ABST
    Figure CN114170124B_ABST
Patent Text Reader

Abstract

The present application discloses an image synthesis method, an electronic device and a readable storage medium, which belong to the field of image processing technology. The specific implementation scheme includes: selecting n frames of images from each video stream of m video streams, and taking the n frames of images of each video stream as a group of image sets to obtain m groups of image sets; selecting a frame of image from each image set of the m groups of image sets to obtain m frames of images, and performing face superposition and fusion on the m frames of images to obtain a first image; determining the amount of facial expression change corresponding to the m groups of image sets based on feature analysis of the m groups of image sets; and superimposing the first image using the amount of facial expression change to obtain a composite image. In this way, based on the universality of facial features, faces in multiple video streams can be continuously fused to obtain continuously dynamically changing expressions, thereby improving the expression generation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to an image synthesis method, an electronic device and a readable storage medium. Background Art

[0002] At present, dynamic expression generation methods are usually based on a single picture or a single frame of an image, and are implemented based on a deep learning neural network, and the output result is a static picture. In this case, since the application effect of the neural network is often affected by the image quality and application scenario, the existing image synthesis effect will be poor. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide an image synthesis method, an electronic device and a readable storage medium to solve the problem of poor image synthesis effect in the prior art.

[0004] In a first aspect, an image synthesis method is provided, which is applied to an electronic device, comprising:

[0005] Selecting n frames of images from each of the m video streams, and taking the n frames of images of each video stream as a group of image sets, to obtain m groups of image sets; wherein m is an integer greater than 1, and n is an integer greater than 1;

[0006] Selecting one frame of image from each of the m sets of image sets to obtain m frames of image, and performing face superposition and fusion on the m frames of image to obtain a first image;

[0007] Determine the amount of change in facial expression corresponding to the m groups of images based on feature analysis of the m groups of images;

[0008] The first image is overlaid with the amount of change in facial expression to obtain a composite image.

[0009] In a second aspect, an image synthesis device is provided, which is applied to an electronic device, including:

[0010] A first selection module is used to select n frames of images from each of the m video streams, and use the n frames of images of each video stream as a group of image sets to obtain m groups of image sets; wherein m is an integer greater than 1, and n is an integer greater than 1;

[0011] A second selection module is used to select one frame of image from each of the m sets of image sets to obtain m frames of image;

[0012] A fusion module, used for performing face superposition and fusion on the m frames of images to obtain a first image;

[0013] A determination module, used for determining the amount of change of facial expression corresponding to the m groups of images based on feature analysis of the m groups of images;

[0014] The processing module is used to perform superposition processing on the first image using the amount of change in the facial expression to obtain a composite image.

[0015] In a third aspect, an electronic device is provided, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0016] In a fourth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] In the embodiment of the present application, n frames of images can be selected from each video stream of m video streams, and the n frames of images of each video stream are used as a group of image sets to obtain m groups of image sets; one frame of image can be selected from each image set of the m groups of image sets to obtain m frames of image, and the m frames of image can be superimposed and fused with faces to obtain a first image; based on the feature analysis of the m groups of image sets, the amount of facial expression change corresponding to the m groups of image sets can be determined; the first image can be superimposed and processed using the amount of facial expression change to obtain a composite image. Thus, based on the universality of facial features, faces in multiple video streams can be continuously fused to obtain continuously dynamically changing expressions, thereby improving the expression generation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flow chart of an image synthesis method provided by an embodiment of the present application;

[0019] Figure 2 is a schematic diagram of the image synthesis effect in the embodiment of the present application;

[0020] Figure 3 is a schematic diagram of identification of characteristic points in an embodiment of the present application;

[0021] Figure 4 is a structural schematic diagram of an image synthesis device provided in an embodiment of the present application;

[0022] Figure 5 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0025] The image synthesis method, electronic device and readable storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0026] See also Figure 1 , Figure 1 is a flow chart of an image synthesis method provided in an embodiment of the present application, the method is applied to electronic devices such as Figure 1 As shown, the method comprises the following steps:

[0027] Step 11: select n frames of images from each of the m video streams, and use the n frames of images of each video stream as a group of image sets to obtain m groups of image sets.

[0028] In this embodiment, m is an integer greater than 1. n is an integer greater than 1. In some embodiments, the value of n can be selected as the frame rate of the video stream*0.125.

[0029] Optionally, since human facial expressions generally last for a short time, generally about 0.1 seconds, in order to capture expressions more accurately, multiple consecutive frames of images can be selected as a group of image sets, that is, the n frames of images in each of the above image sets are continuous frame images.

[0030] Step 12: Select one frame of image from each of the m image sets to obtain m frames of image, and perform face superposition and fusion on the m frames of image to obtain a first image.

[0031] In this embodiment, in order to accurately characterize the initial amount of facial expression, the first frame of image can be selected from each image set of m sets of image sets to obtain m frames of image, and the m frames of image are superimposed and fused with faces to obtain the first image. The first image can be understood as an average face and can be represented by a vector of its feature points. For the method of superimposing and fusion of faces on m frames of image, the existing image fusion method can be used, and this is not limited.

[0032] In some embodiments, before performing face superposition and fusion, face region annotation may be performed first, such as annotating the selected m*n frame images, identifying the face, and delineating a rectangular face region; then feature point identification and annotation are performed in the annotated face region, such as annotating the selected m*n frame images, so as to perform image processing based on the annotated feature points. For example, for the delineated rectangular face region, the facial features and contours may be further annotated, mainly the key positions of the face such as the face outline, eyebrows, eyes, lips, nose, etc. may be annotated. For example, in order to accurately capture changes in facial expressions, the following may be annotated: 10 feature points for the left eye, 10 feature points for the right eye, 8 feature points for the left eyebrow, 8 feature points for the right eyebrow, 8 feature points for the nose, 18 feature points for the mouth, 20 feature points for the contour, 8 feature points for the forehead, 3 feature points for each of the left and right cheekbones, for a total of 96 feature points.

[0033] Step 13: Based on the feature analysis of the m groups of image sets, determine the amount of change in facial expression corresponding to the m groups of image sets.

[0034] Optionally, when determining the amount of change in facial expression, operations such as translation, rotation, and scaling may be performed on the images in each group of images to convert facial features in different images into a unified mapping space, making the facial features comparable and quantifying the degree of facial change.

[0035] Step 14: Use the facial expression variation to perform superposition processing on the first image to obtain a composite image.

[0036] In this step, when the first image is superimposed using the facial expression variation, the facial feature vector of the first image can be Based on the face expression change b, the fusion face is generated. .

[0037] The image synthesis method of the embodiment of the present application can select n frames of images from each video stream of m video streams, and use the n frames of images of each video stream as a group of image sets to obtain m groups of image sets; select one frame of image from each image set of the m groups of image sets to obtain m frames of image, and perform face superposition and fusion on the m frames of image to obtain a first image; determine the amount of facial expression change corresponding to the m groups of image sets based on feature analysis of the m groups of image sets; and use the amount of facial expression change to superimpose the first image to obtain a synthesized image. Thus, based on the universality of facial features, faces in multiple video streams can be continuously fused to obtain continuously dynamically changing expressions, thereby improving the expression generation effect. Furthermore, the scheme in the present application also has good versatility and portability, and the effect is generally applicable to various face fusion scenarios.

[0038] It should be pointed out that the scenarios applicable to the embodiments of the present application include but are not limited to video conferencing, directing or live broadcasting, etc., to achieve the interaction of multiple video participants. A person can affect the fused expression through the change of expression, that is, everyone is a contributor to the final expression, thereby achieving multi-person interaction, dynamically changing expressions, and increasing fun and interactivity. In addition, the embodiments of the present application can also be suitable for face-changing scenarios, by adjusting parameters, generating multiple intermediate result images between the original image and the target image.

[0039] For example, when this solution is applied to video conferencing, the facial expression information of all participants can be extracted, the expression changes of each participant can be calculated, and the degree of expression changes of all participants can be superimposed and fused to jointly generate a facial expression, which will continue to change with time. Everyone can influence the fused facial expression at any time, thereby achieving interaction among participants and reflecting the overall atmosphere of the meeting.

[0040] For example, see Figure 2 As shown, if there are 4 video streams, among the corresponding 4 groups of images, 2 groups of images have facial expressions with obvious smiles, and 2 groups of images have facial expressions with calmer facial expressions. Then, through the image processing process in this scheme, an image that integrates the facial expressions of the 4 groups of images can be generated, and the face in the image has a smiling expression.

[0041] In the embodiments of the present application, since different angles, distances, posture changes, etc. will cause differences in the faces in the images, alignment based on facial features can be used to convert the facial features in different images into the same mapping space, making the facial features comparable and quantifying the degree of change of the face.

[0042] Optionally, the above process of determining the amount of change in facial expression corresponding to the m groups of image sets based on feature analysis of the m groups of image sets may include: first, for each group of image sets in the m groups of image sets, calculating the feature vectors of the n frames of images in each group of image sets after alignment, obtaining n feature vectors, and using the n feature vectors to calculate the amount of change in facial expression for each group of image sets; then, using the amount of change in facial expression for each group of image sets, calculating the amount of change in facial expression corresponding to the m groups of image sets.

[0043] Optionally, when aligning images, feature points can be rotated, scaled, horizontally translated, and vertically translated to the same mapping space without changing the distribution of feature points, so that feature points are comparable, feature changes are quantified, and interference caused by non-shape factors is reduced. The above calculation of the feature vectors of n frames of images in each set of images after alignment can obtain n feature vectors that may include:

[0044] The following iterative process is performed cyclically until the average value of the feature vectors of the n frames of images in each group of images converges, and n feature vectors of the n frames of images when the average value converges are obtained:

[0045] S1: Rotate and scale the n frames in each image set to obtain the first eigenvector X i '; Among them, X i '=M(S i ,θ i )[X i ], X i Represents the feature vector of the i-th frame image among the n frames of images, i = {1,…,n}. θ i represents the rotation angle, S i Indicates the scaled size, represents the horizontal coordinate of the kth feature point in the i-th frame image, represents the ordinate of the kth feature point in the i-th frame image; k = {1,…,a}, a represents the number of feature points in the i-th frame image. i , based on the two-dimensional position information of the feature points of the i-th frame image, the feature points can be connected into vectors in order: Represents the coordinate value of the kth feature point in the i-th frame image. For example, if the i-th frame image has 96 feature points, then n is equal to 96, and the i-th frame image is represented by a 96*2-dimensional vector.

[0046] S2: Perform translation transformation on n frames of images in each set of images to obtain the second eigenvector Z i ;in, Represents the average value of the feature vectors of n frames of images calculated in the previous iteration process, Indicates the horizontal offset of the kth feature point in the i-th frame image. represents the vertical offset of the kth feature point in the i-th frame image. It should be noted that in the initial iteration process, Z i =X1-X i '-M i , X1 represents the feature vector of the reference image. The reference image can be any one of the n frames. Preferably, since the expression change needs to be quantified, X1 can be the first frame in the corresponding group.

[0047] S3: Calculate the average value of the feature vectors of n frames of images in each set of images in this round of iteration in, W represents the weight matrix of each feature point in the i-th frame image, E i Represents the feature vector of the i-th frame image after this round of iterative alignment.

[0048] It should be noted that the condition for the average value of the feature vectors of n frames of images in each set of images to converge may be that the number of iterations reaches a preset value, or that the change in the average value in adjacent iterations is less than a preset value, and there is no limitation on this. i , S i , and The specific value of E i For θ i , S i , and Take the derivatives separately and set the derivatives to 0 to find the corresponding values.

[0049] Optionally, the weight matrix W can be expressed as follows:

[0050] in, Represents the weight of the kth feature point in the i-th frame image; Represents the R between the n frames of images in the image set where the i-th frame image is located kj The variance, R kj represents the distance between the kth feature point and the jth feature point in the i-th frame image, j = {1,…,a}, that is, R kj It represents the distance between the kth feature point and other feature points, and a represents the number of feature points in the i-th frame image.

[0051] Optionally, if the n eigenvectors are E1, E2, ..., E n , the above process of calculating the facial expression variation of each image set using n feature vectors may include:

[0052] S1: Calculate the average vector of the n eigenvectors and the covariance matrix S; where

[0053] S2: Calculate multiple eigenvalues ​​of the covariance matrix S;

[0054] S3: Using the average vector and the multiple eigenvalues, and calculate the facial expression variation of each image set.

[0055] In some embodiments, the multiple eigenvalues ​​of the covariance matrix S are arranged in descending order and can be expressed as λ1, λ2, ..., λ q , q represents the number of eigenvalues ​​of the covariance matrix S, and the corresponding eigenvalue vector P q =(λ1,λ2,...,λ q ).

[0056] Optionally, when calculating the facial expression variation of each set of images, the i-th frame image can be simplified as: Among them, b q is a vector containing q parameters, Let b m =P q b q , then the i-th frame image can be expressed as b m It is the variation of facial expression in this set of images.

[0057] Since the first eigenvalues ​​of the covariance matrix S mainly affect the change in facial expression, in order to reduce the amount of calculation, the formula can be used Select t eigenvalues ​​from q eigenvalues ​​to calculate the corresponding facial expression change. f v Indicates precision, and its value range is (0, 1). For example, the possible value is 0.92.

[0058] Furthermore, after selecting t eigenvalues, when calculating the facial expression variation of each image set, the i-th frame image can be simplified as: Among them, b t is a vector containing t parameters, Let b m =P q b q , then the i-th frame image can be expressed as b m It is the variation of facial expression in this set of images.

[0059] Optionally, when the facial expression change amount of each image set is used to calculate the facial expression change amount corresponding to the m image sets, the following formula can be used to calculate the facial expression change amount b corresponding to the m image sets:

[0060]

[0061] Among them, b i represents the facial expression change of the i-th image set in the m-image set, α i represents the weight of the facial expression change in the i-th image set. For example, In some embodiments, the default But it is not limited to this and can be adjusted according to actual conditions. Different people have different weights. The greater the weight, the greater the impact on the final synthesis effect.

[0062] In an embodiment of the present application, in order to optimize the image synthesis effect, after obtaining the synthesized image, the relationship between adjacent feature points can be comprehensively considered on the basis of the synthesized image F to supplement local features and enrich the details of the feature points, so that the image feature representation is more accurate and smoother.

[0063] Optionally, after obtaining the synthesized image, the image synthesis method in this embodiment further includes:

[0064] Calculate the local features of each feature point in the synthesized image;

[0065] The local features of each feature point are superimposed on the composite image to obtain the final composite image.

[0066] Optionally, when calculating the local features of each feature point in the composite image, for the i-th feature point in the composite image, the pixel vector v1 between the i-th feature point and the i+1-th feature point, and the pixel vector v2 between the i-th feature point and the i-1-th feature point can be calculated first; wherein, i = {2,…,a}, a represents the number of feature points in the composite image; then, according to the distance between the i-th feature point and the i+1-th feature point and the distance between the i-th feature point and the i-1-th feature point, the composite vector v of the pixel vector v1 and the pixel vector v2 is calculated, wherein the influence of feature points with a long distance is smaller and the influence of feature points with a short distance is greater; finally, the grayscale values ​​of the pixels contained in the composite vector v are derived to obtain the local features of the i-th feature point.

[0067] In some embodiments, taking the ith feature point and the i+1th feature point as an example, the pixel point vector between the ith feature point and the i+1th feature point can be represented by the coordinates of a preset number of pixel points between the ith feature point and the i+1th feature point. The preset number can be pre-set based on actual needs and is not limited to this.

[0068] For example, see Figure 3 As shown in the figure, for feature point i, the two adjacent feature points before and after it are the i-1th feature point and the i+1th feature point. 2m+1 pixels can be selected in the direction of the vertical midline between the i-th feature point and the i+1th feature point to form the pixel vector v1=[x1,x2,...,x 2m+1 ], and select 2m+1 pixels in the direction of the vertical midline between the ith feature point and the i-1th feature point to form the pixel vector v2=[y1,y2,...,y 2m+1 ].

[0069] Furthermore, when merging v1 and v2, the distance d between the i-th feature point and the i+1-th feature point can be used. i-1,i and the distance d between the i-th feature point and the i-1-th feature point i,i+1 , using the following formula Get the synthetic vector v. Derivative the grayscale value of the pixel contained in the synthetic vector v, and get the local feature δ i Similarly, after obtaining the local features of all other feature points in the composite image, the obtained local features are superimposed on the corresponding feature points to obtain the final composite image.

[0070] It should be noted that the image synthesis method provided in the embodiment of the present application can be executed by an image synthesis device or a control module in the image synthesis device for executing the image synthesis method. In the embodiment of the present application, the image synthesis device provided in the embodiment of the present application is described by taking the image synthesis device executing the image synthesis method as an example.

[0071] See also Figure 4 , Figure 4 is a schematic diagram of the structure of an image synthesis device provided in an embodiment of the present application, and the device is applied to electronic devices such as Figure 4 As shown, the image synthesis device 40 may include:

[0072] The first selection module 41 is used to select n frames of images from each of the m video streams, and use the n frames of images of each video stream as a group of image sets to obtain m groups of image sets; wherein m is an integer greater than 1, and n is an integer greater than 1;

[0073] A second selection module 42 is used to select one frame of image from each of the m sets of image sets to obtain m frames of image;

[0074] A fusion module 43 is used to perform face superposition fusion on the m frames of images to obtain a first image;

[0075] A determination module 44, configured to determine the amount of change in facial expression corresponding to the m groups of images based on feature analysis of the m groups of images;

[0076] The processing module 45 is used to perform superposition processing on the first image using the facial expression change amount to obtain a composite image.

[0077] Optionally, the determining module 44 includes:

[0078] A first calculation unit is used to calculate, for each image set in the m image sets, feature vectors of n frames of images in each image set after alignment to obtain n feature vectors, and calculate the amount of facial expression change in each image set using the n feature vectors;

[0079] The second calculation unit is used to calculate the facial expression change amounts corresponding to the m groups of image sets by using the facial expression change amounts of each group of image sets.

[0080] Optionally, the first calculation unit is specifically used to: loop through the following iterative process until the average value of the feature vectors of the n frames of images in each group of images converges, and obtain n feature vectors of the n frames of images when the average value converges:

[0081] Rotate and scale the n frames of image to obtain the first eigenvector X i '; Among them, X i '=M(S i ,θ i )[X i ], X i A feature vector representing the i-th frame image among the n frames of images, i={1,…,n}; θ i represents the rotation angle, S i Indicates the scaled size, represents the horizontal coordinate of the kth feature point in the i-th frame image, represents the ordinate of the kth feature point in the i-th frame image; k={1,…,a}, where a represents the number of feature points in the i-th frame image;

[0082] Perform translation transformation on the n frames of images to obtain the second eigenvector Z i ;in, Represents the average value of the feature vectors of n frames of images calculated in the previous iteration process, represents the horizontal offset of the kth feature point in the i-th frame image, represents the vertical offset of the k-th feature point in the i-th frame image;

[0083] Calculate the average value of the feature vectors of the n frames of images in this round of iteration in, W represents the weight matrix of each feature point in the i-th frame image, E i Represents the feature vector of the i-th frame image after this round of iterative alignment.

[0084] Optionally, W is as follows:

[0085] in, Represents the weight of the kth feature point in the i-th frame image; represents the R between n frames of images in the image set where the i-th frame image is located kj The variance, R kj represents the distance between the kth feature point and the jth feature point in the i-th frame image, j={1,…,a}.

[0086] Optionally, the first calculation unit is specifically used to: calculate the average vector and covariance matrix of the n eigenvectors, calculate multiple eigenvalues ​​of the covariance matrix, and use the average vector and the multiple eigenvalues ​​to calculate the change in facial expression of each group of images.

[0087] Optionally, the second calculation unit is specifically used to calculate the facial expression change amount b corresponding to the m groups of image sets using the following formula:

[0088]

[0089] Among them, b i represents the facial expression change of the i-th image set in the m-image set, α i The weight representing the amount of change in facial expression of the i-th image set.

[0090] Optionally, the image synthesis device 40 further includes:

[0091] A calculation module, used for calculating the local features of each feature point in the composite image;

[0092] The processing module 45 is further configured to: superimpose the local features of each feature point onto the composite image to obtain a final composite image.

[0093] Optionally, the calculation module is specifically used to: for the i-th feature point in the composite image, calculate the pixel point vector v1 between the i-th feature point and the i+1-th feature point, and calculate the pixel point vector v2 between the i-th feature point and the i-1-th feature point; wherein, i = {2,…,a}, a represents the number of feature points in the composite image; calculate the composite vector v of the pixel point vector v1 and the pixel point vector v2 based on the distance between the i-th feature point and the i+1-th feature point and the distance between the i-th feature point and the i-1-th feature point; and derive the grayscale values ​​of the pixels contained in the composite vector v to obtain the local features of the i-th feature point.

[0094] Optionally, the second selection module 42 is specifically used to: select the first frame image from each group of image sets to obtain m frames of images.

[0095] Optionally, the n frames of images in each group of images are: continuous frame images.

[0096] The image synthesis device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and the non-mobile electronic device can be a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0097] The image synthesis device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0098] The image synthesis device 40 of the embodiment of the present application can realize the above Figure 1 The various processes of the method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0099] Optional, such as Figure 5As shown, an embodiment of the present application also provides an electronic device 50, including a processor 51, a memory 52, and a program or instruction stored in the memory 52 and executable on the processor 51. When the program or instruction is executed by the processor 51, each process of the above-mentioned image synthesis method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0100] The embodiment of the present application also provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned image synthesis method embodiment can be implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0101] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0102] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0103] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0104] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0105] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image synthesis method, applied to electronic equipment, characterized in that: include: Selecting n frames of images from each of the m video streams, and taking the n frames of images of each video stream as a group of image sets, to obtain m groups of image sets; wherein m is an integer greater than 1, and n is an integer greater than 1; Selecting one frame of image from each of the m sets of image sets to obtain m frames of image, and performing face superposition and fusion on the m frames of image to obtain a first image; Determine the amount of change in facial expression corresponding to the m groups of images based on feature analysis of the m groups of images; The first image is overlaid with the amount of change in facial expression to obtain a composite image.

2. The method according to claim 1, characterized in that Determining the amount of change in facial expression corresponding to the m groups of images based on feature analysis of the m groups of images includes: For each image set in the m image sets, calculate the feature vectors of n frames of images in each image set after alignment to obtain n feature vectors, and calculate the amount of facial expression change in each image set using the n feature vectors; The facial expression changes of each image set are used to calculate the facial expression changes corresponding to the m image sets.

3. The method according to claim 2, characterized in that The step of calculating the aligned feature vectors of the n frames of images in each group of images to obtain n feature vectors includes: The following iterative process is performed cyclically until the average value of the feature vectors of the n frames of images in each group of images converges, and n feature vectors of the n frames of images when the average value converges are obtained: Rotate and scale the n frames of image to obtain the first eigenvector X i '; Among them, X i '=M(S i ,θ i )[X i ], X i A feature vector representing the i-th frame image among the n frames of images, i={1,…,n}; θ i represents the rotation angle, S i Indicates the scaled size, represents the horizontal coordinate of the kth feature point in the i-th frame image, represents the ordinate of the kth feature point in the i-th frame image; k={1,…,a}, where a represents the number of feature points in the i-th frame image; Perform translation transformation on the n frames of images to obtain the second eigenvector Z i ;in, Represents the average value of the feature vectors of n frames of images calculated in the previous iteration process, represents the horizontal offset of the kth feature point in the i-th frame image, represents the vertical offset of the k-th feature point in the i-th frame image; Calculate the average value of the feature vectors of the n frames of images in this round of iteration in, W represents the weight matrix of each feature point in the i-th frame image, E i Represents the feature vector of the i-th frame image after this round of iterative alignment.

4. The method according to claim 3, characterized in that The W is as follows: in, Represents the weight of the kth feature point in the i-th frame image; represents the R between n frames of images in the image set where the i-th frame image is located kj The variance, R kj represents the distance between the kth feature point and the jth feature point in the i-th frame image, j={1,…,a}.

5. The method according to claim 2, characterized in that: The step of calculating the amount of change in facial expression of each image set by using the n feature vectors includes: Calculate the mean vector and covariance matrix of the n eigenvectors; Calculating a plurality of eigenvalues ​​of the covariance matrix; The average vector and the multiple eigenvalues ​​are used to calculate the variation of facial expressions in each group of image sets.

6. The method according to claim 2, characterized in that The step of calculating the facial expression change amounts corresponding to the m groups of image sets by using the facial expression change amounts of each group of image sets comprises: The following formula is used to calculate the facial expression change b corresponding to the m sets of images: Among them, b i represents the facial expression variation of the i-th image set in the m image sets, α i The weight representing the amount of change in facial expression of the i-th image set.

7. The method according to claim 1, characterized in that After obtaining the composite image, the method further includes: Calculating a local feature of each feature point in the composite image; The local features of each feature point are superimposed on the composite image to obtain a final composite image.

8. The method according to claim 7, characterized in that The calculating the local feature of each feature point in the synthesized image comprises: For the i-th feature point in the synthetic image, calculate the pixel point vector v1 between the i-th feature point and the i+1-th feature point, and calculate the pixel point vector v2 between the i-th feature point and the i-1-th feature point; wherein i={2,…,a}, and a represents the number of feature points in the synthetic image; Calculate a composite vector v of the pixel point vector v1 and the pixel point vector v2 according to the distance between the i-th feature point and the i+1-th feature point and the distance between the i-th feature point and the i-1-th feature point; The grayscale values ​​of the pixels included in the synthetic vector v are derived to obtain the local features of the i-th feature point.

9. An electronic device, characterized in that: The method comprises a processor, a memory and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the image synthesis method as described in any one of claims 1 to 8.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the image synthesis method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Face image generation method and device, electronic equipment and storage medium

    CN113222876A

  • Stereoscopic image printing system

    JP1997127622A