Method, device and storage medium for matching animation skeletons based on topological structure
By determining the topological structure in the static image and matching it with the animation skeleton library, the problem of inconsistent dynamic display caused by unsuitable animation skeletons is solved, a more coordinated dynamic display effect is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202310111239.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-02-03
AI Technical Summary
In existing methods for converting static images into dynamic images, the matched animation skeleton information files are inappropriate, resulting in an uncoordinated dynamic display of static images and a poor user experience.
By obtaining the outer contour of the target object in the static image, its topological structure is determined, and based on the topological structure, it is matched with the animation skeleton in the animation skeleton library, and the animation skeleton is driven to perform deformation processing to achieve dynamic display.
Improves the coordination of dynamic display of static images and enhances the user experience.
Smart Images

Figure CN116228940B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method, device and storage medium for matching animation skeletons based on topological structures. Background Art
[0002] With the widespread adoption and application of smart electronic devices such as mobile phones and tablets, especially with the continuous upgrade of camera hardware and the maturity of facial recognition technology, more and more users prefer to use mobile phones for taking photos, and mobile phone photography has gradually replaced still cameras. Although existing mobile phone camera technologies can perform some simple photo processing, such as beautifying and adjusting the color of people in photos, or blurring the background, the photos processed by these image processing technologies are still static images with poor interactivity, which is far from enough to meet people's entertainment needs.
[0003] To enable the dynamic display of static images captured by smart electronic devices such as mobile phones, the applicant submitted an invention patent application with application number 201611088517.6 to the State Intellectual Property Office on May 31, 2017. This application discloses a method and apparatus for converting static images into dynamic images. In this application, an electronic image is first acquired, and the contour features of the object in the electronic image are extracted to obtain the object's inner contour image. Then, based on topological structure analysis, an animation skeleton information file corresponding to the inner contour image is obtained. Based on the inner contour image and the corresponding animation skeleton information file, the target object in the electronic image is converted into a vector model. Finally, based on the vector model and the corresponding inner contour image, a vector model with facial features is obtained, and the static image is driven by facial expression or limbs.
[0004] However, this solution does not disclose how to analyze the topological structure, and the matched animation skeleton information file is sometimes not very suitable, resulting in poor overall coordination when dynamically displaying static images and a low user experience.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] The embodiments of the present application provide a method, apparatus, and storage medium for matching animation skeletons based on topological structures, so as to at least solve the technical problem of incoordination of static images in dynamic display caused by inappropriate matching of animation skeletons.
[0007] According to one aspect of an embodiment of the present application, a method for matching animation skeletons based on topological structure is provided, comprising: acquiring a static image and extracting the outer contour of a target object in the static image from the static image; determining the topological structure of the image within the outer contour, and matching the image within the outer contour with animation skeletons in an animation skeleton library based on the topological structure, wherein the matched animation skeletons are used to dynamically display the image within the outer contour.
[0008] According to another aspect of an embodiment of the present application, a device for matching animation skeletons based on topological structure is also provided, including: a contour acquisition module, configured to acquire a static image and extract the outer contour of a target object in the static image from the static image; a topological analysis module, configured to determine the topological structure of the image within the outer contour, and based on the topological structure, match the image within the outer contour with animation skeletons in an animation skeleton library, wherein the matched animation skeletons are used to dynamically display the image within the outer contour.
[0009] In an embodiment of the present application, the topological structure of the image within the outer contour of the target object in the static image is determined, and based on the topological structure, the image within the outer contour is matched with the animation skeleton in the animation skeleton library, thereby solving the technical problem of the incoordination of the dynamic display of the static image caused by the unsuitability of the matched animation skeleton. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0011] Figure 1 is a flow chart of a method for dynamically displaying a static image according to an embodiment of the present application;
[0012] Figure 2 is a flowchart of another method for dynamically displaying a static image according to an embodiment of the present application;
[0013] Figure 3 is a flowchart of another method for dynamically displaying a static image according to an embodiment of the present application;
[0014] Figure 4 is a flow chart of a dynamic display method capable of replacing a static image of a face according to an embodiment of the present application;
[0015] Figure 5 is a flow chart of a method for performing face-changing processing on a dynamically displayed target object according to an embodiment of the present application;
[0016] Figure 6Schematic diagram of the optical flow of pixels in two adjacent frames of images according to an embodiment of the present application;
[0017] Figure 7 2 is a schematic diagram of the optical flow of pixels in two adjacent frames of pictures with an integrated window according to an embodiment of the present application;
[0018] Figure 8 is a flowchart of a method for matching animation skeletons based on topological structures according to an embodiment of the present application;
[0019] Figure 9 is a flowchart of another method for matching animation skeletons based on topological structures according to an embodiment of the present application;
[0020] Figure 10 is a structural schematic diagram of a dynamic display device for static images according to an embodiment of the present application;
[0021] Figure 11 is a structural diagram of a device for matching animation skeletons based on topological structures according to an embodiment of the present application;
[0022] Figure 12 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] Example 1
[0026] According to an embodiment of the present application, a method for dynamically displaying a static image is provided, such as Figure 1 As shown, the method includes:
[0027] Step S102 : acquiring a static image, and extracting the outer contour of a target object in the static image from the static image to obtain outer edge position data of the outer contour of the target object.
[0028] A static image can be a photo taken by an imaging device or an image created by graphic software or a graphic tool, and the format can be a frame image in a video, a JPG picture, etc.
[0029] In some embodiments, multiple target sequence values corresponding to each of the at least two edges of the static image can be obtained based on the length of each edge; then, the probability that each of the multiple target sequence values corresponds to the outer contour is determined; finally, based on the determined probability, the outer contour of the target object is extracted from the static image.
[0030] In other embodiments, contour features of the target object in a static image may be extracted to obtain an image within the outer contour of the target object. For example, contour feature extraction, image edge detection, and color space feature mapping operations may be performed on the static image to obtain outer edge position data of the outer contour of the target object.
[0031] Step S104: Match the outer edge position data with the animation skeleton in the animation skeleton library, and bind the outer edge position data with the matched animation skeleton.
[0032] In some embodiments, a directional cutting method may be used to match the outer edge position data with each animation skeleton in the animation skeleton library.
[0033] First, the animation skeleton is preprocessed. For example, based on the outer edge position data, the area and geometric shape of the image within the outer contour are estimated; based on the estimated area, the animation skeleton in the animation skeleton library is scaled, and based on the estimated geometric shape, the animation skeleton in the animation skeleton library is rotated.
[0034] The image within the outline is then segmented based on the direction, and the resulting sub-images are matched. For example, the image within the outline is divided into multiple sub-images based on multiple preset directions. Then, for each of the multiple directions, a determination is made as to whether the outer edge position data of the sub-image corresponding to that direction matches the position data of the processed animation skeleton in that direction.
[0035] In other embodiments, a method of analyzing a topological structure may be used to match the outer edge position data with each animation skeleton in an animation skeleton library.
[0036] For example, first, based on the outer edge position data and the image within the outline, a three-dimensional model corresponding to the image within the outline is constructed; then, a continuous function is defined on the triangular mesh of the three-dimensional model, and the function value of each vertex on the triangular mesh is calculated using the continuous function; then, vertices with the same function value and located on the same connected component are classified into one category and used as a new node quotient set; finally, a topological structure is generated based on the new node quotient set, and based on the topological structure, it is matched with the animation bones in the animation bone library.
[0037] Step S106 : driving the bound animation skeleton and deforming the image within the outer contour based on the motion trajectory of the bound animation skeleton to dynamically display the target object in the static image.
[0038] First, pre-deformation is performed. For example, an energy map of the image within the outer contour is calculated. Based on the energy map, the minimum energy line of the image within the outer contour is found. Next, based on the minimum energy line and the motion trajectory, each pixel within the image within the outer contour is moved, and the optical flow field of each pixel after the movement is calculated. Finally, based on the optical flow field, the motion vector of each pixel within the image within the outer contour is calculated to pre-deform the image within the outer contour.
[0039] Afterwards, the pre-deformed image is divided into grids, and the vertices of the divided grids are used as control vertices to perform global deformation on the pre-deformed image using similarity transformation.
[0040] In this embodiment, the outer edge position data of the target object is matched with the animation skeleton in the animation skeleton library, and the outer edge position data is bound to the matched animation skeleton; the bound animation skeleton is driven, and the image within the outer contour is deformed based on the motion trajectory of the bound animation skeleton, thereby solving the technical problem of uncoordinated dynamic display caused by not deforming the image within the contour.
[0041] Example 2
[0042] According to an embodiment of the present application, another method for dynamically displaying a static image is provided, such as Figure 2 As shown, the method includes:
[0043] Step S202: Acquire a static image.
[0044] A static image can be a photo taken by an intelligent electronic device such as a mobile phone, tablet computer, or camera. It can also be an image created using graphics software or a drawing tool. It can also be a photo obtained from the photo album of the intelligent electronic device or a picture downloaded from the Internet. In some examples, a static image can also be a frame image from a video. For example, a static image corresponding to an eye can be a frame image of a closed eye or a frame image of an open eye.
[0045] Step S204: extracting the outer contour of the target object from the static image.
[0046] In some embodiments, detection can be performed on a static image to extract the outer contour of a target object, such as a person, animal, plant, or object. For example, edge detection and color space feature mapping can be performed on the static image to obtain the outer contour of the target object. The edge detection method will be described in detail below and will not be repeated here.
[0047] In some other embodiments, a user input instruction may be received, and the outer contour of the target object may be selected from the static image based on the user input instruction. For example, the outer contour of the target object may be obtained by the user defining a range in the static image by swiping a finger on a touch screen or by using a drawing tool to select a range in the static image.
[0048] The target object in this embodiment is not limited to the overall target in the static image, such as people, animals, plants, objects, etc., but can also be one or some feature areas in the static image. For example, when the static image contains a person, the feature area can be the person's eyes, mouth, nose, ears, hands, feet, torso, etc., or it can be a combination of the above parts, for example, a combination of eyes, mouth, nose, and ears.
[0049] Afterwards, based on the extracted outer contour, outer edge position data of the outer contour is obtained.
[0050] Step S206: Bind the image within the outline to the selected animation skeleton.
[0051] Traverse each animation bone in the animation bone library, adjust each animation bone by moving, rotating and scaling to match the image within the outline.
[0052] In some embodiments, when matching the image within the outer contour and the animated skeleton, a layer of object voxels of the image within the outer contour is peeled off while maintaining the topology and local elongation of the image within the outer contour. The image within the outer contour is then divided into sub-images in six directions, for example, north, south, east, west, up, and down. A sub-iteration is performed for each direction to determine whether the outer edge position data of the animated skeleton in that direction matches the outer edge position data of the image within the outer contour in that direction, wherein the voxels in the sub-images in the corresponding directions are processed in parallel.
[0053] In other embodiments, the image within the outer contour can be first segmented so that the simplicity of any voxel is independent of the object configuration of any other voxel from the same sub-image. Eight subfields are then defined within the 3D cube grid, and their topology-preserving properties are established. An iterative parallel algorithm is then used to calculate whether the distance transform of the image within the outer contour matches the distance transform of the scaled animated skeleton in each of the eight subfields.
[0054] Step S208 , driving the animation skeleton to perform corresponding deformation processing on the image within the outer contour to generate a dynamic image.
[0055] First, it is necessary to find and calculate the minimum energy line of the image within the outer contour and pre-deform the image within the outer contour. For example, the image within the outer contour can be enlarged into a rectangular image. Calculate the energy map of the rectangular image; based on the energy map, find the minimum energy line of the rectangular image. Each pixel in the image within the outer contour can move its position based on the motion trajectory of the animation skeleton. The pixel points of the image within the outer contour can be enhanced using the Seam mapping algorithm. Assuming that the position movement of the pixel point is I'(x, y), then its optical flow field u(x, y)-I'(x, y) can be calculated. In this way, after the image within the outer contour is converted into a locally twisted image, the motion vector of each pixel is calculated.
[0056] Next, the rectangular image is divided into a grid. After recording the partitions, each pixel is deformed through the motion field to produce a local twisted image. In the motion field, the deformation is reversed while recording the new position of each vertex and storing it in the grid.
[0057] Using the new positions v'(x, y) of these mesh vertices as control vertices, we use a similarity transformation to globally deform the image within the outer contour. The control vertices for the deformation are v'(x, y), and the positions after deformation are v(x, y). This gives us the global image deformation result.
[0058] In this embodiment, the optical flow field of each pixel of the image within the outer contour is calculated to pre-deform the image within the outer contour, and after the pre-deformation, global deformation is performed to achieve the purpose of more coordinated dynamic display of the target object in the static image, thereby solving the technical problem of poor user experience caused by the inability to coordinate dynamic display of static images in the existing technology.
[0059] Example 3
[0060] According to an embodiment of the present application, another method for dynamic display of static images is provided, such as Figure 3 As shown, the method includes:
[0061] Step S302: Acquire a static image and extract the outer contour of the target object from the static image.
[0062] In this embodiment, an edge detection method is used to extract the outer contour of the target object from a static image.
[0063] First, based on the length of each of at least two edges in a static image, multiple target sequence values corresponding to the edge are obtained. After obtaining the static image, multiple target sequence values corresponding to the edge are obtained based on the length of each edge in the static image. For example, the edge lengths of the at least two edges in the static image are discretized to obtain multiple target sequence values. The multiple target sequence values can be any value between 0 and x, where x is the edge length value.
[0064] Next, a target tracking feature is obtained according to the static image, where the target tracking feature represents a probability that each target sequence value among a plurality of target sequence values corresponding to at least two edges corresponds to an outer contour position of the target object in the static image.
[0065] The edge of the outer contour is the result of discontinuity in grayscale values. This discontinuity can be detected by convolution using a spatial differential operator to further obtain target tracking features. Specifically, a filter is used to filter out the noise in the outer contour image to improve the performance of related edge detection. Then, the first-order derivative and second-order derivative of the intensity of the image within the outer contour after filtering out the noise are used to perform edge detection. For example, the neighborhood intensity change value of each pixel in the image within the outer contour is first determined, so that the pixel points with obvious changes in intensity value can be highlighted. There will be many points with relatively large gradient amplitudes in the image within the outer contour. At this time, the edge position can be judged by sub-pixel resolution, and the orientation of the edge can also be judged to obtain a probability distribution prediction of the outer contour position of the target object in the static image corresponding to each target sequence value.
[0066] When detecting multiple target sequence values for the longer edge of at least two edges, as many probability distribution predictions as possible are performed on the static image to improve the accuracy of the probability distribution represented by the obtained target tracking feature, thereby ultimately improving the accuracy of the target detection feature.
[0067] In this embodiment, multiple target sequence values corresponding to at least two edges of a static image are extracted, and then the target sequence values are matched with position features of the target object and probability distribution prediction is performed to obtain target tracking features, thereby more accurately extracting the outer contour of the target object.
[0068] Step S304: generating a three-dimensional model based on the image within the outer contour.
[0069] Obtain the image within the outer contour and its depth information. Depth information is used to describe the grayscale or color of each pixel in the image. The depth information of the corresponding scene is generally described by the same grayscale size to describe the grayscale of the image. The grayscale value of each pixel in the grayscale image describes the depth value of the corresponding scene, i.e., the depth information.
[0070] This embodiment calculates the depth of each pixel in the image within the outer contour using depth information. This calculated depth information is then used to construct a 3D model. For example, the depth value of each pixel in the image within the outer contour is obtained based on the depth information in the depth map. A 3D modeling algorithm is then used to map the depth value of each pixel in the depth information to coordinates in a 3D coordinate system to obtain a 3D image. The 3D model is then generated by performing 3D modeling on the 3D image.
[0071] The image within the outer contour and the corresponding depth information together determine the overall shape of the 3D model in 3D space. To make the constructed 3D model more accurate, it is necessary to render the constructed 3D model. The rendering method can adopt the existing methods, so it will not be described here.
[0072] Step S306: Bind the three-dimensional model to the selected animation skeleton.
[0073] A continuous function is defined on the triangular mesh of the three-dimensional model, and the function value of each vertex on the triangular mesh is calculated using the continuous function based on the three-dimensional coordinates of each vertex. Vertices with the same function value and located on the same connected component are then grouped together as a new node quotient set. Finally, a topological structure is generated based on the new node quotient set, and each animation skeleton in the animation skeleton library is traversed to match the topological structure. Before matching the topological structure, the animation skeleton can also be rotated and scaled to better match the topological structure.
[0074] Step S308 , driving the animation skeleton to perform corresponding deformation processing on the image within the outer contour to generate a dynamic image.
[0075] The deformation processing method is similar to that in Examples 1 and 2 and will not be described again here.
[0076] In this embodiment, while driving the animation skeleton, an audio file adapted to the motion trajectory of the animation skeleton is also played.
[0077] Audio files are generated directly from user input, such as voice or character input. Character information refers to the characters input by the user. After using animation skeletons to drive the target object in a static image, the character information corresponding to the input character can be extracted from the static image and converted into a corresponding audio file. When the static image is dynamically displayed, the corresponding audio file can be played to create a more interesting display effect.
[0078] This embodiment achieves the purpose of creating an audio file for a dynamic image and can satisfy the rendering effect of a sound element during the dynamic display of a static image.
[0079] Example 4
[0080] According to an embodiment of the present application, another method for dynamic display of static images is provided, such as Figure 4 As shown, the method includes:
[0081] Step S402: Acquire a static image and identify a target object in the static image.
[0082] For example, identifying a target object in a static image, which can be a body part or face of a person or character. The static image can be a photo taken by an intelligent electronic device or an image created by graphics software or drawing tools.
[0083] Edge detection and color space feature mapping are performed on the static image to obtain outline feature data of a person image or person; and the character or avatar is obtained based on the outline feature data. In other embodiments, image processing operations such as edge detection or color space mapping can be performed on the static image to obtain outline feature data of a target person or person in the static image. In this way, the outer contour of the target object can be extracted based on the identified outline feature data.
[0084] Step S404: matching the animation skeleton based on the outer contour of the target object.
[0085] Based on topological structure analysis, the animation skeleton corresponding to the image within the outline is obtained. Based on the outer edge position data of the outer outline, the specific animation skeleton information file is matched through topological structure analysis. The image within the outer outline is bound to the corresponding animation skeleton information file.
[0086] In some embodiments, a three-dimensional model corresponding to the image within the outline can be constructed based on the outer edge position data and the image within the outline; a continuous function is defined on the triangular mesh of the three-dimensional model, and the function value of each vertex on the triangular mesh is calculated using the continuous function according to the three-dimensional coordinates of each vertex; vertices with the same function value and located on the same connected component are classified into one category and used as a new node quotient set; a topological structure is generated based on the new node quotient set, and based on the topological structure, it is matched with the animation bones in the animation bone library.
[0087] In some other embodiments, other topology analysis methods may be used. These topology analysis methods will be described in detail below and will not be repeated here.
[0088] Step S406: driving the animation skeleton based on the driving instruction.
[0089] In some embodiments, the image within the outer contour of a static image can also be converted into a vector model based on the image within the outer contour and the information file of the corresponding animation skeleton according to actual driving needs. For example, the image within the outer contour is triangulated to obtain a new vector model with animation skeleton data. A facial feature vector model is obtained from the vector model and the corresponding image within the outer contour. Based on the vector model and the image within the relevant outer contour, information such as the facial features of the person is identified, the facial feature information is triangulated, and a facial feature vector model is obtained. The facial feature vector model is driven by facial expression drive or limb drive.
[0090] In some embodiments, the driving instructions can include sound signals, facial expressions, and body movements. For example, the driving instructions can be sounds, facial expressions, and body movements input by the user. For another example, when a user uses a mobile phone to take a photo, the user's various facial expressions and facial movement characteristics are used as driving instructions for a static photo on the phone screen. In this way, the static image on the phone can be moved in a certain manner. This can be used to replace different images on a mobile phone or to decorate a specific image to replace the picture on a computer screen. In this embodiment, by introducing driving instructions, the effect of moving a static image through facial expressions or gestures can be achieved.
[0091] Step S408: performing face-changing processing on the dynamically displayed target object.
[0092] For example, if the identified target object is a person or an animal, the face of the target object can be replaced. Of course, in other embodiments, other parts of the face can also be replaced, such as legs, hands, eyes, etc.
[0093] Figure 5 is a flow chart of a method for performing face-changing processing on a dynamically displayed target object according to an embodiment of the present application. Figure 5 As shown, the method includes the following steps:
[0094] Step S4082: extract facial motion trajectory data.
[0095] After dynamically displaying a static image to generate a dynamic video, two adjacent frames within the dynamic video can be obtained: the current frame and the next frame, also referred to as the first and second frames. Feature points are obtained from the current frame through focus detection. Afterwards, the feature points are initialized, a successful feature point tracking flag is set, and the feature points are plotted. The KTL is then tracked using sparse optical flow to obtain the number of tracked feature points. The tracked moving feature points are then connected in a vector, and lost and stationary feature points are removed. Valid feature points are saved and the tracking trajectory is plotted.
[0096] like Figure 6 As shown in FIG, in two adjacent frames I and J, there is movement of pixels, that is, the position of the pixels in the current frame will change slightly in the next frame. This change is the displacement vector, which is the optical flow of the pixels.
[0097] In order to calculate the optical flow, it is necessary to determine whether the following three conditions are met between adjacent frames: constant brightness between adjacent frames, short distance movement, and spatial consistency, that is, the pixels of the same image have the same movement.
[0098] First, determine whether the brightness of two adjacent frames I and J in a video is the same within the integration window w. That is, whether I(x, y, t) = J(x', y', t + τ) within the integration window w. Only when the brightness between two adjacent frames is constant can the KLT algorithm find the pixel.
[0099] Next, we determine whether there is spatial consistency between adjacent frames. That is, for the same window, are all pixel offsets equal? On the integration window w, all (x, y) points are shifted in the same direction by (dx, dy), resulting in (x', y'). This means that the (x, y) point at time t is (x+dx, y+dy) at time t+τ. Therefore, the matching problem can be reduced to minimizing the vector of the difference function ε to find its minimum.
[0100] refer to Figure 6, calculate the displacement vector d of the pixel point, let u = [ux uy]T, represent the position of the pixel point, then the new position of the pixel point in the next frame can be expressed as v = u + d = [ux + dx uy + dy]T, where u represents the position of the pixel point in the current frame, v represents the new position of the pixel point in the next frame, ux and uy represent the horizontal and vertical coordinates of the pixel point in the current frame respectively, T represents time, dx, dy represent the displacement of the pixel point on the horizontal axis and the vertical axis in the current frame and the next frame respectively.
[0101] The displacement vector d is calculated using the vector that minimizes the difference function ε:
[0102]
[0103] Among them, I(x,y) represents the brightness of the pixel in the current frame, and J(x+dx,y+dy) represents the brightness of the pixel in the next frame. A circle with a length of w is preset around the pixel. x Width w y The neighborhood of x +1)*(2w y +1). All pixels in the current frame’s integration window are squared with all pixels in the next frame’s integration window that have been displaced. Then, the sum of the differences is calculated. When the obtained minimized difference function is the smallest, the displacement vector d can be obtained, as shown in the following example: Figure 7 shown.
[0104] By using the above method, the difference function is minimized to obtain the displacement vector, and based on the obtained displacement vector, the motion trajectory of the face in the video can be tracked.
[0105] Step S4084: Calculate the lens motion trajectory.
[0106] By analyzing the continuous frames of dynamic videos, tracking the movement of key pixels, and using the perspective principle to calculate the lens movement trajectory of the animation skeleton.
[0107] Step, S4086, performs face replacement.
[0108] Based on the lens motion trajectory, the coordinates of the pixels of the facial image in the dynamic video are fused with the world coordinates, and these pixels are replaced with three-dimensional materials. In some embodiments, after the video is fused, error detection can also be performed to remove the images that were falsely detected.
[0109] In this embodiment, the face is replaced as an example. In other embodiments, other parts such as the torso may also be replaced.
[0110] This embodiment can directly replace the face in the dynamic video through the above method. In this way, the face in the user's mobile phone selfie can be replaced with the face of the character in the cartoon animation, or the face of the cartoon character can be replaced with the user's own face, thereby enhancing the user's interest and improving the user experience.
[0111] Example 5
[0112] The following describes in detail the method of matching animation skeletons using the topological structure analysis method. In this embodiment, the topological structure of the image within the outer contour is determined by associating the image within the outer contour with the manifold boundary component and tracking the changes between the interior and the gap, so that a more suitable animation skeleton can be matched for the target object. Figure 8 As shown, the method includes the following steps:
[0113] Step S802: Create a contour map based on the image within the outer contour and pre-process the contour map.
[0114] A contour map is created for the image within the outer contour based on a preset mapping function. This mapping function reflects the general connectivity of the manifold topology, and its domain is simply connected. The shape of the contour map is completely defined by the mapping function itself. A simply connected subdomain is defined within the created contour map.
[0115] By creating, merging, or deleting level set components in the contour graph corresponding to the presence of scalar field critical points in the contour graph, each connected component of the level set of the contour graph's scalar field is contracted to a single point, forming monotonic paths that connect points to each other such that no point belongs to the contour of any component critical point. Level set components are constructed in this way. The number of level set components varies, but the genus of the level set does not.
[0116] The isosurfaces of the scalar field change genus at critical values of the scalar field, and all saddles of the contour map are encoded by analyzing the level set components, enriching the contour map with further information of all topological changes of the level set.
[0117] Step S804: Expand and discretize the pre-processed contour image.
[0118] The triangular mesh of the contour graph manifold is expanded, with f representing a region of the triangular mesh that contains the codomain of the mapping function defined on the surface. Each region is defined as either a regular region or a critical region based on the number and value of components along its boundary. Critical regions are categorized as maximum, minimum, and saddle regions and correspond to nodes of the contour graph. Arcs between nodes are then detected through the expansion process of the critical regions.
[0119] Since all points of the inverse image of a pixel in a region are equivalent in terms of expansion, all points of the inverse image of a pixel in a region can be shrunk to the same point in the quotient space, resulting in a discrete space. Connecting discrete points that share the same mapping function value yields a discrete contour map.
[0120] In the prior art, the genus of a surface with a boundary is set to the genus of a closed surface obtained by enclosing each boundary component with a disk. This effectively encloses some boundary components. This embodiment extends the contour map to surfaces with an arbitrary number of boundary components, thereby enabling the representation of the surface via a finite level set of a given mapping function, thereby generating an accurate topological structure.
[0121] Step S806: Acquire the topological structure through a multi-resolution slicing method.
[0122] The image is first extracted at the minimum required resolution, and then a multiresolution representation is created using adjacency rules in ascending order. Specifically, topology control is not performed during image extraction. Instead, the contour image is scanned with a set of parallel planes, generating a set of slices formed by the set of mesh elements bounded by two adjacent isosurfaces. Each connected component of a slice is determined by the intersection of an isosurface with a set of slice planes.
[0123] Next, a level set map is constructed. In the level set map of the triangulated surface, each contour is visualized via its centroid. To automatically select source points, a heuristic method is used to determine the slicing direction. Seed points are located based on a multi-scale curvature evaluation, for example, by using a set of intersection curves between the input surface and a set of spheres of increasing radius centered at the mesh vertices. The seed points are then sequentially connected using the wavefront traversal distance defined for simplicial complexes. The number of seed points and the selected curvature scale determine the complexity of the level set map.
[0124] Finally, the topological structure is extracted. For each pair of contours on adjacent level sets in the level set graph, a weight function is defined that depends on the average distance between the vertices of the two different contours. This weight function is used to determine the connection between two vertices. In this way, critical points are identified and classified by analyzing each vertex. Once all critical points have been detected, all vertices are processed according to the increasing value of the mapping function.
[0125] In related technologies, when analyzing topological structures through local adjustments or perturbations, artifacts that do not correspond to any shape features are introduced, leading to incorrect interpretation of the shape. However, the embodiments of the present application use the semantic features of the model and introduce discrete structures to analyze the topological structure, thereby improving the accuracy of the topological structure.
[0126] Example 6
[0127] According to an embodiment of the present application, another method for matching animation skeletons using a topological structure analysis method is provided, such as Figure 9 As shown, the method includes the following steps:
[0128] Step S902 : acquiring a static image, and extracting the outer contour of the target object in the static image from the static image.
[0129] According to the length of each of at least two edges of the static image, a plurality of target sequence values corresponding to each edge are obtained; the probability that each target sequence value in the plurality of target sequence values corresponds to the outer contour is determined; and based on the determined probability, the outer contour of the target object is extracted from the static image.
[0130] Step S904 , determining the topological structure of the image within the outer contour, and matching the image within the outer contour with animation skeletons in an animation skeleton library based on the topological structure, wherein the matched animation skeletons are used to dynamically display the image within the outer contour.
[0131] First, a contour map is created for the image within the outer contour based on a preset mapping function, and the contour map is preprocessed. For example, each connected component of the level set of the contour map's scalar field is collapsed to a single point to form a component of the level set. The components of the level set are analyzed to encode all saddles of the contour map, obtaining all topological variations of the level set. Based on this topological variation information, the contour map details are constructed to preprocess the contour map. The mapping function is used to reflect the connectivity of the manifold topology of the image within the outer contour.
[0132] Next, the pre-processed contour image is expanded and discretized, and the contour image after expansion and discretization is analyzed using a multi-resolution slicing method to obtain the topological structure of the image within the outer contour.
[0133] For example, the triangular mesh of the manifold of the preprocessed contour map is expanded to obtain the expanded contour map; all points of the inverse image of the pixels within the area of each triangular mesh are mapped to the same point in the quotient space to obtain a discrete space; and the points in the discrete space that share the same mapping function value are connected to obtain the contour map after discrete processing.
[0134] After discretization, the multi-resolution slicing method is used to scan the expanded and discretized contour map using a set of parallel planes to obtain a set of slices; a heuristic method is used to determine the respective directions of the set of slices, and multi-scale curvature evaluation is used to locate the respective seed points of the set of slices; the respective seed points are connected to construct a level set graph, and the topological structure is obtained based on the level set graph.
[0135] For example, a weight function is set for each pair of contours on adjacent level sets in the level set graph, which is the average distance between the vertices of each pair of contours. The weight function is used to determine the connection between the vertices of each pair of contours to obtain the topological structure. Finally, the obtained topological structure is matched with each animation skeleton in the animation skeleton library.
[0136] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0137] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0138] Example 7
[0139] According to an embodiment of the present application, a dynamic display device for static images is also provided, such as Figure 10 As shown, the device includes: an acquisition module 102 , a matching module 104 and a driving module 106 .
[0140] The acquisition module 102 is configured to acquire a static image, and extract the outer contour of a target object in the static image from the static image to obtain outer edge position data of the outer contour of the target object.
[0141] The matching module 104 is used to match the outer edge position data with the animation skeleton in the animation skeleton library, and bind the outer edge position data with the matched animation skeleton.
[0142] The driving module 106 is used to drive the bound animation skeleton and deform the image within the outer contour based on the motion trajectory of the bound animation skeleton to dynamically display the target object in the static image.
[0143] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments 1 to 4, and this embodiment will not be described in detail here.
[0144] Example 8
[0145] According to an embodiment of the present application, a device for matching animation skeletons based on topological structures is also provided. Figure 11 As shown, the apparatus includes a contour acquisition module 112 and a topology analysis module 114 .
[0146] The contour acquisition module 112 is configured to acquire a static image and extract the outer contour of the target object in the static image from the static image;
[0147] The topology analysis module 114 is configured to determine the topological structure of the image within the outer contour, and based on the topological structure, match the image within the outer contour with animation bones in an animation bone library, wherein the matched animation bones are used to dynamically display the image within the outer contour.
[0148] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments 5 and 6, and this embodiment will not be repeated here.
[0149] Example 9
[0150] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 12 As shown, the electronic device includes:
[0151] The electronic device includes a processor 291 and a memory 292; a communication interface 293, and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via bus 294. The communication interface 293 can be used for information transmission. The processor 291 can call logic instructions in the memory 292 to execute the methods of embodiments 1 to 6 above.
[0152] In addition, the logic instructions in the memory 292 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0153] Memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present application. Processor 291 executes the software programs, instructions, and modules stored in memory 292 to perform functional applications and data processing, thereby implementing the methods in the above-mentioned method embodiments.
[0154] Memory 292 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Memory 292 may also include high-speed random access memory and non-volatile memory.
[0155] Example 10
[0156] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium may be located in at least one network device among a plurality of network devices in the virtual network.
[0157] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.
[0158] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments 1 to 6, and this embodiment will not be described in detail here.
[0159] An embodiment of the present application further provides a computer program product, including a computer program, which is used to implement the method described in any embodiment when executed by a processor.
[0160] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0161] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0162] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0163] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0164] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0165] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0166] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for matching animation skeletons based on topological structure, characterized in that: include: Acquire a static image, and extract the outer contour of the target object in the static image from the static image; Determining a topological structure of the image within the outer contour, and matching the image within the outer contour with an animation skeleton in an animation skeleton library based on the topological structure, wherein the matched animation skeleton is used to dynamically display the image within the outer contour; Determining the topological structure of the image within the outer contour includes: creating a contour map for the image within the outer contour based on a preset mapping function, and preprocessing the contour map, wherein the mapping function is used to reflect the connectivity of the manifold topology of the image within the outer contour; expanding and discretizing the preprocessed contour map, and analyzing the expanded and discretized contour map using a multi-resolution slicing method to obtain the topological structure of the image within the outer contour; Among them, preprocessing the contour map includes: shrinking each connected component of the level set of the scalar field of the contour map to a point to constitute the components of the level set; encoding all saddles of the contour map by analyzing the components of the level set to obtain all topological changes of the level set, and constructing the details of the contour map based on the information of the topological changes to preprocess the contour map.
2. The method according to claim 1, characterized in that Expanding and discretizing the pre-processed contour map includes: Expanding the triangular mesh of the manifold of the preprocessed contour map to obtain the expanded contour map; Map all points of the inverse image of the pixels within the area of each triangular mesh to the same point in the quotient space to obtain a discrete space; Points in the discrete space that share the same mapping function value are connected to obtain the contour map after discretization processing.
3. The method according to claim 1, characterized in that The contour image after expansion and discretization is analyzed using a multi-resolution slicing method to obtain a topological structure of the image within the outer contour, including: Using the multi-resolution slicing method, a set of parallel planes are used to scan the expanded and discretized contour image to obtain a set of slices; determining respective orientations of the set of slices using a heuristic method, and locating respective seed points of the set of slices using a multi-scale curvature evaluation; The seed points are connected to construct a level set graph, and the topological structure is obtained based on the level set graph.
4. The method according to claim 3, characterized in that Obtaining the topological structure based on the level set graph includes: Setting a weight function of the average distance between vertices of each pair of contours on adjacent level sets in the level set graph; The weight function is used to determine the connection between the vertices of each pair of contours to obtain the topological structure.
5. The method according to claim 1, wherein Extracting the outer contour of the target object in the static image from the static image includes: obtaining, according to a length of each of at least two edges of the static image, a plurality of target sequence values corresponding to each edge; determining a probability that each target sequence value in the plurality of target sequence values corresponds to the outer contour; Based on the determined probability, an outer contour of the target object is extracted from the static image.
6. The method according to claim 1, wherein Matching the image within the outer contour with the animation skeleton in the animation skeleton library, including: constructing a three-dimensional model corresponding to the image within the outline based on the image within the outline; defining a continuous function on the triangular mesh of the three-dimensional model, and calculating the function value of each vertex on the triangular mesh using the continuous function according to the three-dimensional coordinates of each vertex; The vertices with the same function value and located on the same connected component are classified into one category and used as a new node quotient set; A topological structure is generated based on the new node quotient set, and is matched with animation bones in an animation bone library based on the topological structure.
7. A device for matching animation skeletons based on topological structure, characterized in that: include: a contour acquisition module, configured to acquire a static image and extract an outer contour of a target object in the static image from the static image; a topology analysis module configured to determine a topological structure of the image within the outer contour, and based on the topological structure, match the image within the outer contour with an animation skeleton in an animation skeleton library, wherein the matched animation skeleton is used to dynamically display the image within the outer contour; The topology analysis module is further configured to: establish a contour map for the image within the outer contour based on a preset mapping function, and preprocess the contour map, wherein the mapping function is used to reflect the connectivity of the manifold topology of the image within the outer contour; expand and discretize the preprocessed contour map, and analyze the expanded and discretized contour map using a multi-resolution slicing method to obtain the topological structure of the image within the outer contour; In which, the topology analysis module is further configured to: shrink each connected component of the level set of the scalar field of the contour map to a point to constitute the components of the level set; encode all saddles of the contour map by analyzing the components of the level set to obtain all topological changes of the level set, and construct the details of the contour map based on the information of the topological changes to preprocess the contour map.
8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for converting static image into dynamic image
CN106791032A
Animation migration method and device, equipment and storage medium
CN113313794A