Image processing method and device, electronic equipment and computer readable storage medium

CN116029948BActive Publication Date: 2026-08-07FACE CUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FACE CUTE CO LTD
Filing Date
2021-10-25
Publication Date
2026-08-07

Smart Images

  • Figure CN116029948B_ABST
    Figure CN116029948B_ABST
Patent Text Reader

Abstract

An image processing method, an image processing device, an electronic device and a computer readable storage medium. The image processing method comprises: obtaining at least one first virtual sub-image; processing a selected virtual sub-image in the at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image; in response to detecting a detection object, obtaining current feature information of the detection object, the current feature information being used to indicate a current state of the detection object, the detection object comprising a plurality of target features; based on the current feature information, determining movement information of a plurality of feature points in an initial virtual image, the initial virtual image being obtained by superimposing the first virtual sub-image and the second virtual sub-image; and driving the plurality of feature points to move in the initial virtual image according to the movement information to generate a current virtual image corresponding to the current state. The method can reduce the design difficulty and driving difficulty of the virtual image, and improve the design efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to an image processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With the rapid development of the internet, virtual avatars are widely used in emerging fields such as live streaming, short videos, and games. The application of virtual avatars not only enhances the fun of human-computer interaction but also brings convenience to users. For example, on live streaming platforms, broadcasters can use virtual avatars to broadcast without showing their faces. Summary of the Invention

[0003] At least one embodiment of this disclosure provides an image processing method, comprising: acquiring at least one first virtual sub-image, each of the at least one first virtual sub-image corresponding to one of a plurality of target features; processing a selected virtual sub-image among the at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image; in response to detecting a detection object, acquiring current feature information of the detection object, the current feature information indicating the current state of the detection object, the detection object including a plurality of target features; determining movement information of a plurality of feature points in an initial virtual image based on the current feature information, the initial virtual image being obtained by superimposing the at least one first virtual sub-image and the at least one second virtual sub-image; and driving the plurality of feature points to move in the initial virtual image according to the movement information to generate a current virtual image corresponding to the current state.

[0004] For example, in one embodiment of the image processing method provided in this disclosure, the method further includes: obtaining a fill image; and superimposing at least one first virtual sub-image, at least one second virtual sub-image, and the fill image to generate an initial virtual image.

[0005] For example, in one embodiment of the image processing method provided in this disclosure, the method further includes: obtaining depth information of each first virtual sub-image and each second virtual sub-image in an initial virtual image; and superimposing at least one first virtual sub-image and at least one second virtual sub-image according to the depth information of each first virtual sub-image and each second virtual sub-image to generate an initial virtual image.

[0006] For example, in an image processing method provided in an embodiment of this disclosure, processing a selected virtual sub-image in at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image includes: performing morphological processing on the selected virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image.

[0007] For example, in an image processing method provided in one embodiment of this disclosure, the selected virtual sub-image includes the sclera image of the detected object, and the method further includes:

[0008] Along the central axis of the sclera image, the sclera image is divided into a first sclera sub-image and a second sclera sub-image. The direction of the central axis is parallel to the length direction of the eye of the detected object. The first sclera sub-image is located on the side of the central axis away from the mouth of the detected object, and the second sclera sub-image is located on the side of the central axis closer to the mouth. Morphological processing is performed on the selected virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image. This includes: performing morphological processing on the first sclera sub-image and the second sclera sub-image respectively to obtain the first sub-image and the second sub-image; and filling the first sub-image according to the color and texture of the detected object's face to obtain the upper eyelid image of the detected object, and filling the second sub-image to obtain the lower eyelid image of the detected object.

[0009] For example, in an image processing method provided in an embodiment of this disclosure, processing a selected virtual sub-image in at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image includes: splitting the selected virtual sub-image to obtain a plurality of second virtual sub-images.

[0010] For example, in an image processing method provided in an embodiment of this disclosure, the selected virtual sub-image includes the mouth image of the detected object, and the selected virtual sub-image is split to obtain multiple second virtual sub-images, including: splitting the mouth image into two second virtual sub-images from the widest position of the mouth of the detected object, one of the two second virtual sub-images being an upper lip image and the other being a lower lip image.

[0011] For example, in an embodiment of the image processing method provided in this disclosure, the method further includes: calculating the maximum deformation value of multiple feature points; and determining the movement information of multiple feature points in the initial virtual image based on the current feature information, including: determining the movement information of multiple feature points in the initial virtual image based on the current feature information and the maximum deformation value.

[0012] For example, in an image processing method provided in an embodiment of this disclosure, calculating the maximum deformation value of multiple feature points includes: determining the maximum deformation curve that a first part of the feature points conforms to; determining the position coordinates of each feature point in the first part of the feature points; and substituting the position coordinates of each feature point in the first part of the feature points into the maximum deformation curve to obtain the maximum deformation value of the first part of the feature points.

[0013] For example, in an image processing method provided in an embodiment of this disclosure, calculating the maximum deformation value of multiple feature points includes: calculating the difference between the extreme position coordinates and the reference position coordinates of a second part of the feature points, where the extreme position coordinates are the coordinates of the second part of the feature points in the initial virtual image; and using the difference as the maximum deformation value of each of the second part of the feature points.

[0014] For example, in an image processing method provided in one embodiment of this disclosure, determining the movement information of multiple feature points in an initial virtual image based on current feature information and maximum deformation value includes: determining the current state value of the current state relative to the reference state based on the current feature information; and calculating the movement information of multiple feature points in the initial virtual image based on the current state value and the maximum deformation value.

[0015] For example, in an image processing method provided in an embodiment of this disclosure, calculating the current state value and the maximum deformation value to determine the movement information of multiple feature points in the initial virtual image includes: performing a multiplication operation on the current state value and the maximum deformation value to obtain the movement information of multiple feature points in the initial virtual image.

[0016] For example, in an image processing method provided in one embodiment of this disclosure, determining the current state value of the current state relative to a reference state based on current feature information includes: obtaining a mapping relationship between feature information and state value; and determining the current state value of the current state relative to a reference state based on the mapping relationship and current feature information.

[0017] For example, in an image processing method provided in one embodiment of this disclosure, obtaining the mapping relationship between feature information and state value includes: obtaining multiple samples, each sample including the correspondence between sample feature information and sample state value; and constructing a mapping function based on the correspondence included in each of the multiple samples, the mapping function representing the mapping relationship between feature information and state value.

[0018] For example, in an image processing method provided in one embodiment of this disclosure, the sample feature information includes target feature information and second feature information, and the sample state value includes a first value corresponding to the target feature information and a second value corresponding to the second feature information. Based on the correspondence, a mapping function is constructed, including: constructing a system of linear equations; and solving the system of linear equations according to the target feature information, the first value, the second feature information, and the second value to obtain the mapping function.

[0019] For example, in an image processing method provided in an embodiment of this disclosure, the second virtual sub-image includes an upper lip image and a lower lip image, the first part of feature points includes feature points in the upper lip image and feature points in the lower lip image, and determining the maximum deformation curve that the first part of feature points conforms to among a plurality of feature points includes: respectively determining the first maximum deformation curve that the feature points in the upper lip image conform to and the second maximum deformation curve that the feature points in the lower lip image conform to.

[0020] For example, in an image processing method provided in one embodiment of this disclosure, the feature points in the upper lip image and the feature points in the lower lip image correspond one-to-one. The feature points in the upper lip image and the feature points in the lower lip image include n columns, and the first maximum deformation curve is y1 = (x′-n). 2 The second largest deformation curve is y2 = c - (x′ - n). 2 x' represents the x'-axis coordinates of the feature points in the upper lip image and the lower lip image, with the widest point of the mouth being the x'-axis. y1 represents the distance the feature points in the upper lip image move away from the lower lip, and y2 represents the distance the feature points in the lower lip image move away from the upper lip.

[0021] For example, in an image processing method provided in one embodiment of this disclosure, the selected sub-image includes the sclera image of the detection object, the second virtual sub-image includes the upper eyelid image and the lower eyelid image, the reference position coordinates are the coordinates of the reference point corresponding to the second part of the feature points on the central axis of the sclera image, the direction of the central axis is parallel to the length direction of the eye of the detection object, the second part of the feature points includes feature points of the upper eyelid image and feature points of the lower eyelid image, and the calculation of the difference between the extreme position coordinates of the second part of the feature points and the reference position coordinates among multiple feature points includes: calculating the difference between the extreme position coordinates of the feature points in the upper eyelid image and the reference point corresponding to the feature points in the upper eyelid image on the central axis; and calculating the difference between the extreme position coordinates of the feature points in the lower eyelid image and the reference point corresponding to the feature points in the lower eyelid image on the central axis.

[0022] For example, in an image processing method provided in one embodiment of this disclosure, the motion information includes a motion distance, and according to the motion information, multiple feature points are driven to move in an initial virtual image, including: driving multiple feature points to move a motion distance toward a preset reference position.

[0023] For example, in an image processing method provided in an embodiment of this disclosure, a filling image is used to fill gaps in an overlay sub-image. The overlay sub-image is obtained by superimposing two second virtual sub-images. The gap refers to the position between the two second virtual sub-images in the overlay sub-image. According to the movement information, multiple feature points are driven to move in the initial virtual image, including: according to the movement information, multiple feature points are driven to move in the initial virtual image, and the shape of the filling image is changed according to the size of the gap, so that the filling image is adapted to the gap to generate the current virtual image.

[0024] For example, in an image processing method provided in one embodiment of this disclosure, the filling image includes an oral cavity image, and the two second virtual sub-images include an upper lip image and a lower lip image.

[0025] At least one embodiment of this disclosure provides an image processing apparatus, comprising: an acquisition unit configured to acquire at least one first virtual sub-image, each of the at least one first virtual sub-image corresponding to one of a plurality of target features; a processing unit configured to process a selected virtual sub-image among the at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image; a detection unit configured to acquire current feature information of a detected object in response to detecting a detected object, the current feature information indicating the current state of the detected object, the detected object including a plurality of target features; a determination unit configured to determine movement information of a plurality of feature points in an initial virtual image based on the current feature information, the initial virtual image being obtained by superimposing at least one first virtual sub-image and at least one second virtual sub-image; and a driving unit configured to drive the plurality of feature points to move in the initial virtual image according to the movement information to generate a current virtual image corresponding to the current state.

[0026] At least one embodiment of this disclosure provides an electronic device including a processor; a memory including one or more computer program modules; the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the image processing method provided in any embodiment of this disclosure.

[0027] At least one embodiment of this disclosure provides a computer-readable storage medium for storing non-transitory computer-readable instructions that, when executed by a computer, can implement the image processing method provided in any embodiment of this disclosure. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0029] Figure 1A A flowchart of an image processing method provided by at least one embodiment of the present disclosure is shown;

[0030] Figure 1B The illustration shows a schematic diagram of a plurality of first virtual sub-images provided in some embodiments of the present disclosure;

[0031] Figure 1C The illustrations illustrate some embodiments of the present disclosure that provide support for... Figure 1B A schematic diagram of the second virtual sub-image obtained by processing the image 110 of the middle mouth area;

[0032] Figure 1D The illustration shows a schematic diagram of processing a selected virtual sub-image to obtain a second virtual sub-image, according to some embodiments of the present disclosure;

[0033] Figure 1E A schematic diagram illustrating the process of obtaining a complete mouth image using a filled image, according to at least one embodiment of the present disclosure, is shown.

[0034] Figure 2A A flowchart is shown, illustrating a method for calculating the maximum deformation value of multiple feature points according to at least one embodiment of this disclosure;

[0035] Figure 2B A schematic diagram of the maximum deformation curve provided in at least one embodiment of the present disclosure is shown;

[0036] Figure 2C A flowchart is shown showing another method for calculating the maximum deformation value of multiple feature points provided by at least one embodiment of the present disclosure;

[0037] Figure 2D This illustration shows a schematic diagram of calculating the difference between the extreme position coordinates and the reference position coordinates of a feature point in the upper eyelid, according to at least one embodiment of the present disclosure.

[0038] Figure 3 A flowchart is shown showing a method for determining the movement information of multiple feature points in an initial virtual image based on current feature information and maximum deformation value, according to at least one embodiment of the present disclosure.

[0039] Figure 4 A schematic diagram illustrating the effect of an image processing method provided in at least one embodiment of this disclosure is shown.

[0040] Figure 5 A schematic block diagram of an image processing apparatus provided in at least one embodiment of the present disclosure is shown;

[0041] Figure 6 A schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure is shown;

[0042] Figure 7 A schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure is shown; and

[0043] Figure 8 A schematic diagram of a computer-readable storage medium provided in at least one embodiment of the present disclosure is shown. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0045] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0046] Virtual avatars can be driven in real-time based on the actions and expressions of objects detected by electronic devices (e.g., users). The design and driving of virtual avatars are both highly complex. Creating a 3D virtual avatar requires character art design, model creation, and animation rigging. Further driving and presentation involve motion capture technology, holographic hardware technology, augmented reality (AR) technology, virtual reality (VR) technology, and driver development, resulting in a long production cycle, significant implementation and driving difficulties, and high costs. Creating a 2D virtual avatar requires professional concept art design according to the requirements of different concept art design platforms. During driving, each frame of animation is drawn to form actions and expressions, and material transformations are also required on specific rendering software such as Live2D, making implementation and driving also highly complex.

[0047] At least one embodiment of this disclosure provides an image processing method, an image processing apparatus, an electronic device, and a computer-readable storage medium. The image processing method includes: acquiring at least one first virtual sub-image, each of the at least one first virtual sub-image corresponding to one of a plurality of target features; processing a selected virtual sub-image from the at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image; in response to detecting a target object, acquiring current feature information of the target object, the current feature information indicating the current state of the target object, the target object including the plurality of target features; determining movement information of a plurality of feature points in an initial virtual image based on the current feature information; and driving the plurality of feature points to move within the initial virtual image according to the movement information to generate a current virtual image corresponding to the current state, the initial virtual image being obtained by superimposing the at least one first virtual sub-image and the at least one second virtual sub-image. This image processing method can reduce the design difficulty of virtual avatars, improve design efficiency, and reduce the driving difficulty of virtual avatars, making virtual avatars easier to implement and drive.

[0048] Figure 1A A flowchart of an image processing method provided by at least one embodiment of the present disclosure is shown.

[0049] like Figure 1A As shown, the method may include steps S10 to S50.

[0050] Step S10: Obtain at least one first virtual sub-image, each of which corresponds to one of a plurality of target features.

[0051] Step S20: Process the selected virtual sub-image in at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image.

[0052] Step S30: In response to the detection of a target object, obtain the current feature information of the target object. The current feature information is used to indicate the current state of the target object. The target object includes multiple target features.

[0053] Step S40: Based on the current feature information, determine the movement information of multiple feature points in the initial virtual image. The initial virtual image is obtained by superimposing at least one first virtual sub-image and at least one second virtual sub-image.

[0054] Step S50: Based on the movement information, drive multiple feature points to move in the initial virtual image to generate the current virtual image corresponding to the current state.

[0055] This embodiment can process the first virtual sub-image to obtain the second virtual sub-image. Therefore, it is only necessary to draw some of the virtual sub-images required for the initial virtual image, instead of drawing all the virtual sub-images required to generate the initial virtual image, thereby improving the efficiency of designing virtual images and reducing the design difficulty.

[0056] In the embodiments of this disclosure, by acquiring the current state of the detected object in real time and driving the movement of feature points in the initial virtual image according to the current feature information of the current state, the generation and driving of the virtual image can be realized. Therefore, the embodiments of this disclosure do not require the use of animation skeleton binding technology, holographic hardware technology or specific drawing software (such as Live2D) to realize the generation and driving of the virtual image, which reduces the difficulty of realizing and driving the virtual image.

[0057] For step S10, for example, the first virtual sub-image is drawn in advance by the designer using graphic design tools such as Photoshop and stored in the storage unit. For example, according to the layer names preset by the designer, the corresponding images are drawn in the layers, so that the layers can be retrieved and driven according to the layer names. For example, at least one first virtual sub-image is obtained by reading the PSD file of the Photoshop graphic design tool from the storage unit.

[0058] For example, the first virtual sub-image is an image of the target features. For example, the target features are facial features such as upper eyelashes, lower eyelashes, sclera of the eyes, and pupils. Accordingly, the first virtual sub-image can be an image of the upper eyelashes, lower eyelashes, sclera of the eyes, mouth, and pupils, etc.

[0059] In some embodiments of this disclosure, the first virtual sub-image is, for example, an image of a target feature drawn by a designer when it is in an extreme state. For example, the upper eyelash image is an image of the upper eyelashes when the eyes are open to their maximum extent. For example, the mouth image is an image of the mouth when it is open to its maximum extent.

[0060] Figure 1B The illustration shows a schematic diagram of a plurality of first virtual sub-images provided in some embodiments of the present disclosure.

[0061] like Figure 1B As shown, the multiple first virtual sub-images include a mouth image 110, an upper eyelash image 120, a pupil image 130, and a sclera image 140.

[0062] It is necessary to understand that Figure 1B Only a portion of the first virtual sub-images is shown, not all of them. For example, multiple first virtual sub-images also include images of hair, nose, etc. Designers can draw the required first virtual sub-images according to their actual needs.

[0063] For step S20, in some embodiments of this disclosure, the second virtual sub-image associated with the selected virtual sub-image may be a local image of the target features corresponding to the selected virtual sub-image. In this embodiment, step S20 may include splitting the selected virtual sub-image to obtain multiple second virtual sub-images.

[0064] For example, the virtual sub-image is selected as image 110 of the mouth, and the second virtual sub-image associated with image 110 of the mouth can be an image of the upper lip and an image of the lower lip. In this embodiment, step S20 may include splitting the mouth image into two second virtual sub-images from the widest position of the mouth of the detected object, one of the two second virtual sub-images being an image of the upper lip and the other being an image of the lower lip.

[0065] Figure 1C The illustrations illustrate some embodiments of the present disclosure that provide support for... Figure 1B A schematic diagram of the second virtual sub-image obtained by processing the image 110 of the middle mouth area.

[0066] like Figure 1C As shown, the mouth image 110 is divided into a second virtual sub-image 111 and a second virtual sub-image 112 along the widest part of the mouth, that is, along the position of the dashed line AA'.

[0067] The second virtual sub-image 111 is the upper lip image, and the second virtual sub-image 112 is the lower lip image.

[0068] In some embodiments of this disclosure, the second virtual sub-image associated with the selected virtual sub-image may be an image of other target features located around the target feature corresponding to the selected virtual sub-image.

[0069] In some embodiments of this disclosure, step S20 may include performing morphological processing on the selected virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image.

[0070] For example, if the virtual sub-image is selected as the image 140 of the sclera, and the upper and lower eyelids are target features located around the sclera, the image 140 of the sclera is processed to obtain the images of the upper and lower eyelids.

[0071] Figure 1D The illustration shows a schematic diagram of processing a selected virtual sub-image to obtain a second virtual sub-image, according to some embodiments of the present disclosure.

[0072] exist Figure 1D In the illustrated embodiment, the selected virtual sub-image includes an image of the sclera of the detected object. In this embodiment, the image processing method further includes: splitting the sclera image into a first sclera sub-image and a second sclera sub-image along the central axis of the sclera image, wherein the direction of the central axis is parallel to the length direction of the detected object's eye, the first sclera sub-image is located on the side of the central axis away from the detected object's mouth, and the second sclera sub-image is located on the side of the central axis closer to the mouth.

[0073] like Figure 1D As shown, the sclera image is split into a first sclera sub-image 141 and a second sclera sub-image 142 along the central axis BB' of the alpha channel of the sclera image.

[0074] In step S20, morphological processing is performed on the first eye white sub-image 141 and the second eye white sub-image 142 to obtain the first sub-image and the second sub-image, and the first sub-image is filled according to the color and texture of the detected object's face to obtain the upper eyelid image 143 of the detected object, and the second sub-image is filled to obtain the lower eyelid image 144 of the detected object.

[0075] For example, the first eye white sub-image 141 and the second eye white sub-image 142 are processed in reverse, and small holes are filtered out (e.g., opening, closing operations, or hole filling in morphology) to obtain the first sub-image and the second sub-image. The first sub-image and the second sub-image are the upper eyelid layer and the lower eyelid layer, respectively. Then, the texture and color of the upper and lower eyelids are extracted from the corresponding positions of the face layer, and the upper eyelid layer and the lower eyelid layer are filled respectively to obtain the upper eyelid image 143 and the lower eyelid image 144.

[0076] For step S30, for example, the detected object can be obtained through a detection device (camera, infrared device, etc.). For example, the detected object can be an object to be virtualized, such as a living being like a person or a pet. For example, the detected object can be a user who needs to be driven by a virtual avatar, such as a streamer who has selected the virtualization function provided by the live streaming platform.

[0077] In some embodiments of this disclosure, the detection object includes a feature to be virtualized, which includes a target feature. For example, the feature to be virtualized is a feature in the detection object that needs to be virtualized, and the target feature is a key part of the detection object. The key part is the part that needs to be driven in the virtual image obtained after the feature to be virtualized. For example, the feature to be virtualized of the detection object may include, but is not limited to, cheeks, shoulders, hair, and various facial features. For example, among the features to be virtualized, cheeks and shoulders are parts that do not need to be driven in the virtual image, while eyebrows, upper eyelashes, lower eyelashes, mouth, upper eyelids, and lower eyelids need to be driven in the virtual image. Therefore, the target feature may include, but is not limited to, eyebrows, upper eyelashes, lower eyelashes, mouth, upper eyelids, and lower eyelids.

[0078] For example, in response to the camera detecting a target object, the current feature information of the target object is acquired. For example, multiple facial landmarks are used to locate and detect the target object, thereby acquiring the current feature information of the target object. For example, a 106-facial landmark detection algorithm or a 280-facial landmark detection algorithm can be used, or other applicable algorithms can be used; the embodiments disclosed herein are not limited in this regard.

[0079] In some embodiments of this disclosure, the current feature information includes, for example, facial movement information and body posture information of the detected object. This current feature information is used to indicate the current state of the detected object. For example, the current feature information can indicate the degree of eye opening, mouth opening, etc.

[0080] In some embodiments of this disclosure, the current feature information includes comparison information between the target features of the current state and the reference features of the detected object, and the comparison information remains unchanged when the distance of the detected object relative to the image acquisition device used to detect the detected object changes.

[0081] For example, when the distance between the detected object and the camera varies, as long as the object's actions or expressions remain unchanged, these current feature information will not change. This can alleviate the shaking of the virtual image caused by changes in the distance between the detected object and the camera.

[0082] For example, reference features can be faces, eyes, etc. For example, the current feature information can include the ratio h1 / h0 of the height h1 from the eyebrows to the eyes to the height h0 of the face, the ratio h2 / h0 of the height h2 of the upper and lower eyelids to the height h0 of the face, the ratio h3 / k0 of the distance h3 from the pupil to the outer corner of the eye to the width k0 of the eye, and the ratio s1 / s0 of the area s1 of the mouth to the area s0 of the face. That is, the height of the eyebrows is represented by the ratio h1 / h0, the degree of eye opening is represented by the ratio h2 / h0, the position of the pupil is represented by the ratio h3 / k0, and the degree of mouth opening is represented by the ratio s1 / s0.

[0083] For step S40, the initial virtual image may be generated in response to the detection of a detected object by superimposing at least one first virtual sub-image and at least one second virtual sub-image. For example, each first virtual sub-image and each second virtual sub-image serve as layers of the initial virtual image, and at least one first virtual sub-image and at least one second virtual sub-image are superimposed to obtain the initial virtual image. Multiple feature points in the initial virtual image are feature points in each virtual sub-image (including the first virtual sub-image and the second virtual sub-image).

[0084] like Figure 1A As shown, in addition to steps S10 to S50, the image processing method may also include steps S60 and S70. For example, steps S60 and S70 may be performed between steps S20 and S30.

[0085] Step S60: Obtain the filled image.

[0086] Step S70: Overlay at least one first virtual sub-image, at least one second virtual sub-image, and a fill image to generate an initial virtual image.

[0087] This embodiment uses a filling image to fill the first virtual sub-image and the second virtual sub-image, making the initial virtual image generated by the first virtual sub-image and the second virtual sub-image more vivid, and thus the current virtual image is also more vivid and lifelike.

[0088] For step S60, for example, the image to be filled is an oral cavity image, and the oral cavity image, the upper lip image 111, and the lower lip image 112 are superimposed to obtain a complete mouth image of the initial virtual image.

[0089] For step S70, for example, the gap between the first virtual sub-image and the second virtual sub-image is filled with a filling image to generate an initial virtual image.

[0090] In other embodiments of this disclosure, step S70 may be, in response to detecting a detected object, superimposing at least one first virtual sub-image, at least one second virtual sub-image, and a fill image to generate an initial virtual image.

[0091] Figure 1E A schematic diagram is shown illustrating the process of obtaining a complete mouth image using a filled image, according to at least one embodiment of the present disclosure.

[0092] like Figure 1E As shown, the filling image is the oral cavity image 113. The oral cavity image 113 is superimposed with the upper lip image 111 and the lower lip image 112 to obtain the complete mouth image 150 of the initial virtual image.

[0093] In some embodiments of this disclosure, the image processing method may further include, in addition to the foregoing steps, obtaining the depth information of each first virtual sub-image and each second virtual sub-image in the initial virtual image, and superimposing at least one first virtual sub-image and at least one second virtual sub-image according to the depth information of each first virtual sub-image and each second virtual sub-image to generate the initial virtual image.

[0094] This embodiment sets depth information for each virtual sub-image (including the first virtual sub-image and the second virtual sub-image), so that the initial virtual image obtained by superimposing multiple virtual sub-images has a three-dimensional effect, thereby driving the movement of feature points in the initial virtual image to obtain a virtual image that also has a three-dimensional effect.

[0095] In some embodiments of this disclosure, the depth information of each first virtual sub-image and each second virtual sub-image can be preset by the designer and stored in a storage unit. For example, the depth information is read from the storage unit, and in response to the detection of a detection object, at least one first virtual sub-image and at least one second virtual sub-image are superimposed according to the depth information of each first virtual sub-image and each second virtual sub-image to generate an initial virtual image.

[0096] For example, the depth information of each virtual sub-image includes the depth value of each virtual sub-image and the depth value of the feature points in each virtual sub-image.

[0097] In some embodiments of this disclosure, each virtual sub-image corresponds to one of a plurality of virtual features to be virtualized of a detection object. These virtual features include target features. In a direction perpendicular to the face of the detection object, the depth value of each virtual sub-image is proportional to a first distance, and the depth value of a feature point in each virtual sub-image is also proportional to the first distance, which is the distance between the virtual feature corresponding to the virtual sub-image and the eyes of the detection object. That is, in a direction perpendicular to the face of the detection object, the depth value of each virtual sub-image and the depth value of a feature point in each virtual sub-image are proportional to the distance between the virtual sub-image and the eyes of the detection object.

[0098] It should be understood that the distance in this disclosure can be a vector, meaning the distance can be either positive or negative.

[0099] For example, multiple virtual sub-images include a nose virtual sub-image corresponding to the nose of the detected object and a shoulder virtual sub-image corresponding to the shoulder of the detected object. In the direction perpendicular to the face of the detected object, the distance between the virtual feature (nose) and the eyes of the detected object is -f1, and the distance between the virtual feature (shoulder) and the eyes of the detected object is f2, where both f1 and f2 are greater than 0. Therefore, the depth value of the nose virtual sub-image is less than the depth value of the eye virtual sub-image, and the depth value of the eye virtual sub-image is less than the depth value of the shoulder virtual sub-image. This embodiment makes the nose of the virtual image appear in front of the eyes and the shoulders appear behind the eyes.

[0100] For example, for each virtual sub-image, the virtual sub-image is divided into multiple bounding boxes. The bounding box coordinates of each bounding box are extracted, and depth values ​​are set for the top-left and bottom-right vertices of each bounding box based on a first distance. Each vertex of the bounding box can represent a feature point. The depth value of a vertex of the bounding box is proportional to the distance between that vertex and the eye of the detected object.

[0101] For example, for the target feature eyebrow, the virtual sub-image of the eyebrow is divided into three equal rectangles. The top-left and bottom-right corner vertices of each of the three rectangles are extracted and their bounding box coordinates (X, Y) in the image coordinate system are obtained. X is the horizontal coordinate and Y is the vertical coordinate. A depth value Z is set for each vertex to obtain the coordinates (X, Y, Z) of each vertex. The W coordinate and four texture coordinates (s, t, r, q) are also added. For example, each vertex has 8 dimensions. The number of vertex groups is 2 rows and 4 columns (2, 4), totaling 8. The first row is (X, Y, Z, W), and the second row is (s, t, r, q). The image coordinate system is, for example, with the width direction of the virtual sub-image as the X-axis, the height direction of the virtual sub-image as the Y-axis, and the bottom-left corner of the virtual sub-image as the origin.

[0102] Compared to the virtual sub-image of the eyebrows, the virtual sub-image of the posterior hair (the part of the hair furthest from the eyes of the detected object) is divided into more rectangles. In the virtual sub-image of the posterior hair, the depth value (i.e., z-coordinate) of the middle region is smaller than the depth values ​​(i.e., z-coordinate) of the sides, forming a curved shape of the posterior hair.

[0103] In some embodiments of this disclosure, the image processing method may include, in addition to steps S10 to S50, calculating the maximum deformation value of multiple feature points, thereby step S40 includes determining the movement information of multiple feature points in the initial virtual image based on the current feature information and the maximum deformation value. For example, the calculation of the maximum deformation value of multiple feature points may be performed before step S40.

[0104] The following is combined with Figure 2A , 2B 2C, 2D and Figure 3 At least one embodiment of step S40 is described.

[0105] Figure 2A A flowchart is shown of a method for calculating the maximum deformation value of multiple feature points according to at least one embodiment of the present disclosure.

[0106] like Figure 2A As shown, the method may include steps S210 to S230.

[0107] Step S210: Determine the maximum deformation curve that the first part of the feature points in the multiple feature points conform to.

[0108] Step S220: Determine the position coordinates of each feature point in the first part of the feature points.

[0109] Step S230: Substitute the position coordinates of each feature point in the first part of the feature points into the maximum deformation curve to obtain the maximum deformation value of the first part of the feature points.

[0110] This embodiment obtains the maximum deformation value through the maximum deformation curve, which avoids missing some feature points located at the edges of the virtual sub-image, thus making the current virtual image more complete and realistic. For example, if feature points at the edges (e.g., the corners of the mouth) of the upper and lower lip images are missed, causing the feature points at the edges (e.g., the corners of the mouth) to be unable to be driven, then when the mouth of the detected object is closed, the corners of the mouth in the current virtual image will not be closed, resulting in an incomplete and unrealistic virtual image.

[0111] For step S210, in some embodiments of this disclosure, the first set of feature points are a subset of multiple feature points. These first set of feature points are multiple feature points that can be fitted into a curve when the target feature is at its maximum deformation. When the target feature is at its maximum deformation, the curve fitted by the first set of feature points is the maximum deformation curve.

[0112] For example, the second virtual sub-image includes an upper lip image and a lower lip image. Feature points in the upper lip image can be fitted into a curve when the mouth is fully open, and feature points in the lower lip image can also be fitted into a curve when the mouth is fully open. Therefore, the first set of feature points can include feature points in the upper lip image and feature points in the lower lip image. In this embodiment, step S210 includes: determining the first maximum deformation curve that the feature points in the upper lip image conform to and the second maximum deformation curve that the feature points in the lower lip image conform to, respectively.

[0113] Figure 2B A schematic diagram of the maximum deformation curve provided in at least one embodiment of the present disclosure is shown.

[0114] like Figure 2B As shown, the scenario includes an upper lip image 111 and a lower lip image 112. The upper lip image 111 and lower lip image 112 are images corresponding to the mouth when it is opened to its maximum extent. The upper lip image 111 includes multiple feature points, and the lower lip image 112 includes multiple feature points.

[0115] The feature points in the upper lip image correspond one-to-one with the feature points in the lower lip image, and the feature points in both images comprise n columns. The first maximum deformation curve is:

[0116] y1=(x′-n) 2 ,

[0117] The second maximum deformation curve is:

[0118] y2 = c - (x' - n) 2 ,

[0119] x' represents the x'-axis coordinates of feature points in the upper lip image and the lower lip image. Figure 2B In the mouth image shown, the widest position of the mouth is taken as the x' axis. In the above formula, y1 is the distance that the feature point in the upper lip image moves away from the lower lip, and y2 is the distance that the feature point in the lower lip image moves away from the upper lip.

[0120] In some embodiments of this disclosure, n is an odd number.

[0121] In some embodiments of this disclosure, the maximum deformation curve conforming to the first part of the feature points can be obtained by the designer through pre-fitting based on multiple samples.

[0122] For step S220, the position coordinates of multiple feature points in the upper lip image and the position coordinates of multiple feature points in the lower lip image are determined. For example, the X coordinates of multiple feature points in the upper lip image are extracted from the coordinates (X, Y, Z, W) described above, and the X coordinates are transformed into... Figure 2B The x-coordinate in the mouth image shown.

[0123] For step S230, for example, the x' axis coordinates of multiple feature points on the upper lip include x' 上1 ~x' 上n , will x' 上1 ~x' 上n Substitute these values ​​into y1=(x′-n) respectively. 2 The maximum deformation value y1 of each feature point on the upper lip is obtained.

[0124] For example, the x-axis coordinates of feature points on the lower lip include x' 下1 ~x' 下n , will x' 下1 ~x' 下n Substituting these values ​​into y² = c - (x′ - n) 2 The maximum deformation value y2 of each feature point on the lower lip is obtained.

[0125] Figure 2C A flowchart is shown of another method for calculating the maximum deformation value of multiple feature points provided by at least one embodiment of the present disclosure.

[0126] like Figure 2C As shown, the method may include steps S240 to S250.

[0127] Step S240: Calculate the difference between the extreme position coordinates and the reference position coordinates of the second part of the feature points among multiple feature points. The extreme position coordinates are the coordinates of the second part of the feature points in the initial virtual image.

[0128] Step S250: Use the difference as the maximum deformation value of each feature point in the second part.

[0129] This embodiment directly determines the maximum deformation value based on the extreme position coordinates and the reference position coordinates, which is simple to calculate and easy to implement.

[0130] For step S240, the second part of the feature points are, for example, other feature points besides the first part of the feature points among multiple feature points. For example, the feature points in the upper eyelash image, upper eyelid image, and lower eyelid image are the second part of the feature points. For the feature points in the upper eyelash image, upper eyelid image, and lower eyelid image, the maximum deformation value can be determined according to steps S240 to S250.

[0131] For example, the extreme position coordinates of the second part of the feature points refer to the coordinates of the second part of the feature points in the direction of feature point movement, while the reference position coordinates refer to the coordinates of the reference point in the direction of movement. For example, the movement direction of the upper eyelash is perpendicular to the width direction of the eye, and the extreme position coordinates of the feature points on the upper eyelash can be the coordinates of the feature points of the upper eyelash in the direction perpendicular to the width direction of the eye. For example, if the width direction of the eye is the first direction, then the movement direction of the upper eyelash is the second direction. The first direction is, for example, the reference direction mentioned above. Figure 2B The x' axis direction is consistent with the x-axis direction in the image coordinate system.

[0132] For example, in the second part of the feature points, the reference point of the feature point is a point on the reference line that has the same X-axis coordinate as the feature point.

[0133] For example, for feature points on the upper eyelash image, upper eyelid image, and lower eyelid image, the reference line is the central axis of the sclera layer (e.g., Figure 1D The reference coordinates of the feature point are the coordinates of the reference point corresponding to the feature point on the central axis of the white of the eye layer in the second direction.

[0134] For example, the selected sub-image includes the sclera image of the detected object, the second virtual sub-image includes the upper eyelid image and the lower eyelid image, and the reference position coordinates are the coordinates of the reference point corresponding to the second part of the feature points on the central axis of the sclera image. The second part of the feature points includes feature points of the upper eyelid image and feature points of the lower eyelid image. In this embodiment, step S240 includes: calculating the difference between the extreme position coordinates of the feature points in the upper eyelid image and the reference point corresponding to the feature point on the central axis; and calculating the difference between the extreme position coordinates of the feature points in the lower eyelid image and the reference point corresponding to the feature point on the central axis.

[0135] Figure 2D This illustration shows a schematic diagram of calculating the difference between the extreme position coordinates and the reference position coordinates of a feature point in the upper eyelid, according to at least one embodiment of the present disclosure.

[0136] like Figure 2D As shown, the coordinates of feature point Q0 in the upper eyelid image are (x0, y0). Q0 ), x0 is the feature point Q0 in Figure 2D The coordinates shown are on the X-axis (eye width direction), y Q0 For feature point Q0 in Figure 2D The coordinates on the Y-axis (direction of movement) are shown. The reference point for feature point Q0 is reference point n1 on the central axis of the sclera layer, and the coordinates of n1 are (x0, y0). n0 If the extreme position coordinates of feature point Q0 are y, then the difference between the extreme position coordinates and the reference position coordinates is y. Q0 -y n0.

[0137] Similarly, for example, the coordinates of feature point Q'0 in the lower eyelid image are (x0, y0). Q’0 The reference point for feature point Q'0 is reference point n1 on the central axis of the sclera layer, and the coordinates of n1 are (x0, y0). n0 If the extreme position coordinates of feature point Q'0 are y, then the difference between the extreme position coordinates and the reference position coordinates is y. n0 -y Q’0 .

[0138] For example, the coordinates of feature point P1 in the upper eyelash image are (x1, y1). p1 The reference point for feature point P1 is reference point m1 on the central axis of the sclera layer, and the coordinates of m1 are (x1, y1). m1 If the extreme position coordinates of feature point P1 are y, then the difference between the extreme position coordinates and the reference position coordinates is y. p1 -y m1 .

[0139] The above y Q0 y Q’0 y n0 y p1 and y m1 All are greater than or equal to 0.

[0140] For step S250, for example, y p1 -y m1 This represents the maximum deformation value of feature point P1.

[0141] Figure 3 A flowchart is shown of a method for determining the movement information of multiple feature points in an initial virtual image based on current feature information and maximum deformation value, according to at least one embodiment of the present disclosure.

[0142] like Figure 3 As shown, the method may include steps S310 to S320.

[0143] Step S310: Determine the current state value relative to the reference state based on the current feature information.

[0144] Step S320: Calculate the current state value and the maximum deformation value to determine the movement information of multiple feature points in the initial virtual image.

[0145] The current state value can be a parameter used to reflect the relationship between the current feature information and the reference feature information when the detected object is in the reference state.

[0146] For example, the reference state is eyes open to their maximum extent, the target feature is eyelashes, and the current state value can be a parameter that reflects the relationship between the current position of the eyelashes and the position of the eyelashes when the eyes are open to their maximum extent.

[0147] In some embodiments of this disclosure, step S310 may include obtaining the mapping relationship between feature information and state value, and determining the current state value of the target feature relative to the reference state based on the mapping relationship and the current feature information.

[0148] For example, obtaining the mapping relationship between feature information and state values ​​may include: obtaining multiple samples, each sample including the correspondence between sample feature information of the target feature and sample state value; and constructing a mapping function based on the correspondence, the mapping function representing the mapping relationship between feature information and state value.

[0149] For example, sample feature information includes target feature information and second feature information, and sample state values ​​include a first value corresponding to the target feature information and a second value corresponding to the second feature information. Constructing the mapping function includes constructing a system of linear equations, and substituting the target feature information, the first value, the second feature information, and the second value into the system of linear equations respectively, and solving the system of linear equations to obtain the mapping function.

[0150] For example, the linear equation system is a system of two linear equations in two variables. The target feature information is the position coordinates Y0 of the feature point on the eyelash image in the direction of movement when the eye is fully open (with a first value of 0). The second feature information is the position coordinates Y1 of the feature point on the eyelash image in the direction of movement when the eye is closed (with a second value of 1). The target feature information and the second feature information can be the result of statistical calculations on multiple samples. A system of two linear equations in two variables is constructed based on (Y0, 0) and (Y1, 1), and the mapping function u = av + b is obtained by solving the system of two linear equations in two variables. a and b are the values ​​obtained by solving the system of two linear equations in two variables, v is the current feature information, and u is the current state value. For example, substituting the current position coordinates v (i.e., the coordinates on the Y-axis) of each feature point on the eyelash into the mapping function u = av + b, the current state value u corresponding to the current feature information v is obtained.

[0151] In some other embodiments of this disclosure, the mapping relationship between feature information and state value can also be a mapping relationship table, and this disclosure does not limit the representation of the mapping relationship.

[0152] In some embodiments of this disclosure, step S320 includes multiplying the current state value and the maximum deformation value to obtain the movement information of multiple feature points in the initial virtual image.

[0153] For example, for feature point P1 in the upper eyelash, if the current state value is 0.5, then the movement distance of feature point P1 is 0.5 and y. p1 -y m1 The product of.

[0154] Return to reference Figure 1A In some embodiments of this disclosure, the movement information includes the movement distance. In this embodiment, step S50 includes: driving multiple feature points to move a movement distance toward a preset reference position.

[0155] For example, for a mouth image, the reference position is the widest part of the mouth. Multiple feature points of the upper lip are driven to move a distance from the initial position toward the widest part of the mouth. The initial position is the position of the feature points in the upper lip image when the mouth is opened to the maximum extent, in the initial virtual image when the mouth is opened to the maximum extent.

[0156] For example, for the upper eyelid, the reference position is the central axis of the sclera. Multiple feature points of the upper eyelid are driven to move a distance from their initial positions toward the central axis of the sclera. The initial position is the position of the feature points of the upper eyelid in the upper eyelid image when the eyes are opened to the maximum extent, and the eyes are opened to the maximum extent in the initial virtual image.

[0157] In some embodiments of this disclosure, a filling image is used to fill gaps in an overlay sub-image, which is obtained by overlaying two second virtual sub-images. The gap refers to the position between two second virtual sub-images in the overlay sub-image. In this embodiment, step S50 includes: driving multiple feature points to move in the initial virtual image according to the movement information, and changing the shape of the filling image according to the gap size so that the filling image matches the gap to generate the current virtual image.

[0158] Based on the movement information, driving multiple feature points to move in the initial virtual image can be done in the manner described above, and will not be repeated here.

[0159] like Figure 1C and 1E As shown, the filled image is, for example, an oral cavity image 113, and the two second virtual sub-images are an upper lip image 111 and a lower lip image 112. As feature points in the upper lip image 111 and lower lip image 112 move, the shape of the gap between them changes. The oral cavity image 113 is stretched or compressed to change its shape, such that it fills the gap between the upper lip image 111 and lower lip image 112, thus presenting a complete mouth image to the detection object.

[0160] In the above embodiments of this disclosure, the movement information of the target feature is the movement information relative to the reference position. Therefore, the feature points of the target feature are driven to move toward the reference position without the need to draw the final limit position to which the feature points of the target feature have moved, thereby improving the efficiency of designing virtual images.

[0161] In other embodiments of this disclosure, references above are made to... Figure 1A In addition to steps S10 to S70, the method shown may also include: calculating pitch angle, yaw angle and roll angle based on current feature information; calculating rotation matrix based on pitch angle, yaw angle and roll angle; driving feature points in the initial virtual image to move based on movement information; and controlling the rotation of the initial virtual image based on the rotation matrix to generate a current virtual image corresponding to the current state.

[0162] For example, if the detection target is a user, the pitch, yaw, and roll angles are calculated based on 280 key points of the user's face. The pitch angle is the angle of rotation of the face around a first axis, the yaw angle is the angle of rotation of the face around a second axis, and the roll angle is the angle of rotation of the face around a third axis. The first axis is perpendicular to the height direction of the face, the second axis is parallel to the height direction of the face, and the third axis is perpendicular to the first and second axes. The pitch, yaw, and roll angles are calculated based on the current feature information. The calculation of the rotation matrix based on the pitch, yaw, and roll angles can be referenced from relevant algorithms in related technologies, and will not be elaborated upon in this disclosure.

[0163] In this embodiment, while driving the feature points in the initial virtual image to move according to the movement information, the rotation of the initial virtual image is controlled according to the rotation matrix to generate the current virtual image corresponding to the current state.

[0164] For example, in response to the head rotation of the detected object, the current virtual image displayed by the electronic device changes from a virtual image of the detected object's front face to a virtual image of the detected object's side face after the head rotation.

[0165] In this embodiment, the rotation of the initial virtual image can be controlled according to the rotation matrix, making the current virtual image more realistic and vivid.

[0166] Figure 4 A schematic diagram illustrating the effect of an image processing method provided in at least one embodiment of this disclosure is shown.

[0167] like Figure 4 As shown in the schematic diagram, the detection object is shown in state 401 at the first moment and state 402 at the second moment.

[0168] The effect diagram also includes the current virtual image 403 of the detected object displayed in the electronic device at the first moment, and the current virtual image 404 of the detected object displayed in the electronic device at the second moment.

[0169] like Figure 4 As shown, the eyes of the detected object are open and its mouth is closed at the first moment. Correspondingly, the eyes of the detected object in the current virtual image 403 displayed in the electronic device are also open and its mouth is closed.

[0170] like Figure 4 As shown, at the second moment, the eyes of the object being detected are closed and the mouth is open. Correspondingly, at this time, the eyes of the object being detected in the current virtual image 404 displayed in the electronic device are also closed and the mouth is also open.

[0171] Figure 5 A schematic block diagram of an image processing apparatus 500 provided in at least one embodiment of the present disclosure is shown.

[0172] For example, such as Figure 5 As shown, the image processing apparatus 500 includes an acquisition unit 510, a processing unit 520, a detection unit 530, a determination unit 540, and a driving unit 550.

[0173] The acquisition unit 510 is configured to acquire at least one first virtual sub-image, wherein each of the at least one first virtual sub-image corresponds to one of a plurality of target features. The acquisition unit 510 may, for example, perform... Figure 1A Step S10 is described.

[0174] Processing unit 520 is configured to process a selected virtual sub-image from at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image. Processing unit 520 may, for example, perform... Figure 1A Step S20 is described.

[0175] The detection unit 530 is configured to acquire current feature information of the detected object in response to detection of the detected object. The current feature information indicates the current state of the detected object, which includes multiple target features. For example, the detection unit 530 may execute... Figure 1A Step S30 is described.

[0176] The determining unit 540 is configured to determine the movement information of multiple feature points in an initial virtual image based on current feature information. The initial virtual image is obtained by superimposing at least one first virtual sub-image and at least one second virtual sub-image. For example, the determining unit 540 may perform... Figure 1A Step S40 is described.

[0177] The driving unit 550 is configured to drive multiple feature points to move within an initial virtual image based on movement information, thereby generating a current virtual image corresponding to the current state. The driving unit 550 may, for example, perform... Figure 1A Step S50 is described.

[0178] For example, the acquisition unit 510, processing unit 520, detection unit 530, determination unit 540, and driving unit 550 can be hardware, software, firmware, or any feasible combination thereof. For example, the acquisition unit 510, processing unit 520, detection unit 530, determination unit 540, and driving unit 550 can be dedicated or general-purpose circuits, chips, or devices, or they can be a combination of a processor and memory. The embodiments of this disclosure do not limit the specific implementation of the above-mentioned units.

[0179] It should be noted that in the embodiments of this disclosure, each unit of the image processing device 500 corresponds to each step of the aforementioned image processing method. For the specific functions of the image processing device 500, please refer to the relevant description of the image processing method, which will not be repeated here. Figure 5 The components and structures of the image processing apparatus 500 shown are merely exemplary and not limiting. The image processing apparatus 500 may also include other components and structures as needed.

[0180] At least one embodiment of this disclosure also provides an electronic device including a processor and a memory, the memory including one or more computer program modules. The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the image processing method described above. This electronic device can reduce the design difficulty of virtual avatars, improve design efficiency, and reduce driving difficulty.

[0181] Figure 6 This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. For example... Figure 6 As shown, the electronic device 800 includes a processor 810 and a memory 820. The memory 820 stores non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 810 executes the non-transitory computer-readable instructions, which, when executed by the processor 810, can perform one or more steps in the image processing method described above. The memory 820 and the processor 810 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0182] For example, processor 810 may be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For example, the central processing unit (CPU) may be an x86 or ARM architecture. Processor 810 may be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 800 to perform desired functions.

[0183] For example, memory 820 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and processor 810 may run one or more computer program modules to implement various functions of electronic device 800. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0184] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 800 can be referred to the description of the image processing method above, and will not be repeated here.

[0185] Figure 7 This is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. The electronic device 900 is, for example, suitable for implementing the image processing method provided in the embodiments of the present disclosure. The electronic device 900 may be a terminal device, etc. It should be noted that... Figure 7 The illustrated electronic device 900 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0186] like Figure 7 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 910, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 920 or a program loaded from a storage device 980 into a random access memory (RAM) 930. The RAM 930 also stores various programs and data required for the operation of the electronic device 900. The processing device 910, the ROM 920, and the RAM 930 are interconnected via a bus 940. An input / output (I / O) interface 950 is also connected to the bus 940.

[0187] Typically, the following devices can be connected to I / O interface 950: input devices 960 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 970 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 980 including, for example, magnetic tapes, hard disks, etc.; and communication devices 990. Communication device 990 allows electronic device 900 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 7 An electronic device 900 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 900 may alternatively implement or have more or fewer devices.

[0188] For example, according to embodiments of this disclosure, the image processing method described above can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the image processing method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 990, or installed from a storage device 980, or installed from a ROM 920. When the computer program is executed by the processing device 910, the functions defined in the image processing method provided by embodiments of this disclosure can be implemented.

[0189] At least one embodiment of this disclosure also provides a computer-readable storage medium for storing non-transitory computer-readable instructions that, when executed by a computer, can implement the image processing method described above. Using this computer-readable storage medium can reduce the design difficulty of virtual avatars, improve design efficiency, and reduce driving complexity.

[0190] Figure 8 This is a schematic diagram of a storage medium provided for some embodiments of this disclosure. For example... Figure 8 As shown, the storage medium 1000 is used to store non-transitory computer-readable instructions 1010. For example, when the non-transitory computer-readable instructions 1010 are executed by a computer, one or more steps in the image processing method described above can be performed.

[0191] For example, the storage medium 1000 can be used in the aforementioned electronic device 800. For example, the storage medium 1000 can be... Figure 6 The memory 820 in the illustrated electronic device 800. For example, a description of the storage medium 1000 can be found here. Figure 6 The corresponding description of the memory 820 in the illustrated electronic device 800 will not be repeated here.

[0192] The following points need to be explained:

[0193] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0194] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0195] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.

Claims

1. An image processing method, comprising: At least one first virtual sub-image is acquired, wherein each of the at least one first virtual sub-image corresponds to one of a plurality of target features; The selected virtual sub-image in the at least one first virtual sub-image is processed to obtain at least one second virtual sub-image associated with the selected virtual sub-image; In response to the detection of a target object, the current feature information of the target object is obtained, wherein the current feature information is used to indicate the current state of the target object, and the target object includes the plurality of target features; Based on the current feature information, the movement information of multiple feature points in the initial virtual image is determined, wherein the initial virtual image is obtained by superimposing the at least one first virtual sub-image and the at least one second virtual sub-image; and Based on the movement information, the plurality of feature points are driven to move in the initial virtual image to generate a current virtual image corresponding to the current state; Wherein, the multiple target features are multiple body parts; the detection object is a movable object; the current feature information of the detection object includes the current action or posture information of the detection object; The processing of the selected virtual sub-image includes splitting or morphological processing; Based on the current feature information, determine the movement information of multiple feature points in the initial virtual image, including: determining the movement information of the multiple feature points based on the current feature information and the maximum deformation value of the multiple feature points; The determination of the movement information of the multiple feature points based on the current feature information and the maximum deformation value of the multiple feature points includes: Based on the current feature information, a current state value relative to a reference state is determined, wherein the current state value is a parameter reflecting the relationship between the current feature information and the reference feature information when the detected object is in the reference state; and The current state value and the maximum deformation value of the plurality of feature points are multiplied to obtain the movement information of the plurality of feature points.

2. The method according to claim 1, further comprising: Get the filled image; as well as The at least one first virtual sub-image, the at least one second virtual sub-image, and the fill image are superimposed to generate the initial virtual image.

3. The method according to claim 1, further comprising: The depth information of each first virtual sub-image and each second virtual sub-image in the initial virtual image is obtained, and the at least one first virtual sub-image and the at least one second virtual sub-image are superimposed according to the depth information of each first virtual sub-image and each second virtual sub-image to generate the initial virtual image.

4. The method according to any one of claims 1 to 3, wherein, Processing a selected virtual sub-image from the at least one first virtual sub-image to obtain the at least one second virtual sub-image associated with the selected virtual sub-image includes: The selected virtual sub-image is subjected to morphological processing to obtain the at least one second virtual sub-image associated with the selected virtual sub-image.

5. The method according to claim 4, wherein, The selected virtual sub-image includes the sclera image of the detected object, and the method further includes: The sclera image is divided into a first sclera sub-image and a second sclera sub-image along its central axis. The direction of the central axis is parallel to the length direction of the subject's eye. The first sclera sub-image is located on the side of the central axis away from the subject's mouth, and the second sclera sub-image is located on the side of the central axis closer to the mouth. Performing morphological processing on the selected virtual sub-image to obtain the at least one second virtual sub-image associated with the selected virtual sub-image includes: Morphological processing is performed on the first eye leukocyte image and the second eye leukocyte image respectively to obtain a first sub-image and a second sub-image; and Based on the color and texture of the detected object's face, the first sub-image is filled to obtain the upper eyelid image of the detected object, and the second sub-image is filled to obtain the lower eyelid image of the detected object.

6. The method according to any one of claims 1 to 3, wherein, Processing a selected virtual sub-image from the at least one first virtual sub-image to obtain the at least one second virtual sub-image associated with the selected virtual sub-image includes: The selected virtual sub-image is split to obtain multiple second virtual sub-images.

7. The method according to claim 6, wherein, The selected virtual sub-image includes the mouth image of the detected object. The selected virtual sub-image is split to obtain multiple second virtual sub-images, including: The mouth image of the detected object is split into two second virtual sub-images from the widest position of the mouth. One of the two second virtual sub-images is the upper lip image, and the other is the lower lip image.

8. The method according to any one of claims 1 to 3, further comprising: Calculate the maximum deformation value of the plurality of feature points.

9. The method according to claim 8, wherein, Calculating the maximum deformation value of the plurality of feature points includes: Determine the maximum deformation curve that the first portion of the feature points among the plurality of feature points conform to; Determine the position coordinates of each feature point in the first part of the feature points; and Substitute the position coordinates of each feature point in the first part of the feature points into the maximum deformation curve to obtain the maximum deformation value of the first part of the feature points.

10. The method according to claim 8, wherein, Calculating the maximum deformation value of the plurality of feature points includes: Calculate the difference between the extreme position coordinates and the reference position coordinates of the second portion of feature points among the plurality of feature points, wherein the extreme position coordinates are the coordinates of the second portion of feature points in the initial virtual image; and The difference is taken as the maximum deformation value of each feature point in the second part.

11. The method according to claim 1, wherein, Based on the current feature information, the current state value relative to the reference state is determined, including: Obtain the mapping relationship between feature information and state values; and Based on the mapping relationship and the current feature information, the current state value relative to the reference state is determined.

12. The method according to claim 11, wherein, Obtaining the mapping relationship between the feature information and the state value includes: Acquire multiple samples, where each sample includes the correspondence between sample feature information and sample state values; and Based on the correspondence included in each of the plurality of samples, a mapping function is constructed, wherein the mapping function represents the mapping relationship between the feature information and the state value.

13. The method according to claim 12, wherein, The sample feature information includes target feature information and second feature information, and the sample state value includes a first value corresponding to the target feature information and a second value corresponding to the second feature information. Based on the aforementioned correspondence, the mapping function is constructed, including: Constructing a system of linear equations; and The mapping function is obtained by solving the linear equation system based on the target feature information and the first value, the second feature information and the second value.

14. The method according to claim 9, wherein, The second virtual sub-image includes an upper lip image and a lower lip image, and the first portion of feature points includes feature points from the upper lip image and feature points from the lower lip image. Determining the maximum deformation curve that the first portion of the feature points conform to among the plurality of feature points includes: The first maximum deformation curve that the feature points in the upper lip image conform to and the second maximum deformation curve that the feature points in the lower lip image conform to are determined respectively.

15. The method according to claim 14, wherein, The feature points in the upper lip image and the feature points in the lower lip image correspond one-to-one. The feature points in the upper lip image and the feature points in the lower lip image comprise n columns. The first maximum deformation curve is The second maximum deformation curve is , Wherein, x' is the x'-axis coordinate of the feature point in the upper lip image and the feature point in the lower lip image, the widest position of the mouth is the x'-axis, y1 is the distance the feature point in the upper lip image moves away from the lower lip, y2 is the distance the feature point in the lower lip image moves away from the upper lip, and c is the maximum distance in the y-axis direction between the feature point in the upper lip image and the feature point in the lower lip image when the mouth is opened to its maximum extent.

16. The method of claim 10, wherein, The selected virtual sub-image includes the sclera image of the detected object, and the second virtual sub-image includes the upper eyelid image and the lower eyelid image. The reference position coordinates are the coordinates of the reference point on the central axis of the sclera image corresponding to the second part of the feature points, and the direction of the central axis is parallel to the length direction of the detected object's eye. The second part of the feature points includes feature points of the upper eyelid image and feature points of the lower eyelid image. Calculating the difference between the extreme position coordinates and the reference position coordinates of the second part of the feature points among the plurality of feature points includes: Calculate the extreme position coordinates of the feature points in the upper eyelid image and the difference between the reference point on the midline and the feature point in the upper eyelid image; and Calculate the extreme position coordinates of the feature points in the lower eyelid image and the difference between the reference point on the central axis and the feature point in the lower eyelid image.

17. The method according to claim 1, wherein, The movement information includes the distance traveled. Based on the movement information, the plurality of feature points are driven to move within the initial virtual image, including: The plurality of feature points are driven to move the distance toward a preset reference position.

18. The method according to claim 2, wherein, The filling image is used to fill the gaps in the overlay sub-image, which is obtained by superimposing two second virtual sub-images. The gap refers to the position between the two second virtual sub-images in the overlay sub-image. Based on the movement information, the plurality of feature points are driven to move within the initial virtual image, including: Based on the movement information, the plurality of feature points are driven to move in the initial virtual image, and the shape of the filling image is changed according to the gap size so that the filling image adapts to the gap to generate the current virtual image.

19. The method according to claim 18, wherein, The filled image includes an oral cavity image, and the two second virtual sub-images include an upper lip image and a lower lip image.

20. An image processing apparatus, comprising: The acquisition unit is configured to acquire at least one first virtual sub-image, wherein each of the at least one first virtual sub-image corresponds to one of a plurality of target features; The processing unit is configured to process a selected virtual sub-image from the at least one first virtual sub-image to obtain at least one second virtual sub-image associated with the selected virtual sub-image; The detection unit is configured to acquire current feature information of the detected object in response to the detection of the detected object, wherein the current feature information is used to indicate the current state of the detected object, and the detected object includes the plurality of target features; The determining unit is configured to determine the movement information of multiple feature points in an initial virtual image based on the current feature information, wherein the initial virtual image is obtained by superimposing the at least one first virtual sub-image and the at least one second virtual sub-image; and The driving unit is configured to drive the plurality of feature points to move in the initial virtual image according to the movement information, so as to generate a current virtual image corresponding to the current state; Wherein, the multiple target features are multiple body parts; the detection object is a movable object; the current feature information of the detection object includes the current action or posture information of the detection object; The processing of the selected virtual sub-image includes splitting or morphological processing; The determining unit is further configured to: determine the movement information of the multiple feature points based on the current feature information and the maximum deformation value of the multiple feature points; The determining unit is further configured to: determine the current state value of the current state relative to the reference state based on the current feature information, wherein the current state value is a parameter used to reflect the relationship between the current feature information and the reference feature information when the detected object is in the reference state; and perform a multiplication operation on the current state value and the maximum deformation value of the plurality of feature points to obtain the movement information of the plurality of feature points.

21. An electronic device, comprising: processor; Memory, which includes one or more computer program instructions; The one or more computer program instructions are stored in the memory and, when executed by the processor, implement the image processing method according to any one of claims 1-19.

22. A computer-readable storage medium that non-transitoryly stores computer-readable instructions, wherein, The image processing method according to any one of claims 1-19 is implemented when the computer-readable instructions are executed by a processor.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN109584152A

  • Image processing method and device, terminal equipment and storage medium

    CN110390704A