Virtual image interaction method, virtual image interaction device, medium and electronic device

By obtaining point cloud data in real images, detecting the collision between face and hand areas, and adjusting the deformation of virtual images, the problem of poor interaction between virtual images and users is solved, and higher correlation and interactivity are achieved.

CN115712339BActive Publication Date: 2025-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110956454.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-19
Publication Date
2025-08-26
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

In the prior art, virtual images are difficult to adjust according to the actual interaction of users, user interaction experience is poor, and virtual images have low correlation with users.

Method used

By acquiring real images, extracting point cloud data of face and hand, determining the face and hand area, and in response to the collision between the face area and the hand area, the virtual image is deformed and adjusted based on the collision results.

Benefits of technology

It improves the relevance and interactivity between virtual images and users, increases the fun of interaction between users and virtual images, and ensures the accuracy of collision detection and image adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712339B_ABST
    Figure CN115712339B_ABST
Patent Text Reader

Abstract

The present disclosure provides a virtual image interaction method, a virtual image interaction device, a computer-readable storage medium, and an electronic device, relating to the field of computer technology. The virtual image interaction method includes: acquiring a real image; extracting facial point cloud data from the real image and determining a facial region based on the facial point cloud data; extracting hand point cloud data from the real image and determining a hand region based on the hand point cloud data; and in response to a collision between the facial region and the hand region, deforming and adjusting the virtual image corresponding to the real image based on the collision result. The present disclosure can deform and adjust the virtual image based on the collision between the face and the hand in the real image, thereby improving the user's sense of interaction with the virtual image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a virtual image interaction method, a virtual image interaction device, a computer-readable storage medium, and an electronic device. Background Art

[0002] With the development of the Internet, in order to meet the personalized needs of users in various application scenarios, many terminals or applications provide users with customized virtual image services, that is, users can set their own virtual images according to their preferences or needs, and use them in scenarios such as chatting, video calls or making short videos.

[0003] In existing technologies, user interaction with avatars is typically achieved through interaction with a terminal device. For example, a user can press the avatar's face on the terminal device's display interface to cause the avatar's face to deform, or push the edges of the avatar's face to create a slimmer, more distorted effect. However, existing technologies rarely offer solutions that can adaptively adjust the avatar's state based on the user's actual facial and hand interactions. This results in a low level of relevance between the avatar and the user, impacting the user's interactive experience with the avatar. Summary of the Invention

[0004] The present disclosure provides a virtual image interaction method, a virtual image interaction device, a computer-readable storage medium and an electronic device, thereby at least to some extent improving the problem in the prior art that the virtual image is difficult to adjust according to the user's actual interaction situation, resulting in a poor user interaction experience.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0006] According to a first aspect of the present disclosure, a virtual image interaction method is provided, comprising: acquiring a real image; extracting facial point cloud data from the real image, and determining a facial region based on the facial point cloud data; extracting hand point cloud data from the real image, and determining a hand region based on the hand point cloud data; and in response to a collision between the facial region and the hand region, performing a deformation adjustment on the virtual image corresponding to the real image based on the collision result.

[0007] According to a second aspect of the present disclosure, a virtual image interaction device is provided, comprising: a real image acquisition module for acquiring a real image; a first area determination module for extracting face point cloud data from the real image and determining a face area based on the face point cloud data; a second area determination module for extracting hand point cloud data from the real image and determining a hand area based on the hand point cloud data; and a virtual image adjustment module for performing deformation adjustment on the virtual image corresponding to the real image based on a collision result in response to a collision between the face area and the hand area.

[0008] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the virtual image interaction method and possible implementation methods of the above-mentioned first aspect are implemented.

[0009] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor, wherein the processor is configured to execute the executable instructions to perform the virtual image interaction method and possible implementations thereof according to the first aspect.

[0010] The technical solution disclosed in this disclosure has the following beneficial effects:

[0011] Acquire a real image; extract facial point cloud data from the real image and determine a face region based on the facial point cloud data; extract hand point cloud data from the real image and determine a hand region based on the hand point cloud data; and in response to a collision between the facial region and the hand region, deform and adjust the virtual image corresponding to the real image based on the collision result. This exemplary embodiment provides a new virtual image interaction method that, based on the actual collision between the user's face and hand in the real image, deforms and adjusts the virtual image accordingly, so that the virtual image can truly and accurately reflect the user's actual state, improve the relevance and interactivity between the user and the virtual image, and provide high interest. Furthermore, this exemplary embodiment creates facial point cloud data and hand point cloud data from the real image and deforms and adjusts the virtual image based on the collision result between the facial region and the hand region generated by the facial point cloud data and the hand point cloud data. The point cloud data can more accurately reflect the actual state of the face and hand, further ensuring the accuracy and effectiveness of collision detection and virtual image adjustment.

[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0014] Figure 1 A schematic diagram showing a system architecture in this exemplary embodiment;

[0015] Figure 2 A structural diagram of an electronic device according to this exemplary embodiment is shown;

[0016] Figure 3 A flowchart showing a virtual image interaction method in this exemplary embodiment is shown;

[0017] Figure 4 A schematic diagram illustrating an adjustment of a virtual image in this exemplary embodiment is shown;

[0018] Figure 5 A sub-flowchart showing a method for interacting with a virtual character in this exemplary embodiment;

[0019] Figure 6 Another sub-flowchart of a virtual character interaction method according to this exemplary embodiment is shown;

[0020] Figure 7 A schematic diagram showing a network architecture in this exemplary embodiment;

[0021] Figure 8 Another sub-flowchart of a virtual character interaction method in this exemplary embodiment is shown;

[0022] Figure 9 A flowchart illustrating another virtual image interaction method in this exemplary embodiment;

[0023] Figure 10 A structural diagram of a virtual image interaction device in this exemplary embodiment is shown. DETAILED DESCRIPTION

[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0025] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0026] In view of one or more of the above problems, exemplary embodiments of the present disclosure provide a virtual character interaction method. Figure 1 FIG. 1 shows a system architecture diagram of the operating environment of this exemplary embodiment. Figure 1 As shown, the system architecture 100 may include a user terminal 110 and a server 120, which communicate with each other via a network. The server 110 is a backend server that provides Internet services; the user terminal 110 may include, but is not limited to, electronic devices such as smartphones, tablets, AR (Augmented Reality) / VR (Virtual Reality) devices, game consoles, and wearable devices.

[0027] It should be understood that Figure 1 The number of each device is only exemplary. According to the implementation requirements, any number of user terminals can be set, or the server can be a cluster formed by multiple servers.

[0028] The virtual image interaction method provided by the embodiment of the present disclosure can be executed by the user terminal 110. For example, after the user terminal 110 obtains the real image, the face point cloud data and the hand point cloud data are directly generated according to the real image, and subsequent collision detection and deformation adjustment of the virtual image are performed based on the face point cloud data and the hand point cloud data; it can also be executed by the server 120. For example, after the user terminal 110 obtains the real image, it is uploaded to the server 120, and the server 120 extracts the point cloud data, performs collision detection and deformation adjustment of the virtual image according to the real image, and finally returns the deformation adjustment result to the user terminal 110; the processing process can also be jointly performed by the user terminal 110 and the server 120 according to actual needs, for example, the face point cloud processing process and the hand point cloud processing process are executed separately, etc. The present disclosure does not make specific restrictions on this.

[0029] The exemplary embodiment of the present disclosure provides an electronic device for implementing a virtual character interaction method, which may be Figure 1 The user terminal 110 and the server 120 in the embodiment of the present invention are configured to: the electronic device at least comprises a processor and a memory, wherein the memory is used to store executable instructions of the processor, and the processor is configured to execute the virtual image interaction method by executing the executable instructions.

[0030] Below is Figure 2 The structure of the above electronic device is exemplarily described by taking the mobile terminal 200 in FIG. 1 as an example. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 2 The construction in can also be applied to fixed type equipment.

[0031] like Figure 2 As shown, the mobile terminal 200 may specifically include: a processor 210, an internal memory 221, an external memory interface 222, a USB (Universal Serial Bus) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 271, a receiver 272, a microphone 273, an earphone interface 274, a sensor module 280, a display screen 290, a camera module 291, an indicator 292, a motor 293, a button 294 and a SIM (Subscriber Identification Module) card interface 295, etc.

[0032] The processor 210 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit). The encoder may encode (i.e., compress) image or video data, and the decoder may decode (i.e., decompress) the image or video bitstream data to restore the image or video data.

[0033] In some implementations, the processor 210 may include one or more interfaces, and may be connected to other components of the mobile terminal 200 through different interfaces.

[0034] The internal memory 221 can be used to store computer executable program code, which includes instructions. The internal memory 221 may include volatile memory, non-volatile memory, etc. The processor 210 executes various functional applications and data processing of the mobile terminal 200 by running instructions stored in the internal memory 221 and / or instructions stored in a memory provided in the processor.

[0035] The external memory interface 222 can be used to connect an external memory, such as a Micro SD card, to expand the storage capacity of the mobile terminal 200. The external memory communicates with the processor 210 through the external memory interface 222 to implement data storage functions, such as storing music, video and other files.

[0036] The USB interface 230 is an interface that complies with USB standard specifications and can be used to connect a charger to charge the mobile terminal 200 , or to connect headphones or other electronic devices.

[0037] The charging management module 240 is used to receive charging input from the charger. While charging the battery 242, the charging management module 240 can also power the device through the power management module 241. The power management module 241 can also monitor the status of the battery.

[0038] The wireless communication function of the mobile terminal 200 can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modulation and demodulation processor and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the mobile terminal 200. The wireless communication module 260 can provide wireless communication solutions including WLAN (Wireless Local Area Networks) (such as Wi-Fi (Wireless Fidelity) network), BT (Bluetooth), GNSS (Global Navigation Satellite System), FM (Frequency Modulation), NFC (Near Field Communication, short-range wireless communication technology), IR (Infrared, infrared technology) and the like applied on the mobile terminal 200.

[0039] The mobile terminal 200 can implement a display function and display a user interface through the GPU, display screen 290, and AP. The mobile terminal 200 can implement a shooting function through the ISP, camera module 291, encoder, decoder, GPU, display screen 290, and AP, and can also implement an audio function through the audio module 270, speaker 271, receiver 272, microphone 273, headphone jack 274, and AP.

[0040] The sensor module 280 may include a depth sensor 2801 , a pressure sensor 2802 , a gyroscope sensor 2803 , an air pressure sensor 2804 , etc., to implement different sensing detection functions.

[0041] Indicator 292 can be an indicator light that can be used to indicate charging status, power level changes, messages, missed calls, notifications, etc. Motor 293 can generate vibration prompts and can also be used for touch vibration feedback, etc. Buttons 294 include a power button, volume button, etc.

[0042] The mobile terminal 200 may support one or more SIM card interfaces 295 for connecting to SIM cards to implement functions such as calls and data communications.

[0043] Figure 3 An exemplary process of a virtual character interaction method is shown, which can be executed by the user terminal 110 or the server 120, including the following steps S310 to S340:

[0044] Step S310: Acquire a real image.

[0045] Here, a real image refers to an image containing the face or limbs of a real person, such as an image containing the user's face or limbs captured after the user turns on the front or rear camera of the terminal device, or a portrait image captured by the user in the past. In this exemplary embodiment, the real image can be acquired in real time through the camera, or can be acquired from a historical photo album, or can be downloaded from the cloud, etc. In addition, the real image can include multiple images, such as real-time acquisition of the user's video stream data, selecting multiple images from the video stream as real images for processing, or directly capturing multiple real images at the same time through the same camera or different cameras configured in the terminal device, etc. The perspectives of different real images can be different, for example, the face can be the main shooting subject, or the limbs can be the main shooting subject, etc., and this disclosure does not specifically limit this.

[0046] Step S320: extracting facial point cloud data from the real image and determining the facial area based on the facial point cloud data.

[0047] In this exemplary embodiment, the real image may include an image of the face area, an image of the limb area, or an image including the face area and the limb area, wherein the image including the face area and the limb area can reflect the positional relationship and motion information between the face and the limbs, and the combination of the image of the face area and the image of the limb area can also reflect the positional relationship and motion information between the face and the limbs. Based on the key points or feature points of the face area in the real image, facial point cloud data of the face area can be generated. The facial point cloud data refers to a set of three-dimensional feature points established based on the key points of facial features such as eyebrows, eyes, nose, mouth and facial contours. The facial point cloud data can reflect information such as the shape or posture of the face, such as the fatness, length, height of cheekbones, and the state of the mouth or eyes opening and closing, etc.

[0048] Specifically, this exemplary embodiment can extract facial point cloud data from real images in a variety of ways. For example, a pre-trained neural network model or a specific algorithm can be used to process the real image to extract facial feature points. Furthermore, facial point cloud data can be established by determining the shape or posture parameters of the facial feature points. Finally, the corresponding facial region is determined by the distribution of point clouds at different locations in the facial point cloud data. The facial region can be an irregular graphic region, such as a region connected based on the outermost point cloud, or a regular graphic region, such as a rectangular region containing all point clouds. In addition, the facial region can be a connected region, such as a region composed of all point clouds, or a non-connected region. For example, the facial point cloud can be divided into multiple non-connected regions based on the point clouds of the eyes, cheeks, mouth, etc. This exemplary embodiment can characterize the face area by determining a face bounding box, which refers to a face bounding box generated based on all or part of the point cloud in the face point cloud data. For example, a face bounding box is generated based on the point cloud at the boundary in the face point cloud data. This exemplary embodiment can use the face bounding box to represent the face area to perform the subsequent collision detection process.

[0049] Step S330: extracting hand point cloud data from the real image and determining the hand area based on the hand point cloud data.

[0050] Among them, hand point cloud data refers to a set of three-dimensional feature points established based on the key points or feature points of the hand area in the real image, such as the hand joints, elbow joints, or the contours of the hand and arm. Through the hand point cloud data, information such as the shape or posture of the hand can be reflected, such as the size and fatness of the hand, whether the arm is bent, whether the palm is in a fist or open state, etc.

[0051] In this exemplary embodiment, the hand point cloud data extraction process can first extract the feature vectors of the hand joints based on the real image. For example, the real image is used as input data and processed using a pre-trained neural network to obtain the feature vectors of the hand joints. Then, the hand shape or posture parameters corresponding to each hand joint are determined, and hand point cloud data is established based on each hand joint and its corresponding hand shape and posture parameters. Finally, the corresponding hand area is determined by the distribution of point clouds at different positions in the hand point cloud data. The hand area can be an irregular graphic area, such as an area connected based on the outermost point cloud, or a regular graphic area, such as a rectangular area containing all point clouds. In addition, the hand area can be a connected area, such as an area of ​​the entire hand composed of all point clouds, or a non-connected area, such as a point cloud divided into multiple non-connected areas based on the point clouds of the index finger, middle finger, etc. This exemplary embodiment can characterize the hand area by determining a hand bounding box, which refers to a hand bounding box generated based on all or part of the point cloud in the hand point cloud data. For example, a hand bounding box is generated based on the point cloud at the boundary in the hand point cloud data. This exemplary embodiment can use the face bounding box to represent the hand area to perform the subsequent collision detection process.

[0052] Step S340 , in response to the collision between the face region and the hand region, the virtual image corresponding to the real image is deformed and adjusted based on the collision result.

[0053] In practical applications, if a face and a hand collide in a real image, the face will deform. Correspondingly, the virtual image of the real face should also deform in accordance with the real face to ensure the virtual image's authenticity and interactivity. Therefore, this exemplary embodiment can adaptively adjust the deformation of the virtual image corresponding to the real face by determining whether the face and hand regions collide based on the real image. Specifically, considering that general object collisions can be processed as rectangular collisions, this exemplary embodiment can determine whether the face and hand regions collide by performing collision detection on the face and hand bounding boxes. Collision detection essentially detects whether the face and hand bounding boxes overlap. Specifically, this can be determined by comparing the distance and width of the center points of the face and hand bounding boxes in the x- and y-axis directions. When a collision is detected between the face and hand bounding boxes, it is determined that the face and hand regions have collided. The virtual image corresponding to the real image can then be deformed based on the collision result. The collision result may include collision parameters or collision time between the face bounding box and the hand bounding box, such as collision speed, acceleration, collision degree, or duration of a collision action during the collision. This exemplary embodiment can adjust the virtual image's deformation based on the collision result to simulate the effect of a collision between a real face and a hand. Specifically, collisions can be divided into instantaneous collisions and continuous collisions. Instantaneous collisions may include parameters such as collision speed, acceleration, and force. Based on the collision parameters of the instantaneous collision, the virtual image can be deformed, such as causing the face to sag, changes in facial curves, or a tilt of the head to one side. A continuous collision typically also begins with an instantaneous collision. Based on the parameters of the instantaneous collision, the virtual image can be deformed. Furthermore, whether to maintain the virtual image's deformed state can be determined based on the positional relationship between the face and hand regions, or the collision duration. For example, if a finger pokes the cheek, causing it to sag, and then the finger stops moving and remains in contact with the cheek, the collision speed between the finger and the face is zero, and the virtual image's face can still be displayed in a sag state. This exemplary embodiment can control the deformation process of the virtual image through specific collision parameters in the collision result. Further, based on other parameters, such as the positional relationship between the hand area and the face area, it is determined whether the virtual image maintains the deformed state or restores the state before deformation.

[0054] Figure 4 A schematic diagram showing adjustment of a virtual image in this exemplary embodiment is shown. Figure 4 (a) represents the state of the virtual image before adjustment. When the hand area collides with the face area, such as when the hand area collides with the mouth area in the face area, the virtual image can be deformed. Figure 4As shown in (b), the position and shape of the mouth area change due to the hand collision.

[0055] In an exemplary embodiment, the number of the real images may be multiple frames. For example, a camera may be used to collect video stream data of the user's face and hands in real time, and multiple frames of real images may be determined from the video stream for performing deformation adjustment processing on the virtual image. In step S340, deformation adjustment of the virtual image corresponding to the real image based on the collision result may include:

[0056] Determine collision motion information based on the collision result in each frame of the real image, where the collision motion information includes at least one of a collision speed, a collision acceleration, and a collision position;

[0057] The virtual image corresponding to the real image is deformed and adjusted according to the collision motion information.

[0058] That is, this exemplary embodiment can perform face and hand collision detection on each frame of real images based on multiple frames of real images, and determine the dynamic collision process of the face and hand between different image frames. For example, the process of the user's hand touching the face and causing the face to undergo maximum deformation may involve multiple frames of images. By analyzing each frame of these multiple frames of real images, the collision degree of the hand-face collision event can be more accurately determined to accurately adjust the virtual image.

[0059] Collision motion information may include collision velocity, collision acceleration, and collision location. Based on this collision motion information, the avatar can determine the appropriate deformation process, such as the degree of facial concavity, head tilt, and change in facial curves. This exemplary embodiment can simulate and implement specific collision events between a face and a hand using Unity (a game engine).

[0060] In summary, in this exemplary embodiment, a real image is acquired; facial point cloud data is extracted from the real image, and a facial region is determined based on the facial point cloud data; hand point cloud data is extracted from the real image, and a hand region is determined based on the hand point cloud data; and in response to a collision between the facial region and the hand region, a deformation adjustment is performed on the virtual image corresponding to the real image based on the collision result. On the one hand, this exemplary embodiment provides a new virtual image interaction method that, based on the actual collision between the user's face and hand in the real image, performs corresponding deformation adjustment on the virtual image, so that the virtual image can truly and accurately reflect the user's actual state, improve the relevance and interactivity between the user and the virtual image, and at the same time, provide high interest. On the other hand, this exemplary embodiment creates facial point cloud data and hand point cloud data from the real image, and performs deformation adjustment on the virtual image based on the collision result between the facial region and the hand region generated by the facial point cloud data and the hand point cloud data. The point cloud data can more accurately reflect the actual state of the face and hand, further ensuring the accuracy and effectiveness of collision detection and virtual image adjustment.

[0061] In an exemplary embodiment, Figure 5 As shown, in the above step S320, extracting facial point cloud data from the real image may include the following steps:

[0062] Step S510, extracting facial feature points from the real image;

[0063] Step S520, determining a cost function by reducing the dimensionality of facial feature points into a face shape feature space and an expression feature space, and determining face shape parameters and expression parameters by optimizing the cost function;

[0064] Step S530: Generate facial point cloud data according to the face shape parameters and expression parameters.

[0065] Before extracting facial feature points from a real image, to ensure the validity of the real image and improve the accuracy of image processing, this exemplary embodiment may first perform image enhancement on the real image to reduce interference caused by interference factors in the real image to the facial area and enhance image details. Specifically, image enhancement may include filtering the real image using filtering methods such as sliding average filtering, median filtering, and arithmetic mean filtering, or performing image sharpening or contrast enhancement on the real image. After image enhancement, the process of extracting facial feature points may be performed. This exemplary embodiment may use Adaboost (classifier) ​​to process the real image after the image enhancement process to extract the pixel region where the face is located in the real image, and then extract the facial feature points from this pixel region. Facial feature points are key points that can reflect the structure or characteristics of a face, such as key points that can reflect the position or shape of parts such as eyebrows, eyes, nose, and mouth, or key points that reflect the contour or position of facial structure.

[0066] After determining the facial feature points, this exemplary embodiment can reduce the dimensionality of the facial feature points to a face shape feature space and an expression feature space to construct a face reconstruction cost function. By optimizing this cost function, the face shape parameters and expression parameters of the face are determined. The face shape parameters refer to parameters that can reflect the shape of the face, such as fatness, thinness, length, etc., and the expression parameters refer to parameters that can reflect the posture of various parts of the face, such as the opening and closing of the mouth, the opening and closing of the eyes, the height of the eyebrows, and other expression-related parameters. Specifically, the face reconstruction cost function can be expressed by the following formula:

[0067]

[0068] Among them, α i , β i are the face shape parameters and expression parameters of the face, and are the parameters to be optimized. By optimizing the cost function, a stable α can be determined. i , β i As the final determined facial parameters and expression parameters; s is the scale factor of the camera, which refers to a parameter that can reflect the relationship between the focal length of the camera and the distance between the person; R is the external parameter of the camera, which can include a rotation matrix and a translation vector. In this exemplary embodiment, the rotation matrix and the translation matrix can be combined and substituted into the cost function as the external parameters of the camera; P is the projection matrix of the camera; s i and e i are the PCA (Principal Component Analysis) parameters of face shape and facial expression, respectively. The PCA parameters refer to the main characteristic components of face shape data and facial expression data; X is the 68 feature points of the face extracted; δ iThe deviation of the principal component; v i is the PCA coefficient; is an average face point cloud model, which can be determined by calculating the average value of historical or pre-acquired face point cloud calculations of different people (such as people of different genders, nationalities, and ages); i represents the identifier of different facial feature points.

[0069] By optimizing the above cost function, the face shape parameters and expression parameters of the facial feature points can be determined. Furthermore, three-dimensional face point cloud data can be generated based on the face shape parameters and expression parameters of the facial feature points.

[0070] In an exemplary embodiment, the above step 510 may include the following steps:

[0071] Obtain reference facial feature points in a reference image of a real image;

[0072] Determine a confidence region in the real image based on the position of the reference facial feature points in the reference image and the motion information between the reference image and the real image;

[0073] Extract facial feature points within the confidence region.

[0074] To ensure the accuracy and stability of the extracted facial feature points, this exemplary embodiment can optimize and determine the facial feature points in the real image based on the reference facial feature points in the reference image. The real image can be the current frame, and the reference image can be the previous frame relative to the current frame. The reference facial feature points are the facial feature point information in the previous frame, and can be obtained through processing using a pre-trained neural network model.

[0075] Typically, there may be a certain correspondence between the pixels of two adjacent frames. This exemplary embodiment can use an optical flow algorithm to determine the correspondence between the previous frame, i.e., the reference image, and the current frame, i.e., the real image, based on the changes in pixels in the time domain and the correlation between adjacent frames, thereby calculating the motion information of the corresponding objects between the real image and the reference image. Therefore, this exemplary embodiment can first estimate the facial feature points in the real image based on the position of the reference facial feature points in the reference image and the determined motion information, and then determine a confidence region based on the estimated facial feature points. Furthermore, within this confidence region, facial feature points with higher accuracy are extracted. For example, if the estimated facial feature point coordinates are (1, 1), then after re-extracting the facial feature points by determining the confidence region, the coordinates of the facial feature points can be accurately (1.2, 1.5). This exemplary embodiment can extract facial feature point coordinates with higher accuracy by first roughly determining the facial feature points and then optimizing them based on the confidence region, thereby ensuring the accuracy of the extracted facial feature points.

[0076] The specific facial feature point extraction method can adopt the Dlib (face feature detection algorithm) algorithm to perform facial feature point detection on the confidence area, and extract 68 facial feature points and their corresponding coordinates from it. Among them, each part of the face can correspond to multiple facial feature points. For example, eyebrows can be represented by 5 facial feature points, eyes can be represented by 6 facial feature points, and so on.

[0077] In an exemplary embodiment, after determining the facial shape parameters and the expression parameters in step S520, the avatar interaction method may further include:

[0078] The face shape of the avatar is adjusted according to the face shape parameters, and / or the expression of the avatar is adjusted according to the expression parameters.

[0079] In this exemplary embodiment, regardless of whether or not the avatar requires deformation adjustment, once the facial shape and expression parameters of the user's face are determined, the avatar's facial shape can be adjusted in real time based on the facial shape parameters, or the avatar's expression can be adjusted based on the expression parameters, or both can be adjusted based on the facial shape and expression parameters. By adjusting the avatar's facial shape and expression in real time, the avatar can be kept consistent with the user's real face, thereby enhancing the authenticity of the avatar.

[0080] In an exemplary embodiment, Figure 6 As shown, in the above step S330, extracting hand point cloud data from the real image may include the following steps:

[0081] Step S610, extracting feature vectors from the real image, and generating a confidence map of the hand joint points based on the feature vectors;

[0082] Step S620, optimizing the initial hand shape parameters and the initial hand posture parameters using the confidence map to obtain target hand shape parameters and target hand posture parameters;

[0083] Step S630 , generating hand point cloud data according to the target hand shape parameters and the target hand posture parameters.

[0084] Among them, the hand joints refer to key points that can represent the structure or characteristics of the hand, such as the key points of the finger joints, the key points of the arm or elbow joints, etc. This exemplary embodiment can pre-train a neural network, take the real image as input data, and process it through the neural network to extract high-dimensional feature vectors that can represent the hand joints from the real image. The feature vector can be used as the initial vector of the hand joints, and then processed by another neural network and a fully connected network to generate a confidence map of the hand joints. The confidence map of the hand joints refers to an image that can represent the position and association relationship of the hand joints, which can include the result of whether each hand joint is detected, the position information of the detected hand joints, and the connection relationship or connection probability between the detected hand joints and other hand joints.

[0085] Hand shape parameters refer to data that can reflect the fatness, size, and shape of the hand, and posture parameters refer to parameters that can reflect the state of the hand, such as clenching a fist, spreading out, bending some fingers, and other postures. This exemplary embodiment can predetermine an initial hand shape parameter and an initial hand posture parameter, and then optimize the initial hand shape parameter and the initial hand posture parameter based on the determined confidence map to obtain the final target hand shape parameter and target hand posture parameter. Finally, the hand point cloud data is determined based on the target hand shape parameter and the target hand posture parameter through a specific algorithm or model. For example, the target hand shape parameter and the target hand posture parameter can be processed using a MANO (parametric model for the hand) model to estimate the hand posture and determine the hand point cloud data.

[0086] In an exemplary embodiment, in step S610, extracting a feature vector from a real image may include:

[0087] The real image is processed using the convolutional layer in the residual neural network to extract the feature vector.

[0088] Figure 7A schematic diagram of the architecture of the hand point cloud data generation network model in this exemplary embodiment is shown. After acquiring the real image 710, the real image 710 can be convolved by the convolution layer in the residual neural network 720, for example, the real image can be convolved by the convolution layer in the ResNet (Deep residual network) network to extract the feature vectors of the hand joint points in the real image.

[0089] It should be noted that, in addition to the above-mentioned ResNet network, VGG (Visual Geometry Group, Computer Vision Group) network, MobileNet network and EfficientNet network can also be used to extract the feature vectors of the hand joint points. This disclosure does not make specific limitations on this.

[0090] In an exemplary embodiment, in step S610, generating a confidence map of the hand joint points based on the feature vector may include:

[0091] The feature vector is processed using convolutional neural network and fully connected network to obtain the confidence map of the hand joint points.

[0092] This exemplary embodiment can process the feature vectors through a multi-layer convolutional neural network and a fully connected network to determine the confidence map of the hand joint points. Figure 7 As shown, the feature vectors of the hand joints can be processed based on the first convolutional neural network 730 with a 5-layer structure and the first fully connected network 740 with a 2-layer structure to determine the connection relationship of each hand joint point. Then, the connection relationship of the joint points is processed by the second convolutional neural network 750 with a 5-layer structure and the second fully connected network 760 with a 2-layer structure to obtain a confidence map of the hand joint points. Among them, the first convolutional neural network 730 and the second convolutional neural network 750 can be the same or different, and the first fully connected network 740 and the second fully connected network 760 can be the same or different. The specific network structure can be customized according to actual needs, and this disclosure does not make specific limitations on this.

[0093] In an exemplary embodiment, Figure 8 As shown, the above step S620 may include the following steps:

[0094] Step S810, using the initial hand shape parameters as current hand shape parameters, and using the initial hand posture parameters as current hand posture parameters;

[0095] Step S820, executing an iterative process of inputting the confidence map, the current hand shape parameters, and the current hand posture parameters into a fully connected network to update the current hand shape parameters and the current hand posture parameters;

[0096] Step S830: When the iterative process meets the preset conditions, the current hand shape parameters are used as target hand shape parameters, and the current hand posture parameters are used as target hand posture parameters.

[0097] Furthermore, the above step S820 may include the following steps:

[0098] Input the confidence map, current hand shape parameters and current hand posture parameters into the fully connected network to output the state increment;

[0099] Update the current hand shape parameters and current hand posture parameters according to the state increment.

[0100] This exemplary embodiment can input the hand joint confidence map and the set initial hand shape parameters and initial hand posture parameters as input data into a fully connected network for processing, so as to iterate the current hand shape parameters and the current hand posture parameters to obtain the target hand shape parameters and the target hand posture parameters. Specifically, the confidence map, the current hand shape parameters and the current hand posture parameters can be input into the fully connected network, and the state increments of the hand posture parameters and the hand shape parameters can be output first, wherein the fully connected network can be composed of a three-layer fully connected layer network for outputting the state increment. This exemplary embodiment can adopt the network structure of a three-layer Encode-fc3 network, and update the current hand shape parameters and the current hand posture parameters according to the state increment output by the final network until the preset conditions of the iterative process are met, that is, the current hand shape parameters can be used as the target hand shape parameters, and the current hand posture parameters can be used as the target hand posture parameters. Among them, the preset conditions of the iterative process refer to the judgment conditions for whether to continue to execute the update of the current hand shape parameters and the current hand posture parameters. When this exemplary embodiment adopts three consecutive incremental networks to implement the iterative process, the preset conditions are to complete three iterative calculations. When the three fully connected layers respectively complete the output of three state increments, and update the current hand shape parameters and the current hand posture parameters according to the third output state increment, the final target hand shape parameters and target hand posture parameters can be determined. It should be noted that this exemplary embodiment can also use a non-iterative network to determine the target hand shape parameters and target hand posture parameters, and the preset conditions can also be other conditions. For example, when the error is less than a preset threshold or the state increment reaches a certain level, it can be considered that the preset conditions are met, etc. This disclosure does not make specific limitations on this.

[0101] by Figure 7The network shown gives an example, in which the joint confidence map, the current hand shape parameters and the current hand posture parameters ω0 can be input into a fully connected network 770 comprising three fully connected layers, and the fully connected network can include a first fully connected layer 771, a second fully connected layer 772 and a third fully connected layer 773. Since the hand posture parameters have a high degree of nonlinearity, each fully connected layer in Encode_fc3 outputs a state increment. When the first fully connected layer 771 outputs a state increment △ω0, the current hand shape parameters and the current hand posture parameters ω1=ω0+△ω0 can be updated according to the state increment; when the second fully connected layer 772 outputs a state increment △ω1, the current hand shape parameters and the current hand posture parameters ω2=ω1+△ω1 can be updated according to the state increment; when the third fully connected layer 773 outputs a state increment △ω2, the current hand shape parameters and the current hand posture parameters ω3=ω2+△ω2 can be updated according to the state increment, and the current hand shape parameters are used as the target hand shape parameters, and the current hand posture parameters are used as the target hand posture parameters to determine the final target hand shape parameters and target hand posture parameters 780. In addition, during the network training process, in order to increase the generalization ability of the network, this exemplary embodiment can also add a Dropout layer between the fully connected layers. The specific network structure can be adjusted and set according to actual needs, and this disclosure does not make specific limitations on this.

[0102] Figure 9 A flowchart of another virtual image interaction method in this exemplary embodiment is shown, which may specifically include the following steps:

[0103] Step S910, obtaining a real image;

[0104] Step S920, extracting facial feature points from the real image;

[0105] Step S930, determining facial shape parameters and expression parameters by optimizing the cost function;

[0106] Step S940, generating face point cloud data based on the face shape parameters and expression parameters, and determining the hand area based on the hand point cloud data;

[0107] Step S950, adjusting the face shape of the avatar according to the face shape parameters, and / or adjusting the expression of the avatar according to the expression parameters, so as to drive the display of the corresponding avatar;

[0108] Step S960: extracting hand point cloud data from the real image, and determining the hand area based on the hand point cloud data;

[0109] Step S970, performing collision detection on the face area and the hand area;

[0110] Step S980: Adjust the virtual image's shape according to the collision result between the face area and the hand area.

[0111] It should be noted that the above steps are merely illustrative, and the present disclosure does not specifically limit the order of the above steps. For example, the process of determining the hand area in steps S920 to S940 can be performed simultaneously with the process of determining the hand area in step S960. In addition, after determining the facial shape parameters and expression parameters, step S950 can be executed immediately or periodically as needed to update the virtual image's facial shape and / or expression. In addition, to ensure the real-time performance of the overall algorithm, this exemplary embodiment can run the process of generating facial point cloud data on a CPU (central processing unit), and the process of generating hand point cloud data on a DSP (digital signal processor), etc.

[0112] The exemplary embodiment of the present disclosure also provides a virtual image interaction device. Figure 10 As shown, the virtual image interaction device 1000 may include: a real image acquisition module 1010, used to acquire a real image; a first area determination module 1020, used to extract face point cloud data from the real image, and determine the face area based on the face point cloud data; a second area determination module 1030, used to extract hand point cloud data from the real image, and determine the hand area based on the hand point cloud data; a virtual image adjustment module 1040, used to respond to a collision between the face area and the hand area, and perform deformation adjustment on the virtual image corresponding to the real image based on the collision result.

[0113] In an exemplary embodiment, the first area determination module includes: a facial feature point extraction unit, used to extract facial feature points from a real image; a cost function optimization unit, used to determine the cost function by reducing the facial feature points to a facial shape feature space and an expression feature space, and to determine the facial shape parameters and expression parameters by optimizing the cost function; and a facial point cloud data generation unit, used to generate facial point cloud data based on the facial shape parameters and expression parameters.

[0114] In an exemplary embodiment, the facial feature point extraction unit includes: a reference facial feature point acquisition subunit, used to obtain reference facial feature points in a reference image of the real image; a confidence region determination subunit, used to determine a confidence region in the real image based on the position of the reference facial feature points in the reference image and the motion information between the reference image and the real image; and a facial feature point determination subunit, used to extract facial feature points within the confidence region.

[0115] In an exemplary embodiment, the avatar interaction device further includes: a face shape and expression adjustment module for adjusting the face shape of the avatar according to the face shape parameters, and / or adjusting the expression of the avatar according to the expression parameters.

[0116] In an exemplary embodiment, the second area determination module includes: a confidence map generation unit, which is used to extract feature vectors from real images and generate a confidence map of hand joint points based on the feature vectors; a parameter optimization unit, which is used to optimize the initial hand shape parameters and the initial hand posture parameters using the confidence map to obtain target hand shape parameters and target hand posture parameters; and a hand point cloud data generation unit, which is used to generate hand point cloud data based on the target hand shape parameters and the target hand posture parameters.

[0117] In an exemplary embodiment, the confidence map generation unit includes: a vector extraction subunit, configured to process the real image using a convolutional layer in a residual neural network to extract a feature vector.

[0118] In an exemplary embodiment, the confidence map generation unit includes: a vector processing subunit, which is used to process the feature vector using a convolutional neural network and a fully connected network to obtain a confidence map of the hand joint points.

[0119] In an exemplary embodiment, the parameter optimization unit includes: an initial parameter determination subunit, which is used to use the initial hand shape parameters as the current hand shape parameters and the initial hand posture parameters as the current hand posture parameters; an iterative process updating subunit, which is used to execute an iterative process of inputting the confidence map, the current hand shape parameters and the current hand posture parameters into a fully connected network to update the current hand shape parameters and the current hand posture parameters; and a target parameter determination subunit, which is used to use the current hand shape parameters as the target hand shape parameters and the current hand posture parameters as the target hand posture parameters when the iterative process meets preset conditions.

[0120] In an exemplary embodiment, the iterative process update subunit includes: a state increment acquisition subunit, which is used to input the confidence map, current hand shape parameters and current hand posture parameters into a fully connected network to output a state increment; and a parameter update subunit, which is used to update the current hand shape parameters and current hand posture parameters according to the state increment.

[0121] In an exemplary embodiment, the number of the above-mentioned real images is multiple frames; the virtual image adjustment module includes: a collision motion information determination unit, which is used to determine the collision motion information based on the collision detection results in each frame of the real image, and the collision motion information includes at least one of the collision speed, collision acceleration, and collision position; a virtual image deformation adjustment subunit, which is used to perform deformation adjustment on the virtual image corresponding to the real image according to the collision motion information.

[0122] The specific details of each part of the above device have been described in detail in the method part of the implementation method, so they will not be repeated here.

[0123] The exemplary embodiments of the present disclosure further provide a computer-readable storage medium, which can be implemented in the form of a program product, including program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above. For example, Figure 3 、 Figure 5 、 Figure 6 、 Figure 8 or Figure 9 The program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0124] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0125] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0126] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0127] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0128] It will be appreciated by those skilled in the art that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented as the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which may be collectively referred to herein as a "circuit", "module" or "system". Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to encompass any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and implementation are intended to be exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0129] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A virtual image interaction method, characterized in that: include: Get real images; Extracting facial point cloud data from the real image, and determining a facial area based on the facial point cloud data; Extracting hand point cloud data from the real image, and determining a hand area based on the hand point cloud data; In response to a collision between the face region and the hand region, performing deformation adjustment on the virtual image corresponding to the real image based on a collision result; The number of the real images is multiple frames; The deforming and adjusting the virtual image corresponding to the real image based on the collision result includes: Determining collision motion information based on a collision result in each frame of real image, wherein the collision motion information includes at least one of a collision speed, a collision acceleration, and a collision position; The virtual image corresponding to the real image is deformed and adjusted according to the collision motion information.

2. The method according to claim 1, characterized in that The extracting of face point cloud data from the real image includes: Extracting facial feature points from the real image; Determining a cost function by reducing the dimension of the facial feature points to a face shape feature space and an expression feature space, and determining face shape parameters and expression parameters by optimizing the cost function; Generate facial point cloud data based on the facial shape parameters and the expression parameters.

3. The method according to claim 2, characterized in that The extracting facial feature points from the real image includes: Obtaining reference facial feature points in a reference image of the real image; Determining a confidence region in the real image according to positions of the reference facial feature points in the reference image and motion information between the reference image and the real image; The facial feature points are extracted within the confidence region.

4. The method according to claim 2, characterized in that After determining the facial shape parameters and the expression parameters, the method further includes: The face shape of the virtual image is adjusted according to the face shape parameters, and / or the expression of the virtual image is adjusted according to the expression parameters.

5. The method according to claim 1, characterized in that The extracting hand point cloud data from the real image includes: Extracting feature vectors from the real image, and generating a confidence map of hand joint points based on the feature vectors; Optimizing initial hand shape parameters and initial hand posture parameters using the confidence map to obtain target hand shape parameters and target hand posture parameters; Hand point cloud data is generated according to the target hand shape parameters and the target hand posture parameters.

6. The method according to claim 5, characterized in that The extracting a feature vector from the real image comprises: The real image is processed using a convolutional layer in a residual neural network to extract the feature vector.

7. The method according to claim 5, characterized in that Generating a confidence map of the hand joint points according to the feature vector includes: The feature vector is processed using a convolutional neural network and a fully connected network to obtain a confidence map of the hand joint points.

8. The method according to claim 5, characterized in that The step of optimizing the initial hand shape parameters and the initial hand posture parameters using the confidence map to obtain target hand shape parameters and target hand posture parameters includes: The initial hand shape parameters are used as the current hand shape parameters, and the initial hand posture parameters are used as the current hand posture parameters; executing an iterative process of inputting the confidence map, the current hand shape parameters, and the current hand posture parameters into a fully connected network to update the current hand shape parameters and the current hand posture parameters; When the iterative process meets the preset conditions, the current hand shape parameters are used as the target hand shape parameters, and the current hand posture parameters are used as the target hand posture parameters.

9. The method according to claim 8, characterized in that The executing an iterative process of inputting the confidence map, the current hand shape parameters, and the current hand posture parameters into a fully connected network to update the current hand shape parameters and the current hand posture parameters includes: Inputting the confidence map, current hand shape parameters, and current hand posture parameters into a fully connected network to output a state increment; The current hand shape parameters and the current hand posture parameters are updated according to the state increment.

10. A virtual image interaction device, characterized in that: include: A real image acquisition module, used to acquire real images; A first region determination module is configured to extract facial point cloud data from the real image and determine a facial region based on the facial point cloud data; a second region determination module, configured to extract hand point cloud data from the real image and determine a hand region based on the hand point cloud data; a virtual image adjustment module, configured to, in response to a collision between the face region and the hand region, perform deformation adjustment on the virtual image corresponding to the real image based on a collision result; The number of the real images is multiple frames; and the deformation adjustment of the virtual image corresponding to the real image based on the collision result is configured as follows: Determining collision motion information based on a collision result in each frame of real image, wherein the collision motion information includes at least one of a collision speed, a collision acceleration, and a collision position; The virtual image corresponding to the real image is deformed and adjusted according to the collision motion information.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. An electronic device, characterized in that: include: processor; a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 9 by executing the executable instructions.

Citation Information

Patent Citations

  • Simulation writing brush force measuring device based on micro electromechanical system (MEMS)

    CN103605431A

  • Method for rapidly monitoring high face rockfill dam extrusion sidewall deformation

    CN106225707A

  • Virtual figure expression driving method and system

    CN107945255A