Video image processing method and device, electronic equipment and storage medium

CN116630487BActive Publication Date: 2026-09-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210126470.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-10
Publication Date
2026-09-04
Estimated Expiration
2042-02-10

AI Technical Summary

Technical Problem

[0003]目前,视频拍摄过程中,通过控制虚拟对象进行特效显示变得越来越普遍,然而,现有的特效显示技术,仅能在视频拍摄过程中,显示单一动画特效,从而导致呈现出的特效显示效果存在一定的局限性

Benefits of technology

[0018] The technical solution of this disclosure responds to special effect triggering operations, displays a target virtual object model, and acquires an image to be processed including the target object. It then determines the facial image in the image to be processed, enabling the determination of at least one overlay animation effect based on the facial image. Furthermore, the overlay animation effect is overlaid onto the target virtual object model, ultimately resulting in a target video frame for display. This solves the problem in existing video image processing technologies where only a single animation effect can be triggered, and only one effect can be selected for playback during the effect playback process. It enables multiple animation effects to be played simultaneously, enriching the effect display. Moreover, determining the subsequent overlay animation effect based on the target object's facial image not only enhances the richness and interest of the video image but also strengthens the interactive effect with the user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630487B_ABST
    Figure CN116630487B_ABST
Patent Text Reader

Abstract

The method comprises: in response to a special effect triggering operation, displaying a target virtual object model and collecting a to-be-processed image comprising a target object; wherein the target virtual object model is played according to a pre-set basic animation special effect; determining at least one superimposed animation special effect triggered according to a face image in the to-be-processed image; superimposing the at least one superimposed animation special effect on the target virtual object model to obtain a target video frame and displaying the target video frame. The technical scheme of the embodiment of the present disclosure realizes determination of a subsequent superimposed animation special effect according to facial expression changes of a user, and can simultaneously play multiple animation special effects, thereby not only enriching a special effect display effect but also enhancing an interaction effect with the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to a video image processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of internet technology, more and more applications have entered users' lives, especially a series of software that can shoot short videos, which are very popular among users.

[0003] Currently, it is becoming increasingly common to control virtual objects to display special effects during video shooting. However, existing special effects display technologies can only display single animation effects during video shooting, which results in certain limitations in the presented special effects display effects. Summary of the Invention

[0004] This disclosure provides a video image processing method, apparatus, electronic device, and storage medium to achieve the superimposed playback of multiple animation effects.

[0005] In a first aspect, embodiments of this disclosure provide a video image processing method, including:

[0006] In response to a special effects trigger, the target virtual object model is displayed, and an image to be processed, including the target object, is acquired; the target virtual object model plays according to pre-set basic animation effects.

[0007] Based on the facial image in the image to be processed, determine at least one overlay animation effect to be triggered;

[0008] Overlay at least one overlay animation effect onto the target virtual object model to obtain the target video frame and display it.

[0009] Secondly, embodiments of the present invention also provide a video image processing apparatus, the apparatus comprising:

[0010] The image acquisition module is used to display the target virtual object model in response to special effects triggering operations, and to acquire the image to be processed, including the target object; wherein, the target virtual object model plays according to the pre-set basic animation effects;

[0011] The overlay animation effect determination module is used to determine at least one overlay animation effect to be triggered based on the facial image in the image to be processed;

[0012] The target video frame display module is used to overlay at least one overlay animation effect onto the target virtual object model to obtain and display the target video frame.

[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0014] One or more processors;

[0015] Storage device for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the video image processing method as described in any of the embodiments of this disclosure.

[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video image processing method as described in any of the embodiments of this disclosure.

[0018] The technical solution of this disclosure responds to special effect triggering operations, displays a target virtual object model, and acquires an image to be processed including the target object. It then determines the facial image in the image to be processed, enabling the determination of at least one overlay animation effect based on the facial image. Furthermore, the overlay animation effect is overlaid onto the target virtual object model, ultimately resulting in a target video frame for display. This solves the problem in existing video image processing technologies where only a single animation effect can be triggered, and only one effect can be selected for playback during the effect playback process. It enables multiple animation effects to be played simultaneously, enriching the effect display. Moreover, determining the subsequent overlay animation effect based on the target object's facial image not only enhances the richness and interest of the video image but also strengthens the interactive effect with the user. Attached Figure Description

[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0020] Figure 1 This is a schematic flowchart of a video image processing method provided in Embodiment 1 of this disclosure;

[0021] Figure 2 This is a schematic flowchart of a video image processing method provided in Embodiment 2 of this disclosure;

[0022] Figure 3 This is a schematic flowchart of a video image processing method provided in Embodiment 3 of this disclosure;

[0023] Figure 4 This is a schematic diagram of the structure of a video image processing apparatus provided in Embodiment 4 of this disclosure;

[0024] Figure 5 This is a schematic diagram of an electronic device structure provided in Embodiment 5 of this disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] Before introducing this technical solution, an illustrative example of its application scenarios can be provided. This disclosed technical solution can be applied to any scenario requiring special effects display or processing. For example, it can be applied during video shooting to apply special effects to the subject being filmed, resulting in a target special effects image for display. It can also be applied during still image shooting, for instance, when images are captured using a terminal device's built-in camera and then processed into special effects images for display. In this embodiment, the added special effects can be vertical jump, left punch, or overlay effects, etc. In this implementation, the target object can be a user or various animals captured in the video.

[0031] Example 1

[0032] Figure 1 This is a schematic flowchart of a video image processing method provided in Embodiment 1 of this disclosure. This embodiment is applicable to any scenario supported by the Internet for displaying or processing special effects, and is used to achieve simultaneous playback of multiple animation effects. The method can be executed by a video image processing device, which can be implemented in the form of software and / or hardware, or optionally by an electronic device, such as a mobile terminal, a PC, or a server.

[0033] like Figure 1 As shown, the method includes:

[0034] S110, In response to the special effect trigger operation, display the target virtual object model and acquire the image to be processed, including the target object.

[0035] It should be noted that the apparatus for executing the video image processing method provided in this embodiment can be integrated into application software that supports video image processing functions, and this software can be installed on an electronic device, optionally a mobile terminal or a PC. The application software can be a type of software for image / video processing; specific application software will not be detailed here, as long as it can achieve image / video processing. Alternatively, it can be a specially developed application program for adding and displaying special effects, or it can be integrated into a corresponding page, allowing users to add special effects through the integrated page on a PC.

[0036] The target virtual object model can be an animated model displayed on the interface, waiting to be controlled to perform a certain action. Basic animation effects can be pre-set for each virtual object model. These basic animation effects are pre-defined primitive animation effects; for example, the primitive animation effect could be at least one of dancing, running, or walking. The basic animation effects can change depending on the animation scene in which the target virtual object model is located, and the target virtual object model will then play according to the pre-set basic animation effects. For example, when the animation scene is a stage scene, the basic animation effect could be dancing, and the target virtual object model could be a cartoon character model that is dancing.

[0037] The image to be processed can be any image that needs to be processed. This image can be captured by a terminal device. Terminal devices can refer to electronic products with image capture capabilities, such as cameras, smartphones, and tablets. In practical applications, terminal devices are equipped with front-facing cameras, rear-facing cameras, or other camera devices. Correspondingly, the shooting methods can include selfies and still images. When a special effects operation is triggered, and depending on the user's selected shooting method (e.g., selfie), the presence of a target object within the field of view can be detected. When a target object is detected within the field of view of the terminal device, a video frame image from the current terminal device can be captured as the image to be processed. During image capture, if the target object is not detected in the video frame image captured by the current terminal device, no further processing will be performed. Alternatively, if the target object in the image to be processed is stationary, the virtual target object model will continuously play according to pre-set basic animation effects until a change in the target object is detected. Correspondingly, the target object can be any object in the frame whose posture or position information can change, such as a user or a pet.

[0038] It should be noted that when acquiring images containing target objects, the video frames corresponding to the captured video can be processed. For example, a target object corresponding to the captured video can be preset. When the image corresponding to the video frame is detected to contain the target object, the image corresponding to that video frame can be used as the image to be processed, so that the target object in each video frame can be tracked and processed with special effects.

[0039] It should also be noted that the number of target objects in the same shooting scene can be one or more. Whether it is one or more, the technical solution provided in this disclosure can be used to determine the special effects display video image.

[0040] In practical applications, the acquisition of images of the target object is usually initiated only when certain special effects are triggered. These special effects triggering operations include at least one of the following: triggering special effects props corresponding to the target virtual object model; or detecting facial images within the field of view.

[0041] One approach involves pre-setting controls to trigger special effects props. When a user triggers a control, a special effects prop display page pops up on the screen, showing multiple special effects props. The user can trigger the special effects prop corresponding to the target animation. If the special effects prop corresponding to the target virtual animation model is triggered, it indicates that a special effects trigger operation has been performed. Another implementation method is that the terminal device's camera has a certain field of view. When the facial image of the target object is detected within the field of view, it indicates that a special effects trigger operation has been performed. For example, a user can be pre-set as the target object. When the facial image of the user is detected within the field of view, it can be determined that a special effects trigger operation has been performed. Alternatively, the facial image of the target object can be pre-stored in the terminal device. When several facial images are detected within the field of view, if the facial image of the preset target object is detected among these facial images, it can be determined that a special effects trigger operation has been performed. This allows the terminal device to track the facial image of the target object and further acquire the image to be processed, including the target object.

[0042] S120. Based on the facial image in the image to be processed, determine at least one overlay animation effect to be triggered.

[0043] Understandably, the triggered overlay animation effects can be determined based on the facial image in the image to be processed. Accordingly, when the state information of the facial features of the target object in the image to be processed changes, different animation effects may be triggered. For example, when the target object's mouth is detected to be open, the triggered overlay animation effect could be jumping up and down.

[0044] The target virtual object model will play according to pre-set basic animation effects in different virtual scenes. If the facial features in the image to be processed change, the target virtual object model will add other animation effects on top of the original basic animation effects. These subsequently added animation effects can be overlaid as superimposed animation effects. At least one superimposed animation effect can be multiple animation effects simultaneously superimposed on the target virtual object model, such as jumping up and down, right punch, and left punch. If the facial features in the facial image do not change, the target virtual object model will continue to play according to the basic animation effects until a change in the facial features in a subsequently acquired image to be processed is detected, thus determining the triggered superimposed animation effect.

[0045] Specifically, a control for stopping shooting can be preset. When the user triggers an effect, the system can start processing each captured image and generate video frame images. When the stop shooting control is triggered, the system can generate the target video based on all previously generated video frame images.

[0046] S130. Overlay at least one overlay animation effect onto the target virtual object model to obtain the target video frame and display it.

[0047] As mentioned earlier, the target virtual object model will play according to the pre-set basic animation effects in different virtual scenes. After determining at least one superimposed animation effect corresponding to the target object's facial image, the determined superimposed animation effect can be superimposed with the basic animation effect of the target virtual object model, so that the target virtual object model can execute the basic animation effect while executing the superimposed animation effect, and use the currently displayed video frame image as the target video frame for display.

[0048] The technical solution of this disclosure responds to special effect triggering operations, displays a target virtual object model, and acquires an image to be processed including the target object. It then determines the facial image in the image to be processed, enabling the determination of at least one overlay animation effect based on the facial image. Furthermore, the overlay animation effect is overlaid onto the target virtual object model, ultimately resulting in a target video frame for display. This solves the problem in existing video image processing technologies where only a single animation effect can be triggered, and only one effect can be selected for playback during the effect playback process. It enables multiple animation effects to be played simultaneously, enriching the effect display. Moreover, determining the subsequent overlay animation effect based on the target object's facial image not only enhances the richness and interest of the video image but also strengthens the interactive effect with the user.

[0049] Example 2

[0050] Figure 2 This is a flowchart illustrating a video image processing method according to Embodiment 2 of this disclosure. Based on the foregoing embodiments, steps S110 and S120 are further refined. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0051] like Figure 2 As shown, the method specifically includes the following steps:

[0052] S210. In response to the special effect triggering operation, retrieve the target virtual object model corresponding to the special effect triggering operation, and control the target virtual object model to play according to the basic animation special effects.

[0053] In practical applications, different virtual object models can be pre-set. When a user triggers an effect, the system can retrieve the corresponding virtual object model based on the user's basic registration information to serve as the user's target virtual object model. The retrieved target virtual object model will then play according to pre-set basic animation effects. Alternatively, when a user triggers an effect, a special effects display page will pop up on the terminal device's screen. This page includes several virtual object models, allowing the user to choose according to their preferences. The selected virtual object model will then be used as the target virtual object model, and its corresponding basic animation effects will be played. The advantage of this setup is that it allows for personalized configuration based on user preferences, enhancing the interactive experience to some extent.

[0054] S220: Acquire images of the target object based on a camera device deployed on a terminal device.

[0055] Optionally, the terminal device can be a mobile terminal, such as a mobile phone or tablet computer, or a fixed terminal such as a PC. Correspondingly, the camera device deployed on the terminal device can be a built-in camera installed inside the terminal device, such as a front-facing camera or a rear-facing camera; or it can be an external camera on the terminal device, capable of 360° rotation, such as a rotating camera; or it can be other camera devices used to achieve image acquisition functions, etc., and this disclosure does not specifically limit these embodiments.

[0056] Optionally, based on the acquisition of the image to be processed by the camera device, an input device such as a touch screen or physical button in the terminal device can be used to input a camera device start command to control the camera device on the terminal device to enter video image shooting mode and acquire the image to be processed; or, a camera device start control can be pre-set in the terminal device, and when the user triggers the control, the corresponding camera device can be turned on and the image to be processed can be acquired; or, the video image shooting mode of the camera device can be started in other ways to realize the acquisition function of the image to be processed, etc., and the embodiments of this disclosure do not specifically limit this.

[0057] Specifically, when a user triggers a special effect, the system retrieves the target virtual object model corresponding to the special effect trigger and captures the image to be processed, including the target object, through a camera device on the terminal device, so that subsequent operations can be performed on the acquired image to be processed.

[0058] S230. Use the target virtual object model as the foreground image and the image to be processed as the background image.

[0059] In practical applications, during the image processing, to allow users to more clearly capture the special effects performed by the target virtual object model, the target virtual object model can be used as the foreground image, and the image to be processed as the background image. The advantage of this setup is that it allows users to clearly understand the corresponding overlay animation effects triggered when the state information of different parts of the face in the image to be processed changes, as well as the effect display of the target virtual object model. It also gives users a more immersive experience when using the video image processing application, enhancing their sense of participation.

[0060] S240. Based on the image segmentation model, determine the facial images in the image to be processed.

[0061] Generally, the target virtual object model needs to execute corresponding special effects actions based on the facial expression changes of the target object in the image to be processed. Accordingly, the facial expression changes of the target object need to be determined based on the facial image in the image to be processed. Based on this, the facial image in the image to be processed can be determined based on an image segmentation model. Optionally, the facial features include at least two of the following: the left eye, right eye, left eyebrow, right eyebrow, nose, and mouth.

[0062] The image segmentation model can be a pre-trained neural network model used to segment the target image. Optionally, the image segmentation model can be composed of at least one network structure selected from convolutional neural networks, recurrent neural networks, and deep neural networks; this disclosure does not specifically limit this aspect.

[0063] In this embodiment, the image segmentation model can be trained based on the sample image to be processed and the facial region annotation image. The facial region annotation image can be a ground truth image, which can be used as a basis for evaluating subsequent prediction results. The specific training process of the image segmentation model can be as follows: acquire the sample image set to be processed, input the sample image set to be processed into the image segmentation model to be trained, output the initial training result, determine the loss result based on the initial training result and the facial region annotation image, and adjust the model parameters in the image segmentation model to be trained based on the loss result and the preset loss function corresponding to the image segmentation model to be trained, obtaining the corresponding adjustment result. In this embodiment, the convergence of the preset loss function corresponding to the image segmentation model to be trained can be used as the training objective. Based on this, it can be understood that when it is determined that the preset loss function has not converged, it indicates that the adjustment result does not meet the requirements of model training, and it is necessary to continue inputting the sample image set to be processed to train the model. When it is determined that the preset loss function has converged, it indicates that the adjustment result meets the requirements of model training, and the trained image segmentation model is obtained.

[0064] S250. Determine multiple key points to be processed for at least one part of the facial image, and determine the trigger parameters for at least one part of the facial image based on the multiple key points to be processed.

[0065] The "at least one part" can include one or more parts. For more precise control, the "at least one part" can be multiple parts in the facial image. Optionally, when determining the key points to be processed for one part, only the key point information of certain specific parts can be focused on, such as the eyes, eyebrows, or mouth; when determining the key points to be processed for multiple parts, different parts correspond to different trigger parameters, and the key point information to be processed for different parts can be determined separately, and the corresponding trigger parameters can be determined based on the key point information to be processed.

[0066] In some implementations, to determine the change information of each part of the target object's facial image in the image to be processed, it is necessary to determine the key point information around each part of the facial image. These key points can be used as multiple key points to be processed. By determining the coordinate changes of these multiple key points, the change information of each part of the facial image can be determined, so that trigger parameters for each part of the facial image can be determined based on the key point information. The trigger parameters can be parameter information corresponding to different animation effects triggered by different movements of the key points to be processed in each part. Optionally, the trigger parameters include overlay animation effect parameters corresponding to each part.

[0067] Optionally, multiple key points to be processed in at least one part of the facial image are determined, and triggering parameters for at least one part of the facial image are determined based on the multiple key points to be processed, including: determining multiple key points to be processed in at least one part of the facial image based on a key point recognition algorithm; determining feature information of at least one part by processing the key points to be processed in at least one part; and determining corresponding triggering parameters based on each feature information.

[0068] The keypoint recognition algorithm can be a pre-set algorithm used to identify keypoints around various parts of a facial image. Based on the displacement changes of different parts of the facial image, the keypoint recognition algorithm identifies keypoints around those parts that are undergoing relative changes and identifies them as multiple keypoints to be processed for that part. For example, when a target object opens or closes its eyes, the relative positions of multiple keypoints around the eyes in its facial image will change. Through the keypoint recognition algorithm, these multiple keypoints around the eyes can be identified as keypoints to be processed, so that the relative changes of the corresponding parts can be determined by calculating the coordinate information of these keypoints.

[0069] The feature information of each part can be used to display the current state of each part. For example, for the eyes, the feature information can be whether they are open or closed; for the mouth, the feature information can be whether it is open or closed, and can also include information on the degree of opening.

[0070] In practical applications, after identifying multiple key points to be processed in at least one part of a facial image, the changes in the feature information of that part can be calculated based on the changes in the positional information of the key points. For example, for the eyes, the positional changes of the upper and lower eyelids can be detected; if they are closer together, it can be determined that the eyes were closed. For the eyebrows, the positional changes of the brow peak key point can be detected; if the position moves upward, it can be determined that the eyebrows are currently raised. For the mouth, the positional changes of the upper and lower lips can be detected; if the relative distance between the upper and lower lips increases, it can be determined that the mouth is currently open. Furthermore, the special effect parameters triggered by each part can be determined based on the feature information of each part.

[0071] In some implementations, by processing the key points of each part to be processed, the feature information of each part is determined. Movement information can be determined based on the positional changes of key points in two adjacent images to be processed. For example, a point among multiple key points corresponding to the eye can be used as a reference point. The positional information of this reference point in two adjacent images to be processed is determined, and the positional offset is determined according to the distance formula between the two points. This positional offset is then used as movement information. If the movement information meets a preset condition (optionally, the preset condition is the movement distance), its feature information can be determined so that the triggered special effect parameters can be determined based on the feature information.

[0072] Optionally, based on the feature information of at least one part, the corresponding triggering parameters are determined, including: determining the triggering parameters corresponding to each feature information according to a pre-established parameter mapping table.

[0073] In this embodiment, the trigger parameters corresponding to the feature information of different parts of the face image are different. Correspondingly, the trigger parameters corresponding to different feature information of the same part are also somewhat different. The trigger parameters can be used to characterize the state change information of a certain part of the face image. For example, when the part of the face image is the mouth, its corresponding trigger parameter can be the information of the degree of mouth opening; when the part is the eyebrows, its corresponding trigger parameter can be the information of the height of eyebrow raising, etc.

[0074] A pre-established correspondence between the feature information of each part and its corresponding trigger parameters can be created, and a corresponding parameter mapping table can be established based on this correspondence. The mapping table includes the trigger parameters corresponding to each feature, and these trigger parameters correspond to overlaid animation effects. Accordingly, based on the feature information of each part, the corresponding trigger parameters are determined, and thus the overlaid animation effects corresponding to these trigger parameters can be determined.

[0075] S260. Based on at least one trigger parameter, determine at least one superimposed animation effect.

[0076] Understandably, based on the changes in key points of various parts of the facial image, corresponding trigger parameters are determined. To enable the target virtual object model to execute corresponding animation effects according to the changes in various parts of the target object's facial image, at least one overlay animation effect can be determined based on each trigger parameter. It should be noted that the number of overlay animation effects can be determined based on the changes in various parts of the target object's facial image in the image to be processed. For example, when it is detected that the user is simultaneously opening their mouth and blinking to the left, the corresponding overlay animation effects could be two: a switching loop animation and a left fist swing.

[0077] Optionally, based on at least one trigger parameter, at least one overlay animation effect is determined, including: determining the corresponding overlay animation effect according to at least one trigger parameter, and determining the amplitude information and duration information of the overlay animation effect, so as to display the corresponding overlay animation effect based on the amplitude information and duration information.

[0078] The amplitude information of the overlay animation effect can be considered as the intensity information of the target virtual object model when performing the corresponding animation effect. The amplitude information of the overlay animation effect can correspond to the change amplitude of a corresponding part in the target object's facial image. When the change amplitude of a certain part in the target object's facial image increases, the amplitude information of the corresponding overlay animation effect can also increase, that is, the intensity information of the target virtual object model when performing the corresponding animation effect increases. For example, when the target object is in a state of open mouth, the corresponding overlay animation effect is a switching loop animation; as the target object's mouth opens wider, the speed of the switching loop animation can increase, and so on. The duration information of the overlay animation effect can be considered as the duration of the target virtual object model performing the corresponding animation effect. For example, when the overlay animation effect is a left punch, its duration information is the duration of the target virtual object performing the left punch action, and so on.

[0079] In practical applications, the correspondence between trigger parameters and overlay animation effects, as well as the correspondence between the amplitude information and duration of the overlay animation effects, can be established in advance. A corresponding mapping table can be established so that after determining each trigger parameter, the corresponding overlay animation effect can be determined based on each trigger parameter, and the amplitude information and duration information of the overlay animation effect during execution can be determined so that the display interface can display the corresponding overlay animation effect based on the amplitude information and duration information.

[0080] For example, when two or more parts of the target object's face image change, the corresponding overlay animation effects can be blended and superimposed onto the target virtual object model. Different blending ratios can be applied based on the magnitude of the changes in different parts, allowing the target virtual object model to play the desired animation according to the blending ratio. For instance, when it is detected that a user is simultaneously opening their mouth and blinking their left eye, the corresponding overlay animation effects are a switching loop animation and a left fist swing. The target virtual object model can execute both of these overlay animation effects simultaneously. Furthermore, when it is detected that the user's mouth opening becomes wider, the amplitude of the corresponding overlay animation effect will also increase accordingly, meaning the switching loop animation speed becomes faster. Alternatively, when it is detected that the user blinks their left eye for a longer period, the duration of the corresponding overlay animation effect will also increase accordingly, meaning the target virtual object model remains in a left fist swing state. Accordingly, the blending ratio of the overlay animation effects can be determined based on the changes in different parts of the target object's face image, allowing the target virtual object model to play different blended animation effects according to different blending ratios. The advantage of this setup is that it allows for the creation of a unique set of animation effects by freely blending changes in the target object's facial expressions, and the target virtual object model can be controlled to play the blended animation effects.

[0081] It should be noted that when the key point information of two or more parts of the target object's facial image changes, the two parts can be superimposed. That is, the target virtual object model can execute two or more superimposed animation effects at the same time.

[0082] S270. Overlay at least one overlay animation effect onto the target virtual object model to obtain the target video frame and display it.

[0083] The technical solution of this disclosure responds to special effect triggering operations, retrieves the target virtual object model corresponding to the special effect triggering operation, and controls the target virtual object model to play according to basic animation special effects. Simultaneously, it acquires an image to be processed including the target object based on a camera device on a terminal device. Then, based on an image segmentation model, it determines the facial image in the image to be processed and determines multiple key points to be processed for at least one part of the facial image. Based on the key points to be processed, it determines the trigger parameters for at least one part of the facial image, so that at least one superimposed animation special effect can be determined according to each trigger parameter. Furthermore, the superimposed animation special effect is superimposed onto the target virtual object model, ultimately obtaining and displaying the target video frame. This achieves the goal of controlling the corresponding virtual object model to simultaneously execute and display multiple special effect actions based on the facial expression changes of the target object in the image to be processed, enriching the special effect display effect. Moreover, by segmenting the facial image in the image to be processed based on the image segmentation model, it can more accurately capture changes in each part of the facial image, thereby achieving the effect of accurately triggering the corresponding animation special effects.

[0084] Example 3

[0085] Figure 3 This is a schematic flowchart of a video image processing method provided in Embodiment 3 of this disclosure. Based on the foregoing embodiments, S130 is further refined, and the specific implementation method can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0086] like Figure 3 As shown, the method specifically includes the following steps:

[0087] S310, in response to the special effect trigger operation, displays the target virtual object model and acquires the image to be processed, including the target object.

[0088] S320. Based on the facial image in the image to be processed, determine at least one overlay animation effect to be triggered.

[0089] S330. Overlay at least one superimposed animation effect with the basic animation effect of the target virtual object model to obtain the target virtual object model that performs the target effect, and display it.

[0090] In this embodiment, the target virtual object model will have its corresponding basic animation effects depending on the virtual scene in which it is located. Therefore, the target effects may include the basic animation effects of the target virtual object model and at least one superimposed animation effect.

[0091] Specifically, based on the changes in key points of various parts of the facial image in the image to be processed, at least one superimposed animation effect can be determined. The determined superimposed animation effect is superimposed with the current basic animation effect of the target virtual object model to obtain a target virtual object model that executes both the basic animation effect and the superimposed animation effect. The target video frame determined based on the current target effect display parameters can be displayed and played.

[0092] For example, the target video frame may include a target virtual object model for performing the target effect and a target object, where the target object is a background image and the target virtual object model is a foreground image.

[0093] Based on the above technical solution, when acquiring the image to be processed, including the target object, the method further includes: determining the relative position information between the target object and the camera device, so as to adjust the display position information of the target virtual object model in the target video frame based on the relative position.

[0094] Generally, when a camera on a terminal device captures an image of a target object, there is a certain distance between the target object and the camera. This distance information can be used as relative position information. When the target virtual object model is displayed on the terminal device's screen, its position changes based on the movement of the target object. Correspondingly, when capturing the image, the relative position information between the target object and the camera is determined simultaneously. Based on this relative position information, the display position of the target virtual object model in the target video frame is adjusted so that the image in the target video frame uses the target object as the background image, controlling the target virtual object model in the foreground image to execute corresponding animation effects.

[0095] The technical solution of this disclosure responds to special effect triggering operations, displays a target virtual object model, and acquires a to-be-processed image including the target object. Then, based on the facial image in the to-be-processed image, it determines at least one overlay animation effect to be triggered. Furthermore, it overlays the overlay animation effect with the basic animation effect of the target virtual object to obtain and display the target virtual object model that executes the target effect. This enables multiple animation effects to be played simultaneously in the same video frame image, enriching the special effect display effect.

[0096] Example 4

[0097] Figure 4 This is a schematic diagram of the structure of a video image processing apparatus provided in Embodiment 4 of this disclosure, as shown below. Figure 4As shown, the device includes: an image acquisition module 410, an overlay animation effect determination module 420, and a target video frame display module 430.

[0098] The image acquisition module 410 is used to display the target virtual object model and acquire the image to be processed, including the target object, in response to the special effect triggering operation. The target virtual object model plays according to the pre-set basic animation special effects. The superimposed animation special effect determination module 420 is used to determine at least one superimposed animation special effect triggered based on the facial image in the image to be processed. The target video frame display module 430 is used to superimpose at least one superimposed animation special effect on the target virtual object model to obtain the target video frame and display it.

[0099] Based on the above technical solutions, the image acquisition module 410 to be processed includes a special effects trigger operation setting unit.

[0100] The special effects triggering operation setting unit is used to trigger the special effects props corresponding to the target virtual object model; facial images are included in the detected field of view area.

[0101] Based on the above technical solutions, the image acquisition module 410 includes a virtual object model retrieval unit and an image acquisition unit.

[0102] The Unreal Object Model Retrieval Unit is used to retrieve the target virtual object model corresponding to the special effect triggering operation, and control the target virtual object model to play according to the basic animation special effects;

[0103] The image acquisition unit is used to acquire images of the target object based on a camera device deployed on a terminal device.

[0104] Based on the above technical solutions, after acquiring the image to be processed, including the target object, the device further includes a foreground image and background image determination module.

[0105] The foreground and background image determination module is used to use the target virtual object model as the foreground image and the image to be processed as the background image.

[0106] Based on the above technical solutions, the superimposed animation effect determination module 420 includes a facial image determination unit, a trigger parameter determination unit, and a superimposed animation effect determination unit.

[0107] The face image determination unit is used to determine the face image in the image to be processed based on the image segmentation model;

[0108] The trigger parameter determination unit is used to determine multiple key points to be processed in at least one part of the facial image, and to determine the trigger parameters of at least one part of the facial image based on the multiple key points to be processed.

[0109] The overlay animation effect determination unit is used to determine at least one overlay animation effect based on at least one trigger parameter.

[0110] Based on the above technical solutions, the trigger parameter determination unit includes a key point determination subunit, a feature information determination subunit, and a trigger parameter determination subunit.

[0111] The key point determination subunit is used to determine multiple key points to be processed in at least one part of a facial image based on a key point recognition algorithm.

[0112] The feature information determination subunit is used to determine the feature information of at least one part by processing the key points to be processed in at least one part.

[0113] The trigger parameter determination subunit is used to determine the corresponding trigger parameters based on the feature information of at least one part.

[0114] Based on the above technical solutions, the trigger parameter determination subunit is also used to determine the trigger parameters corresponding to the feature information of at least one part according to a pre-established parameter mapping relationship table; wherein, the mapping relationship table includes the trigger parameters corresponding to each feature information, and the trigger parameters correspond to the superimposed animation effects.

[0115] Based on the above technical solutions, the superimposed animation effect determination unit is also used to determine the corresponding superimposed animation effect according to at least one trigger parameter, and to determine the amplitude information and duration information of the superimposed animation effect, so as to display the corresponding superimposed animation effect based on the amplitude information and duration information.

[0116] Based on the above technical solutions, the facial image includes at least two of the following: the left eye, the right eye, the left eyebrow, the right eyebrow, the nose, and the mouth.

[0117] Correspondingly, the triggering parameters include the superimposed animation effect parameters corresponding to each part.

[0118] Based on the above technical solutions, the target video frame display module 430 is further used to overlay at least one superimposed animation effect with the basic animation effect of the target virtual object model to obtain the target virtual object model that performs the target effect, and then display it.

[0119] The target video frame includes a target virtual object model for performing the target effect and a target object; the target object is a background image, and the target virtual object model is a foreground image.

[0120] Based on the above technical solutions, the device further includes: a relative position information determination module.

[0121] The relative position information determination module is used to determine the relative position information between the target object and the camera device, so as to adjust the display position information of the target virtual object model in the target video frame based on the relative position.

[0122] The technical solution of this disclosure responds to special effect triggering operations, displays a target virtual object model, and acquires an image to be processed including the target object. It then determines the facial image in the image to be processed, enabling the determination of at least one overlay animation effect based on the facial image. Furthermore, the overlay animation effect is overlaid onto the target virtual object model, ultimately resulting in a target video frame for display. This solves the problem in existing video image processing technologies where only a single animation effect can be triggered, and only one effect can be selected for playback during the effect playback process. It enables multiple animation effects to be played simultaneously, enriching the effect display. Moreover, determining the subsequent overlay animation effect based on the target object's facial image not only enhances the richness and interest of the video image but also strengthens the interactive effect with the user.

[0123] The video image processing apparatus provided in this disclosure can execute the video image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0124] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0125] Example 5

[0126] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of this disclosure. Refer to the following... Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0127] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 506 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0128] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0130] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0131] The electronic device provided in this embodiment and the video image processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0132] Example 6

[0133] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video image processing method provided in the above embodiments.

[0134] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0136] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0137] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0138] In response to a special effects trigger operation, a target virtual object model is displayed, and an image to be processed, including the target object, is acquired; wherein the target virtual object model is played according to pre-set basic animation effects;

[0139] Based on the facial image in the image to be processed, determine at least one overlay animation effect to be triggered;

[0140] The at least one overlay animation effect is overlaid on the target virtual object model to obtain the target video frame and display it.

[0141] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0143] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0144] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0146] According to one or more embodiments of this disclosure, [Example 1] a video image processing method is provided, the method comprising:

[0147] In response to a special effects trigger, the target virtual object model is displayed, and an image to be processed, including the target object, is acquired; the target virtual object model is played according to pre-set basic animation effects.

[0148] Based on the facial image in the image to be processed, determine at least one overlay animation effect to be triggered;

[0149] Overlay at least one overlay animation effect onto the target virtual object model to obtain the target video frame and display it.

[0150] According to one or more embodiments of this disclosure, [Example 2] provides a video image processing method, which further includes:

[0151] Optionally, the special effects triggering operation includes at least one of the following:

[0152] Trigger the special effects props corresponding to the target virtual object model;

[0153] Facial images are included in the detected visual field area.

[0154] According to one or more embodiments of this disclosure, [Example 3] provides a video image processing method, which further includes:

[0155] Optionally, displaying the target virtual object model and acquiring the image to be processed, including the target object, includes:

[0156] Retrieve the target virtual object model corresponding to the special effect triggering operation, and control the target virtual object model to play according to the basic animation special effect;

[0157] Images of the target object are acquired using camera devices deployed on terminal devices.

[0158] According to one or more embodiments of this disclosure, [Example 4] provides a video image processing method, which, after acquiring the image to be processed including the target object, further includes:

[0159] Optionally, the target virtual object model can be used as the foreground image, and the image to be processed can be used as the background image.

[0160] According to one or more embodiments of this disclosure, [Example 5] provides a video image processing method, which further includes:

[0161] Optionally, determining at least one overlay animation effect to be triggered based on the facial image in the image to be processed includes:

[0162] Based on the image segmentation model, the facial image in the image to be processed is determined;

[0163] Determine multiple key points to be processed for at least one part of the facial image, and determine trigger parameters for at least one part of the facial image based on the multiple key points to be processed;

[0164] Based on at least one trigger parameter, determine at least one overlay animation effect.

[0165] According to one or more embodiments of this disclosure, [Example Six] provides a video image processing method, which further includes:

[0166] Optionally, determining multiple key points to be processed for at least one part of the facial image, and determining trigger parameters for at least one part of the facial image based on the multiple key points to be processed, includes:

[0167] Based on the key point recognition algorithm, multiple key points to be processed in at least one part of the facial image are determined;

[0168] By processing the key points of at least one part, the feature information of at least one part is determined.

[0169] Based on the feature information of at least one part, the corresponding triggering parameters are determined.

[0170] According to one or more embodiments of this disclosure, [Example Seven] provides a video image processing method, which further includes:

[0171] Optionally, determining the corresponding triggering parameters based on feature information of at least one part includes:

[0172] Based on a pre-established parameter mapping table, determine the trigger parameters corresponding to the feature information of at least one part;

[0173] The mapping table includes trigger parameters corresponding to each feature information, and the trigger parameters correspond to the superimposed animation effects.

[0174] According to one or more embodiments of this disclosure, [Example Eight] provides a video image processing method, which further includes:

[0175] Optionally, determining at least one overlay animation effect based on at least one trigger parameter includes:

[0176] Based on at least one trigger parameter, a corresponding overlay animation effect is determined, and the amplitude information and duration information of the overlay animation effect are determined, so as to display the corresponding overlay animation effect based on the amplitude information and duration information.

[0177] According to one or more embodiments of this disclosure, [Example Nine] provides a video image processing method, which further includes:

[0178] Optionally, the facial image includes at least two of the following: the left eye, the right eye, the left eyebrow, the right eyebrow, the nose, and the mouth.

[0179] Correspondingly, the triggering parameters include the superimposed animation effect parameters corresponding to each part.

[0180] According to one or more embodiments of this disclosure, [Example 10] provides a video image processing method, which further includes:

[0181] The step of overlaying at least one superimposed animation effect onto the target virtual object model to obtain a target video frame and displaying it includes:

[0182] Optionally, the at least one superimposed animation effect is superimposed on the basic animation effect of the target virtual object model to obtain a target virtual object model that performs the target effect, and then displayed.

[0183] The target video frame includes a target virtual object model for performing the target effect and a target object; the target object is a background image, and the target virtual object model is a foreground image.

[0184] According to one or more embodiments of this disclosure, [Example 11] provides a video image processing method, which further includes:

[0185] Optionally, when acquiring the image to be processed, which includes the target object, the following may also be included:

[0186] The relative position information between the target object and the camera device is determined so as to adjust the display position information of the target virtual object model in the target video frame based on the relative position.

[0187] According to one or more embodiments of this disclosure, [Example Twelve] provides a video image processing apparatus, the apparatus comprising:

[0188] The image acquisition module is used to display the target virtual object model and acquire the image to be processed, including the target object, in response to the special effect triggering operation; wherein the target virtual object model plays according to the pre-set basic animation special effects;

[0189] The overlay animation effect determination module is used to determine at least one overlay animation effect to be triggered based on the facial image in the image to be processed;

[0190] The target video frame display module is used to overlay the at least one overlay animation effect onto the target virtual object model to obtain and display the target video frame.

[0191] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0192] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0193] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video image processing method, characterized in that, include: In response to a special effect trigger operation, a target virtual object model is displayed, and an image to be processed, including the target object, is acquired; wherein, the target virtual object model is played according to a pre-set basic animation effect, and the target virtual object model is an animated model displayed on the display interface, waiting to be controlled to perform an action; Based on the facial image in the image to be processed, at least one overlay animation effect is determined to be triggered, wherein the number of overlay animation effects is determined according to the changes in various parts of the facial image of the target object in the image to be processed; The at least one overlay animation effect is overlaid on the target virtual object model to obtain the target video frame and display it.

2. The method according to claim 1, characterized in that, The special effects triggering operation includes at least one of the following: Trigger the special effects props corresponding to the target virtual object model; Facial images are included in the detected visual field area.

3. The method according to claim 1, characterized in that, The process of displaying the target virtual object model and acquiring the image to be processed, including the target object, includes: Retrieve the target virtual object model corresponding to the special effect triggering operation, and control the target virtual object model to play according to the basic animation special effect; Images of the target object are acquired using camera devices deployed on terminal devices.

4. The method according to claim 1, characterized in that, After acquiring the image to be processed, which includes the target object, the process further includes: The target virtual object model is used as the foreground image, and the image to be processed is used as the background image.

5. The method according to claim 1, characterized in that, The step of determining at least one overlay animation effect to be triggered based on the facial image in the image to be processed includes: Based on the image segmentation model, the facial image in the image to be processed is determined; Determine multiple key points to be processed for at least one part of the facial image, and determine trigger parameters for at least one part of the facial image based on the multiple key points to be processed; Based on at least one trigger parameter, determine at least one overlay animation effect.

6. The method according to claim 5, characterized in that, The step of determining multiple key points to be processed in at least one part of the facial image, and determining trigger parameters for at least one part of the facial image based on the multiple key points to be processed, includes: Based on the key point recognition algorithm, multiple key points to be processed in at least one part of the facial image are determined; By processing the key points of at least one part, the feature information of at least one part is determined. Based on the feature information of at least one part, the corresponding triggering parameters are determined.

7. The method according to claim 6, characterized in that, The determination of corresponding triggering parameters based on feature information of at least one part includes: Based on a pre-established parameter mapping table, determine the trigger parameters corresponding to the feature information of at least one part; The mapping table includes trigger parameters corresponding to each feature information, and the trigger parameters correspond to the superimposed animation effects.

8. The method according to claim 6, characterized in that, The determination of at least one superimposed animation effect based on at least one trigger parameter includes: Based on at least one trigger parameter, a corresponding overlay animation effect is determined, and the amplitude information and duration information of the overlay animation effect are determined, so as to display the corresponding overlay animation effect based on the amplitude information and duration information.

9. The method according to claim 6, characterized in that, The facial image includes at least two of the following: the left eye, the right eye, the left eyebrow, the right eyebrow, the nose, and the mouth. Correspondingly, the triggering parameters include the superimposed animation effect parameters corresponding to each part.

10. The method according to claim 1, characterized in that, The step of overlaying at least one superimposed animation effect onto the target virtual object model to obtain a target video frame and displaying it includes: The at least one superimposed animation effect is superimposed on the basic animation effect of the target virtual object model to obtain the target virtual object model that performs the target effect, and then displayed. The target video frame includes a target virtual object model for performing the target effect and a target object; the target object is a background image, and the target virtual object model is a foreground image.

11. The method according to claim 1, characterized in that, When acquiring an image to be processed, including the target object, the process also includes: The relative position information between the target object and the camera device is determined so as to adjust the display position information of the target virtual object model in the target video frame based on the relative position.

12. A video image processing apparatus, characterized in that, include: The image acquisition module is used to respond to special effect triggering operations, display the target virtual object model, and acquire the image to be processed, including the target object; wherein, the target virtual object model plays according to the pre-set basic animation effects, and the target virtual object model is an animated model displayed on the display interface, waiting to be controlled to perform actions; The overlay animation effect determination module is used to determine at least one overlay animation effect to be triggered based on the facial image in the image to be processed, wherein the number of overlay animation effects is determined according to the changes in various parts of the facial image of the target object in the image to be processed; The target video frame display module is used to overlay the at least one superimposed animation effect onto the target virtual object model to obtain and display the target video frame.

13. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video image processing method as described in any one of claims 1-11.

14. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video image processing method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Video record method and device, terminal and storage medium

    CN108833818A