Image processing method, device, storage medium and computer system

The image processing method recognizes the position change data of the face area, drives the special effects animation to match the anchor's movement, solves the problems of high hardware costs and complex operations in the existing technology, and improves the experience of running live broadcast.

CN113920167BActive Publication Date: 2025-09-05GUANGZHOU BOGUAN TELECOMM TECH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111284542.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-09-05
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

The existing technology relies on mobile phone gravity sensors for running detection in live running broadcasts, resulting in high hardware costs and complex operation processes, making it difficult to accurately match special effects animations and anchor sports.

Method used

The face area in the image frame sequence is identified by the image processing method, position change data is collected to determine the motion data, and the playback of the target special effect animation is driven to match it with the motion.

Benefits of technology

It reduces hardware costs, simplifies the operation process, and improves the accuracy and realism of matching special effects animations with face areas, improving the interactive fun of running live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920167B_ABST
    Figure CN113920167B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, device, storage medium and computer system, which relate to the field of computer technology. The method includes: responding to a trigger operation of adding a target special effect animation, obtaining an image frame sequence; identifying the face area corresponding to the target object in the image frame sequence; collecting position change data of the face area in a plurality of consecutive image frames, and determining the motion data of the target object based on the position change data; driving the playback of the target special effect animation based on the motion data, so that the played target special effect animation is adapted to the motion of the target object. The present disclosure can determine the motion data of the user object by detecting the position change of the face area in the video, and drive the playback of the target special effect animation based on the motion data, making the target special effect animation more realistic and enhancing the fun of live interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an image processing method, an image processing device, a storage medium, and a computer system. Background Art

[0002] With the development of live broadcast business, the live broadcast projects of anchors in the live broadcast room are becoming more and more diverse. In the running live broadcast project, it is necessary to identify whether the anchor is running and how fast he is running.

[0003] Currently, in related technical solutions, when a host broadcasts a running live, the host's running is generally detected by the gravity sensor built into the mobile phone.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The purpose of the present disclosure is to provide an image processing method, device, storage medium and computer system, thereby improving the live broadcast experience of running at least to a certain extent.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, an image processing method is provided, comprising: obtaining an image frame sequence in response to a trigger operation of adding a target special effects animation; identifying a facial region corresponding to a target object in the image frame sequence; collecting position change data of the facial region in a plurality of consecutive image frames, and determining motion data of the target object based on the position change data; and driving the playback of the target special effects animation based on the motion data, so that the played target special effects animation is adapted to the motion of the target object.

[0008] In an exemplary embodiment of the present disclosure, the collecting of position change data of the face area in a continuous plurality of image frames includes: selecting a target key point in the face area, collecting coordinate data of the target key point in a continuous plurality of image frames; and determining the position change data of the face area in a continuous plurality of image frames based on the coordinate data.

[0009] In an exemplary embodiment of the present disclosure, the motion data includes the change amplitude, change number and movement distance of the face area in multiple consecutive image frames, and the coordinate data includes a first coordinate value; determining the motion data of the target object based on the position change data includes: calculating the difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determining the change number based on the positive and negative changes of the difference; calculating the movement distance of the target key point based on the first coordinate value of the target key point in two consecutive image frames, and determining the change amplitude based on the movement distance; determining the movement distance based on the movement distance of the target key point in multiple consecutive image frames.

[0010] In an exemplary embodiment of the present disclosure, the motion data includes a violent motion state, and the method further includes: if the number of changes within a preset time is greater than or equal to a first threshold, the change amplitude is greater than or equal to a second threshold, and the moving distance is greater than or equal to a third threshold, then it is determined that the target object is in a violent motion state; and when the target object is in a violent motion state, the playback of the target special effects animation is triggered.

[0011] In an exemplary embodiment of the present disclosure, the method further includes: filtering the position change data whose difference value is less than an error threshold.

[0012] In an exemplary embodiment of the present disclosure, driving the playback of the target special effects animation based on the motion data includes: determining the motion rate of the target object through the number of changes corresponding to the target object; configuring the playback rate of the target special effects animation based on the motion rate, and driving the playback of the target special effects animation based on the playback rate.

[0013] In an exemplary embodiment of the present disclosure, the method further includes: obtaining a preset target special effects animation, the target special effects animation including a background animation and a special effects three-dimensional model; replacing the background area in the image frame sequence with the background animation; and displaying the face area in the target area of ​​the special effects three-dimensional model.

[0014] According to a second aspect of the present disclosure, an image processing device is provided, comprising: an image acquisition module for responding to a trigger operation of adding a target special effects animation and acquiring an image frame sequence; a face region recognition module for identifying a face region corresponding to a target object in the image frame sequence; a motion data determination module for collecting position change data of the face region in a plurality of consecutive image frames, and determining motion data of the target object based on the position change data; and a special effects animation playback module for driving the playback of the target special effects animation based on the motion data, so that the played target special effects animation is adapted to the motion of the target object.

[0015] In an exemplary embodiment of the present disclosure, the motion data determination module can be used to: select target key points in the face area, collect coordinate data of the target key points in multiple consecutive image frames; and determine the position change data of the face area in multiple consecutive image frames based on the coordinate data.

[0016] In an exemplary embodiment of the present disclosure, the motion data may include the change amplitude, number of changes and movement distance of the face area in multiple consecutive image frames, and the coordinate data may include a first coordinate value; the motion data determination module may be used to: calculate the difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determine the number of changes based on the positive and negative changes of the difference; calculate the movement distance of the target key point based on the first coordinate values ​​of the target key point in two consecutive image frames, and determine the change amplitude based on the movement distance; determine the movement distance based on the movement distance of the target key point in multiple consecutive image frames.

[0017] In an exemplary embodiment of the present disclosure, the motion data may include a violent motion state, and the image processing device may be used to: determine that the target object is in a violent motion state if the number of changes within a preset time is greater than or equal to a first threshold, the change amplitude is greater than or equal to a second threshold, and the movement distance is greater than or equal to a third threshold; and trigger the playback of the target special effects animation when the target object is in a violent motion state.

[0018] In an exemplary embodiment of the present disclosure, the image processing apparatus may be configured to filter position change data whose difference is less than an error threshold.

[0019] In an exemplary embodiment of the present disclosure, the special effects animation playback module can be used to: determine the movement rate of the target object through the number of changes corresponding to the target object; configure the playback rate of the target special effects animation according to the movement rate, and drive the playback of the target special effects animation according to the playback rate.

[0020] In an exemplary embodiment of the present disclosure, the image processing device can be used to: obtain a preset target special effects animation, wherein the target special effects animation includes a background animation and a special effects three-dimensional model; replace the background area in the image frame sequence with the background animation; and display the face area in the target area of ​​the special effects three-dimensional model.

[0021] According to a third aspect of the present disclosure, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned image processing method is implemented.

[0022] According to a fourth aspect of the present disclosure, there is provided a computer system, comprising:

[0023] processor; and

[0024] a memory for storing executable instructions of the processor;

[0025] The processor is configured to perform the above-mentioned image processing method by executing the executable instructions.

[0026] In the image processing method provided by an embodiment of the present disclosure, when a trigger operation for adding a target special effect animation is detected, an image frame sequence corresponding to the live video can be obtained, and then the facial area corresponding to the target object in the image frame sequence can be identified, and the position change data of the facial area in multiple consecutive image frames can be collected. Then, the motion data of the target object can be determined based on the position change data, and finally, the playback of the target special effect animation can be driven based on the motion data, so that the played target special effect animation is adapted to the motion of the target object. On the one hand, the motion data can be determined by the position change data of the facial area in the image frame sequence, and there is no need to use other devices with gravity sensors to achieve motion detection, which effectively reduces the hardware cost during live broadcast, simplifies the operational process of running live broadcast, and improves the live broadcast experience of running live broadcast; on the other hand, by detecting the motion data generated by the facial area to drive the playback of the target special effect animation, it can effectively improve the matching accuracy of the target special effect animation and the facial area, improve the realism of the target special effect animation, and enhance the interactive fun.

[0027] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present disclosure, and together with the description, serve to explain the principles of the present disclosure. It is apparent that the drawings described below are merely some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0029] Figure 1 A schematic diagram schematically illustrates a flow chart of an image processing method in an exemplary embodiment of the present disclosure;

[0030] Figure 2 A schematic diagram schematically illustrates the positions of target key points in a face region in an exemplary embodiment of the present disclosure;

[0031] Figure 3 A schematic diagram of a model for representing a human face region in an exemplary embodiment of the present disclosure is schematically shown;

[0032] Figure 4 A schematic diagram schematically illustrates a process of determining motion data of a target object in an exemplary embodiment of the present disclosure;

[0033] Figure 5 A schematic diagram schematically illustrates a flow chart of determining the motion state of a target object in an exemplary embodiment of the present disclosure;

[0034] Figure 6 A schematic diagram of a process for driving target special effect animation playback in an exemplary embodiment of the present disclosure is shown schematically;

[0035] Figure 7 A schematic diagram of a process for achieving target special effect animation playback in an exemplary embodiment of the present disclosure is shown schematically;

[0036] Figure 8 A schematic diagram schematically illustrates a human face area and a target special effect animation display area in an exemplary embodiment of the present disclosure;

[0037] Figure 9 A schematic diagram schematically illustrating another face area and a target special effect animation display area in an exemplary embodiment of the present disclosure;

[0038] Figure 10 A schematic diagram schematically illustrates an image processing apparatus in an exemplary embodiment of the present disclosure;

[0039] Figure 11 A schematic diagram schematically illustrates the composition of a computer system in an exemplary embodiment of the present disclosure;

[0040] Figure 12 The following schematically shows the composition of a storage medium in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0042] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0043] In this example embodiment, an image processing method is first provided, which can be applied to terminal devices, such as smart phones, tablet computers, desktop computers and other electronic devices, and correspondingly, the image processing device is also set in the terminal device; of course, those skilled in the art will understand that the image processing method in the present disclosure can also be applied to servers, for example, a live broadcast server implemented in a live broadcast on the Internet, and correspondingly, the image processing device is set in the server.

[0044] The following uses the terminal device to execute this method as an example to illustrate. Figure 1 As shown, the image processing method may include the following steps:

[0045] Step S110, in response to a triggering operation of adding a target special effect animation, obtaining an image frame sequence;

[0046] Step S120, identifying a face region corresponding to a target object in the image frame sequence;

[0047] Step S130, collecting position change data of the face region in a plurality of consecutive image frames, and determining motion data of the target object based on the position change data;

[0048] Step S140 , driving the target special effect animation to be played according to the motion data, so that the played target special effect animation is adapted to the motion of the target object.

[0049] The image processing method provided in this example embodiment can obtain the image frame sequence corresponding to the live video when detecting the trigger operation of adding the target special effect animation, and then can identify the facial area corresponding to the target object in the image frame sequence, and collect the position change data of the facial area in multiple consecutive image frames, and then can determine the motion data of the target object based on the position change data, and finally can drive the playback of the target special effect animation based on the motion data, so that the played target special effect animation is adapted to the motion of the target object. On the one hand, the motion data can be determined by the position change data of the facial area in the image frame sequence, without the need to use other devices with gravity sensors to achieve motion detection, effectively reducing the hardware cost during live broadcast, while simplifying the operating process of running live broadcast and improving the live broadcast experience of running live broadcast; on the other hand, by detecting the motion data generated by the facial area to drive the playback of the target special effect animation, it can effectively improve the matching accuracy of the target special effect animation and the facial area, improve the realism of the target special effect animation, and enhance the interactive fun.

[0050] Below, each step of the image processing method in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.

[0051] The application scenarios of this solution can be:

[0052] Currently, with the development of live broadcast business, the gameplay of the anchor in the live broadcast room is becoming more and more diverse. Some gameplays require identifying whether the anchor is running and how fast he is running, and then playing special effects at the corresponding speed. In the above gameplay, it is necessary to detect the anchor's running speed. During the live broadcast of the anchor's video, the face is usually in the video screen, and the body and feet may be outside the video screen or blocked. By using the image processing method disclosed in the present invention, the anchor's live video is used as the video to be detected in this solution, and the image frame sequence in the anchor's live video is obtained. After identifying the area where the anchor's face is located in each frame of the image, the anchor's motion data can be determined based on the position change of the area where the anchor's face is located in the image frame sequence. After obtaining the anchor's motion data, the playback speed of the selected target special effect animation can be determined based on the motion data, and the target special effect animation can be played according to the determined playback speed in the graphical user interface of the anchor client.

[0053] In step S110 , in response to a triggering operation of adding a target special effect animation, an image frame sequence is acquired.

[0054] In an example embodiment, the target special effect animation refers to a pre-set animation for user selection and that can be added to a video. For example, the target special effect animation can be a special effect animation composed of a cartoon three-dimensional model and a background image, wherein the cartoon three-dimensional model can include an area for displaying part of the video content, thereby realizing the fusion of the video content and the cartoon three-dimensional model. Of course, the target special effect animation can also be a virtual decoration added to the video and changes as the video content changes. This example embodiment does not specifically limit this.

[0055] A trigger operation refers to a pre-set operation for adding a target special effect animation to a video. For example, a special effect list consisting of multiple special effect animations can be provided on a graphical user interface. The trigger operation can be a click operation on the target special effect animation in the special effect list, or it can be a drag operation of dragging the target special effect animation in the special effect list to the video area.

[0056] Of course, the trigger operation can also be other operations that can add target special effects animations to the video. For example, during a live broadcast, the trigger operation can be an operation in which the host object on the host side makes a specific body gesture to add the target special effects animation corresponding to the specific gesture, or it can be an operation in which the user object on the audience side inputs voice data to add the corresponding target special effects animation through the keywords contained in the voice data. This example embodiment does not impose any special limitations on the trigger operation.

[0057] When a trigger operation for adding a target special effect animation is detected, the image frame sequence corresponding to the target video can be acquired. The target video can be a video pre-stored in the terminal device or a real-time video displayed during a live broadcast. This example embodiment does not specifically limit this. Specifically, the target video can be a live video. For example, in a live broadcast scenario, the target video can be a live video captured by a camera on the PC corresponding to the host object. The camera can be a camera integrated on the PC or an external camera connected via a USB interface. This example embodiment is not limited to this.

[0058] In this example embodiment, the method for obtaining an image frame sequence based on a target video can be: when the target video is a real-time live video, each frame of the live video can be read from the host object in the camera used for live broadcast through a Python image acquisition tool and stored to obtain an image frame sequence; When the target video is a video stored in a memory, the target video can be decoded by the ffmpeg video codec tool in the opencv framework to obtain each frame of the target video, and then each frame of the image can be formed into an image frame sequence in chronological order. Among them, opencv is a cross-platform computer vision and machine learning software library released based on the BSD license (open source), which can run on Linux, Windows, Android and Mac OS operating systems. Of course, the method for obtaining the image frame sequence here is only a schematic example and should not cause any special limitation to this example embodiment.

[0059] It is understandable that the image processing method in this embodiment can also be executed by a server. When the server executes the method, the image frame sequence can be obtained in the following manner:

[0060] The first method: Taking the live broadcast scenario as an example, the PC used by the anchor transmits the live video to the server via wired or wireless means, and the server parses the live video to obtain an image frame sequence. The wired method can be a broadband transmission method, and the wireless method can include but is not limited to: 4G network transmission method, 5G network transmission method, etc.

[0061] The second method: The PC can read the video to be detected from its own storage system and transmit the read video to be detected to the server. The server can parse and process the received video to be detected to obtain an image frame sequence.

[0062] The third method: The server can pre-store the video to be tested in a storage system. When the terminal device sends a trigger instruction to add the target special effect animation, the server can directly extract the video to be tested from its own storage and parse it to obtain the image frame sequence. The server's storage system can be cloud storage, MySQL database, etc.

[0063] In step S120, a face region corresponding to a target object in the image frame sequence is identified.

[0064] In an example embodiment, the target object refers to the key content contained in the image frame. For example, the target object can be a person object contained in the image frame, or an animal object, a movable object object, etc. contained in the image frame. This example embodiment does not specifically limit this.

[0065] The face area refers to the area of ​​interest corresponding to the target object in the image frame. For example, when the target object is a person object, the face area can be the area corresponding to the facial features of the person object. It is easy for those skilled in the art to understand that the face area can also refer to other types of areas of interest. For example, if the target object is an animal object, the face area can be the contour area corresponding to the animal object. This example embodiment does not specifically limit this.

[0066] In this example embodiment, the image frame sequence may include one or more target objects. Accordingly, the number of face regions detected in the image frames may be one or more, which is not particularly limited in this example embodiment.

[0067] Alternatively, taking the target object as a person object as an example, when the number of person objects is one, the method for identifying the face region in each image frame can be: grayscale conversion of each frame image, then converting the grayscale image into a binary image by grayscale threshold screening, performing contour extraction on the binary image, excluding small contours to obtain the area with the largest contour, which can be considered as the face region, or further setting the pixel values ​​of other areas outside the face region in the binary image to 0. In this case, the face position of the target object, i.e., the face region, can be determined from the image frame. Of course, when the number of face regions is one, the method for extracting the position of the face region in the image frame can also be: HOG (Histogram of Oriented Gradient) feature extraction method, MATLAB-based PCA (Principal Component Analysis) face recognition and other image processing methods, etc., which are not specifically limited in this example embodiment.

[0068] Optionally, taking the target object as a human object as an example, when there are multiple human objects, a target detection algorithm can be used to perform face region detection on the multiple human objects in the graphic frame sequence. That is, in each image frame of the graphic frame sequence, different faces in the image frame can be selected by multiple rectangular boxes, and the area corresponding to the rectangular box is the face area corresponding to the target object in the image frame sequence. The target detection algorithm can be a multi-task convolutional neural network (MTCNN) for face detection, or a FaceNET model for face recognition. Of course, it can also be other types of detection methods that can detect and identify face areas, and this example embodiment does not specifically limit this.

[0069] By determining the number of target objects and using different face detection algorithms based on the number, the efficiency of face area detection can be effectively improved, thereby enhancing the system performance in real-time video scenarios such as live video.

[0070] In step S130 , position change data of the face region in a plurality of consecutive image frames is collected, and motion data of the target object is determined based on the position change data.

[0071] In an example embodiment, the position change data refers to the position change of the face area in multiple consecutive image frames. Generally, the size of each image frame in the image frame sequence is fixed, so a coordinate system can be established in each image frame to represent the position of the face area in the image frame.

[0072] When the position of the face area is represented by coordinates, a key point can be determined in the face area, such as the midpoint of the face area or the vertex of the face area, and the position change data of the face area in the continuous multiple image frames can be represented by the coordinate change of the key point in the face area in the continuous multiple image frames.

[0073] For example, taking the midpoint of the face area as the key point of the face area, assuming that the coordinates of the midpoint of the face area in the first frame image are (0, 1), and the coordinates of the midpoint of the face area in the second frame image are (0, 5), then it can be considered that the face area moves relatively upward, and the data changing from the coordinates (0, 1) to (0, 5) can be considered as the position change data of the face area in multiple consecutive image frames; of course, in actual application, other data calculated from the coordinate data, such as the distance to the origin, can also be used as the position change data of the face area, and this example embodiment does not make any special limitations on this.

[0074] It can be understood that the coordinate data in this embodiment is only a schematic example. The change in the distance data of the face area from the edges of the image frames in multiple consecutive image frames can also be used as the position change data. Other data that can represent the displacement change of the face area can also be used as the position change data of the face area in multiple consecutive image frames. This example embodiment is not limited to this.

[0075] Motion data refers to data that is converted from position change data of a facial region in a plurality of consecutive image frames and is used to characterize the motion posture of a target object. For example, the motion data may be the number of times the facial region moves up and down in a plurality of consecutive image frames. If the number of times the facial region moves up and down in a plurality of consecutive image frames increases, it can be considered that the target object corresponding to the facial region is performing continuous motion, such as walking, running, or nodding. The motion data may also be the amplitude of the up and down movement of the facial region in a plurality of consecutive image frames. If the amplitude of the up and down movement of the facial region in a plurality of consecutive image frames is large, it can be considered that the target object corresponding to the facial region is more likely to be running. If the amplitude of the up and down movement of the facial region in a plurality of consecutive image frames is small, it can be considered that the target object corresponding to the facial region is more likely to be walking or nodding. Of course, the motion data here is merely an illustrative example. The motion data may also be other data that can be converted from position change data and is used to characterize the motion posture of the target object. This exemplary embodiment does not specifically limit this.

[0076] Optionally, motion data can also represent the target object's motion state. For example, the motion data can represent strenuous motion such as running or walking, mild motion such as nodding or shaking, or static motion such as standing still. It is understood that in a live video, when a person is standing still, the position of the person's facial region in the live video remains essentially horizontal. When the person is walking, the position of the person's facial region in the live video will move up and down at a certain frequency depending on the speed of the person's walking. When the person is running, the position of the person's facial region in the live video will move up and down at a certain frequency depending on the speed of the person's running. The frequency and amplitude of the up and down movement of the facial region during running are greater than those during walking. Therefore, this solution can determine the target object's motion state by obtaining the frequency and amplitude of the up and down movement of the target object's facial region in the acquired image frame sequence as a judgment basis.

[0077] In an example embodiment, the continuous multi-frame image frame can be any continuous segment of image frames in the image frame sequence, or can be all continuous image frames in the image frame sequence, and this example embodiment does not specifically limit this. Specifically, the image frame corresponding to the moment of responding to the triggering operation of adding the target special effect animation can be used as the first image frame, and the image frames corresponding to the duration of the target special effect animation and the first image frame can be used to form the continuous multi-frame image frame.

[0078] In one exemplary embodiment, target key points may be selected within a facial region, and coordinate data for the target key points in multiple consecutive image frames may be collected. Position change data for the facial region within the multiple consecutive image frames may then be determined based on the coordinate data. The target key points may refer to feature points screened within the facial region. For example, if the target object is a person, feature points corresponding to the left eye, right eye, or nose within the facial region may be used as target key points, and changes in the positions of these key points may be used to represent changes in the entire facial region. Of course, marker points may also be pre-set within the facial region as target key points, and this exemplary embodiment does not impose any specific limitations on this.

[0079] By determining the position of the target key points in the face region, the position change of the face region in the image frame sequence can be indirectly determined, which can improve the accuracy of the face region position change detection.

[0080] In this exemplary embodiment, the face region in each frame of image may be estimated using a face pose estimation algorithm, and a coordinate system may be constructed based on the size of the image frame to obtain an xy coordinate dataset of the face region.

[0081] For example, the position change data of the face region in multiple consecutive image frames can be obtained by:

[0082] The first way, refer to Figure 2 As shown, a feature point can be selected in the facial region and its coordinate data across multiple consecutive image frames can be determined to obtain corresponding position change data for the facial region. Feature points can be features such as the user's eyes 202, ears 201, nose 203, or mouth 204, or they can be specially marked locations 205 within the facial region. For example, a red dot can be selected from the cheek of a person within the facial region and used as feature point 205. The coordinate data of the feature point across multiple image frames can then be obtained and transformed to obtain position change data for the facial region.

[0083] The second method, refer to Figure 3As shown, the facial region can be virtualized as an identification model, and the position change of facial region 301 in the image frame sequence can be determined by the changes in the coordinate data of the identification model in multiple consecutive image frames. For example, the facial region can be virtualized as a triangle 302, and the position change data of the facial region in the image frame sequence can be determined by determining the average coordinate data change of multiple endpoints of triangle 302 in the image frame sequence. Alternatively, the face can be virtualized as a circle, and the average coordinate data change of each coordinate point on the circle's edge can be directly determined to determine the position change data of the facial region in the image frame sequence. By virtualizing the entire facial region as a simple graphic model and then representing the position change data of the entire facial region based on the coordinate data changes of each coordinate point on the graphic model, compared to the solution of determining a single key point, the accuracy is higher and the error is smaller.

[0084] In an optional embodiment, the movement trajectory of the face can be determined based on the position change of the face area in multiple consecutive image frames, wherein the movement trajectory is a position curve formed by connecting the positions of the face area corresponding to different time points in the multiple consecutive image frames. The movement trajectory can be considered as a coordinate change curve obtained by converting the coordinate data of the face area. Specifically, two adjacent frames can be obtained from the image frame sequence, and the positions of the face areas in the two frames are marked respectively. The two positions are connected to form the trajectory of the face in the two frames. All image frames in the image frame sequence are processed separately to identify the position of the face area in each frame. The positions of the face area in each frame are connected to form the change trajectory of the face area in the image frame sequence. The change trajectory is the position change of the face area in the multiple consecutive image frames.

[0085] It is understood that when the target object is standing, the movement trajectory can be displayed as a relatively smooth horizontal line; when the target object is walking, the movement trajectory can be displayed as a curve with less fluctuation; when the target object is running, the movement trajectory can be displayed as a curve with greater fluctuation. When the target object is running, the head of the target object moves up once and down once with each step. It is understood that the number of times the direction of the curve changes per unit time can be obtained based on the movement trajectory; the more times the number of changes, the more times the target object's head floats up and down, that is, the more steps the target object has taken. When the number of fluctuations per unit time exceeds a preset threshold, it can be determined that the target object is currently running.

[0086] For example, experiments have shown that a person's head moves up and down at least twice per second when running, and less than twice per second when walking. Based on the drawn movement trajectory, it is determined whether the number of curve direction changes within one second is less than four. If it is not less than four, the target subject's head movement frequency per unit time is greater than two times per second. Therefore, based on the movement trajectory analysis, the target subject is running. If the number of curve direction changes is less than four, the target subject is walking during the current period. This is merely an illustrative example and should not be construed as limiting this exemplary embodiment.

[0087] In an example embodiment, the motion data may include at least the change amplitude, number of changes and movement distance of the face area in multiple consecutive image frames, and the coordinate data of the target key point corresponding to the face area may include at least a first coordinate value, wherein the first coordinate value may be the vertical coordinate in the coordinate data or the horizontal coordinate in the coordinate data. The specific selection of the horizontal coordinate or the vertical coordinate as the first coordinate value can be customized according to actual conditions. This example embodiment is not limited to this. When this embodiment is used to detect the motion state of the target object, the first coordinate value in this embodiment may refer to the vertical coordinate of the coordinate data.

[0088] Specifically, step S130 can be performed by Figure 4 The steps in the above are used to determine the motion data of the target object based on the position change data of the face area in multiple consecutive image frames. Figure 4 Specifically, it may include:

[0089] Step S410, calculating the difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determining the number of changes according to the positive and negative changes of the difference;

[0090] Step S420, calculating the moving distance of the target key point according to the first coordinate values ​​of the target key point in two consecutive image frames, and determining the change amplitude according to the moving distance;

[0091] Step S430 : determining the moving stroke according to the moving distance of the target key point in a plurality of consecutive image frames.

[0092] The number of changes refers to the frequency at which the face region moves up and down in the image frames. For example, in the image from the first frame to the second frame, the coordinate data corresponding to the face region changes from (0, 1) to (0, 5). In this case, the first coordinate values ​​can be 1 and 5, and the difference between the first coordinate values ​​is 4. It can be considered that the face region of the target object has moved upward once. In the image from the second frame to the third frame, the coordinate data corresponding to the face region changes from (0, 5) to (0, 1). In this case, the first coordinate values ​​can be 5 and 1, and the difference between the first coordinate values ​​is -4. It can be considered that the face region of the target object has moved downward once. Thus, in the image from the first frame to the third frame, the difference between the first coordinate values ​​of the target key point in the face region changes from positive to negative once. In this case, it can be considered that the face region has moved up and down once, that is, the target object has completed a step of running or walking, and the number of changes is 1. Of course, this is only an illustrative example and should not impose any special limitations on this exemplary embodiment.

[0093] The change amplitude refers to the moving distance of the face area when it moves up and down in the image frame. The moving distance can be the maximum moving distance of the face area when it moves up and down in multiple image frames, or it can be the average moving distance of the face area when it moves up and down in multiple image frames. For example, in the image from the first frame to the second frame, the coordinate data corresponding to the face area changes from (0, 1) to (0, 5). At this time, the first coordinate values ​​can be 1 and 5, and the moving distance is 4; in the image from the second frame to the third frame, the coordinate data corresponding to the face area changes from (0, 5) to (0, 2). At this time, the first coordinate values ​​can be 5 and 2, and the moving distance is 3. At this time, the change amplitude of the face area can be the maximum moving distance of 4, or the average moving distance of 3.5. Of course, this is only an illustrative example and should not impose any special limitations on this exemplary embodiment.

[0094] The movement distance refers to the total movement distance of the face region when it moves up and down in the image frame. For example, in the image from the first frame to the second frame, the coordinate data corresponding to the face region changes from (0, 1) to (0, 5). In this case, the first coordinate values ​​can be 1 and 5, and the movement distance is 4. In the image from the second frame to the third frame, the coordinate data corresponding to the face region changes from (0, 5) to (0, 2). In this case, the first coordinate values ​​can be 5 and 2, and the movement distance is 3. In this case, the movement distance of the face region is 7. Of course, this is merely an illustrative example and should not impose any special limitations on this exemplary embodiment.

[0095] In an optional embodiment, the movement trajectory of the facial area can be determined in advance based on the position change data of the facial area in multiple consecutive image frames. It should be noted that, if the host runs in place during the live broadcast, the analysis points of this solution are the amplitude of the up and down changes in the host's facial area in the video, the number of changes, and the movement distance. Therefore, the position change of the target key point along the y-axis direction in the movement trajectory can be used as the position change data of the target key point, so as to determine the position change data of the facial area in the image frame sequence.

[0096] In this embodiment, the movement trajectory can be a curve formed by connecting the changes in the coordinate data of the facial area in the y-axis direction with time as the horizontal axis; based on the movement trajectory, a difference sequence corresponding to the movement trajectory can be obtained; first, all peak values ​​and valley values ​​can be obtained from the movement trajectory, and all peak values ​​and valley values ​​can be arranged in chronological order to form a sequence, and then the difference between the latter value and the former value in the sequence is calculated in order, and the differences are arranged in sequence to form a difference sequence; based on the difference sequence, when the sign of the difference changes twice (i.e., from positive to negative or from negative to positive), the number of changes in the facial area is increased by 1, corresponding to the number of steps the user runs plus 1; the difference with the largest absolute value in the difference sequence is obtained, and the absolute value of the difference is used as the change amplitude; the sum of the absolute values ​​of the differences within unit time is calculated to obtain the movement distance.

[0097] For example, based on the movement trajectory drawn based on the position change of the face area, the position change data of the face area of ​​the target object within 1 second can be obtained. For example, the sequence of all peaks and valleys is: {1, 10, 1, 9, 1, 12, 2, 10, 2}. Based on this sequence, the difference sequence that can be obtained is: {10-1, 1-10, 9-1, 1-9, 12-1, 2-12, 10-2, 2-10}, that is, {9, -9, 8, -8, 11, -10, 8, -8}. From the difference sequence, it can be seen that the number of changes in the face area is 8 times, which proves that the number of steps taken by the target object in 1 second is 4; the change amplitude of the face area is 11 (in this case, the maximum movement distance is taken as the change amplitude), and the movement distance of the face area is 71. Of course, this is only an illustrative example and should not impose any special limitations on this example embodiment.

[0098] Specifically, position change data can be filtered for data where the difference in the first coordinate values ​​of a target key point between two consecutive image frames is less than an error threshold. For example, the error threshold can be 5 or 4, or other thresholds. The specific setting can be customized based on actual circumstances, and this exemplary embodiment is not limited thereto. That is, when the difference in the first coordinate values ​​of a target key point is less than the error threshold, the data can be considered error data or cheating data in live running broadcasts. For example, in live running broadcasts, this can be used to prevent users from cheating by standing and nodding, thereby eliminating errors in facial pose estimation.

[0099] In an example embodiment, the motion data may also include a vigorous motion state, which may be a running state, a fast walking state, a jumping state, etc. Specifically, it is possible to determine whether the target object is in a vigorous motion state by using the motion data such as the change amplitude, the number of changes, and the movement distance corresponding to the target object. Figure 5 The steps in the above code are used to determine whether the target object is in a state of intense exercise. Figure 5 Specifically, it may include:

[0100] Step S510: If the number of changes within a preset time is greater than or equal to a first threshold, the change amplitude is greater than or equal to a second threshold, and the movement distance is greater than or equal to a third threshold, then it is determined that the target object is in a state of intense exercise; and

[0101] Step S520 , when the target object is in a state of intense motion, triggering the playing of the target special effects animation.

[0102] Among them, the first threshold refers to the threshold used to judge whether the number of changes of the target object reaches the standard of strenuous exercise state. For example, the first threshold can be 4 times / second. At this time, it can be considered that the number of changes of the target object is greater than or equal to 4 times / second, that is, the step frequency of the target object reaches 2 steps / second. It can be considered that the number of changes of the target object reaches the standard of strenuous exercise state. Of course, the first threshold can also be 6 times / second. It can be determined according to the actual situation or the situation of the target object. This example embodiment does not make any special limitations on this.

[0103] The second threshold refers to the threshold used to determine whether the change amplitude of the target object reaches the standard of a strenuous exercise state. For example, the second threshold can be 10. At this time, it can be considered that the change amplitude of the target object is greater than or equal to 10, that is, when the target object takes a step, the moving distance of the face area in the image frame sequence is greater than or equal to 10. It can be considered that the change amplitude of the target object reaches the standard of a strenuous exercise state. Of course, the second threshold can also be 12. It can be customized according to the size of the image frame or other actual conditions. This example embodiment does not make any special limitations on this.

[0104] The third threshold value refers to the threshold value used to determine whether the target object's moving distance reaches the standard of a strenuous exercise state. For example, the third threshold value can be 20. At this time, it can be considered that the target object's moving distance is greater than or equal to 20, that is, the total moving distance of the target object's face area in the image frame sequence reaches 20. It can be considered that the target object's moving distance reaches the standard of a strenuous exercise state. Of course, the third threshold value can also be 30. It can be customized according to the size of the image frame or other actual conditions. This example embodiment does not make any special limitations on this.

[0105] In this exemplary embodiment, limiting conditions are introduced to further accurately determine the motion state of the target object, which may include at least limiting the number of changes, limiting the amplitude of changes, and limiting the movement range.

[0106] In a unit of time, the target object can only be judged to be in a state of intense exercise (i.e., running, fast walking, jumping on the spot, etc.) when the number of steps (i.e., obtained by converting the number of changes), the amplitude of change (i.e., the amplitude of each step), and the movement distance all meet preset conditions (i.e., the first threshold, the second threshold, and the third threshold).

[0107] For example, if data is counted every second and three conditions are met at the same time, the user is considered to be running: the number of steps of the target object is greater than or equal to 2, that is, the number of changes is greater than or equal to 4, the change amplitude of each step of the target object is greater than 10, and the movement distance of the target object is greater than 20.

[0108] In this example embodiment, after responding to the trigger operation of adding a target special effects animation, the target special effects animation can be set to a static state first, and after determining that the target object is in a state of intense motion, the playback of the target special effects animation can be triggered. That is, assuming that during a live broadcast of running, after adding a target special effects animation, the target special effects animation will not be played immediately. Instead, after detecting that the host starts running (i.e., a state of intense motion), the target special effects animation is triggered to start playing, thereby achieving an interactive effect of driving the target special effects animation through the host's running action.

[0109] In an example embodiment, the Figure 6 To achieve the goal of driving the special effects animation according to the motion data, refer to Figure 6 Specifically, it may include:

[0110] Step S610, determining the movement speed of the target object according to the number of changes corresponding to the target object;

[0111] Step S620 : configuring the play rate of the target special effect animation according to the motion rate, and driving the play of the target special effect animation according to the play rate.

[0112] Among them, the motion rate of the target object refers to the motion rate of the target object obtained by converting the number of changes. For example, assuming that in a live broadcast of running, the host can actually run, and the motion rate obtained according to the number of changes is the real motion rate of the host. Alternatively, the host can run in situ in the live broadcast room, and the motion rate obtained according to the number of changes is the estimated motion rate of the host. Specifically, the number of changes corresponding to the host can be obtained by the method in this embodiment. If the number of changes can be 4 times / second, the motion rate corresponding to the host can be 2 steps / second (or 2 meters / second). Of course, this is only a schematic example and should not cause any special limitation to this example embodiment.

[0113] The playback rate refers to the refresh frequency of the target special effects animation displayed on the graphical user interface. For example, the movement rate of the target object can be 2 steps / second, then the playback rate of the target special effects animation can be 20 frames / second, the movement rate of the target object can also be 3 steps / second, then the playback rate of the target special effects animation can be 30 frames / second. By matching the playback rate of the target special effects animation with the movement rate of the target object, the target special effects animation can be visually changed in real time following the movement of the target object, thereby improving the realism of the target special effects animation and enhancing the interactive experience.

[0114] In an optional embodiment, if the position change data of the facial area of ​​the target object is a moving trajectory, the number of steps taken by the user in unit time can be determined through the moving trajectory, that is, if the number of times the peak-to-peak-to-valley alternation of the moving trajectory changes is 2 times, the number of steps taken by the user increases by one step. When the user is in a running state, the number of steps taken by the user in unit time can be determined by the number of times the direction of the moving trajectory changes in unit time; for example, if the peak-to-peak-to-valley alternation of the moving trajectory changes 4 times in unit time, it proves that the user has run 2 steps in unit time; if the peak-to-peak-to-valley alternation of the moving trajectory changes 6 times in unit time, it proves that the user has run 3 steps in unit time.

[0115] In an example embodiment, the Figure 7 To play the target special effects animation, refer to Figure 7 Specifically, it may include:

[0116] Step S710: obtaining a preset target special effect animation, wherein the target special effect animation includes a background animation and a special effect three-dimensional model;

[0117] Step S720, replacing the background area in the image frame sequence with the background animation;

[0118] Step S730: Display the human face area in the target area of ​​the special effect three-dimensional model.

[0119] The target special effects animation may include at least a background animation and a special effects 3D model. For example, the special effects 3D model may be a virtual 3D cartoon character or a 3D virtual decorative object. This example embodiment does not impose any specific restrictions on the type of special effects 3D model. The background animation may be a racetrack background or a landscape background. This example embodiment does not impose any specific restrictions on the content of the background animation.

[0120] For example, assuming that in a live broadcast scenario, the method of playing the target special effects animation can be to replace the background of the live video with the background animation of the target special effects animation, such as the track background animation, and replace the host's body part with a special effects three-dimensional model, such as a 3D cartoon character. The facial area corresponding to the 3D cartoon character can be a blank area, which can be used to display the facial area detected in the image frame sequence. The 3D cartoon character will move with the movement of the host's facial area. As the host runs, the 3D cartoon character will perform a running animation, and the running speed of the 3D cartoon character (i.e., the playback rate) is proportional to the frequency of the up and down movement of the host's facial area. At the same time, the track background animation will also change with the running speed of the 3D cartoon character, such as visually showing that the track content moves backward. The faster the host runs, the faster the 3D cartoon character animation plays, and the faster the track background animation moves.

[0121] In this example implementation, during a live video broadcast, the host's face is usually within the video frame, while their body and feet may be outside the video frame. The host's live video is captured by a camera and displayed in a graphical user interface. The PC transmits the live video captured by the camera to a server via broadband transmission. The server uses the host's live video as the video to be detected, obtains a sequence of image frames from the live video, and then identifies the area in each frame of the image frame sequence where the host's face is located. The host's motion state can then be determined based on the positional changes in the area where the host's face is located in the image frame sequence. The motion state can include the host's running speed. After obtaining the host's motion state, the playback speed of the special effects can be configured according to pre-set configuration rules, and the special effects can be played according to the playback speed of the special effects in the graphical user interface of the host client.

[0122] It can respond to a user's instruction to add a target special effect animation triggered in the graphical user interface, and configure the display position of the target special effect animation in the graphical user interface based on the position of the face in the graphical user interface.

[0123] For example, when starting a running game, the server sends a game start command to the client used by the user. After receiving the game start command, the client needs to click on the client's graphical user interface to indicate whether to enter the running game. For example, the client's graphical user interface can be provided with an "Accept" button and a "Reject" button. When the user clicks the "Accept" button, it is considered that the user agrees to enter the running game. The server responds to the user's confirmation and enters the running game, displaying a target special effect animation in the user's graphical user interface. If the user clicks the "Reject" button, the server considers that the user refuses to participate in the running game. The server will respond to the rejection command sent by the user and will not respond to the user command issued by the client.

[0124] Specifically, the display position of the target special effect animation in the graphical user interface based on the face position can be configured in the following ways:

[0125] The first method: Reference Figure 8 As shown in , when the server responds to the user's instruction, it can divide the user's graphical interface into a user's face display area 801 and a target special effect animation display area 802, and the two areas do not intersect.

[0126] The second method: Reference Figure 9 As shown in , when the server responds to the user's instructions, it can identify the display position of the face area in the graphical user interface 902, and then control the client's graphical user interface 902 to assemble a target special effect animation (such as a special effect three-dimensional model in the target special effect animation) around the outside of the user's head. For example, the server can assemble a 3D cartoon character image 903 in the graphical user interface, deduct the face of the 3D cartoon character image 903 to form a blank area, and the user's face area 901 is displayed in the blank area.

[0127] Specifically, there are two ways to play the target special effect animation in the graphical user interface:

[0128] The first method: When the server responds to the user's instruction, the server sends an assembly instruction to the client, and the client executes the assembly instruction to assemble the target special effect animation, and plays it in the graphical user interface when it detects that the target object is in a state of intense motion.

[0129] The second method: store the target special effects animation in the client in advance. When the server responds to the user's instruction, the server sends an instruction to call the target special effects animation to the client. The client executes the instruction to call the target special effects animation, retrieves the target special effects animation from the client storage, and plays the target special effects animation in the client's graphical user interface when it detects that the target object is in a state of intense motion.

[0130] Specifically, after receiving a user instruction from the client, the server can capture the video to be detected. For example, after the host clicks the "Accept" button on the client, the server will respond to the host's instruction by capturing the camera feed from the host's terminal as a sequence of image frames. It is understood that capturing image frames can be done in more ways than just the above. For example, after responding to the user instruction, the server can also retrieve a video recording of the user running in place or on a treadmill from memory and analyze the movement trajectory of the target object's face in the video recording to determine the user's motion state.

[0131] The embodiments disclosed herein provide an application scenario in which a host broadcasts a game. When a server is the executing entity, during a live broadcast, after the host's terminal device receives a game start command from the server, the host sends a start detection command to the server via the terminal device. That is, after receiving the start detection command, the server responds to the host's start detection command and assembles and generates a target special effects animation (such as a special effects three-dimensional model in the target special effects animation) around a human face area on the live broadcast screen. The target special effects animation may include a 3D or 2D cartoon character image, or a 3D or 2D cartoon animal. To enhance the fun, the face of the virtual model may be removed, and the host's face area may be filled in the face position of the 3D or 2D cartoon character image. The playback rate of the 3D or 2D cartoon character image is configured according to the host's running speed, and the server controls the 3D or 2D cartoon character image to play in a graphical user interface according to the playback rate.

[0132] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0133] The present disclosure also provides an image processing device, referring to Figure 10 As shown, the image processing device 1000 may include an image acquisition module 1010, a face region recognition module 1020, a motion data determination module 1030, and a special effects animation playback module 1040.

[0134] The image acquisition module 1010 is used to respond to the triggering operation of adding a target special effect animation and acquire an image frame sequence;

[0135] The face region recognition module 1020 is used to recognize the face region corresponding to the target object in the image frame sequence;

[0136] The motion data determination module 1030 is used to collect position change data of the face region in a plurality of consecutive image frames, and determine the motion data of the target object based on the position change data;

[0137] The special effect animation playing module 1040 is used to drive the playing of the target special effect animation according to the motion data, so that the played target special effect animation is adapted to the motion of the target object.

[0138] In an exemplary embodiment of the present disclosure, the motion data determination module 1030 may be configured to:

[0139] Selecting a target key point in the face area and collecting coordinate data of the target key point in a plurality of consecutive image frames;

[0140] Position change data of the face region in a plurality of consecutive image frames is determined according to the coordinate data.

[0141] In an exemplary embodiment of the present disclosure, the motion data may include a change amplitude, a change number, and a movement distance of the face region in a plurality of consecutive image frames, and the coordinate data may include a first coordinate value;

[0142] The motion data determination module 1030 may be configured to:

[0143] Calculating a difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determining the number of changes according to a positive or negative change of the difference;

[0144] Calculating a moving distance of the target key point according to the first coordinate values ​​of the target key point in two consecutive image frames, and determining the change amplitude according to the moving distance;

[0145] The moving stroke is determined according to the moving distance of the target key point in a plurality of consecutive image frames.

[0146] In an exemplary embodiment of the present disclosure, the motion data may include a strenuous motion state, and the image processing apparatus 1000 may be configured to:

[0147] If the number of changes within a preset time is greater than or equal to a first threshold, the amplitude of the change is greater than or equal to a second threshold, and the movement distance is greater than or equal to a third threshold, then it is determined that the target object is in a state of intense exercise; and

[0148] When the target object is in a state of intense motion, the playing of the target special effects animation is triggered.

[0149] In an exemplary embodiment of the present disclosure, the image processing apparatus 1000 may be used to:

[0150] The position change data whose difference is less than the error threshold is filtered.

[0151] In an exemplary embodiment of the present disclosure, the special effects animation playback module 1040 may be used to:

[0152] Determining the movement speed of the target object according to the number of changes corresponding to the target object;

[0153] The play rate of the target special effect animation is configured according to the motion rate, and the play of the target special effect animation is driven according to the play rate.

[0154] In an exemplary embodiment of the present disclosure, the image processing apparatus 1000 may be used to:

[0155] Obtaining a preset target special effect animation, wherein the target special effect animation includes a background animation and a special effect three-dimensional model;

[0156] replacing the background area in the image frame sequence by the background animation;

[0157] The face area is displayed in the target area of ​​the special effect three-dimensional model.

[0158] The specific details of each module in the above-mentioned image processing device have been described in detail in the corresponding image processing method, and therefore will not be repeated here.

[0159] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0160] In an exemplary embodiment of the present disclosure, a computer system capable of implementing the above method is also provided.

[0161] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented as a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0162] Refer to the following Figure 11 1100 according to this embodiment of the present invention will be described. Figure 11 The computer system 1100 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0163] like Figure 11 As shown, computer system 1100 is implemented as a general-purpose computing device. Components of computer system 1100 may include, but are not limited to, at least one processing unit 1110, at least one storage unit 1120, and a bus 1130 connecting various system components (including storage unit 1120 and processing unit 1110).

[0164] The storage unit stores program codes, which can be executed by the processing unit 1110, so that the processing unit 1110 performs the steps of various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 1110 can perform the following steps: Figure 1 In step S110 shown, in response to a trigger operation of adding a target special effects animation, an image frame sequence is obtained; in step S120, the face area corresponding to the target object in the image frame sequence is identified; in step S130, position change data of the face area in a plurality of consecutive image frames is collected, and motion data of the target object is determined based on the position change data; in step S140, the playback of the target special effects animation is driven based on the motion data, so that the played target special effects animation is adapted to the motion of the target object.

[0165] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 11201 and / or a cache memory unit 11202 , and may further include a read-only memory unit (ROM) 11203 .

[0166] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205, such program modules 11205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0167] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0168] Computer system 1100 can also communicate with one or more external devices 1001 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with computer system 1100, and / or any device that enables computer system 1100 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication can occur via input / output (I / O) interface 1150. Furthermore, computer system 1100 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1160. As shown, network adapter 1160 communicates with other modules of computer system 1100 via bus 1130. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with computer system 1100, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0169] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0170] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the aforementioned methods of this specification. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section of this specification.

[0171] refer to Figure 12 , a program product 1200 for implementing the above-described method according to an embodiment of the present invention is described. The program product 1200 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0172] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0173] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0174] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0175] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0176] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0177] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0178] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that: include: In response to the triggering operation of adding the target special effect animation, the image frame sequence is obtained; Identifying a face region corresponding to a target object in the image frame sequence; Collecting position change data of the face region in a plurality of consecutive image frames of the image frame sequence, and determining motion data of the target object based on the position change data; The motion data includes the number of changes of the face region in a plurality of consecutive image frames; determining the motion data of the target object based on the position change data includes: calculating the difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determining the number of changes based on the positive and negative changes of the difference; wherein the target key point is a key point selected in the face region; The movement rate of the target object is determined by the number of changes corresponding to the target object; the playback rate of the target special effect animation is configured according to the movement rate, and the playback of the target special effect animation is driven according to the playback rate, so that the played target special effect animation is adapted to the movement of the target object. The playback rate is the screen refresh frequency of the target special effect animation on the graphical user interface.

2. The image processing method according to claim 1, wherein: The collecting of position change data of the face region in a plurality of consecutive image frames includes: Selecting a target key point in the face area and collecting coordinate data of the target key point in a plurality of consecutive image frames; Position change data of the face region in a plurality of consecutive image frames is determined according to the coordinate data.

3. The image processing method according to claim 2, wherein: The motion data includes a change amplitude, a change number, and a movement distance of the face region in a plurality of consecutive image frames, and the coordinate data includes a first coordinate value; Determining the motion data of the target object according to the position change data includes: Calculating a difference between the first coordinate values ​​of the target key point in two consecutive image frames, and determining the number of changes according to a positive or negative change of the difference; Calculating a moving distance of the target key point according to the first coordinate values ​​of the target key point in two consecutive image frames, and determining the change amplitude according to the moving distance; The moving stroke is determined according to the moving distance of the target key point in a plurality of consecutive image frames.

4. The image processing method according to claim 3, wherein: The exercise data includes a strenuous exercise state, and the method further includes: If the number of changes within a preset time is greater than or equal to a first threshold, the amplitude of the change is greater than or equal to a second threshold, and the movement distance is greater than or equal to a third threshold, then it is determined that the target object is in a state of intense exercise; and When the target object is in a state of intense motion, the playing of the target special effects animation is triggered.

5. The image processing method according to claim 4, characterized in that The method further comprises: The position change data whose difference is less than the error threshold is filtered.

6. The image processing method according to claim 1, wherein: The method further comprises: Obtaining a preset target special effect animation, wherein the target special effect animation includes a background animation and a special effect three-dimensional model; replacing the background area in the image frame sequence by the background animation; The human face area is displayed in the target area of ​​the special effect three-dimensional model.

7. An image processing device, characterized in that include: An image acquisition module is used to respond to a trigger operation of adding a target special effect animation and acquire an image frame sequence; A face region recognition module, configured to recognize a face region corresponding to a target object in the image frame sequence; a motion data determination module, configured to collect position change data of the facial region in a plurality of consecutive image frames of the image frame sequence, and determine motion data of the target object based on the position change data; wherein the motion data includes the number of changes of the facial region in the plurality of consecutive image frames; determining the motion data of the target object based on the position change data comprises: calculating a difference between first coordinate values ​​of a target key point in two consecutive image frames, and determining the number of changes based on a positive or negative change of the difference; wherein the target key point is a key point selected from the facial region; A special effects animation playback module is used to determine the movement rate of the target object through the number of changes corresponding to the target object; configure the playback rate of the target special effects animation according to the movement rate, and drive the playback of the target special effects animation according to the playback rate, so that the played target special effects animation is adapted to the movement of the target object. The playback rate is the screen refresh frequency of the target special effects animation on the graphical user interface.

8. A storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the image processing method according to any one of claims 1 to 6 is implemented.

9. A computer system, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the image processing method according to any one of claims 1 to 6 by executing the executable instructions.

Citation Information

Patent Citations

  • Video special effect adding method and device, terminal equipment and storage medium

    CN109618183A