Video processing method, device, storage medium and program product
Patent Information
- Application Number
- CN202611260689.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-22
AI Technical Summary
然而在视频中目标对象需要形变处理的部位运动时,呈现出不同的形态,因此,对该部位进行形变处理导致相邻帧的背景扭曲区域、扭曲程度和扭曲方向不一致,这就会引起背景抖动,造成感观上的眩晕感
[0011]在本申请实施例中,根据形变区域在当前视频帧中的空间属性信息确定目标形态参数,量化待处理部位在形变过程中引发背景抖动的风险程度,并根据目标形态参数对初始形变强度进行衰减处理得到目标形变强度。由于背景抖动的根源在于形变区域的空间属性异常导致相邻帧间形变场的不连续,而目标形态参数直接表征了这种空间异常所对应的抖动风险程度,因此,通过对初始形变强度施加与该风险程度相匹配的衰减处理,可使作用于待处理部位的形变强度随空间属性的异常程度自适应降低,从而削弱因空间属性不稳定引发的帧间形变场差异,而缓解由于待处理部位的形态异常导致的背景抖动。
Smart Images

Figure CN122802640A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a video processing method, device, storage medium, and program product. Background Technology
[0002] In applications such as live streaming, short videos, and film special effects, adjusting the appearance of target objects like people in videos can not only present a more aesthetically pleasing and harmonious visual effect, but also compensate for perspective distortion caused by shooting distance, restore spatial proportions that are closer to natural human vision, and enhance the geometric realism of the image. Deformation technology is one of the key means to achieve this goal; it can adjust the overall or partial structure of an object, including changing its size, position, and angle.
[0003] The essence of deformation is image distortion and deformation technology. When a target object in the foreground of an image is distorted, the background is also distorted. However, when the part of the target object that needs to be deformed moves in a video, it presents different shapes. Therefore, deforming this part will cause the background distortion area, degree of distortion, and direction of distortion to be inconsistent between adjacent frames. This will cause background jitter and create a sense of dizziness. Summary of the Invention
[0004] This application provides a video processing method, device, storage medium, and program product to alleviate background jitter when processing the deformation of target objects in a video.
[0005] In a first aspect, embodiments of this application provide a video processing method, including: Identify the parts of the target object that require deformation treatment; Determine the deformation region of the part to be processed in the current video frame, and the deformation region corresponds to the initial deformation intensity; Based on the spatial attribute information of the deformed region in the current video frame, the target shape parameters of the part to be processed are determined; the target shape parameters characterize the risk level of background jitter caused by the part to be processed during the deformation process. Based on the target morphological parameters, the initial deformation intensity is attenuated to obtain the target deformation intensity; wherein, the attenuation magnitude of the initial deformation intensity attenuation is positively correlated with the risk level characterized by the target morphological parameters; Based on the deformation region and the target deformation intensity, the part to be processed in the current video frame is subjected to deformation processing to obtain the target video frame after deformation processing of the current video frame.
[0006] Optionally, before attenuating the initial deformation intensity according to the target morphological parameters, the method further includes: Obtain the first position information of the deformed region in the current video frame and the second position information of the historical deformed region of the part to be processed in the historical video frames; Based on the first location information and the second location information, determine the positional change of the deformed region relative to the historical deformed region; Based on the positional change, determine the target movement amplitude of the part to be processed; The step of attenuating the initial deformation intensity according to the target shape parameters to obtain the target deformation intensity includes: Based on the target shape parameters and the target motion amplitude, the initial deformation intensity is attenuated to obtain the target deformation intensity.
[0007] Secondly, embodiments of this application also provide a video processing method, including: Receives a video stream uploaded by a live streaming device; the video stream includes an image of the host in the live streaming room. Using the video processing method provided in the first aspect, the part of the video stream that needs to be deformed is deformed to obtain the target video stream; The target video stream is transmitted to the terminal device corresponding to the live broadcast room.
[0008] Thirdly, embodiments of this application also provide an electronic device, including: a memory and a processor; wherein the memory is used to store a computer program; The processor is coupled to the memory for executing the computer program to perform the steps in the first and / or second aspects of the method described above.
[0009] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps in the methods of the first and / or second aspects described above.
[0010] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the methods of the first and / or second aspects described above.
[0011] In this embodiment, target morphological parameters are determined based on the spatial attribute information of the deformed region in the current video frame. The risk level of background jitter caused by the deformation of the part to be processed during the deformation process is quantified. The initial deformation intensity is then attenuated based on the target morphological parameters to obtain the target deformation intensity. Since the root cause of background jitter lies in the discontinuity of the deformation field between adjacent frames due to the abnormal spatial attributes of the deformed region, and the target morphological parameters directly characterize the jitter risk level corresponding to this spatial abnormality, by applying an attenuation process to the initial deformation intensity that matches the risk level, the deformation intensity acting on the part to be processed can be adaptively reduced with the degree of spatial attribute abnormality, thereby weakening the difference in the deformation field between frames caused by the instability of spatial attributes and alleviating the background jitter caused by the morphological abnormality of the part to be processed. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 These are partial video frames extracted from the video before and after the target object in the video underwent arm-slimming deformation processing. Figure 2 A flowchart illustrating the video processing method provided in this application embodiment; Figure 3 A schematic diagram of the distribution of key points on the arm provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the process of determining the deformation region provided in an embodiment of this application; Figure 5 This is a schematic diagram illustrating the effect of an abnormal arm aspect ratio provided in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the effect of the arm going out of bounds in an embodiment of this application. Figure 7 This is a schematic diagram illustrating the conflict between the arm-slimming and torso-slimming effects provided in the embodiments of this application; Figure 8 A schematic diagram illustrating the functional relationship between the aspect ratio and attenuation factor of the rectangular deformation region provided in an embodiment of this application; Figure 9 This is a schematic diagram showing the effect of using the initial deformation intensity and the target deformation intensity (obtained by attenuating the initial deformation intensity after the attenuation factor determined by the aspect ratio) to perform arm-slimming deformation on the target object, as provided in the embodiments of this application. Figure 10 A schematic diagram illustrating the functional relationship between the percentage of the deformed region in a video frame and the attenuation factor, as provided in the embodiments of this application; Figure 11This is a schematic diagram showing the effect of using initial deformation intensity and target deformation intensity (obtained by attenuating the initial deformation intensity after using the attenuation factor determined by the attenuation factor of the deformation area in the video frame) to perform arm-slimming deformation on the target object, as provided in the embodiments of this application. Figure 12 A schematic diagram illustrating the functional relationship between the angle between the first deformation direction vector and the second deformation direction vector and the attenuation factor, provided for embodiments of this application; Figure 13 A flowchart illustrating another video processing method provided in an embodiment of this application; Figure 14 A schematic diagram illustrating the functional relationship between the motion amplitude of the part to be processed and the attenuation factor, provided for an embodiment of this application; Figure 15 This is a schematic diagram showing the effect of using initial deformation intensity and target deformation intensity (obtained by attenuating the initial deformation intensity after using the attenuation factor determined by the attenuation factor of the target motion amplitude of the part to be processed) to perform arm-slimming deformation on the target object according to the embodiments of this application. Figure 16 A schematic diagram illustrating the principle of the image liquefaction algorithm provided in this application embodiment; Figure 17 This is a flowchart illustrating the process of using the video processing method provided in this application to slim arms; Figure 18 This is a schematic diagram illustrating the before-and-after effects of using the video processing method provided in this application to slim arms; Figure 19 This is a schematic diagram of the structure of the video processing system provided in the embodiments of this application; Figure 20 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. The various neural network models involved in the embodiments of this application (such as keypoint detection models, etc.) comply with relevant laws and standards.
[0015] The core idea of video stream-based human body deformation algorithms is as follows: Key points of the target object are obtained from the input video frame; these key points are smoothed; the smoothed key points are used to obtain the local region of the target object to be processed within the video frame; and the pixel positions of this local region are changed to achieve deformation of the target object's desired area. Common image local deformation methods include image stretching, triangulation, and image liquefaction, among which image liquefaction is often used for human body deformation tasks where key points are relatively sparse.
[0016] Since the essence of deforming the target object is to change the position of the local image pixels where the part to be processed is located, the change in the position of the local image pixels not only distorts the part to be processed in the foreground of the image, but also changes the relative positional relationship between the pixels in the background area, causing the background to be distorted after deformation.
[0017] When a part of a target object that needs to be deformed moves, it will present different shapes. For example, when a human arm moves, it will present different extension states in the video frame, such as the arm straight, the arm bent, the arm raised, and the arm lowered; sometimes the arm will even go out of bounds, that is, it will be outside the video frame. These complex shapes will cause the background distortion area, distortion degree and distortion direction of adjacent frames to be inconsistent when deforming the arm, thus causing background shaking and causing visual dizziness.
[0018] For example, such as Figure 1 As shown, the deformed part of the target object is the arm, and the deformation effect is to make the arm look thinner. Figure 1 These are partial video frames extracted from the video before and after deformation (i.e., arm contraction deformation) of the target object's arm in the video. The left image shows a schematic diagram of the effect before deformation of the target object's arm in the video; Figure 1 The image on the right shows the effect of slimming down the arm of the target object in the video. Figure 1 The illustrations only depict the human body as the target object, but this does not constitute a limitation.
[0019] from Figure 1It can be seen that before deforming the arm of the target object in the video, the background remains unchanged in different frames, such as... Figure 1 (a) and Figure 1 In the left image of (b), the background remains unchanged. After slimming the arm of the target object in the video, the background of the video frames is also deformed. In scenes where the target object's arm moves continuously and forms complex shapes, the area, size, and direction of background distortion will vary between different frames. For example, Figure 1 (a) and Figure 1 In the right image of (b), the background is distorted, as shown below. Figure 1 (a) and Figure 1 As shown in the red rectangular dashed box in the right image of Figure (b). Figure 1 (a) and Figure 1 In the right image of (b), the background distortion area, size, and direction are also different. In this kind of scene where the part of the target object that needs to be deformed is constantly moving and has a complex shape, the deformation causes the background distortion area, size, and direction of adjacent frames to be inconsistent, which will cause video background jitter and create a sense of dizziness.
[0020] To mitigate background jitter caused by deformation processing in videos, in some embodiments of this application, a target morphological parameter is determined based on the spatial attribute information of the deformed region in the current video frame. This quantifies the risk level of background jitter caused by the deformation process of the part to be processed, and the initial deformation intensity is attenuated according to the target morphological parameter to obtain the target deformation intensity. Since the root cause of background jitter lies in the discontinuity of the deformation field between adjacent frames due to abnormal spatial attributes of the deformed region, and the target morphological parameter directly characterizes the jitter risk level corresponding to this spatial abnormality, by applying an attenuation process to the initial deformation intensity that matches this risk level, the deformation intensity acting on the part to be processed can adaptively decrease with the degree of spatial attribute abnormality, thereby weakening the difference in the deformation field between frames caused by unstable spatial attributes and mitigating the background jitter caused by the morphological abnormality of the part to be processed.
[0021] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0022] Figure 2 This is a schematic flowchart illustrating the video processing method provided in an embodiment of this application. Figure 2 As shown, the method mainly includes the following steps: 201. Determine the parts of the target object that require deformation treatment.
[0023] 202. Determine the deformation region of the part to be processed in the current video frame, and the deformation region corresponds to the initial deformation intensity.
[0024] 203. Based on the spatial attribute information of the deformed region in the current video frame, determine the target morphological parameters of the part to be processed; the target morphological parameters characterize the risk of the part to be processed causing background jitter during the deformation process.
[0025] 204. Based on the target morphological parameters, the initial deformation intensity is attenuated to obtain the target deformation intensity. The attenuation magnitude of the initial deformation intensity is positively correlated with the risk level characterized by the target morphological parameters.
[0026] 205. Based on the deformation area and the target deformation intensity, perform deformation processing on the parts to be processed in the current video frame to obtain the target video frame after deformation processing of the current video frame.
[0027] The video processing method provided in this application can perform offline video processing, such as deformation processing of the image of a target object in a pre-recorded video frame in film and television special effects scenes or short video scenes. The video processing method provided in this application can also perform online processing of real-time acquired video streams, such as receiving real-time acquired video streams from the broadcaster's end in live streaming applications and performing deformation processing on the broadcaster's image in the video stream.
[0028] In the embodiments of this application, the target object refers to a physical object with autonomous movement capabilities, such as a person, animal, or robot. During video capture, the parts of the target object that require deformation processing are not static but undergo shape changes. For example, a human arm may straighten, bend, raise, lower, or fold. To present a more aesthetically pleasing and harmonious visual effect, deformation processing can be applied to individual parts of the target object, such as slimming the arms, legs, waist, head, legs, or overall body shape. This not only presents a more aesthetically pleasing and harmonious visual effect but also compensates for perspective distortion caused by shooting distance, restores spatial proportions closer to natural human vision, and enhances the geometric realism of the image.
[0029] The part of the target object to be processed refers to the part of the target object that needs to be deformed. It can be selected by the target object itself, by the user initiating the video processing request, by the staff of the video processing server, or pre-configured by the video processing server, etc. None of these limitations are made in the embodiments of this application.
[0030] In order to perform deformation processing on the parts of the target object to be processed, Figure 2 In step 201, the area to be processed that requires deformation can be determined. For example, an interface for setting the area to be processed can be provided, through which the user sending the video processing request sets the area to be processed; or, a pre-set identifier of the area to be processed can be obtained, and the area corresponding to the identifier can be determined as the area to be processed.
[0031] A video frame is a video frame within the video of the target object being captured. The current video frame is denoted as video frame t. In film and television special effects scenes or short video scenes, the current video frame can be read from the video to be processed. In live streaming scenes, the current video frame uploaded by the broadcaster's terminal can be received. This current video frame includes the image of the broadcaster in the live streaming room. The broadcaster is the target object. In this embodiment, the electronic device that captures video of the target object can be in a fixed position and uses the same pose to shoot. The electronic device has video capture capabilities, such as a mobile phone, computer, or camera. In this embodiment, the current video frame includes an image of the target object, and the target object has a part to be processed. The current video frame includes a partial image of the part to be processed.
[0032] To perform deformation processing on the area to be deformed, the deformation region to be deformed must be determined from the current video frame, and the required deformation intensity must be determined. The deformation region determines where in video frame t the deformation will be performed, and the deformation intensity determines the magnitude of the deformation applied to that region. Accordingly, in Figure 2 In step 202, the deformation region of the part to be processed in the current video frame t can be determined. This deformation region corresponds to the initial deformation intensity.
[0033] In some embodiments, a deformation region setting component may be provided, which is used by the user sending the video processing request to set the deformation region in the current video frame t. For example, the user can manually mark the deformation region in the current video frame t using the deformation region setting component.
[0034] In other embodiments, the deformation region of the part to be processed in the current video frame t can be determined based on the image coordinates of the target key points of the part to be processed in the current video frame t. This deformation region corresponds to the initial deformation intensity. This implementation automatically determines the deformation region, eliminating the need for manual user intervention during the processing of each video frame, thus improving video processing efficiency. Furthermore, the image coordinates of the target key points of the part to be processed in the current video frame t can accurately determine the position information of the part to be processed within the current video frame. Compared to manual determination of the deformation region by the user, this improves the accuracy of the determined deformation region and enhances the subsequent deformation processing effect.
[0035] The following is an exemplary description of the specific implementation method for determining the deformation region of the part to be processed in the current video frame t, and determining the initial deformation intensity corresponding to the deformation region.
[0036] Specifically, keypoint detection can be performed on the current video frame t to determine the target keypoints of the area to be processed, i.e., their image coordinates within the current video frame t. Target keypoints refer to a set of points that can describe the key locations of the area to be processed. For example, ... Figure 3 As shown, the target key points of the arm include the midpoint of the shoulder, the shoulder boundary point, the midpoint of the elbow, the elbow boundary point, the midpoint of the wrist, and the wrist boundary point. Figure 3 The green dots indicate the location of key target points on the arm.
[0037] In this embodiment, a pre-trained keypoint detection model can be used to detect keypoints in the current video frame t, representing the area to be processed. The keypoint detection model can be a neural network model, and its training data can include images or videos of objects of the same species as the target object, with the keypoint locations pre-annotated. Objects of the same species as the target object refer to objects belonging to the same biological category as the target object. For example, if the target object is a human, then the objects of the same species are also humans; if the target object is a cat, then the objects of the same species are also cats, and so on.
[0038] When training a keypoint detection model, images or videos of similar objects to the target object are used as training samples, and pre-annotated keypoint locations are used as supervision to train the initial model. The loss function can be the difference between the keypoint locations predicted by the keypoint detection model and the pre-annotated keypoint locations, such as mean squared error or cross-entropy. Accordingly, with the objective of minimizing the loss function, images or videos of similar objects to the target object are used as training samples, and pre-annotated keypoint locations are used as supervision to train the initial model of the keypoint detection model until the training epochs reach a set threshold or the keypoint detection model converges.
[0039] Based on a pre-trained keypoint detection model, the current video frame t can be input into the model to detect keypoints of the target object in the current video frame t, obtaining the image coordinates of the target keypoints of the target area in the current video frame t. The image coordinates of the target keypoints of the target area determine the corresponding deformation region of the target area in the current video frame t.
[0040] Furthermore, the initial deformation intensity of the deformation region can be determined based on the image coordinates of the target key points of the area to be processed and the shape of the deformation region that adapts to the area to be processed. The shape of the deformation region in the video of the area to be processed can be determined by the shape of the area to be processed. For example, if the area to be processed is the head, the shape of the deformation region in the video of the area to be processed can be circular. If the area to be processed is the eye, the shape of the deformation region that adapts to the area to be processed can be circular or elliptical. If the area to be processed is the arm or leg, the shape of the deformation region that adapts to the area to be processed can be rectangular. If the area to be processed is the waist, the shape of the deformation region in the video of the area to be processed can be a symmetrical arc, trapezoid, or rectangle, etc.
[0041] This embodiment detects key points of the area to be processed in the current video frame, accurately obtaining the image coordinates of the target key points within the current video frame. The key point-based method for locating deformation regions offers high accuracy and robustness. By utilizing the detected image coordinates of the target key points in the current video frame, combined with a preset or learned adaptive deformation region shape, a deformation region matching the current shape and scale of the area to be processed can be dynamically calculated, helping to improve the naturalness and consistency of the deformation.
[0042] The following section provides an example of how to determine the deformation region of the part to be processed in the current video frame and the initial deformation intensity corresponding to the deformation region, using several parts to be processed as examples.
[0043] In some embodiments, the part to be processed is the human arm. The part to be processed includes the upper arm, and the target deformation region is rectangular. The target key points of the upper arm include the shoulder key point and the midpoint of the elbow. The shoulder key points include the midpoint of the shoulder and the inner and outer boundary points of the shoulder. Accordingly, the rectangular deformation region of the upper arm can be constructed based on the image coordinates of the shoulder key point and the midpoint of the elbow in the current video frame t. Specifically, the image coordinates of the midpoint of the shoulder in the current video frame can be... Extend the first length adjustment coefficient along the direction of the upper arm bones away from the midpoint of the elbow. k Using one times the length of the upper arm, the starting point of the target midline of the rectangular deformation region along the upper arm bone direction is obtained. Correspondingly, the image coordinates of the midpoint of the elbow in the current video frame t can also be... The second length adjustment coefficient is extended along the upper arm bone towards the direction away from the midpoint of the shoulder. k The upper arm length of 2 is used to obtain the endpoint of the target's central axis. Furthermore, the width can be adjusted based on the distance between the inner and outer boundary points of the shoulder and the set width adjustment coefficient. k The product of 3 determines the width of the deformable region of the rectangle. The distance between the inner and outer boundary points of the shoulder can be represented as the image coordinates of the inner and outer boundary points of the shoulder in the current video frame t. and The Euclidean distance between them, i.e. .in, k 1. k 2 and k 3 is a set constant. k 1. k 2 and k The value range of 3 is generally (0, 2).
[0044] Furthermore, it can be based on the starting point of the target's central axis. The endpoint of the target's central axis and the width of the rectangular deformation region Determine the deformation region of the rectangle, such as Figure 4 The deformation area is indicated by the red rectangle. The starting point of the target's central axis is also shown. The endpoint of the target's central axis and the width of the rectangular deformation region The rectangular deformation region of the upper arm is characterized by these features. The starting point of the target's central axis is also considered. End point of the target centerline The distance between them, such as the Euclidean distance, is the length of the deformed region of the rectangle. The starting point of the central axis of the aforementioned target. The endpoint of the target's central axis and the width of the rectangular deformation region , can be represented as: (1); (2); (3); (4).
[0045] In the above equations (1)-(3), The image coordinates of the midpoint of the shoulder in the current video frame t; The image coordinates of the midpoint of the elbow in the current video frame t; and These are the image coordinates of the inner and outer boundary points of the shoulder in the current video frame t, respectively. This represents the Euclidean distance between the inner and outer boundary points of the shoulder; and These represent the start and end points of the target central axis along the direction of the upper arm bone in the deformed region of the rectangle, respectively. Indicates the starting point of the target's central axis. End point of the target centerline The Euclidean distance between them; and These represent the length and width of the deformed region of the rectangle, respectively.
[0046] This embodiment expands the central axis endpoints outward at both ends, allowing the deformation area to fully cover the main deformation influence range of the upper arm muscle group, thus avoiding visual unnaturalness caused by the truncation of the deformation edge.
[0047] Accordingly, in embodiments where the area to be processed includes the upper arm and the target deformation region is rectangular, when determining the initial deformation intensity, the initial deformation intensity of the rectangular deformation region can also be determined based on the relative positions of each pixel within the rectangular deformation region to the target central axis. Specifically, the length of the rectangular deformation region can be determined by the distance between the starting point and the ending point of the target central axis. For any pixel within the rectangular deformation region (defined as a target pixel), the first projection length of the target line connecting the target pixel to the starting point of the target central axis perpendicular to the direction of the upper arm bones, and the second projection length of the target line parallel to the direction of the upper arm bones, can be calculated. Further, the first projection length and the width of the rectangular deformation region can be used as a reference. The vertical intensity component is obtained by normalizing and attenuating the first projected length. Correspondingly, the intensity component can also be determined based on the second projected length and the length of the deformed region of the rectangle. The second projected length is normalized and attenuated to obtain the parallel direction intensity component. Furthermore, the initial deformation intensity can be determined based on the vertical and parallel direction intensity components.
[0048] Specifically, the first projection length can be compared with the width of the deformed region of the rectangle. The ratio of the two projection lengths is used to input a preset vertical attenuation function to normalize the attenuation of the first projection length, thus obtaining the vertical intensity component; alternatively, the second projection length can be compared with the length of the deformed region of the rectangle. The ratio is input to a preset parallel direction attenuation function to normalize the attenuation of the second projection length, thus obtaining the parallel direction intensity component. Furthermore, the product of the vertical direction intensity component and the parallel direction intensity component can be determined as the initial deformation intensity of the target pixel. Using the same method, the initial deformation intensity of each pixel within the rectangular deformation region can be determined.
[0049] The vertical and parallel decay functions are both smooth, monotonically decreasing functions (such as the Sigmoid function or a variation of the Sigmoid function), which makes the initial deformation intensity exhibit a smooth decay from the center to the edge in both the direction perpendicular to the bone and the direction parallel to the bone.
[0050] For example, in some embodiments, the initial deformation intensity can be expressed as: (5); (6); (7).
[0051] In the above equations (5)-(7), Indicates the initial deformation intensity; This represents the first projection length of the target line along the direction perpendicular to the upper arm bones; This represents the second projection length of the target line along the direction parallel to the upper arm bones; This represents the initial deformation intensity function, whose independent variable is... . Indicates the intensity component in the vertical direction; This represents the intensity component in the parallel direction.
[0052] Because the upper arm muscles have an asymmetrical shape that extends longitudinally and narrows laterally along the bone structure, using a single radial attenuation (such as a circular area) would lead to a mismatch in the attenuation rates at both ends of the central axis and the sides, resulting in visual artifacts of edge truncation or overstretching. This embodiment decomposes the pixel position into a first projection length perpendicular to the bone structure and a second projection length parallel to the bone structure. The first and second projection lengths are then independently normalized and attenuated based on their width and length, respectively. This ensures that the contour lines of deformation intensity naturally conform to the aspect ratio of the rectangular area, accurately matching the anatomical distribution characteristics of the upper arm muscles. This ensures that the slimming effect is concentrated in areas with full muscles, avoiding ineffective deformation on joints or areas with thin skin.
[0053] In other embodiments, the area to be processed includes the forearm, and the target deformation region is rectangular. The target key points of the forearm include the midpoint of the elbow, the wrist key point, and the inner and outer boundary points of the elbow. Accordingly, the rectangular deformation region of the forearm can be constructed based on the image coordinates of the midpoint of the elbow and the wrist key point in the current video frame t. Specifically, the image coordinates of the midpoint of the elbow can be extended along the forearm bone direction away from the wrist key point by a third length adjustment coefficient. k Four times the forearm length, the starting point of the first central axis along the forearm bone direction of the rectangular deformation region of the forearm is obtained; the image coordinates of the wrist key points are extended along the forearm bone direction away from the midpoint of the elbow by a fourth length adjustment coefficient. k The lower arm length of 5 is used to obtain the endpoint of the first central axis; the distance between the inner and outer boundary points of the elbow is adjusted with the set lower arm width adjustment coefficient. k The product of 6 is used to determine the width of the rectangular deformation region of the lower arm. The third and fourth length adjustment coefficients can be the same as or different from the first and second length adjustment coefficients of the upper arm, to accommodate the different muscle distribution characteristics of the upper and lower arms.
[0054] Correspondingly, the initial deformation intensity of each pixel within the rectangular deformation region of the lower arm is determined in a similar manner to that of the upper arm. Specifically, the distance between the starting point and the ending point of the first central axis, such as the Euclidean distance, can be determined as the length of the rectangular deformation region of the lower arm. For any pixel within the rectangular deformation region of the lower arm (defined as the first pixel), the third projection length of the line connecting the first pixel to the starting point of the first central axis (defined as the first line) perpendicular to the direction of the lower arm bones, and the fourth projection length of the first line parallel to the direction of the lower arm bones are calculated.
[0055] Furthermore, the third projection length can be normalized and attenuated based on the third projection length and the width of the rectangular deformation region of the lower arm to obtain the vertical intensity component; and the fourth projection length can be normalized and attenuated based on the fourth projection length and the length of the rectangular deformation region of the lower arm to obtain the parallel intensity component. Furthermore, the initial deformation intensity of the first pixel can be determined based on the vertical and parallel intensity components. The initial deformation intensity of each pixel within the rectangular deformation region of the lower arm can be determined using the same method.
[0056] For details on how to determine the initial deformation intensity of each pixel within the deformation region of the lower arm, please refer to the aforementioned implementation method for the initial deformation intensity of each pixel within the deformation region of the upper arm. Since the lower arm has more tendons and thinner muscles near the wrist, the decay rate of the parallel direction decay function for the lower arm can be set to be greater than that for the upper arm. This concentrates the deformation intensity in the more muscular mid-section of the lower arm, enhancing the naturalness of a slimmer lower arm.
[0057] In other embodiments, the area to be processed is the eye, and the target deformation region is elliptical. The target key points of the eye include the inner and outer corners of the left and right eyes, the center points of the left and right pupils, and the key points of the upper and lower eyelids of the left and right eyes. Accordingly, the horizontal span of the left eye can be calculated based on the image coordinates of the inner and outer corners of the left eye in the current video frame t; the vertical span of the left eye can be calculated based on the key points of the upper and lower eyelids. The elliptical deformation region of the left eye is constructed by using the midpoint of the line connecting the inner and outer corners of the left eye as the center, a preset proportion of the horizontal span of the left eye as the major axis of the ellipse, and a preset proportion of the vertical span of the left eye as the minor axis. The elliptical deformation region of the right eye is constructed in the same way as the left eye. Using an ellipse as the shape of the deformation region better conforms to the horizontally extended anatomical contour of the eye, avoiding unnecessary cascading deformation of the eyebrows or cheekbone area.
[0058] Accordingly, for any pixel within the elliptical deformation region (defined as the second pixel), the initial deformation intensity of each pixel can be determined based on the normalized elliptical distance from the second pixel to the center of the ellipse. Specifically, the normalized elliptical distance is obtained by calculating the square root of the normalized coordinates of the second pixel along the major and minor axes of the ellipse. This normalized elliptical distance is then input into a preset radial decay function (such as the sigmoid function) to obtain the initial deformation intensity of the pixel. The initial deformation intensity is greatest when the normalized elliptical distance is 0 (i.e., at the center) and least when the normalized elliptical distance is 1 (i.e., at the ellipse boundary). This radial decay characteristic ensures a smooth transition from the center of the pupil to the outer edges of the eye, avoiding a clear deformation boundary between the eyeball and the sclera, thus conforming to the natural perception of eye enhancement by the human eye.
[0059] In some embodiments, the area to be processed is the leg, and the target deformation region is trapezoidal. The target key points of the leg include hip joint key points, knee joint key points, ankle key points, and the inner and outer boundary points of each joint. Considering the anatomical feature of the human leg gradually narrowing from top to bottom, a trapezoidal shape better fits the leg contour than a rectangle. Accordingly, the thigh segment length can be calculated based on the image coordinates of the hip and knee joint key points in the current video frame t. The thigh segment length is multiplied by the set upper and lower base coefficients to obtain the hip joint lateral width and knee joint lateral width, respectively. The upper base coefficient is greater than the lower base coefficient. Using two points of the hip joint key point offset by half the hip joint lateral width along the thigh normal, and two points of the knee joint key point offset by half the knee joint lateral width along the thigh normal, as the four vertices of the thigh trapezoid, the trapezoidal deformation region of the thigh is constructed. The trapezoidal deformation region of the lower leg is constructed in the same way, using the knee joint lateral width and ankle lateral width as the upper and lower bases.
[0060] Accordingly, if the purpose of the deformation is to slim the legs, the initial deformation intensity of any pixel point (defined as the third pixel point) within the trapezoidal deformation area can be determined based on the vertical distance from the third pixel point to the trapezoid's midline. Specifically, the vertical distance from the third pixel point to the trapezoid's midline (i.e., the line segment connecting the midpoints of the upper and lower bases) can be calculated, and this vertical distance can be divided by the local half-width of the trapezoid at that vertical position to obtain the normalized horizontal distance. The normalized horizontal distance is then input into a preset smooth, monotonically decreasing function to obtain the initial deformation intensity of the third pixel point. This embodiment allows the deformation intensity of the leg-slimming effect to smoothly decay from the midline of the leg towards the sides, and the decay rate is adaptively adjusted according to the width of the trapezoid—the decay gradient is gentler where the thigh is wider and steeper where the calf is narrower, ensuring the visual uniformity of the deformation intensity across the entire leg.
[0061] If the purpose of deformation is to lengthen the legs, the initial deformation intensity can be determined based on the longitudinal distance from the third pixel to the vertical line of the trapezoid's center line (i.e., the horizontal center line), so that the stretching force decreases smoothly from the middle of the legs to the upper and lower ends.
[0062] In other embodiments, the area to be processed is the waist, and the target deformation region is an hourglass shape or a combination of two trapezoids. The key target points for the waist include the midpoint of the lower thoracic rim, the midpoint of the narrowest part of the waist, the midpoint of the upper edge of the pelvis, and the left and right boundary points at each horizontal position. Since the waist has a natural inward curve, neither a rectangle nor a single trapezoid can accurately fit it. Accordingly, an upper waist trapezoid can be constructed based on the image coordinates of the midpoint of the lower thoracic rim and the midpoint of the narrowest part of the waist in the current video frame t. The upper base is the width of the thoracic rim, and the lower base can be the width of the narrowest part of the waist. Further, a lower waist trapezoid can be constructed based on the image coordinates of the midpoint of the narrowest part of the waist and the midpoint of the upper edge of the pelvis. The upper base is the width of the narrowest part of the waist, and the lower base is the width of the pelvis. Further, the upper and lower trapezoids are joined along the narrowest part of the waist to form an hourglass-shaped deformation region of the waist.
[0063] Accordingly, for each pixel within the hourglass-shaped deformation region, the initial deformation intensity of each pixel can be determined based on the normalized lateral distance from each pixel to the waist's central axis. Specifically, for any pixel within the hourglass-shaped deformation region (defined as the fourth pixel), the vertical height position of the fourth pixel can be determined. Then, based on the vertical height position, interpolation is performed in the upper or lower waist trapezoid to obtain the local half-width at that height. The lateral distance from the fourth pixel to the central axis of the hourglass-shaped deformation region is divided by this local half-width to obtain the normalized lateral distance. The normalized lateral distance is input into a preset smooth, monotonically decreasing function to obtain the initial deformation intensity of the fourth pixel. This adaptive normalization mechanism causes the deformation intensity of the waist to decrease relative to the current body width. Regardless of changes in waist curvature, the contour lines of deformation intensity remain parallel to the waistline, avoiding the local over- or under-pressure problems caused by fixed geometric models on non-standard body types.
[0064] The foregoing embodiments only illustrate the implementation method of identifying the deformation region of the part to be processed in the current video frame t and determining the initial deformation intensity of the deformation region, using rectangular, elliptical, trapezoidal, and hourglass-shaped deformation regions as examples, and the parts to be processed as the upper arm, lower arm, eye, leg, and waist. This is not intended to be limiting. Those skilled in the art can select appropriate deformation region shapes and intensity attenuation functions based on the anatomical characteristics and aesthetic requirements of the part to be processed, all of which fall within the protection scope of this application.
[0065] In other embodiments, when determining the initial deformation intensity of the deformation region of the part to be processed in the current video frame t by matching the image coordinates of the target key point of the part to be processed with the shape of the target deformation region adapted to the part to be processed, the image coordinates of the target key point in the current video frame t can also be temporally smoothed to obtain the target image coordinates of the target key point in the current video frame t. Specifically, the image coordinates of the target key point in historical video frames can also be obtained. In the embodiments of this application, for ease of description and distinction, the image coordinates of the target key point in the current video frame t are defined as the first image coordinates; and the image coordinates of the target key point in historical video frames are defined as the second image coordinates. Historical video frames are the N previous historical video frames in the same video frame t. N is a positive integer. For example, if the historical video frame is the previous video frame of the current video frame t. If the number of historical video frames in the current video frame t is less than N, then all video frames before the current video frame are determined as historical video frames. Furthermore, the first image coordinates can be temporally smoothed based on the second image coordinates to obtain the target image coordinates of the target key points in the current video frame.
[0066] Specifically, a smoothing algorithm can be used to smooth the first image coordinates using the second image coordinates. The smoothing algorithm can be Gaussian smoothing or ridge regression smoothing, etc. When using Gaussian smoothing, the second image coordinates corresponding to the previous N historical video frames and the first image coordinates corresponding to the current video frame t are combined to form an image coordinate sequence of length (N+1), based on the temporal order of the previous N historical video frames and the current video frame t. The sequence number i of each image coordinate in the image coordinate sequence is input into a Gaussian function with i as the independent variable and Gaussian weights as the dependent variable to obtain the weight of each image coordinate. Here, i = 0, 1, ..., N. Then, the Gaussian weights of each image coordinate can be normalized to obtain the normalized weights of each image coordinate. Further, using the normalized weights of each image coordinate, a weighted sum of the second image coordinates corresponding to the previous N historical video frames and the first image coordinates corresponding to the current video frame t can be performed to smooth the first image coordinates, obtaining the target image coordinates of the target keypoint in the current video frame t. In this embodiment, a Gaussian smoothing algorithm is used to perform temporal smoothing on key points in the current video frame. Gaussian smoothing weakens the influence of isolated outliers through weighted averaging, making the trajectory smoother and significantly reducing visual "flickering" or "jitter". On the other hand, the image coordinates closer to the current time have a greater weight, and the more recent historical video frames have a greater influence on the current frame, which conforms to the assumption of motion continuity. This makes the smoothed key point sequence more coherent in the time dimension, which is consistent with the human eye's perception of natural motion.
[0067] When using ridge regression smoothing as the smoothing algorithm, the motion trajectory of the part of the target object to be processed can be simulated based on the second image coordinates corresponding to the previous N historical video frames, obtaining a motion trajectory function that changes the image coordinates over time. Then, the relative time between the current video frame t and the previous N historical video frames can be substituted into the motion trajectory function for calculation, yielding the target image coordinates corresponding to the current video frame t. This embodiment models the dynamic motion of the part of the target object to be processed using the image coordinates of the target key points in historical video frames, obtaining high-precision and physically consistent target image coordinates. This makes the target image coordinates of the key points in the current video frame t more coherent in the time dimension, conforming to the human eye's perception of natural motion.
[0068] After obtaining the target image coordinates of the target key points in the current video frame, the deformation region can be determined based on the shape of the deformation region that matches the target image coordinates and the area to be processed. For the specific implementation of determining the deformation region based on the shape of the deformation region that matches the target image coordinates and the area to be processed, please refer to the aforementioned content on determining the deformation region based on the target image coordinates of the target key points in the current video frame and the shape of the deformation region that matches the area to be processed; it will not be repeated here.
[0069] In this embodiment, the second image coordinates corresponding to the historical video frame can be the smoothed image coordinates corresponding to the historical video frame. By using the image coordinates of the target key points in the historical video frame to smooth the image coordinates of the target key points in the current video frame, the temporal smoothing of the key points can be achieved, which can effectively suppress high-frequency noise and improve the temporal stability of the key point trajectory. Therefore, it can alleviate the jitter caused by background deformation due to key point jitter and jumps.
[0070] However, in video streams, the difference in deformation regions between adjacent frames of the liquefaction algorithm depends not only on the target keypoints but also on the motion amplitude of the target object and the deformation intensity of the region to be deformed. Therefore, smoothing only the keypoints of the target object can only alleviate background jitter caused by keypoint jumps, and cannot further alleviate background jitter caused by the complex shape of the region to be processed. When the region to be processed is in a non-standard state such as complex extension, partial out-of-bounds, conflicting with the deformation direction of other parts, or rapid movement, if deformation is still performed according to the initial deformation intensity, it will lead to background jitter.
[0071] To mitigate background jitter caused by the complex shape of the area to be processed, this application introduces a deformation intensity attenuation mechanism based on the spatial attribute information of the deformed region in the current video frame, which will be explained in detail below.
[0072] Accordingly, in Figure 2In step 203, the target morphological parameters of the part to be processed can be determined based on the spatial attribute information of the deformation area of the part to be processed in the current video frame t.
[0073] Spatial attribute information refers to the static spatial configuration attributes determined based on the geometry, size, orientation, and relative positional relationship of the deformed region within the current video frame to other parts or image boundaries. Target morphological parameters are used to characterize the risk of background jitter caused by the deformed part during the deformation process.
[0074] Target morphological parameters can be used to quantify the risk level of background jitter caused by geometric anomalies or spatial layout conflicts in the current video frame for the part to be processed. The target morphological parameters are calculated based on the static spatial configuration attributes of the deformed region within the current video frame, without including cross-frame positional changes or motion trajectory information, thereby decoupling geometric morphological risks from temporal motion risks and enabling differentiated attenuation strategies for different types of risk sources.
[0075] The following section provides an illustrative example of the process for determining the target morphological parameters of the area to be processed, using the specific content of spatial attribute information. Each implementation method corresponds to a specific risk detection dimension, which can capture geometric anomalies that may cause jitter from different angles.
[0076] Implementation Method 1: Since the area to be processed has stable geometric proportions in its normal extended state—for example, when the arm is naturally extended, its length is greater than its width, and the corresponding rectangular deformation area is a long, thin rectangle, meaning its length is much greater than its width—if the area to be processed is in a complex extended state, such as when the arm is bent in front of the body, the lower arm portion will be invisible when the video is captured from behind the body. For example, ... Figure 5 The illustrated human image shows the lower arm obscured by the body, leading to keypoint detection position shifts and abnormal deformation region construction. Applying deformation at the initial intensity under these conditions would increase inter-frame jitter in the background. Therefore, the projection span of the deformation region's boundary contour across multiple preset orthogonal directions can be used as a spatial attribute.
[0077] Accordingly, spatial attribute information may include the projected span of the boundary contour of the deformed region along multiple preset orthogonal directions. These preset orthogonal directions refer to two or more mutually perpendicular reference axes pre-defined in the local coordinate system of the part to be treated. These orthogonal directions are anchored to the anatomical structure or geometric principal axis of the part itself and dynamically adapt to the part's posture and orientation. Because the deformation-sensitive dimensions of each part of the target object have a high degree of direction specificity, if the projected span is measured only in the global coordinate system, when the part rotates, tilts, or adopts a non-standard posture, the projected values of the global axes will not accurately reflect the true geometric extension state of the part, leading to morphological parameter distortion. By binding the orthogonal directions to the part's own structure, it can be ensured that the measurement of the projected span is always along its functional principal axis regardless of the part's posture, thereby obtaining stable and comparable geometric measurements.
[0078] The following uses the arm or robotic arm, leg, eye, and waist as examples to define and explain the preset orthogonal directions for each part. For the arm or robotic arm, multiple preset orthogonal directions may include: a first direction parallel to the skeletal orientation of the arm and a second direction perpendicular to the skeletal orientation of the arm. For the leg, multiple preset orthogonal directions may include: a direction parallel to the skeletal orientation of the leg and a direction perpendicular to the skeletal orientation of the leg. For the eye, multiple preset orthogonal directions may include: a direction parallel to the line connecting the left and right corners of the eye, and a direction perpendicular to the line connecting the left and right corners of the eye and located in the frontal plane of the face. For the waist, multiple preset orthogonal directions may include: a direction parallel to the tangent of the spine at the waist and a direction perpendicular to the aforementioned tangent and located in the frontal projection plane of the torso; and so on. The preset orthogonal directions for each part can be calculated in real time through target key points and are dynamically updated according to the posture of the part, rather than being predefined fixed angles. This adaptive orthogonal orientation setting mechanism enables the projection span ratio in the target morphological parameters to maintain a stable physical meaning under any complex posture, providing a reliable and consistent geometric risk measurement basis for subsequent deformation intensity attenuation.
[0079] Accordingly, the target ratio of the projected span of the boundary contour in multiple preset orthogonal directions can be calculated, and this target ratio can be used as the target morphological parameter. This implementation method is mainly used to measure the risk of morphological abnormalities in the parts to be processed (such as arms, legs, etc.) in complex extended states.
[0080] In embodiments where the area to be processed is the arm, under normal extended posture, the rectangular deformation region of the arm generally appears as a slender rectangle with a length greater than its width. However, during actual arm deformation, when the arm is in a backward, twisted, or unnatural extended posture, the fluctuation range of key point detection increases, causing the aspect ratio of the constructed rectangular deformation region to deviate from the normal human body proportions. For example, as... Figure 5As shown, the aspect ratio of the deformed area in the lower left arm is abnormal; that is, the shorter side within the blue rectangle is the longer side. The longer side is the width. This abnormal aspect ratio indicates inaccurate keypoint localization, and continued deformation could easily cause background jitter. Therefore, the aspect ratio of the deformed rectangular region can be defined as a specific form of the target ratio, and its calculation formula can be expressed as: (8).
[0081] in, The aspect ratio of the deformed region of the rectangle. This represents the length of the longer side of the deformed region of the rectangle. This represents the width of the shorter side of the deformed region of the rectangle. When... When the deformation deviates significantly from the preset normal range, it indicates that the arm is in a complex extended state. In this case, it is necessary to reduce the initial deformation intensity to weaken the deformation effect on the deformation area and alleviate the background shaking caused by the deformation.
[0082] This implementation quantifies the degree of geometric distortion in the deformable region by the ratio of the projection span. It can automatically identify morphological anomalies in complex poses where keypoint detection is unreliable, alleviate deformation field distortion and background jitter caused by incorrect keypoint localization, and help improve the robustness of the algorithm in non-standard poses.
[0083] Implementation Method 2: When the area to be processed is completely within the frame, each pixel in the deformed region has valid image data support, and deformation processing can be performed normally. If the area to be processed partially or completely exceeds the frame boundary of the current video frame, for example, when an arm extends outward or a person approaches the edge of the frame, the portion of the deformed region outside the frame will lack pixel information, forming an invalid deformed region. For example, as... Figure 6 As shown, the human's right arm extends beyond the frame, meaning the right arm is out of bounds. The blue rectangular deformation area does not represent the image of the right arm and is considered an invalid region. If deformation is applied at its full intensity to the deformation area containing the invalid region, the background texture at the edge of the image will be forcibly stretched and distorted, resulting in noticeable in-and-out image jitter and visual artifacts. Therefore, the area of the deformation region and the area overlapping between the deformation region and the current video frame can be considered as spatial attribute information.
[0084] Accordingly, spatial attribute information may include: the area of the deformed region and the overlap area between the deformed region and the current video frame. The area of the deformed region refers to the theoretical area of the complete deformed region constructed in the current video frame based on the target key points; the overlap area refers to the actual visible area of the portion where the deformed region intersects with the effective display area of the current video frame. Since the overlap area ratio directly reflects the proportion of effective pixels in the deformed region, it is a direct geometric basis for determining whether the area to be processed is out of bounds.
[0085] Accordingly, a target ratio of the overlapping area of the image to the area of the deformed region can be calculated, and this target ratio can be used as a target shape parameter. This implementation method is mainly used to measure the out-of-bounds state of the part to be processed and the risk of edge jitter caused by it.
[0086] In the embodiment where the part to be treated is the arm, the degree to which the arm exceeds the boundary can be defined as a specific form of the target proportion, and its calculation formula can be expressed as: (9).
[0087] in, This indicates the percentage of the area within the image where the arm is deformed. This represents a rectangular area of the same size as the current video frame. This indicates the deformation area calculated based on key point detection. This represents the actual visible area of the deformed region on the screen, that is, the area of overlap between the deformed region and the current video frame. This represents the area of the deformed region calculated based on key point detection. For example, as... Figure 6 As shown, the right arm, located on the left side of the image, is mostly outside the frame. The blue rectangle represents the complete deformation area calculated based on key points. The area where the blue rectangle intersects with the image border is the overlapping area of the image. .when When the value equals 1, it indicates that the deformed area is completely within the image frame and there is no risk of it going out of bounds; when... When the value is less than 1, it indicates that the deformation region has exceeded its boundary, and The smaller the value, the more severe the deviation from the boundary, and the higher the risk indicated by the target shape parameters. When When the value is below the preset safety threshold, it indicates that the arm is significantly out of bounds. In this case, the initial deformation intensity needs to be reduced to weaken the deformation effect on the deformation area, especially the edge area, and avoid background distortion caused by invalid pixels participating in deformation calculation.
[0088] This implementation method can accurately sense the boundary status of the part to be processed by quantifying the degree of overlap between the deformed area and the image. During the transition phase when the part enters or leaves the image, the deformation intensity is adaptively reduced, which effectively avoids edge distortion and background jitter caused by invalid pixel areas participating in deformation calculation, and helps to ensure the visual stability at the image boundary.
[0089] Implementation Method 3: Since the deformation direction of the part to be processed is usually consistent with its own anatomical contraction direction, for example, the deformation direction of a slimmer arm is perpendicular to the direction of the arm bones, and the deformation direction of a slimmer body contracts inward along the width of the torso. When a single function is activated, the deformation field naturally matches the body structure. If the deformation function of multiple parts is activated simultaneously, and the deformation directions of different parts are significantly different or even opposite, for example, the lateral compression direction of a slimmer arm is inconsistent with the lateral contraction direction of a slimmer body, then the two deformation fields will interfere with, cancel each other out, or superimpose in the same spatial area. If the deformation in conflicting directions is not suppressed, it will lead to unintended body deformation, such as "fatting" the torso when slimming the arms, destroying the overall aesthetic effect and causing background texture disorder. For example, Figure 7 As shown, slimming arms and slimming body are activated simultaneously. The white arrow indicates the lateral compression direction of slimming arms, which is inconsistent with the lateral contraction direction of slimming body, which is indicated by the green arrow. When slimming arms, the body can be "plumped up".
[0090] Therefore, the first deformation direction vector of the deformable region can be considered as a spatial attribute information. Correspondingly, the spatial attribute information may include the first deformation direction vector of the deformable region. Here, the first deformation direction vector refers to the principal direction of the deformation force indicated by the deformable region of the part to be processed. For the arm, this direction is defined as the direction perpendicular to the direction of the arm bones, pointing from the edges of the two long sides of the arm region. The spatial relationship between deformation directions is the core geometric basis for determining whether the deformation functions of multiple parts conflict.
[0091] Correspondingly, the second deformation direction vector of the associated parts of the part to be processed can be obtained. The associated parts refer to adjacent or functionally related parts that have a direct coupling relationship with the part to be processed in terms of anatomical structure, kinematic linkage, or spatial distribution of the deformation field, and whose own deformation state may interfere with, superimpose, or conflict with the deformation effect of the part to be processed. Because the various parts of the human body do not exist in isolation, when multiple parts participate in deformation processing simultaneously, if only the local geometric properties of a single part are considered while ignoring its spatial interaction with surrounding parts, it is very easy to cause unintended interference between deformation fields, leading to disproportionate body shape or disordered background texture. By incorporating the associated parts into the calculation dimension of morphological parameters, deformation risk can be assessed from a global collaborative perspective, rather than being limited to single-point local optimization.
[0092] When the area to be treated is the arm, the associated area is the shoulder or torso. This is because if the lateral compression deformation direction of the slimming arm is inconsistent with the slimming contraction direction of the torso, a deformation field conflict will occur at the junction of the shoulder and arm, causing the torso to be accidentally widened or the shoulder contour to be distorted. When the area to be treated is the eyes, the associated areas are the facial contours or eyebrows. This is because if the vertical deformation of the eyes is excessively extended, it will cause the eyebrows to move downward or pull the skin of the temples, thus disrupting the overall proportion and harmony of the face.
[0093] When the area to be treated is the legs, the associated areas are the buttocks or waist. This is because if the longitudinal deformation of slimming or lengthening the legs does not match the shaping direction of the buttocks and waist area, it will create a force discontinuity at the hip joint, resulting in a stiff connection between the legs and buttocks or a shift in the waistline position. When the area to be treated is the waist, the associated areas are the chest or buttocks. This is because if the horizontal narrowing deformation of the waist goes against the natural width change trend of the ribcage or pelvis, it will create compression artifacts at the junction of the upper and lower parts, disrupting the smooth transition of the hourglass curve.
[0094] The selection of the aforementioned associated parts is not exhaustive. In actual implementation, associated parts can be dynamically determined based on the currently enabled deformation function combination: if only single-part beautification is enabled, the associated parts can be set as default adjacent parts; if multiple-part collaborative beautification is enabled simultaneously, all parts with enabled deformation functions will be automatically identified as associated parts, and the directional angles or spatial overlap relationships between them will be calculated, thereby constructing a multi-part conflict detection system.
[0095] The second deformation direction vector of the associated site refers to a vector, predefined or calculated in real time, representing the principal direction of the deformation force of the associated site itself, in a site that has a spatial interaction relationship with the site to be treated. The specific definition of the second deformation direction vector dynamically adapts to the anatomical structure and deformation function type of the associated site, rather than using a uniform, globally fixed direction. When the associated site is the trunk and the deformation function is slimming, the second deformation direction vector is defined as an inward contraction vector in the width direction of the trunk, i.e., a horizontal vector perpendicular to the longitudinal axis of the spine and pointing towards the midline of the body. When the associated site is the facial contour and the deformation function is face slimming, the second deformation direction vector is defined as an inward narrowing vector in the width direction of the face, i.e., a horizontal vector perpendicular to the midline of the face and pointing towards the bridge of the nose. When the associated site is the buttocks and the deformation function is buttock lifting, the second deformation direction vector is defined as a longitudinal lifting vector in the buttocks, i.e., a vertical vector parallel to the lower segment of the spine and pointing upwards. When the associated part is the eyebrow and the deformation function is to adjust the eyebrow shape, the second deformation direction vector is defined as the tangent direction vector of the eyebrow arch arc, that is, the local curve tangent vector along the natural direction of the eyebrow.
[0096] All the above definitions are anchored to the functional axis of the associated part itself, and are at the same semantic level as the first deformation direction vector of the part to be processed, ensuring that the angle calculation between the two has a clear physical meaning. The method of obtaining the second deformation direction vector can be consistent with the first deformation direction vector, that is, it is calculated in real time through the image coordinates of the target key points of the associated part in the current video frame, and is dynamically updated according to the pose and deformation function configuration of the associated part.
[0097] Furthermore, the target angle between the first deformation direction vector and the second deformation direction vector can be calculated, and this target angle can be used as the target shape parameter. This implementation method is mainly used to measure the risk of directional conflict when deformation functions for multiple parts such as arm slimming and body slimming are activated simultaneously.
[0098] In an embodiment where the area to be treated is the arm and the associated area is the torso, the degree of deviation between the arm's deformation direction and the torso's contraction direction can be defined as a specific form of the target angle, and its calculation formula can be expressed as: (10).
[0099] in, A This represents the angle between the first deformation direction vector (such as the arm deformation direction vector) and the second deformation direction vector (such as the trunk contraction direction vector). Let the vector be the direction of arm deformation. The vector representing the direction of torso contraction is defined as the vector along the width of the torso. "·" indicates the dot product operation. and These represent the magnitudes of the arm deformation direction vector and the torso contraction direction vector, respectively. For example, as... Figure 7 As shown, the white arrow indicates the direction of deformation of the lower right arm. The green arrow indicates the direction of trunk contraction. When the angle between the two A A larger value indicates a significant conflict between the deformation directions of slimming arms and slimming the body. When A When the angle approaches 90 degrees or exceeds the preset conflict threshold, it indicates that the two deformation directions are highly contradictory, and the risk of background shaking caused by the target shape parameters increases. In this case, it is necessary to reduce the initial deformation intensity to weaken the associated effects of arm deformation on the torso area and avoid making the body "fatter" during the arm-slimming operation.
[0100] This implementation method quantifies the spatial relationship between the deformation directions of different parts, which can automatically detect the conflict state of the deformation functions of multiple parts. When the deformation directions are contradictory, the deformation intensity of the part to be processed is reduced in time, avoiding the imbalance of body proportions and background texture disorder caused by mutual interference of deformation fields. This helps to alleviate background jitter and also ensures the overall visual coordination when multiple parts are beautified together.
[0101] By combining the above implementation methods, a single-frame background jitter risk perception mechanism independent of temporal information is constructed by extracting target morphological parameters from independent static dimensions such as geometric distortion, image boundary relationships, and cross-part spatial conflicts. This mechanism can identify various abnormal states that may cause jitter based on the spatial configuration of the current frame without introducing historical frame data, and provides accurate quantitative basis for subsequent deformation intensity attenuation processing, thereby suppressing background jitter risk.
[0102] In practical use, one of the aforementioned embodiments 1-3 can be implemented, or one or more of embodiments 1-3 can be combined. "Multiple" refers to two or more (including two). When multiple combined implementations are selected, the target morphological parameters include multiple target morphological parameters determined by the multiple implementation methods combined. For example, when embodiments 1 and 2 are combined, the target morphological parameters include: the target ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions, and the target proportion of the overlapping area of the image to the area of the deformed region. When embodiments 1 and 3 are combined, the target morphological parameters include: the target ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions, and the target angle between the first deformation direction vector and the second deformation direction vector. When embodiments 2 and 3 are combined, the target morphological parameters include: the target proportion of the overlapping area of the image to the area of the deformed region, and the target angle between the first deformation direction vector and the second deformation direction vector. When implemented in combination with embodiments 1-3, the target morphological parameters include: the target ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions, the target proportion of the overlapping area of the image to the area of the deformed region, and the target angle between the first deformation direction vector and the second deformation direction vector.
[0103] After determining the target morphological parameters of the area to be processed, Figure 2 In step 204, the initial deformation intensity can be attenuated according to the target morphological parameters to obtain the target deformation intensity. The attenuation magnitude of the initial deformation intensity is positively correlated with the risk level represented by the target morphological parameters. That is, the greater the risk level represented by the target morphological parameters, the stronger the attenuation magnitude of the initial deformation intensity.
[0104] In some embodiments, a first mapping relationship between the attenuation factor and the morphological parameter can be preset. The first mapping relationship defines the conversion rule from the target morphological parameter to the attenuation factor. The first mapping relationship can be a function expression, a lookup table, or a piecewise mapping model, which is not limited in this application. The attenuation factor is negatively correlated with the attenuation amplitude. In the first mapping relationship, the attenuation factor ranges from (0,1). Since the target morphological parameter characterizes the relative degree of background jitter risk, its value itself cannot be directly used as an adjustment amount for deformation intensity. If the target morphological parameter is directly calculated with the initial deformation intensity, the adjustment result will be uncontrollable due to the mismatch of parameter dimensions and inconsistent value ranges. Therefore, this embodiment introduces the concept of an attenuation factor as an intermediate conversion variable connecting risk measurement and deformation intensity adjustment. The attenuation factor refers to a dimensionless coefficient used to adjust the initial deformation intensity, and its value range is usually limited to the interval (0,1); the attenuation amplitude refers to the proportion by which the initial deformation intensity is reduced. The attenuation factor is negatively correlated with the attenuation amplitude. That is, the larger the attenuation factor, the smaller the attenuation amplitude and the higher the retained deformation intensity; the smaller the attenuation factor, the larger the attenuation amplitude and the lower the retained deformation intensity. When the attenuation factor approaches 1, it indicates no attenuation, and the target deformation intensity approaches the initial deformation intensity; when the attenuation factor approaches 0, it indicates complete attenuation, and the target deformation intensity approaches 0.
[0105] Based on the first mapping relationship between the attenuation factor and the morphological parameters, the first attenuation factor corresponding to the target morphological parameters can be determined according to the target morphological parameters and the preset first mapping relationship between the attenuation factor and the morphological parameters. Furthermore, the initial deformation intensity can be attenuated according to the first attenuation factor to obtain the target deformation intensity.
[0106] This embodiment uses the attenuation factor as an intermediate transformation variable to transform the abstract risk metric into a controllable intensity adjustment coefficient, thereby ensuring the numerical stability and physical interpretability of the attenuation process. At the same time, the negative correlation between the attenuation factor and the attenuation amplitude is intuitive and makes the parameter tuning process more intuitive and efficient.
[0107] In some embodiments, the target morphological parameter includes: a target ratio of the projected spans of the boundary contour of the deformed region in multiple preset orthogonal directions. The first mapping relationship between the attenuation factor and the ratio of the projected spans of the boundary contour of the deformed region in multiple preset orthogonal directions can be implemented as a smooth, monotonically increasing function (defined as a first monotonically increasing function). The independent variable of the first monotonically increasing function is the ratio of the projected spans of the boundary contour of the deformed region in multiple preset orthogonal directions, and the dependent variable is the attenuation factor.
[0108] The area to be processed has a stable aspect ratio under normal extended state. When the target ratio deviates from the normal range, it indicates that key point detection is unreliable or the area is in a non-standard posture. In this case, the deformation intensity needs to be reduced to suppress jitter. However, the relationship between the target ratio and the risk level is not a simple linear one. When the target ratio is close to the normal range, the risk of causing background jitter is low, and the attenuation factor should be maintained at a high level. When the target ratio deviates significantly from the normal range, the risk of causing background jitter increases sharply, and the attenuation factor needs to decrease rapidly. If a step-like or piecewise linear mapping is used, a sudden change in the attenuation factor will occur at the threshold boundary, causing a jump in deformation intensity between adjacent frames, which will introduce new visual jitter. Therefore, this embodiment uses a smooth first monotonically increasing function as the specific implementation of the first mapping relationship. Here, smooth means that the second derivative of the function is continuous or higher-order derivatives are continuous, ensuring that the curve of the attenuation factor changing with the target ratio has no inflection point or sudden change.
[0109] In this embodiment, when determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and the first mapping relationship between the preset attenuation factor and the shape parameter, the target ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions can be input into a first monotonically increasing function for calculation to obtain the first attenuation factor. .
[0110] In one specific implementation, the first monotonically increasing function can be a modified form of a Sigmoid function. The saturation characteristic of the Sigmoid function maintains high deformation intensity under normal conditions and triggers deformation intensity attenuation under abnormal conditions, thus balancing aesthetics and stability. For example, in an embodiment where the part to be processed is an arm, the deformation region is rectangular, and the ratio of the projection spans of the boundary contour of the deformation region in multiple preset orthogonal directions can be realized as the aspect ratio of the rectangle. First attenuation factor It can be represented as: (11).
[0111] In equation (11), Indicates the first attenuation factor. This indicates the aspect ratio of the deformed region (i.e., the target ratio). and These correspond to the maximum and minimum attenuation factors, respectively. and The values of all are in the range [0,1] and Greater than .like Figure 8 As shown, when When it approaches 0.75, Approaching When the attenuation factor drops to its minimum, the deformation intensity is significantly weakened; when When it approaches 1.25, Approaching The attenuation factor rises to its highest level, while the deformation intensity remains essentially unchanged. This function... The gradient around 1 is steep, precisely corresponding to the critical range of normal aspect ratio, while it tends to saturate at both ends, avoiding oversensitivity caused by extreme outliers.
[0112] To verify the effect of the attenuation factor determined by the ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions provided in this embodiment on background shaking, some embodiments of this application perform arm-slimming deformation on the target object, with the deformed region as a rectangle, and use the calculation method of the deformation region and the target deformation intensity shown in the aforementioned equations (1)-(11) to... Figure 9 The corresponding video shows a person's arm undergoing a slimming deformation; and the deformation intensity attenuation mechanism provided in this embodiment is not used. Figure 9 The corresponding video shows a person's arm undergoing a slimming and deforming process to obtain... Figure 9 The video screenshot shown is a screenshot of the effect. Figure 9 The deformation processing algorithm used is the image liquefaction algorithm. The details of the image liquefaction algorithm will be described in the following examples and will not be repeated here.
[0113] Figure 9 The left-middle figure shows the effect of not using the deformation intensity attenuation mechanism provided in the embodiments of this application. Figure 9 The corresponding video shows the effect of arm slimming and deformation on the person's arm. Figure 9 It is based directly on the initial deformation intensity corresponding to each video frame. Figure 9 The corresponding video shows the effect of slimming and deforming the arms of the person in the video. Figure 9 The right-hand figure shows the deformation region as described by equations (1)-(11) and the target deformation intensity obtained by attenuating the initial deformation intensity using the attenuation factor determined by equation (11). Figure 9 The corresponding video shows the effect of the arm being slender and deformed. Figure 9 This is a schematic diagram illustrating the effect of processing a captured video frame.
[0114] Figure 9 The rectangle in the middle represents the boundary of the deformation region. According to... Figure 9 As can be seen from the processed video frames (a) and (b), as shown... Figure 9 (a) and Figure 9As shown in the elliptical circle on the left of image (b), the arm of the person in the video is slendered based on the initial deformation intensity corresponding to each video frame. Due to the complex shape of the target object's arm movement in the video, the deformation direction, position, and degree of distortion of the background area of the video are different, resulting in temporal background jitter when playing the continuous video. Figure 9 (a) and Figure 9 As shown in the elliptical circle on the right side of (b), the target deformation intensity corresponding to each frame in the video is determined according to the aforementioned formulas (1)-(11). The person in the video is subjected to arm deformation. Since the deformation intensity is small and the background distortion difference is small, the degree of deformation in the background of the video is reduced, and the deformation direction, position and degree of distortion are almost the same, thereby alleviating the background shaking caused by the complex shape of the target object's arm.
[0115] This embodiment achieves a smooth transition of the attenuation factor with the target ratio through a smooth first monotonically increasing function, avoiding inflection point jitter and resulting in a softer deformation transition, which helps eliminate the risk of intensity jumps at the threshold boundary. Combined with... Figure 9 The comparison results show that the mechanism can automatically suppress deformation intensity and reduce the amplitude of background distortion in complex poses where key point detection is unreliable, rather than relying on post-processing smoothing, thus improving the inherent robustness of the algorithm.
[0116] In other embodiments, the target morphological parameter includes a target ratio of the area of the deformed region overlapping with the current video frame to the area of the deformed region. The first mapping relationship between the attenuation factor and the ratio of the area of the deformed region overlapping with the current video frame to the area of the deformed region can be implemented as another smooth, monotonically increasing function (defined as a second monotonically increasing function). The independent variable of the second monotonically increasing function is the ratio of the area of the deformed region overlapping with the current video frame to the area of the deformed region, and the dependent variable is the attenuation factor.
[0117] When the area to be processed is completely within the frame, each pixel within the deformation area has valid image data support, and deformation processing can be performed normally. When the target ratio is less than 1, it indicates that the deformation area is outside the boundary, and the smaller the target ratio, the more severe the boundary violation, and the higher the risk of background jitter. In this case, the deformation intensity needs to be reduced to suppress edge distortion. However, the relationship between the target ratio and the risk level is not a simple linear one. When the target ratio is close to 1, the risk of background jitter is low, and the attenuation factor should be maintained at a high level; when the target ratio is significantly lower than the safety threshold, the risk of background jitter increases sharply, and the attenuation factor needs to decrease rapidly. If a step-like or piecewise linear mapping is used, a sudden change in the attenuation factor will occur at the threshold boundary, causing a jump in deformation intensity during the transition phase of the arm entering and leaving the frame, which in turn introduces new visual jitter. Therefore, this embodiment uses a smooth second monotonically increasing function as the specific implementation of the first mapping relationship. Here, smooth means that the second derivative of the second monotonically increasing function is continuous or higher-order derivatives are continuous, ensuring that the curve of the attenuation factor changing with the target ratio has no inflection point or sudden change.
[0118] In this embodiment, when determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and the first mapping relationship between the preset attenuation factor and the shape parameter, the target proportion of the area overlapping the deformed region with the current video frame to the area of the deformed region can be input into a second monotonically increasing function for calculation to obtain the first attenuation factor. .
[0119] In one specific implementation, the second monotonically increasing function can be a Sigmoid function or a variation of the Sigmoid function. The saturation characteristic of the Sigmoid function maintains high deformation intensity when fully within the frame and triggers deformation intensity attenuation when out of bounds, balancing aesthetic effects and edge stability. For example, in an embodiment where the part to be processed is the arm, the deformation area is rectangular. The ratio of the overlapping area between the deformation area and the current video frame to the area of the deformation area can be used as an indicator of the arm's out-of-bounds degree. The first attenuation factor corresponds to the ratio of the overlapping area between the deformation area and the current video frame to the area of the deformation area. It can be represented as: (12).
[0120] in, The first attenuation factor represents the proportion of the area of the deformed region overlapping with the current video frame to the area of the deformed region. This indicates the proportion of the area of the deformed region that overlaps with the current video frame, which is the area of the deformed region itself; and These correspond to the maximum and minimum attenuation factors, respectively, and their values range from [0,1]. Greater than .like Figure 10 As shown, when When it approaches 0, Approaching When the attenuation factor drops to its minimum, the deformation intensity is significantly weakened; when When it approaches 1, Approaching The attenuation factor rises to its maximum, while the deformation intensity remains essentially unchanged. The center point of this function is set at... At 0.475, the attenuation factor has dropped to the middle value when about half of the arm area goes out of bounds, reflecting an early response strategy to the risk of going out of bounds. At both ends, it tends to saturate, avoiding oversensitivity caused by extreme out-of-bounds values.
[0121] To verify the effect of the attenuation factor determined by the ratio of the overlapping area of the deformation region and the current video frame to the area of the deformation region provided in this embodiment on background jitter, some embodiments of this application perform arm-slimming deformation on the target object, with the deformation region as a rectangle, and use the calculation method of the deformation region and the target deformation intensity shown in the aforementioned equations (1)-(10) and (12) to... Figure 11 The corresponding video shows a person's arm undergoing a slimming deformation; and the deformation intensity attenuation mechanism provided in this embodiment is not used. Figure 11 The corresponding video shows a person's arm undergoing a slimming and deforming process to obtain... Figure 11 The video screenshot shown is a screenshot of the effect. Figure 11 The deformation processing algorithm used is the image liquefaction algorithm. The details of the image liquefaction algorithm will be described in the following examples and will not be repeated here.
[0122] Figure 11 The left-middle figure shows the effect of not using the deformation intensity attenuation mechanism provided in the embodiments of this application. Figure 11 The corresponding video shows the effect of arm slimming and deformation on the person's arm. Figure 11 It is based directly on the initial deformation intensity corresponding to each video frame. Figure 11 The corresponding video shows the effect of slimming and deforming the arms of the person in the video. Figure 11 The right-hand figure shows the deformation region as described by equations (1)-(10) and the target deformation intensity obtained by attenuating the initial deformation intensity using the attenuation factor determined by equation (12). Figure 11 The corresponding video shows the effect of the arm being slender and deformed. Figure 11 This is a schematic diagram illustrating the effect of processing a captured video frame.
[0123] Figure 11 The rectangle in the middle represents the boundary of the deformation region. According to... Figure 11 The processed video frames (a), (b), and (c) show that, as can be seen... Figure 11 (a) Figure 11 (b) and Figure 11 As shown in the elliptical circle on the left of image (c), the arm of the person in the video is slendered based on the initial deformation intensity corresponding to each video frame. Because the degree to which the target object's arm moves out of bounds varies, the deformation direction, position, and degree of distortion of the background area of the video differ, resulting in temporal background jitter during continuous video playback. For example... Figure 11 (a) Figure 11 (b) and Figure 11 As shown in the elliptical circle on the right side of (c), the target deformation intensity corresponding to each frame in the video is determined according to the aforementioned equations (1)-(10) and (12). The person in the video is subjected to arm deformation. Since the deformation intensity is small and the background distortion difference is small, the degree of deformation in the background of the video is reduced, and the deformation direction, position and degree of distortion are almost the same, thereby alleviating the background shaking caused by the complex shape of the target object's arm.
[0124] This embodiment achieves a smooth transition of the attenuation factor with varying degrees of boundary deviation through a smooth, second monotonically increasing function. This avoids inflection point jitter, resulting in a softer transition when the arm enters or leaves the frame, and helps eliminate the risk of intensity jumps at threshold boundaries. Combined with... Figure 11 The comparison results show that the mechanism can automatically suppress the deformation force and reduce the amplitude of edge texture distortion in scenes where the arm part goes out of bounds, rather than relying on post-processing smoothing, thus improving the visual comfort at the edge of the image and the inherent robustness of the algorithm.
[0125] In some embodiments, the target morphological parameters include: a target angle between the first deformation direction vector of the deformed region and the second deformation direction vector of the associated part of the part to be processed. The first mapping relationship between the attenuation factor and the angle between the first and second deformation direction vectors can be implemented as a smooth, monotonically decreasing function (defined as the first monotonically decreasing function). The independent variable of the first monotonically decreasing function is the angle between the first and second deformation direction vectors, and the dependent variable is the attenuation factor.
[0126] When performing multi-part collaborative beautification, if the deformation directions of the part to be processed and the related parts are consistent or orthogonal, and there is no significant interference between the deformation fields, deformation processing can be performed normally. When the target angle increases, it indicates that the spatial contradiction between the two deformation directions intensifies, and the larger the target angle (especially when it is close to π), the more opposite the directions are, and the higher the risk of causing shape proportion imbalance. At this time, the deformation intensity needs to be reduced to suppress the conflict. However, the relationship between the target angle and the risk level is not a simple linear one. When the target angle is less than or equal to π / 2, the risk of causing shape proportion imbalance is low, and the attenuation factor should be maintained at a high level; when the target angle significantly exceeds the orthogonal angle and approaches π, the risk of causing shape proportion imbalance increases sharply, and the attenuation factor needs to decrease rapidly. If a step or piecewise linear mapping is used, a sudden change in the attenuation factor will occur at the threshold boundary, causing a jump in deformation intensity during multi-part collaborative beautification, which will introduce new shape jump artifacts. Therefore, this embodiment uses a smooth first monotonically decreasing function as the specific implementation of the first mapping relationship. Smoothness refers to the continuity of the second derivative or higher derivatives of the function, ensuring that the curve of the attenuation factor changing with the target angle has no inflection point or abrupt change.
[0127] In this embodiment, when determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and the first mapping relationship between the preset attenuation factor and the shape parameter, the target angle between the first deformation direction vector and the second deformation direction vector can be input into the first monotonically decreasing function for calculation to obtain the first attenuation factor. .
[0128] In one specific implementation, the first monotonically decreasing function can be a mirror image of a sigmoid function. The saturation characteristic of a sigmoid function maintains high deformation intensity when directions are aligned or orthogonal, and triggers deformation intensity attenuation when directions severely conflict, thus balancing multi-functional coordination and conflict protection. For example, in an embodiment where the part to be processed is the arm and the associated part is the torso, the first deformation direction vector is the arm deformation direction, and the second deformation direction vector is the torso contraction direction. The target angle can then be represented as the degree of offset between the arm deformation direction and the torso contraction direction. The angle between the first and second deformation direction vectors corresponds to the first attenuation factor. It can be represented as: (13).
[0129] in, Indicates the first attenuation factor. A This represents the angle (in radians) between the first deformation direction vector and the second deformation direction vector. and These correspond to the maximum and minimum attenuation factors, respectively, and their values range from [0,1]. Greater than .like Figure 12 As shown, when A When less than or equal to π / 2, equal The attenuation factor remains at its highest level, and the deformation intensity does not decrease; when A When it approaches π, Approaching The attenuation factor is reduced to a minimum, and the deformation intensity is significantly weakened. The center point of the function is set at... A At 2.35 radians (approximately 135°), strong attenuation is triggered only when the two deformation directions are close to opposite. Full-intensity output is maintained within the orthogonal and lower angle range, reflecting a tolerant threshold design for directional conflicts. At the same time, it tends to saturate at both ends, avoiding oversensitivity caused by small angle fluctuations.
[0130] This embodiment achieves a smooth transition of the attenuation factor with varying directional conflict through a smooth first monotonically decreasing function. This avoids inflection point jitter and makes the transition process of multi-part collaborative beautification more gentle, helping to eliminate the risk of intensity jumps at threshold boundaries. This mechanism can automatically suppress the deformation intensity of the parts to be processed in scenarios with conflicting deformation directions, avoiding local optimization that could disrupt the overall proportions, rather than relying on hard switching of manual rules. This improves the visual coordination and inherent robustness of the algorithm during multi-part collaborative beautification.
[0131] The method for determining the first attenuation factor shown in the foregoing embodiments is merely illustrative. After determining the first attenuation factor, the initial deformation intensity can be attenuated according to the first attenuation factor to obtain the target deformation intensity.
[0132] In some embodiments of this application, there may be one or more target morphological parameters. In an embodiment where there is one target morphological parameter, the target deformation intensity can be determined by multiplying the first attenuation factor corresponding to the target morphological parameter with the initial deformation intensity.
[0133] In other embodiments, there are multiple target morphological parameters, each corresponding to a first attenuation factor. When attenuating the initial deformation intensity based on these multiple target morphological parameters, the first attenuation factor corresponding to each target morphological parameter can be determined according to the first mapping relationship, resulting in multiple first attenuation factors. Then, the initial deformation intensity can be weighted based on the product of these multiple first attenuation factors to attenuate the initial deformation intensity. For example, the product of the multiple first attenuation factors can be multiplied by the initial deformation intensity to attenuate the initial deformation intensity and obtain the target deformation intensity.
[0134] In this study of the deformation processing, the researchers discovered that, in addition to the spatial attribute information of the deformed region within the current video frame causing inter-frame background jitter, the motion of the part to be processed itself also leads to background jitter during deformation. Specifically, when the part to be processed undergoes rapid displacement between adjacent frames, the position of the deformed region changes drastically. If deformation is applied with the same intensity as in a static scene, it will cause significant differences in the background distortion field between adjacent frames, resulting in visible jitter and motion blur during video playback. To address this, this application introduces the target motion amplitude as a temporal risk metric, which, together with the target morphological parameters, constitutes a two-dimensional attenuation decision-making basis.
[0135] In some embodiments, the target motion amplitude of the part to be processed can also be obtained, and the initial deformation intensity can be attenuated according to the target morphological parameters and the target motion amplitude to obtain the target deformation intensity. This mechanism incorporates temporal motion risk and single-frame geometric risk into a unified attenuation decision, which can respond to two independent risk sources and achieve more comprehensive jitter suppression.
[0136] Accordingly, such as Figure 13 As shown in step 144, before attenuating the initial deformation intensity according to the target morphological parameters, the first position information of the deformed region in the current video frame and the second position information of the historical deformed region of the part to be processed in the historical video frames can be obtained. The historical video frame can be the N frames preceding the current video frame t in the same video, where N is a positive integer, such as N=1, 2, ..., or 5, etc. In this embodiment, the historical deformed region is the deformed region of the part to be processed in the historical video frames after smoothing using the deformation region smoothing mechanism provided in this application embodiment. The first position information refers to the coordinates of the reference point used to characterize the spatial position of the deformed region in the current video frame; the second position information refers to the coordinates of the reference point corresponding to the historical deformed region in the historical video frames.
[0137] Furthermore, such as Figure 13 As shown in step 145, the positional change of the deformed region relative to the historical deformed region can be determined based on the first positional information and the second positional information. The positional change refers to the displacement of the first and second positional information in the image coordinate system. This positional change reflects the actual displacement of the part to be processed between adjacent frames.
[0138] To improve the accuracy of the determined positional changes, the previous video frame (t-1) of the current video frame t is generally selected for positional change calculation. Specifically, the aforementioned historical video frames include: the previous video frame (t-1) of the current video frame t; correspondingly, the historical deformation region includes: the historical deformation region of the part to be processed in the previous video frame (t-1) (defined as the target historical deformation region). Accordingly, the positional change between the target historical deformation region and the deformation region can be calculated based on the first positional information of the target historical deformation region in the previous video frame (t-1) and the second positional information of the deformation region in the current video frame t.
[0139] The calculation methods for the positional changes of deformation regions of different shapes differ. The following provides illustrative examples of how to calculate the positional changes of deformation regions of several shapes.
[0140] In embodiments where the area to be processed is a region with a clear skeletal orientation, such as an arm or leg, and the deformed region and the target historical deformed region are rectangular, the first position information includes the first image coordinates of the first projection endpoint in the current video frame; the second position information includes the second image coordinates of the second projection endpoint in the historical video frame. The first projection endpoint is the endpoint of the projection line segment of the deformed region on the central axis of the skeletal orientation of the area to be processed, along the target direction; the second projection endpoint is the endpoint of the projection line segment of the historical deformed region on the central axis, along the target direction. The target direction is parallel to the skeletal orientation. Accordingly, the positional change of the deformed region relative to the historical deformed region can be determined based on the distance between the first image coordinates and the second image coordinates. This distance can be Euclidean distance.
[0141] Since the deformation of parts such as the arms and legs primarily serves aesthetic purposes such as slimming arms and legs or lengthening along the skeletal axis, their effective motion components are concentrated in the longitudinal dimension. If the geometric center of a rectangle is used as the position reference point, when the arm bends or twists, the lateral noise from keypoint detection will directly mix into the calculation of the center coordinates, resulting in an artificially high positional change and triggering unnecessary deformation intensity attenuation. Therefore, this embodiment selects the endpoint of the projection line segment of the deformation region on the central axis of the skeletal axis as the position reference point. The central axis refers to the straight line passing through the center of the deformation region along the skeletal axis of the part to be processed; the projection line segment refers to the line segment intercepted by the boundary of the deformation region on the central axis; the target direction is the direction from the proximal keypoint to the distal keypoint. The projection endpoint can be anchored to the anatomical principal axis of the part, and its displacement purely reflects the longitudinal movement along the skeletal axis, filtering out lateral detection noise and non-motion deformation interference, and is a reference point characterizing the true motion state of the part to be processed. The reason for choosing the endpoint of the projection line segment rather than the starting point or midpoint is that the range of motion at the endpoint (far end) is usually greater than that at the proximal end, making it more sensitive to rapid movements and triggering attenuation protection earlier. At the same time, the endpoint position is less affected by joint bending, and can still maintain stable longitudinal displacement representation ability under complex postures such as arm bending.
[0142] This embodiment uses the endpoint of the central axis projection as a position reference point to anchor the position change calculation of the rectangular deformation area to the anatomical principal axis of the part to be processed. This can effectively filter out the interference of lateral detection noise and non-motion deformation on the motion amplitude measurement. The longitudinal displacement is highly consistent with the actual motion state of the part to be processed, avoiding false attenuation caused by the artificially high target motion amplitude.
[0143] In other embodiments, the deformed region and the target historical deformed region are circular. Accordingly, the first position information of the target historical deformed region in the previous video frame may include the coordinates of the center of the target historical deformed region, i.e., the center coordinates are the image coordinates of the center of the target historical deformed region in the previous video frame. The second position information of the deformed region in the current video frame may include the coordinates of the center of the deformed region, which are the image coordinates of the center of the deformed region in the current video frame. Accordingly, the distance between the center of the deformed region and the center of the target historical deformed region can be calculated based on the center coordinates of the deformed region and the center coordinates of the target historical deformed region, as the positional change between the target historical deformed region and the deformed region. This distance may be Euclidean distance.
[0144] Since circular deformation regions are typically used in isotropic deformation scenarios such as eye magnification and overall head enhancement, the deformation force radiates uniformly from the center outwards, without a dominant skeletal orientation or functional axis. Therefore, the center of the circle, as the geometric symmetry center of the circular region and the origin of the deformation force field, can completely and unbiasedly characterize the overall translational amount of the entire deformation region in any direction. Introducing the center as a positional reference point is due to the rotational symmetry of the circle, which means that any reference point deviating from the center will introduce directional deviations. The center maintains consistent measurement sensitivity in all directions of motion, ensuring that the calculated positional change accurately reflects the actual intensity of the motion regardless of the direction in which the part being processed moves.
[0145] In this embodiment, the center of the circle is used as the position reference point of the circular deformation region, which can make full use of the rotational symmetry of the circle to achieve unbiased motion measurement in all directions; the center of the circle coincides with the origin of the deformation force field radiation, which can ensure the consistency between position change and deformation field temporal offset.
[0146] In other embodiments, the deformed region and the target historical deformed region are elliptical. Accordingly, the first position information of the target historical deformed region in the previous video frame may include the center coordinates of the target historical deformed region, i.e., the center coordinates are the image coordinates of the center of the target historical deformed region in the previous video frame. The second position information of the deformed region in the current video frame may include the center coordinates of the deformed region, which are the image coordinates of the center of the deformed region in the current video frame. Accordingly, the distance between the center of the deformed region and the center of the target historical deformed region can be calculated based on the center coordinates of the deformed region and the center coordinates of the target historical deformed region, as the positional change between the target historical deformed region and the deformed region. This distance may be Euclidean distance.
[0147] Since elliptical deformation regions are typically used in beautification scenarios such as eye enlargement and cheek enhancement, which are directional but not strictly skeletal guided, their major and minor axes correspond to different deformation sensitivity dimensions. However, the overall deformation field still decays outward from the center of the ellipse as the origin of symmetry. The center of the ellipse is introduced as a positional reference point because, although the deformation intensity of the ellipse decays at different rates along its major and minor axes, its positional offset has an equivalent effect on background jitter in all directions. That is, regardless of whether the ellipse moves the same number of pixels along either its major or minor axis, the resulting temporal differences in background distortion are similar. Using the center point avoids confusing the directionality of deformation intensity with the directionality of position measurement, maintaining the physical consistency of motion amplitude calculation.
[0148] This embodiment uses the center of the ellipse as the position reference point of the elliptical deformation region, which preserves the directional deformation characteristics of the ellipse while achieving isotropic position measurement. The center point, as the intersection of the major and minor axes and the origin of the deformation field symmetry, can ensure the accurate correspondence between position changes and the overall offset of the deformation field. It avoids the incorrect transmission of deformation intensity anisotropy to the motion amplitude calculation and ensures the orthogonal decoupling of the two-dimensional attenuation decision.
[0149] In some embodiments, the deformed region and the target historical deformed region are trapezoidal. Accordingly, the second position information of the target historical deformed region in the previous video frame may include the centroid coordinates of the target historical deformed region, i.e., the centroid coordinates are the image coordinates of the centroid of the target historical deformed region in the previous video frame. The first position information of the deformed region in the current video frame may include the centroid coordinates of the deformed region, which are the image coordinates of the centroid of the deformed region in the current video frame. Accordingly, the distance between the centroids of the deformed region and the target historical deformed region can be calculated based on the centroid coordinates of the deformed region and the target historical deformed region, as the positional change between the target historical deformed region and the deformed region. This distance may be Euclidean distance.
[0150] Since trapezoidal deformation regions are typically applied to conical limbs such as the thigh and calf, which are thicker at the top and thinner at the bottom, the widths of the upper and lower bases are unequal, and the geometric center is not equivalent to the center of mass distribution or the equilibrium point of the deformation force field. If the midpoint of the line connecting the midpoints of the upper and lower bases is used as the reference point, the position measurement will be biased towards the wider side due to the asymmetry of the trapezoid, resulting in false lateral displacement components when the limb swings. Therefore, this embodiment selects the centroid (i.e., the center of area) of the trapezoid as the position reference point. The centroid is the equilibrium point of the trapezoidal area distribution, and its position is determined by the length and height of the upper and lower bases, accurately reflecting the center of gravity of the trapezoidal region in space. The background distortion effect of the trapezoidal deformation region is closely related to the area distribution it covers. As the equilibrium point of the area distribution and the position reference point, the displacement of the centroid can accurately characterize the temporal change of the overall drag effect of the entire trapezoidal region on the background. Compared with the geometric center, the centroid has better stability when the aspect ratio of the trapezoid changes, and will not produce drastic reference point jumps due to fluctuations in the width of the upper and lower bases caused by key point detection.
[0151] This embodiment uses the centroid as the positional reference point of the trapezoidal deformation region, which helps to overcome the positional measurement deviation caused by the geometric asymmetry of the trapezoid and ensures the physical accuracy of motion amplitude calculation. As the equilibrium point of area distribution, the centroid's displacement is highly consistent with the overall drag effect of background distortion, improving the effectiveness of motion risk measurement. Compared with the geometric center, the centroid is more robust to aspect ratio fluctuations caused by key point detection noise, avoiding false motion amplitude caused by reference point jumps.
[0152] The calculation methods for the shape of the deformed region and the positional changes between the target historical deformed region and the deformed region shown in the foregoing embodiments are merely illustrative and do not constitute a limitation. Those skilled in the art can select appropriate location reference points and distance calculation methods based on the anatomical features of the part to be treated and the geometric characteristics of the deformed region, all of which fall within the protection scope of this application.
[0153] After determining the positional change of the deformation region relative to the historical deformation region, it can be done as follows: Figure 13 As shown in step 146, the target motion amplitude of the part to be processed is determined based on the position change. The target motion amplitude is positively correlated with the position change; that is, the greater the displacement, the higher the target motion amplitude, indicating more intense motion. In some embodiments, the position change can be determined as the target motion amplitude.
[0154] In other embodiments, since the size of the area to be processed varies significantly under different users and shooting distances, the pixel displacement corresponding to the same physical motion speed may differ by several times. If the original pixel displacement is directly used as the target motion amplitude, the attenuation threshold set for large-sized areas is too sensitive to small-sized areas, while the threshold set for small-sized areas is too insensitive to large-sized areas. To solve this scale-dependent problem, some embodiments of this application use the scale of the historical deformation region as a normalization benchmark. The scale of the historical deformation region refers to the characteristic size of the deformation region of the area to be processed in the previous video frame. For rectangular areas, it can be the length of the long side, the width of the short side, or the square root of the area; for circular areas, it can be the radius or the diameter. This embodiment of the application does not limit this.
[0155] Specifically, a scale normalization coefficient is calculated based on the scale of the historical deformation region. The scale of the deformation region is a parameter used to measure the size of the deformation region.
[0156] Accordingly, in embodiments where the historical deformation region and the deformation region are circular, the average radius of the historical deformation region can be calculated as a scale normalization coefficient. This embodiment uses the average radius of the historical deformation region as the scale normalization coefficient, which is equivalent to performing a low-pass filter on the radius of the deformation region, helping to smooth the changes in deformation intensity. In real-world videos, the scale changes of the target object are usually continuous and slow. Using the radius of the historical deformation region and the average radius of the deformation region aligns with the prior knowledge of motion continuity and is more consistent with real-world scenes.
[0157] In embodiments where the historical deformation region is rectangular, the scale of the historical deformation region is represented by its width. The scale normalization coefficient is equal to the average width of the historical deformation regions corresponding to each of the N historical video frames.
[0158] In embodiments where the historical deformation region is rectangular, the scale of the historical deformation region is represented by its width, and the scale normalization coefficient is set to the average width of the historical deformation regions corresponding to each of the N historical video frames. The width of the rectangular deformation region directly corresponds to the lateral dimension of the part to be processed (such as an arm or leg), and is a core geometric parameter for measuring the spatial occupancy of that part. Normalizing using width as the scale benchmark converts absolute pixel displacement into a relative motion ratio relative to the thickness of the part itself, eliminating measurement biases caused by differences in shooting distance, image resolution, and user body size. This allows the same set of attenuation thresholds to adaptively adapt to the parts to be processed in different scenarios. Using the average width of N historical video frames, rather than the width of a single frame, as the scale normalization coefficient is equivalent to applying a temporal low-pass filter to the scale parameter. This processing effectively smooths the width fluctuation noise generated in a single frame during keypoint detection, avoids abrupt changes in the scale normalization coefficient due to abnormal keypoint positioning in a particular frame, and thus prevents false jumps in the target motion amplitude, ensuring the temporal continuity of attenuation decisions. On the other hand, the lateral dimensions of limbs change slowly over a short time series, exhibiting a high degree of temporal continuity. Using the average width of multiple frames as the normalization benchmark conforms to the physiological law of gradual changes in limb scale, allowing the normalized target motion amplitude to more realistically reflect the actual intensity of motion of the part to be processed, rather than being interfered with by instantaneous detection errors, thus improving the naturalness and visual comfort of the beautification effect.
[0159] In embodiments where the historical deformation region is elliptical, the mean of the square roots of the product of the major and minor axes of the historical deformation region can be calculated as the scale normalization coefficient. The square root of the product of the major and minor axes of the ellipse is on the same dimension as the square root of the area of the ellipse, ensuring scale comparability between different shapes. The deflection of the part to be processed causes a change in the direction of the ellipse's principal axis. As long as the lengths of the major and minor axes of the ellipse are correctly extracted, the product of the major and minor axes remains unchanged, thus ensuring the rotational invariance of the scale normalization coefficient.
[0160] In embodiments where the historical deformation region and the deformation region are trapezoidal, the average length of the upper and lower bases of the trapezoidal historical deformation region can be used as the scale normalization coefficient. The trapezoidal deformation region typically corresponds to conical limb parts such as the thigh and calf, which are thicker at the top and thinner at the bottom. Its width gradually changes along the bone direction. Using the average length of the upper and lower bases is essentially a linear approximation of the lateral span of the trapezoidal region, comprehensively reflecting the overall spatial occupancy from the proximal end to the distal end. This avoids the bias of taking only the upper base, resulting in an oversized scale, or taking only the lower base, resulting in an undersized scale, ensuring that the normalization benchmark accurately matches the true geometric shape of the conical limb. The thickness of the limb changes slowly over a short time series, and the lengths of the upper and lower bases exhibit a highly coordinated trend of change. Using the average length of the upper and lower bases as the scale benchmark preserves the natural scale change information of the trapezoidal region as the limb moves, and smooths out high-frequency noise between frames through averaging. This ensures that the normalized target motion amplitude can not only respond sensitively to real limb scaling, but also effectively suppress interference caused by detection jitter, ensuring the continuity and physiological rationality of the attenuation decision in the temporal dimension.
[0161] In the aforementioned embodiments, the geometric parameters are fused into a physically meaningful equivalent scale in the scale normalization coefficients corresponding to the deformation regions of the aforementioned shapes. Then, temporal smoothing and robust normalization are achieved by using the average value of historical video frames and the current video frame, which conforms to the prior of motion continuity and is more in line with real-world scenarios.
[0162] After obtaining the scale normalization coefficient, it can be used to normalize the positional changes between historical deformation regions and deformation regions to obtain the target motion amplitude of the part to be processed. Specifically, the target motion amplitude of the part to be processed can be obtained by dividing the positional changes between historical deformation regions and deformation regions by the scale normalization coefficient. For example, in the aforementioned embodiment where the historical deformation regions and deformation regions are rectangles, the target motion amplitude can be expressed as: (14).
[0163] in, Indicates the amplitude of the target motion; Indicates the first image coordinates of the first projection endpoint in the current video frame; This represents the second image coordinates of the second projection endpoint within the historical video frame; The average arm width of the previous N frames can be improved. Stability.
[0164] In video, the same body part (such as an arm) can appear different sizes depending on its proximity to or distance from the camera. Directly using pixel-level positional changes (such as a 20-pixel displacement) might represent minute movements in close-up shots but significant movements in distant shots. By introducing a scale normalization coefficient, the positional change can be divided by this coefficient, converting absolute displacement into a relative motion ratio. This ensures that the motion intensity is judged consistently regardless of the distance between the part being processed and the camera, avoiding misjudgments caused by differences in imaging scale.
[0165] After determining the target motion amplitude, such as Figure 13 As shown in step 147, the initial deformation intensity can be attenuated based on the target shape parameters and the target motion amplitude to obtain the target deformation intensity. Regarding... Figure 13 For a description of steps 141-143, please refer to the preceding text. Figure 2 The details of steps 201-203 will not be repeated here.
[0166] This embodiment introduces the target motion amplitude to compensate for the blind spot of purely static morphological parameters in temporal risk perception, and realizes the collaborative detection of single-frame geometric risk and cross-frame motion risk; and by combining the target morphological parameters of a single frame and the motion amplitude of a cross-frame, the initial deformation intensity is attenuated, which can suppress jitter caused by morphological abnormalities and jitter caused by rapid movement, and further alleviate the background jitter caused by the deformation of the part to be processed.
[0167] When attenuating the initial deformation intensity based on the target shape parameters and the target motion amplitude to obtain the target deformation intensity, in some embodiments, the first attenuation factor corresponding to the target shape parameters can be determined based on the target shape parameters and the first mapping relationship between the preset attenuation factor and the shape parameters. For an explanation of this step and its specific implementation, please refer to the relevant content of the foregoing embodiments.
[0168] Accordingly, a second attenuation factor corresponding to the target motion amplitude can be determined based on the target motion amplitude and a preset second mapping relationship between the attenuation factor and the motion amplitude. In the second mapping relationship, the attenuation factor is negatively correlated with the motion amplitude. The attenuation factor is negatively correlated with the attenuation amplitude. The value range of the attenuation factor in the second mapping relationship is [0,1].
[0169] Optionally, the motion amplitude in the mapping relationship between the attenuation factor and the motion amplitude includes multiple motion amplitude ranges, and each motion amplitude range corresponds to a functional expression between the attenuation factor and the motion amplitude. The independent variable of this functional expression is the motion amplitude, and the dependent variable is the attenuation factor. The functional expression corresponding to each motion amplitude range is a monotonically decreasing function. The functional expressions of all motion amplitude ranges show a decreasing trend. Accordingly, the target motion amplitude can be determined from multiple motion amplitude ranges. Further, the target motion amplitude can be input into the functional expression corresponding to the target motion amplitude range for calculation to obtain the second attenuation factor corresponding to the target motion amplitude. The value range of the dependent variable of each functional expression is [0,1].
[0170] In other embodiments, the mapping relationship between the attenuation factor and the motion amplitude can be implemented as a function reflecting this mapping relationship, which is a monotonically decreasing function (defined as a second monotonically decreasing function). The independent variable of the monotonically decreasing function is the motion amplitude, and the dependent variable is the attenuation factor. The second monotonically decreasing function is a smooth monotonically decreasing function; therefore, its second derivative is continuous or its higher-order derivatives are continuous. The dependent variable of the second monotonically decreasing function takes values in the range [0,1]. For example, in some embodiments, the second monotonically decreasing function can be implemented as follows: (15).
[0171] in, This represents the second attenuation factor. , These correspond to the maximum and minimum attenuation factors, respectively. hour, Approaching The deformation intensity approaches the initial deformation intensity; hour, Approaching The deformation intensity weakens to a minimum, and the deformation intensity weakens significantly. Equation (15) reflects the relationship between the attenuation amplitude and the motion amplitude as follows: Figure 14 As shown, when the motion amplitude approaches 0, the attenuation factor approaches the preset maximum attenuation factor. This means the deformation intensity is close to maintaining the initial deformation intensity. When the motion amplitude approaches 1, the attenuation factor approaches the preset minimum attenuation factor. The deformation intensity is significantly reduced, and the deformation intensity of the area to be deformed is significantly reduced.
[0172] Accordingly, the target motion amplitude can be input into the second monotonically decreasing function for calculation to obtain the second attenuation factor corresponding to the target motion amplitude. This embodiment uses a globally smooth monotonically decreasing function to calculate the attenuation factor. Because the second derivative or even higher derivatives of the function are continuous, the attenuation factor changes gradually with the motion amplitude, which can avoid inflection point jitter, making the deformation transition smoother and helping to improve deformation stability.
[0173] To test the effect of the attenuation factor corresponding to the target motion amplitude on mitigating background jitter, the researchers in this application used arm deformation as an example, and used the attenuation factor corresponding to the target motion amplitude of the arm alone to attenuate the deformation intensity. The resulting deformation effect is as follows: Figure 15 As shown. To verify the motion amplitude provided in this embodiment and the effect of the determined attenuation factor on background shaking, some embodiments of this application perform arm-slimming deformation on the target object, with the deformation area as a rectangle, and use the calculation method of deformation area and target deformation intensity shown in the aforementioned equations (1)-(10) and (15) to... Figure 15 The corresponding video shows a person's arm undergoing a slimming deformation; and the deformation intensity attenuation mechanism provided in this embodiment is not used. Figure 15 The corresponding video shows a person's arm undergoing a slimming and deforming process to obtain... Figure 15 The video screenshot shown is a screenshot of the effect. Figure 15 The deformation processing algorithm used is the image liquefaction algorithm. The details of the image liquefaction algorithm will be described in the following examples and will not be repeated here.
[0174] Figure 15 The left-middle figure shows the effect of not using the deformation intensity attenuation mechanism provided in the embodiments of this application. Figure 15 The corresponding video shows the effect of arm slimming and deformation on the person's arm. Figure 15 The left image shows the results directly based on the initial deformation intensity corresponding to each video frame. Figure 15 The corresponding video shows the effect of slimming and deforming the arms of the person in the video. Figure 15 The right-hand figure shows the deformation region as described by equations (1)-(10) and the target deformation intensity obtained by attenuating the initial deformation intensity using the attenuation factor determined by equation (15). Figure 15 The corresponding video shows the effect of the arm being slender and deformed. Figure 15 This is a schematic diagram illustrating the effect of processing a captured video frame.
[0175] Figure 15 The rectangle in the middle represents the boundary of the deformation region. According to... Figure 15 As can be seen from the processed video frames (a) and (b), as shown... Figure 15 (a) and Figure 15As shown in the elliptical circle on the left of (b), the arm-slimming deformation of the person in the video is performed based on the initial deformation intensity corresponding to each video frame. Due to the different range of motion of the target object's arm in the video, the background area of the video (such as...) is affected. Figure 15 Different deformation directions, positions, and degrees of distortion (in the area circled by the ellipse) can cause temporal background jitter during continuous video playback. For example... Figure 15 (a) and Figure 15 As shown in the elliptical circle on the right side of Figure (b), the target deformation intensity corresponding to each frame in the video is determined according to the aforementioned equations (1)-(10) and (15). The person in the video is subjected to arm deformation. Since the deformation intensity is small and the background distortion difference is small, the degree of deformation in the background of the video is reduced, and the deformation direction, position and degree of distortion are almost the same, thereby alleviating the background shaking caused by the complex shape of the target object's arm.
[0176] After obtaining the first attenuation factor corresponding to the target shape parameters and the second attenuation factor corresponding to the target motion amplitude, the initial deformation intensity can be attenuated according to the first and second attenuation factors to obtain the target deformation intensity. This embodiment dynamically adjusts the deformation intensity of the initial deformation intensity through a preset rule, namely "the greater the motion amplitude, the smaller the attenuation factor". Even if the target deformation area is offset, the image distortion is reduced because the deformation intensity, i.e., the force of deformation, is weakened, which can effectively alleviate the visual distortion caused by the lag of the deformation area.
[0177] In some embodiments, there is one target morphological parameter. When attenuating the initial deformation intensity according to the first attenuation factor and the second attenuation factor, the initial deformation intensity can be attenuated by multiplying the first attenuation factor corresponding to the target morphological parameter with the second attenuation factor corresponding to the target motion amplitude. For example, the initial deformation intensity can be multiplied by the product of the first attenuation factor and the second attenuation factor to achieve weighted attenuation of the initial deformation intensity and obtain the target deformation intensity.
[0178] In some embodiments, there are multiple target morphological parameters, each corresponding to a first attenuation factor, resulting in multiple first attenuation factors. Accordingly, when attenuating the initial deformation intensity based on the first and second attenuation factors, a weighted attenuation process can be performed on the initial deformation intensity using the product of multiple first and second attenuation factors. For example, the weighted attenuation process can be achieved by multiplying the product of multiple first and second attenuation factors by the initial deformation intensity to obtain the target deformation intensity. The product of multiple first and second attenuation factors refers to multiplying the product of multiple first attenuation factors by the second attenuation factor. Accordingly, the target deformation intensity can be expressed as: (16).
[0179] In the above equation (16), Indicates the target deformation intensity; Indicates the initial deformation intensity; The attenuation factor corresponding to the aspect ratio of the deformed region; The attenuation factor represents the target proportion of the area of the deformed region that overlaps with the current video frame to the area of the deformed region. The attenuation factor represents the angle between the first deformation direction vector of the deformed region and the second deformation direction vector of the associated part of the part to be processed. This represents the attenuation factor corresponding to the target motion amplitude of the part to be processed.
[0180] This embodiment integrates the first attenuation factor corresponding to each of the various target shape parameters and the second attenuation factor corresponding to the target motion amplitude to attenuate the initial deformation intensity. In complex scenarios where multiple jitter risks exist simultaneously, the attenuation intensity of multiple attenuation factors is automatically superimposed. More comprehensive safety protection can be achieved without the need for additional design of multi-risk collaborative logic. It can reduce the probability of background jitter or shape distortion caused by insufficient single-dimensional attenuation under any abnormal combination, suppress background jitter caused by multiple factors, and further alleviate background jitter.
[0181] After determining the target deformation intensity for deforming the deformation region, such as Figure 2 Step 205 and Figure 13 As shown in step 148, the part to be processed in the current video frame can be deformed according to the deformation area and the target deformation intensity to obtain the target video frame corresponding to the current video frame t.
[0182] This application embodiment determines the target morphological parameters based on the spatial attribute information of the deformed region in the current video frame, quantifies the risk level of background jitter caused by the deformation of the part to be processed during the deformation process, and obtains the target deformation intensity by attenuating the initial deformation intensity according to the target morphological parameters. Since the root cause of background jitter lies in the discontinuity of the deformation field between adjacent frames due to the abnormal spatial attributes of the deformed region, and the target morphological parameters directly characterize the jitter risk level corresponding to this spatial abnormality, by applying an attenuation process to the initial deformation intensity that matches the risk level, the deformation intensity acting on the part to be processed can be adaptively reduced with the degree of spatial attribute abnormality, thereby weakening the difference in the deformation field between frames caused by the instability of spatial attributes and alleviating the background jitter caused by the morphological abnormality of the part to be processed.
[0183] The following describes in detail the implementation method of deforming the part to be processed in the current video frame t based on the deformable region and the target deformable intensity. In some embodiments, an image liquefaction algorithm can be used to deform the part to be processed in the current video frame t based on the deformable region and the target deformable intensity. The image liquefaction algorithm is a local image deformation technique based on displacement mapping. It does not directly modify pixel values, but instead defines a deformable region and maps the coordinates of each pixel in the original image to new coordinates. The deformation process maintains visual continuity and smoothness, avoiding tearing or jaggedness. For a rectangular deformable region, the deformation is determined by the center point, length, and width of the deformable region, as well as the target deformable intensity, following these rules: all pixels within the deformable region are offset according to a normalized distance; pixels closer to the center point have a greater degree of deformation; pixels closer to the boundary have a smaller degree of deformation; when a pixel is located at the boundary, the pixel is not deformed; pixels outside the deformable region are not offset.
[0184] Based on the image liquefaction algorithm, deformation processing of the part to be processed in the current video frame t can be achieved by considering the deformation region and the target deformation intensity. For any pixel P within the deformation region, the new coordinates of pixel P after deformation processing can be calculated based on the characterization parameters of the deformation region, the image coordinates of pixel P, and the target deformation intensity of pixel P. The characterization parameters of the rectangular deformation region include the center point, length, and width of the deformation region.
[0185] Then, the pixel value of pixel P in the original image can be assigned to the pixel corresponding to the new coordinates to obtain the target video frame corresponding to the current video frame t containing the deformed local image, thus realizing the deformation processing of the part to be processed. The above embodiment only takes pixel P in the rectangular deformation area as an example to give the method of determining the pixel value. The same method can be used to obtain the pixel value of any point in the rectangular deformation area.
[0186] The forward transformation rule between the image coordinates of pixels within the deformation region and those before and after deformation processing can be determined based on the shape of the deformation region, its characterization parameters, and the target deformation intensity of each pixel within the deformation region. In the forward transformation rule, the original image coordinates of pixel P, the characterization parameters of the deformation region, and the target deformation intensity of pixel P are known quantities, while the new coordinates of pixel P after deformation processing are quantities to be determined. For example, in some embodiments, such as... Figure 16 As shown, if the deformation region is rectangular, then the characterization parameters of the deformation region include the coordinates of the center point (…). , ),length H and width W. For any pixel P within the deformation region ( x , y ), its new coordinates P' after deformation processingx ', y The solution can be obtained using the following formula: (17); (18); (19).
[0187] In the above equations (17)-(19), ( x , y ) represents the original image coordinates of pixel P before deformation. x ', y ') represents the new coordinates of pixel P after deformation processing. , () represents the coordinates of the center point of the rectangular deformation region. W and H These represent the width and length of the rectangular deformation region, respectively. and The target deformation intensity of pixel P is respectively The components in the horizontal (x-axis) and vertical (y-axis) directions. Based on equations (17)-(19), the new coordinate position P' to which pixel P should be mapped before deformation processing can be calculated. x ', y The preceding steps have determined the coordinates of the center point, length, and width of the deformed region, as well as the target deformation intensity of pixel P. Therefore, these known quantities can be input into equations (17)-(19) above to calculate the new coordinates of pixel P. Then, the pixel value of pixel P in the original image can be assigned to the pixel corresponding to the new coordinates (x', y'), thereby realizing the deformation processing of pixel P. Using the same method, the deformation processing of each pixel in the deformation area can be performed, thereby realizing the deformation processing of the part to be processed and obtaining the target video frame of the current video frame t.
[0188] In a live streaming scenario, after obtaining the target video frame of the current video frame t, the target video frame can be sent to the terminal device corresponding to the live streaming room, such as the terminal device of the viewers in the live streaming room. In an offline video effects processing scenario, after obtaining the target video frame of the current video frame t, the target video frame can be stored, and the video processing method provided in this application embodiment can be used to continue processing the next video frame until the video to be processed is completed.
[0189] To verify the processing effect of the video processing method provided in the embodiments of this application, in some embodiments of this application, the processing effect of the video processing method provided in the embodiments of this application is tested by taking the shrinkage deformation of the target object's arm, i.e., the thinning arm deformation, with the deformation area being a rectangle. Specifically, the video processing process is as follows: Figure 17 As shown, it includes the following steps: S1. Perform arm keypoint detection on the current video frame t to obtain the target keypoints of the target object's arm. The target keypoints include the midpoint of the shoulder, the shoulder boundary point, the midpoint of the elbow, the elbow boundary point, the midpoint of the wrist, and the wrist boundary point.
[0190] S2. Perform temporal smoothing on the image coordinates of the target key points in the current video frame t to obtain the target image coordinates of the target key points in the current video frame t.
[0191] S3. Determine the deformation region of the rectangle based on the target image coordinates of the target key points in the current video frame t. See equations (1)-(4) for the specific calculation method.
[0192] S4. Determine the initial deformation intensity of the rectangular deformation region based on the relative positions of each pixel within the deformation region to the central axis of the target. See equations (5) to (7) for specific calculation methods.
[0193] S5. Determine the aspect ratio of the deformable region of the rectangle based on its length and width. See equation (8) for the specific calculation method.
[0194] S6. Calculate the target ratio of the area overlapping the deformed region with the current video frame to the area of the deformed region. See equation (9) for the specific calculation method.
[0195] S7. Calculate the target angle between the first deformation direction vector of the deformed region and the second deformation direction vector of the torso. See equation (10) for the specific calculation method.
[0196] S8. Based on the distance between the first image coordinates of the first projection endpoint in the current video frame and the second image coordinates of the second projection endpoint in the historical video frame, determine the positional change of the deformed region relative to the target historical deformed region in the previous video frame (t-1). Here, the first projection endpoint is the endpoint of the projection line segment of the deformed region along the target direction on the central axis of the bone orientation of the area to be processed; the second projection endpoint is the endpoint of the projection line segment of the target historical deformed region along the target direction on the central axis. The target direction is parallel to the bone orientation. The target historical deformed region is the deformed region after smoothing.
[0197] S9. Using the average width of the smoothed historical deformation region of the previous N historical video frames of the current video frame, normalize the positional change of the deformation region relative to the target historical deformation region to obtain the target motion amplitude of the target object's arm. The specific calculation methods of steps S8 and S9 are shown in equation (14).
[0198] S10. Calculate the first attenuation factor based on the aspect ratio of the deformed region. For the specific calculation method, please refer to equation (11).
[0199] S11. Calculate the first attenuation factor based on the target ratio of the area overlapping the deformed region with the current video frame to the area of the deformed region. For the specific calculation method, please refer to equation (12).
[0200] S12. Calculate the first attenuation factor based on the target angle between the first deformation direction vector and the second deformation direction vector. For the specific calculation method, please refer to equation (13).
[0201] S13. Calculate the second attenuation factor based on the target motion amplitude. For the specific calculation method, please refer to equation (15).
[0202] S14, Utilization The initial deformation intensity is attenuated to obtain the target deformation intensity.
[0203] S15. Using the image liquefaction algorithm, based on the center point, length, and width of the deformation region and the target deformation intensity of each pixel, perform arm-slimming deformation on the arm of the target object in the current video frame.
[0204] This embodiment utilizes Figure 17 The video processing procedure shown is for Figure 18 The corresponding video shows the target subject's arm undergoing a slimming and deforming process, resulting in the following effect: Figure 18 (a) and Figure 18 As shown in (b). Wherein, Figure 18 The left image shows the result using an image liquefaction algorithm based on... Figure 17 The deformation area obtained in step S3 and the initial deformation intensity of each pixel obtained in step S4 are used to apply a slimmer arm deformation to the arm of the target object in the video. Figure 18 The image on the right is using Figure 17 The deformation area obtained in step S3 and the target deformation intensity obtained in step S14 are used to apply a slimmer arm deformation effect to the arm of the target object in the video. Figure 18 The comparison of the left and right images shows that the video processing method provided in this application can effectively alleviate background deformation in video frames when processing the deformation of the target object, making the background deformation effect of different video frames more consistent, and effectively alleviating background jitter between frames during continuous video playback.
[0205] The video processing method provided in this application can be applied to various scenarios, such as post-production processing of film and television special effects, live streaming, and short video processing. The following example illustrates the video processing procedure provided in this application.
[0206] Figure 19 This is a schematic diagram of the structure of a video processing system in a live streaming scenario. For example... Figure 19 As shown, the video processing system may include: a live streaming device 10, a service device 20, and terminal devices 30 for users in the live streaming room. The live streaming device 10 may be an electronic device with video capture capabilities, such as a mobile phone, computer, camera, or camcorder. The live streaming device 10 can capture the live stream footage of the host in the live streaming room and upload the captured live stream footage to the service device 20 in the form of a video stream. The video stream uploaded by the live streaming device 10 is the current video frame to be processed by the service device 20. This current video frame includes the image of the host in the live streaming room, and the host's image includes the image of the part of the host that needs to be deformed, i.e., a partial image of the part of the host that needs to be processed, such as the host's arms, legs, or upper body.
[0207] Service device 20 refers to a server-side device that performs special effects processing on the video. It can be a single server device, a cloud-based server array, or a virtual machine (VM) or container running within a cloud-based server array. Alternatively, service device 20 can also refer to other computing devices with corresponding service capabilities, such as computers or other terminal devices (running service programs). Service device 20 can receive video streams uploaded by live streaming devices and, using the video processing methods provided in the foregoing embodiments of this application, perform deformation processing on the parts of the video stream to be processed to obtain the target video stream. For specific implementation methods, please refer to the video processing methods provided in the foregoing embodiments. Furthermore, service device 20 can transmit the target video stream to the terminal device 30 corresponding to the live streaming room. Terminal device 30 can play the target video stream. Viewers in the live streaming room can observe live video with minimal or even imperceptible background jitter between frames.
[0208] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 201 and 202 can be device A; or the execution subject of step 201 can be device A, and the execution subject of step 202 can be device B; and so on.
[0209] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 201, 202, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.
[0210] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps in the video processing methods provided in the foregoing embodiments.
[0211] Computer-readable storage media include volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium.
[0212] This application also provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the video processing methods provided in the foregoing embodiments.
[0213] In this application embodiment, the specific implementation form of the computer program product is not limited. In some embodiments, the computer program product may be implemented as an application (APP), a mini-program, a PC client, a program module, a plug-in, an installation package, a software development kit (SDK), an image file of an optical disc, a plug-in, or software in the form of Software as a Service (SaaS), etc., but is not limited to these.
[0214] The computer program product should understand that each or a combination of the above-described method flow can be implemented by a computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0215] Figure 20 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 20 As shown, the electronic device includes a memory 21a and a processor 21b. The memory 21a is used to store computer programs and can be configured to store various other data to support operation on a computing platform. Examples of this data include instructions for any application or method used to operate on the electronic device, data structures, contact data, phonebook data, messages, pictures, videos, etc.
[0216] Processor 21b is coupled to memory 21a and is used to execute computer programs to perform the steps in the video processing methods provided in the foregoing embodiments. Specific implementation details of each step can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0217] In some alternative implementations, such as Figure 20 As shown, the electronic device may also include optional components such as a communication component 21c, a power supply component 21d, a display component 21e, and an audio component 21f. Figure 20 The diagram only shows some components and does not mean that the electronic device must contain them. Figure 20 The inclusion of all components does not imply that an electronic device can only include... Figure 20 The components shown.
[0218] in addition, Figure 20 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the form factor of the electronic device. The electronic device in this embodiment can be a desktop computer, laptop computer, mobile phone, or IoT device; it can also be a traditional server, cloud server, or server cluster, or other server equipment.
[0219] In this embodiment, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0220] In the embodiments of this application, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), or a complex programmable logic device (CPLD); or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited thereto.
[0221] In this embodiment, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device hosting the communication component can access wireless networks based on communication standards, such as 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0222] In embodiments of this application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0223] In this embodiment, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0224] In embodiments of this application, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. For example, in devices with voice interaction capabilities, voice interaction with the user can be achieved through the audio component.
[0225] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0226] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.
[0227] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A video processing method, characterized in that, include: Identify the parts of the target object that require deformation treatment; Determine the deformation region of the part to be processed in the current video frame, and the deformation region corresponds to the initial deformation intensity; Based on the spatial attribute information of the deformed region in the current video frame, the target morphological parameters of the part to be processed are determined; The target morphological parameters characterize the degree of risk of background jitter caused by the deformation of the part to be processed; Based on the target morphological parameters, the initial deformation intensity is attenuated to obtain the target deformation intensity; wherein, the attenuation magnitude of the initial deformation intensity attenuation is positively correlated with the risk level characterized by the target morphological parameters; Based on the deformation region and the target deformation intensity, the part to be processed in the current video frame is subjected to deformation processing to obtain the target video frame after deformation processing of the current video frame.
2. The method according to claim 1, characterized in that, Determining the target shape parameters of the part to be processed based on the spatial attribute information of the deformed region in the current video frame includes: The spatial attribute information includes: the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions, and the target ratio of the projection span of the boundary contour in multiple preset orthogonal directions is calculated as the target morphological parameter. And / or, The spatial attribute information includes: the area of the deformation region and the area of overlap between the deformation region and the current video frame. The target proportion of the overlap area to the area of the deformation region is calculated and used as the target shape parameter. And / or, The spatial attribute information includes: the first deformation direction vector of the deformation region, then the second deformation direction vector of the associated part of the part to be processed is obtained, and the target angle between the first deformation direction vector and the second deformation direction vector is calculated as the target morphological parameter.
3. The method according to claim 1, characterized in that, Before attenuating the initial deformation intensity according to the target shape parameters, the method further includes: Obtain the first position information of the deformed region in the current video frame and the second position information of the historical deformed region of the part to be processed in the historical video frames; Based on the first location information and the second location information, determine the positional change of the deformed region relative to the historical deformed region; Based on the positional change, determine the target movement amplitude of the part to be processed; The step of attenuating the initial deformation intensity according to the target shape parameters to obtain the target deformation intensity includes: Based on the target shape parameters and the target motion amplitude, the initial deformation intensity is attenuated to obtain the target deformation intensity.
4. The method according to claim 3, characterized in that, The historical video frame includes the video frame preceding the current video frame; the historical deformation region includes the target historical deformation region of the part to be processed in the preceding video frame; the position change includes the target position change between the deformation region and the target historical deformation region; Determining the target movement amplitude of the part to be processed based on the position change includes: Calculate the scale normalization coefficient based on the scale of the historical deformation region; The target position change is normalized using the scale normalization coefficient to obtain the target motion amplitude.
5. The method according to claim 3, characterized in that, The deformed region is rectangular; the first position information includes: the first image coordinates of the first projection endpoint in the current video frame; the second position information includes: the second image coordinates of the second projection endpoint in the historical video frame; wherein, the first projection endpoint is the endpoint of the projection line segment of the deformed region on the central axis of the bone orientation of the part to be processed along the target direction; the second projection endpoint is the endpoint of the projection line segment of the historical deformed region on the central axis along the target direction; the target direction is a direction parallel to the bone orientation; Determining the positional change of the deformed region relative to the historical deformed region based on the first position information and the second position information includes: The position change is determined based on the distance between the first image coordinates and the second image coordinates.
6. The method according to claim 1, characterized in that, The step of attenuating the initial deformation intensity according to the target shape parameters to obtain the target deformation intensity includes: Based on the target shape parameters and the first mapping relationship between the preset attenuation factor and the shape parameters, the first attenuation factor corresponding to the target shape parameters is determined; the attenuation factor is negatively correlated with the attenuation amplitude.
7. The method according to claim 3, characterized in that, The step of attenuating the initial deformation intensity based on the target shape parameters and the target motion amplitude to obtain the target deformation intensity includes: Based on the target shape parameters and the first mapping relationship between the preset attenuation factor and the shape parameters, the first attenuation factor corresponding to the target shape parameters is determined; the attenuation factor is negatively correlated with the attenuation amplitude. Based on the target motion amplitude and the second mapping relationship between the preset attenuation factor and the motion amplitude, a second attenuation factor corresponding to the target motion amplitude is determined; the attenuation factor is negatively correlated with the motion amplitude. The initial deformation intensity is attenuated based on the first attenuation factor and the second attenuation factor to obtain the target deformation intensity.
8. The method according to claim 6 or 7, characterized in that, The target morphological parameters include: the target ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions; The step of determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and a preset mapping relationship between the attenuation factor and the shape parameter includes: The target ratio is input into a smooth first monotonically increasing function in the mapping relationship for calculation to obtain a first attenuation factor; the independent variable of the first monotonically increasing function is the ratio of the projection span of the boundary contour of the deformed region in multiple preset orthogonal directions, and the dependent variable is the attenuation factor.
9. The method according to claim 6 or 7, characterized in that, The target morphological parameters include: the target ratio of the area of the deformation region overlapping with the current video frame to the area of the deformation region; The step of determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and the preset mapping relationship between the attenuation factor and the shape parameter includes: The target ratio is input into a smooth, monotonically increasing second function in the mapping relationship to calculate the first attenuation factor; the independent variable of the second monotonically increasing function is the ratio of the area of the deformation region to the area of the current video frame, and the dependent variable is the attenuation factor.
10. The method according to claim 6 or 7, characterized in that, The target morphological parameters include: the target angle between the first deformation direction vector of the deformed region and the second deformation direction vector of the associated part of the part to be processed; The step of determining the first attenuation factor corresponding to the target shape parameter based on the target shape parameter and the preset mapping relationship between the attenuation factor and the shape parameter includes: The target angle is input into a smooth first monotonically decreasing function in the mapping relationship for calculation to obtain a first attenuation factor; the independent variable of the first monotonically decreasing function is the angle between the first deformation direction vector and the second deformation direction vector, and the dependent variable is the attenuation factor.
11. The method according to claim 7, characterized in that, The step of determining the second attenuation factor corresponding to the target motion amplitude based on the target motion amplitude and the second mapping relationship between the preset attenuation factor and the motion amplitude includes: The target motion amplitude is input into a smooth, monotonically decreasing function in the second mapping relationship for calculation to obtain a second attenuation factor; the independent variable of the second monotonically decreasing function is the motion amplitude of the part to be processed, and the dependent variable is the attenuation factor.
12. The method according to claim 7, characterized in that, The target morphological parameters are multiple; Multiple target morphological parameters correspond to multiple first attenuation factors; The step of attenuating the initial deformation intensity according to the first attenuation factor and the second attenuation factor to obtain the target deformation intensity includes: The initial deformation intensity is weighted and attenuated by multiplying the plurality of first attenuation factors and the second attenuation factor to obtain the target deformation intensity.
13. The method according to any one of claims 1-7 and 11-12, characterized in that, The step of determining the deformation region of the part to be processed in the current video frame based on the image coordinates of the target key points of the part to be processed in the current video frame includes: Perform key point detection on the part to be processed in the current video frame to determine the first image coordinates of the target key point of the part to be processed in the current video frame; The deformation region is determined based on the shape of the deformation region that matches the first image coordinates and the part to be processed. The initial deformation intensity is determined based on the shape of the deformation region that matches the first image coordinates and the part to be processed.
14. The method according to claim 13, characterized in that, The step of determining the deformation region based on the shape of the deformation region adapted to the first image coordinates and the region to be processed includes: Based on the second image coordinates of the target key point in the historical video frame, the first image coordinates are temporally smoothed to obtain the target image coordinates of the target key point in the current video frame; The deformation region is determined based on the shape of the deformation region that matches the target image coordinates and the part to be processed.
15. The method according to claim 13, characterized in that, The deformable area is rectangular in shape; the part to be processed includes the upper arm; the target key points include the midpoint of the shoulder, the midpoint of the elbow, and the inner and outer boundary points of the shoulder. The step of determining the deformation region based on the shape of the deformation region adapted to the first image coordinates and the region to be processed includes: The first image coordinates of the shoulder midpoint in the current video frame are extended along the upper arm bone direction away from the elbow midpoint by a first length adjustment factor multiple of the upper arm length, to obtain the starting point of the target central axis of the rectangular deformation area along the upper arm bone direction. The first image coordinates of the elbow midpoint are extended along the upper arm bone direction away from the shoulder midpoint by a second length adjustment factor multiple of the upper arm length to obtain the endpoint of the target central axis. The width of the deformable region of the rectangle is determined based on the distance between the inner and outer boundary points of the shoulder and the set width adjustment coefficient. The deformation region of the rectangle is determined based on the starting point of the target centerline, the ending point of the target centerline, and the width of the deformation region of the rectangle.
16. The method according to claim 13, characterized in that, The deformed region is rectangular in shape; the part to be processed includes the upper arm; determining the initial deformation intensity based on the shape of the deformed region adapted to the first image coordinates and the part to be processed includes: The length of the deformable region of the rectangle is determined by the distance between the starting point and the ending point of the target central axis; the target central axis is the central axis of the deformable region of the rectangle along the direction of the upper arm bones. For a target pixel within the deformation region of the rectangle, calculate the first projected length of the line connecting the target pixel to the starting point of the target central axis, which is perpendicular to the skeletal direction of the upper arm, and calculate the second projected length of the line connecting the target pixels, which is parallel to the skeletal direction of the upper arm; the target pixel is any pixel within the deformation region of the rectangle. Based on the first projection length and the width of the deformed region of the rectangle, the first projection length is normalized and attenuated to obtain the vertical intensity component; and based on the second projection length and the length, the second projection length is normalized and attenuated to obtain the parallel intensity component. The initial deformation intensity is determined based on the vertical direction intensity component and the parallel direction intensity component.
17. A video processing method, characterized in that, include: Receives a video stream uploaded by a live streaming device; the video stream includes an image of the host in the live streaming room. Using the video processing method according to any one of claims 1-16, the part of the video stream that needs to be deformed is deformed to obtain the target video stream; The target video stream is transmitted to the terminal device corresponding to the live broadcast room.
18. An electronic device, characterized in that, include: A memory and a processor; wherein the memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-17.
19. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the method according to any one of claims 1-17.
20. A computer program product, characterized in that, Includes a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps of the method according to any one of claims 1-17.