Image processing method and device, equipment and storage medium
By identifying key human body points and facial expressions in image frames, image frames that match specified human postures and facial expressions are automatically selected, solving the problem of low efficiency in manual operation in existing technologies and achieving highly efficient automated screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INTELLINDUST INFORMATION TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Current technologies for capturing exciting moments in sports scenes require manual operation, which is inefficient and cannot automatically select target image frames that meet user needs.
By performing human key point recognition on the image frames to be detected, calculating the matching degree between the key point set and the pre-stored standard key point set, and combining it with facial expression scoring, a final score is generated, and target image frames that conform to the specified human posture and facial expression are automatically selected.
It achieves automated filtering of target image frames that meet user needs, reduces manual operation, improves efficiency, and can accurately identify and select image frames that are close to the specified human posture and match the facial expression.
Smart Images

Figure CN121901449A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology
[0002] In sports and other similar settings, sports enthusiasts have a need to record and share memorable images of their activities. For example, on a tennis court, enthusiasts want to capture images of themselves in a specific pose with appropriate facial expressions during a move, for motion analysis or social media sharing. This specific pose could be a standard serving or receiving stance, etc.
[0003] However, capturing exciting moments usually requires cameramen to take snapshots, followed by manual selection of the best images from the captured footage. Alternatively, the entire motion can be recorded, and then manually reviewed to extract the best shots. Clearly, existing methods rely on manual operation, which is inefficient. Summary of the Invention
[0004] The purpose of this application is to provide an image processing method, apparatus, device, and storage medium to automatically filter out target image frames that meet user needs, thereby reducing user operations. The specific technical solution is as follows:
[0005] In a first aspect, embodiments of this application provide an image processing method, the method comprising:
[0006] For each image frame to be detected in the image sequence, human key points are identified in the human image region of the image frame to be detected to obtain the first set of key points.
[0007] The matching degree between the first set of key points and the pre-stored second set of key points is calculated to obtain the pose score of the image frame to be detected; wherein, the second set of key points includes a standard set of key points, and the standard set of key points includes: human key points of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree;
[0008] The facial score of the image frame to be detected is determined according to the index value of the facial image region in the image frame to be detected relative to the preset index; wherein, the preset index includes: the degree of matching between facial expression and the specified human posture; the facial score of the image frame to be detected is positively correlated with the index value of the preset index;
[0009] The final score of the image frame to be detected is generated by combining the face score and pose score of the image frame to be detected.
[0010] A target image frame is selected from each image frame to be detected; wherein the final score of the target image frame is higher than the final score of the other image frames to be detected.
[0011] Optionally, the matching degree between the first set of key points and the pre-stored second set of key points is calculated to obtain the pose score of the image frame to be detected, including:
[0012] Based on the similarity between each limb key point in the first key point set and the corresponding limb key points in the pre-stored second key point set, the matching degree between the first key point set and the second key point set is calculated.
[0013] Based on the calculated matching degree, the pose score of the image frame to be detected is obtained.
[0014] Optionally, the similarity between any limb keypoint in the first keypoint set and the corresponding limb keypoint in the pre-stored second keypoint set is calculated using the following formula:
[0015] ;
[0016] Wherein, OKS represents similarity, d represents the Euclidean distance between the first coordinate of the limb keypoint in the first keypoint set in the image frame to be detected and the second coordinate of the corresponding limb keypoint in the second keypoint set in the standard image; σ represents the normalization factor corresponding to the limb keypoint; and the σ value of the preset active keypoint among the limb keypoints is less than the σ value of other limb keypoints.
[0017] Optionally, the second set of key points may be multiple; the attitude score is determined based on the highest preset number of matching degrees.
[0018] Optionally, each second key point set further includes a sub-standard key point set, and any sub-standard key point set includes: human body key points of the sub-standard image; the similarity between the human body key points of any sub-standard image and the human body key points of the standard image is greater than a preset similarity threshold.
[0019] The pose score is: the weighted average of the highest preset number of matching degrees; the weight of the matching degree corresponding to the standard image is greater than the weight of the matching degree corresponding to the next standard image.
[0020] Optionally, the method further includes:
[0021] For each selected target image frame, the area ratio of the human image region in that target image frame is detected;
[0022] If the detected area ratio is less than a preset ratio threshold, then calculate the magnification factor required for the human image region in the target image frame to reach the preset ratio threshold.
[0023] The target image frame is virtually magnified according to the calculated magnification factor, and the virtual cropping coordinates are calculated according to the predetermined cropping conditions; wherein, the predetermined cropping conditions are that the cropping meets the preset resolution, preset aspect ratio, and the human image area in the cropped image area conforms to the golden ratio composition principle.
[0024] The virtual cropping coordinates are mapped onto the target image frame to crop the target image frame, and the resulting cropped image is enlarged according to the magnification factor to obtain an image frame for sending to the user terminal.
[0025] Optionally, the obtained cropped image is enlarged according to the magnification factor to obtain an image frame for sending to the user terminal, including:
[0026] The cropped image is then magnified according to the stated magnification factor to obtain an intermediate image;
[0027] A portrait enhancement operation is performed on the face image region of the intermediate image to obtain the enhanced image;
[0028] The human image region in the obtained enhanced image is subjected to a first number of high-definition restorations, and the obtained enhanced image is subjected to a second number of high-definition restorations to obtain an image frame for sending to the user terminal; the first number is greater than the second number.
[0029] Optionally, the preset index further includes image fidelity, and determining the facial score of the image frame to be detected according to the index value of the facial image region in the image frame to be detected relative to the preset index includes:
[0030] The face image region of the image frame to be detected is input into the face quality model to obtain the first score; wherein, the face quality model is used to score the input image based on the image fidelity of the input image;
[0031] The facial image region of the image frame to be detected is input into the expression recognition model to obtain a second score; wherein, the expression recognition model is used to score the input image based on the degree of matching between the facial expression of the input image and the specified human posture;
[0032] The facial score of the image frame to be detected is generated by combining the first score and the second score.
[0033] Secondly, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0034] The recognition module is used to identify human key points in the human image region of each image frame in the image sequence to be detected, and obtain the first set of key points.
[0035] The pose score determination module is used to calculate the matching degree between the first keypoint set and the pre-stored second keypoint set to obtain the pose score of the image frame to be detected; wherein, the second keypoint set includes a standard keypoint set, which includes human keypoints of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree;
[0036] The face score determination module is used to determine the face score of the image frame to be detected according to the index value of the face image region in the image frame to be detected against the preset index; wherein, the preset index includes: image fidelity, and / or, the degree of matching between facial expression and the specified human posture; the face score of the image frame to be detected is positively correlated with the index value of the preset index;
[0037] The final score generation module is used to combine the face score and pose score of the image frame to be detected to generate the final score of the image frame to be detected.
[0038] The selection module is used to select a target image frame from each image frame to be detected; wherein the final score of the target image frame is higher than the final scores of the other image frames to be detected.
[0039] Optionally, the attitude score determination module includes:
[0040] The calculation submodule is used to calculate the matching degree between the first key point set and the second key point set based on the similarity between each limb key point in the first key point set and the corresponding limb key point in the pre-stored second key point set.
[0041] The determination submodule is used to obtain the pose score of the image frame to be detected based on the calculated matching degree.
[0042] Optionally, the similarity between any limb keypoint in the first keypoint set and the corresponding limb keypoint in the pre-stored second keypoint set is calculated using the following formula:
[0043] ;
[0044] Wherein, OKS represents similarity, d represents the Euclidean distance between the first coordinate of the limb keypoint in the first keypoint set in the image frame to be detected and the second coordinate of the corresponding limb keypoint in the second keypoint set in the standard image; σ represents the normalization factor corresponding to the limb keypoint; and the σ value of the preset active keypoint among the limb keypoints is less than the σ value of other limb keypoints.
[0045] Optionally, the second set of key points may be multiple; the attitude score is determined based on the highest preset number of matching degrees.
[0046] Optionally, each second key point set further includes a sub-standard key point set, and any sub-standard key point set includes: human body key points of the sub-standard image; the similarity between the human body key points of any sub-standard image and the human body key points of the standard image is greater than a preset similarity threshold.
[0047] The pose score is: the weighted average of the highest preset number of matching degrees; the weight of the matching degree corresponding to the standard image is greater than the weight of the matching degree corresponding to the next standard image.
[0048] Optionally, the device further includes:
[0049] The first detection module is used to detect the area ratio of the human image region in each selected target image frame.
[0050] The magnification factor determination module is used to calculate the magnification factor required for the human image region in the target image frame to reach the preset ratio threshold if the detected area ratio is less than the preset ratio threshold.
[0051] The virtual magnification module is used to virtually magnify the target image frame according to the calculated magnification factor, and calculate the virtual cropping coordinates according to the predetermined cropping conditions; wherein, the predetermined cropping conditions are that the cropping meets the preset resolution, the preset aspect ratio, and the human image area in the cropped image area conforms to the golden ratio composition principle.
[0052] The cropping module is used to map virtual cropping coordinates onto the target image frame to crop the target image frame, and then enlarge the cropped image according to the magnification factor to obtain an image frame for sending to the user terminal.
[0053] Optionally, the cropping module includes:
[0054] The magnification submodule is used to magnify the obtained cropped image according to the magnification factor to obtain an intermediate image;
[0055] The enhancement submodule is used to perform portrait enhancement operations on the face image region of the intermediate image to obtain an enhanced image;
[0056] The repair submodule is used to perform a first number of high-definition repairs on the human image region in the obtained enhanced image, and to perform a second number of high-definition repairs on the obtained enhanced image, to obtain an image frame for sending to the user terminal; the first number is greater than the second number.
[0057] Optionally, the preset index further includes image fidelity, and the face score determination module includes:
[0058] The first score determination submodule is used to input the face image region of the image frame to be detected into the face quality model to obtain the first score; wherein, the face quality model is used to score the input image based on the image fidelity of the input image;
[0059] The second score determination submodule is used to input the face image region of the image frame to be detected into the expression recognition model to obtain a second score; wherein, the expression recognition model is used to score the input image based on the degree of matching between the facial expression of the input image and the specified human posture;
[0060] The face score generation submodule is used to combine the first score and the second score to generate the face score of the image frame to be detected.
[0061] Thirdly, embodiments of this application provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0062] Memory, used to store computer programs;
[0063] The processor, when executing a program stored in memory, implements the image processing method described in the first aspect above.
[0064] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the image processing method described in the first aspect.
[0065] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the image processing method described in the first aspect above.
[0066] Beneficial effects of the embodiments in this application:
[0067] The solution provided in this application, for each image frame to be detected, calculates the matching degree between a first set of key points and a pre-stored second set of key points to obtain the matching degree between the human posture in the image frame and a specified human posture. Since the posture score is determined based on the matching degree between the human posture in the image frame and the specified human posture, and the higher the matching degree, the higher the posture score, image frames containing human postures close to the specified human posture can be identified from each image frame based on their posture scores. Since the facial score of the image frame to be detected is positively correlated with the value of a preset index, image frames with a high degree of matching between facial expressions and the specified human posture can be identified from each image frame based on their facial scores. Therefore, by combining the facial score and posture score to generate a final score, and selecting target image frames with higher final scores from each image frame, target image frames that are close to the specified human posture and whose facial expressions have a high degree of matching with the specified human posture can be filtered out from each image frame.
[0068] As can be seen, this solution can comprehensively consider the matching degree between human posture and facial expression and the specified human posture, so as to automatically select target image frames from each image frame whose human posture is close to the specified human posture and whose facial expression matches the specified human posture to a high degree, without the need for manual selection, thus reducing user operations.
[0069] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0070] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0071] Figure 1 A flowchart illustrating an image processing method provided in an embodiment of this application;
[0072] Figure 2 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0073] Figure 3 A flowchart illustrating a specific example of an image processing method provided in an embodiment of this application;
[0074] Figure 4 A flowchart illustrating a method for optimizing a person's image in an embodiment of this application;
[0075] Figure 5A schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application;
[0076] Figure 6 A block diagram of an electronic device for implementing the image processing method provided in the embodiments of this application. Detailed Implementation
[0077] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0078] The image processing method provided in this application can be applied to various electronic devices, such as AI (Artificial Intelligence) cameras, personal computers, servers, and other devices with data processing capabilities. In one scenario, the electronic device executing the image processing method can be an image acquisition device, that is, the image acquisition device performs image processing on the acquired image frames locally after acquiring them. Furthermore, it is understood that the image processing method provided in this application can be implemented through software, hardware, or a combination of both.
[0079] like Figure 1 As shown, the image processing method provided in this application embodiment includes steps S101-S105:
[0080] S101, for each image frame to be detected in the image sequence to be detected, human key points are identified in the human image region of the image frame to be detected to obtain the first key point set.
[0081] For example, in practical applications, if a user needs motion capture, an image acquisition device can be used to capture images of the scene in which the user is located. For instance, the image acquisition device can continuously capture images to obtain a video of the user's motion, and use each image frame in the video as the image sequence to be detected. Alternatively, the image acquisition device can also capture images at preset time intervals (e.g., at 1-second intervals) to obtain a sequence of motion photos of the user, which can then be used as the image sequence to be detected.
[0082] For example, after acquiring each image frame to be detected, the image acquisition device can send the image frame to be detected to an electronic device that performs the image processing method provided in this application embodiment, so that the electronic device can detect the human image region and the face image region in each image frame to be detected according to the time sequence of each image frame to be detected acquired by the image acquisition device. This application embodiment does not limit the specific type of image acquisition device; for example, the image acquisition device can be a PTZ camera, a bullet camera, or a mobile phone, etc., that have image acquisition capabilities.
[0083] For example, each image frame to be detected can be input into a pre-trained human face detection model to obtain the human image region and face image region in the image frame to be detected, as output by the human face detection model.
[0084] For example, the initial human face detection model can employ models such as Mask R-CNN (Mask Region-based Convolutional Neural Network) or SSD (Single-Shot Multibox Detector). In practical applications, the initial human face detection model can be trained using a large number of sample images labeled with bounding boxes for human and face regions, enabling the model to learn the ability to locate human and face regions within an image. After training, inputting a frame containing any person into the trained human face detection model will yield the output bounding boxes for human and face detection. The image region containing the human bounding box is the human image region, and the image region containing the face bounding box is the face image region.
[0085] For example, human key points in human image regions can be identified using human key point recognition models such as PoseNet (Pose Estimation Network) and SimplePose (Lightweight Human Pose Estimation Model).
[0086] For example, inputting a human image region into the PoseNet model yields the coordinates and confidence scores of 17 human keypoints output by the model. After obtaining the model's output keypoints, the set containing all human keypoints can be used as the first keypoint set, or the set containing keypoints with confidence scores greater than a preset confidence threshold can be used as the first keypoint set; both are reasonable. For instance, if 5 of the 17 keypoints output by the model have confidence scores less than the preset confidence threshold, these 5 keypoints are considered invalid. In this case, the remaining 12 keypoints with confidence scores not less than the preset confidence threshold can be used to construct the first keypoint set. For example, the preset confidence threshold could be 0.2 or 0.3, etc.
[0087] In one implementation, if the confidence level of more than a preset number of human key points obtained based on any image frame to be detected is less than a preset confidence threshold, the image frame to be detected can be directly removed. That is, no further processing is performed on the image frame to be detected, and the image frame to be detected does not participate in the selection of target image frames.
[0088] Understandably, if the confidence level of more than a preset number of human keypoints is less than a preset confidence threshold, it indicates that there are many human keypoints in the image frame to be detected whose exact location is difficult to determine. In other words, the human image area in the image frame to be detected may be largely occluded, making it difficult for the model to accurately detect most human keypoints. In this case, it can be considered that there are a large number of missing human keypoints in the image frame to be detected, and the image quality of the image frame to be detected is poor, so it can be directly discarded.
[0089] For example, the preset number could be 8 or 10, etc., or the preset number could be determined based on the total number of human keypoints that the model can detect. For example, the preset number could be 1 / 3 or 1 / 2 of the total number of detected human keypoints, etc. If the confidence level of the human keypoints is less than the preset confidence threshold, then the subsequent step S102 can be executed.
[0090] S102, calculate the matching degree between the first key point set and the pre-stored second key point set to obtain the pose score of the image frame to be detected; wherein, the second key point set contains a standard key point set, and any standard key point set contains: human key points of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree.
[0091] In this embodiment, a standard image containing a specified human posture can be pre-selected from images containing different human postures by a relevant technician. The selected standard image is related to the current motion capture scene. The images containing different human postures can be pre-captured by image acquisition devices in the current motion capture scene, or images related to the motion capture scene can be filtered from a network image library. For example, if the motion capture scene is a tennis scene, the technician can select images containing proper tennis serve and / or tennis receive postures from images related to tennis as standard images. If the motion capture scene is a tourist photography scene, the technician can select images containing aesthetically pleasing and graceful human postures from various tourist photos (such as jumping postures) as standard images.
[0092] For example, in practical applications, after selecting a standard image containing a specified human pose, relevant technicians can input the standard image containing the specified human pose into a human keypoint detection model to extract the human keypoints of the standard image containing the specified human pose. The human keypoints of each standard image constitute a set of standard keypoints.
[0093] For example, the second set of key points can be one or more. That is, one or more standard images can be selected in advance, and correspondingly, one or more sets of standard key points can be extracted from the pre-selected standard images, each set of standard key points being a second set of key points.
[0094] In one implementation, step S102, which calculates the matching degree between the first set of key points and the pre-stored second set of key points to obtain the pose score of the image frame to be detected, may include steps A1-A2:
[0095] A1. Based on the similarity between each limb keypoint in the first keypoint set and the corresponding limb keypoint in the pre-stored second keypoint set, calculate the matching degree between the first keypoint set and the second keypoint set.
[0096] A2, based on the calculated matching degree, obtain the pose score of the image frame to be detected.
[0097] Understandably, key points for each human body can include "left eye," "right eye," "nose," "neck," "left elbow," "right elbow," "left hip," "right hip," and so on. When calculating the matching degree between the first set of key points and the second set of key points, for key points representing the same human body position in the first and second sets of key points, the similarity between these key points is calculated. For example, for the "left elbow" key point in the first set of key points, the similarity between this key point and the "left elbow" key points in each of the second set of key points is calculated; for the "right elbow" key point in the first set of key points, the similarity between this key point and the "right elbow" key points in each of the second set of key points is calculated.
[0098] Since posture assessment in motion scenarios primarily focuses on limb movements, key points of the limbs can be used as core key points for posture comparison. Based on the similarity between the corresponding limb key points in the first and second key point sets, the matching degree between the two sets is calculated. For example, the selected limb key points may include: shoulder (left, right), elbow (left, right), wrist (left, right), hip (left, right), knee (left, right), and ankle (left, right).
[0099] In this implementation, each limb key point can be selected from the first key point set, and for each selected limb key point, the similarity between the limb key point in the first key point set and the corresponding limb key point in the second key point set can be calculated.
[0100] For example, the Euclidean distance between the first coordinates of the limb keypoints in the first keypoint set in the image frame to be detected and the second coordinates of the corresponding limb keypoints in the second keypoint set in the standard image can be calculated. Based on the Euclidean distances between the corresponding limb keypoints, the similarity between them can be determined. Then, the matching degree between the first and second keypoint sets can be determined based on the similarity between the corresponding limb keypoints. Here, Euclidean distance is negatively correlated with similarity; the reciprocal of the Euclidean distance can be used as the similarity. After obtaining the similarity between the corresponding limb keypoints, the mean or weighted sum of the similarities between the limb keypoints can be used as the matching degree between the first and second keypoint sets. At this point, the weights of each limb keypoint can be set by relevant technical personnel based on experience. For example, the weights of all limb keypoints can be the same, or the weights of limb keypoints for parts with large movements such as elbows, wrists, knees, and ankles can be set higher than the weights of limb keypoints for stable parts such as shoulders; these are all reasonable approaches.
[0101] For example, in another implementation, the similarity between any limb keypoint in the first keypoint set and the corresponding limb keypoint in the second keypoint set is calculated using the following formula:
[0102] ;
[0103] Wherein, OKS (Object Keypoint Similarity) represents the similarity, d represents the Euclidean distance between the first coordinate of the limb keypoint in the first keypoint set in the image frame to be detected and the second coordinate of the corresponding limb keypoint in the second keypoint set in the standard image; σ represents the normalization factor corresponding to the limb keypoint; and the σ value of the preset active keypoint among the limb keypoints is less than the σ value of other limb keypoints.
[0104] In this implementation, the preset active keypoints are keypoints of parts with large range of motion, such as the elbow, wrist, knee, and ankle. Smaller σ values are set for these active keypoints, making the similarity between them more sensitive to Euclidean distance. This makes the matching of active keypoints more stringent; that is, a smaller Euclidean distance is needed between active keypoints to achieve a higher OKS (Objectives and Search Results). For example, the value of σ ranges from (0,1), the σ value for active keypoints can be set to 0.2, and the σ value for other limb keypoints can be set to 0.5.
[0105] In this implementation, after calculating the similarity between each limb keypoint in the first keypoint set and the second keypoint set, the average of the similarities between the limb keypoints can be used as the matching degree between the first keypoint set and the second keypoint set. For example, if 10 limb keypoints are selected from the first keypoint set, then for each limb keypoint, the similarity between that limb keypoint and the limb keypoint in the second keypoint set can be calculated, resulting in 10 similarity scores. The average of these 10 similarity scores is used as the matching degree between the first keypoint set and the second keypoint set.
[0106] In one implementation, the second keypoint set is multiple, and the pose score is determined based on the highest preset number of matching degrees.
[0107] In this implementation, after calculating the matching degree between the first set of key points and each of the second set of key points, the attitude score can be determined based on the highest preset number of matching degrees. For example, the preset number can be 3 or 5, etc.
[0108] For example, the average of the highest preset number of matching degrees can be used as the pose score of the image frame to be detected. For instance, if 10 sets of second keypoints are pre-stored (preset number 3), the average of the matching degrees corresponding to the 3 second keypoint sets with the highest matching degrees to the first keypoint set can be used as the pose score of the image frame to be detected. It is understood that determining the pose score of the image frame to be detected based on the preset number of matching degrees with the highest matching degrees can improve the stability of pose scoring and reduce the bias caused by accidental matching.
[0109] S103, determine the facial score of the image frame to be detected according to the index value of the face image region in the image frame to be detected relative to the preset index; wherein, the preset index includes: the degree of matching between facial expression and specified human posture; the facial score of the image frame to be detected is positively correlated with the index value of the preset index.
[0110] In this embodiment, the facial score of the image frame to be detected is determined based on the degree of matching between the facial expression in the image frame and the specified human posture, and the higher the degree of matching, the higher the facial score.
[0111] For example, the face image region of the image frame to be detected can be input into the expression recognition model to obtain the index value of the degree of matching between the facial expression and the specified human posture output by the expression recognition model.
[0112] For example, an expression recognition model can be trained as follows: First, obtain a facial image region from a standard image containing a specified human pose, as the first sample; second, obtain facial image regions from images of people containing other human poses, as the second sample. The labels for each sample are set as follows: the first sample for a smiling expression is labeled 1, the second sample is labeled 0, and the first sample for a natural expression is labeled between 0 and 1, for example, 0.5. The first and second samples are input into the initial expression recognition model to obtain the predicted index value for each sample output by the model. Based on the difference between each predicted index value and its corresponding label, the model loss value is calculated. The model is trained by backpropagation that minimizes the model loss value until the model converges, resulting in a trained expression recognition model.
[0113] The facial expression recognition model trained in the above manner can learn facial expressions that match a specified human posture. Among the facial expressions that match a specified human posture, the matching degree of smiling expressions is higher than that of natural expressions.
[0114] For example, the initial facial expression recognition model can use CNN (Convolutional Neural Network) or DNN (Deep Neural Network), and so on.
[0115] In one implementation, the preset index further includes image fidelity. Step S103, which determines the facial score of the image frame to be detected based on the index value of the face image region in the image frame relative to the preset index, may include steps B1-B3:
[0116] B1, input the face image region of the image frame to be detected into the face quality model to obtain the first score; wherein, the face quality model is used to score the input image based on the image fidelity of the input image;
[0117] B2, input the face image region of the image frame to be detected into the expression recognition model to obtain the second score; wherein, the expression recognition model is used to score the input image based on the degree of matching between the facial expression of the input image and the specified human posture;
[0118] B3, combining the first score and the second score, generates the facial score for the image frame to be detected.
[0119] For example, image fidelity may include sharpness and / or occlusion. The higher the sharpness of the face image region, the higher the image fidelity; the higher the occlusion of the face image region, the lower the image fidelity.
[0120] In this implementation, a face image region can be input into a pre-trained face quality model. This model detects the clarity and occlusion of the face in the image region and scores them based on the detected clarity and occlusion, resulting in a first score output by the model. Higher clarity and lower occlusion in the face image region result in a higher first score from the face quality model.
[0121] For example, a face quality model can be trained as follows: Acquire face images with varying sharpness and occlusion levels as sample images. Assign higher-resolution and less occluded sample images a larger label. Input each sample image into the initial face quality model to obtain a predicted score for each image. Based on the difference between each predicted score and its corresponding label, calculate the model loss value. Train the model by minimizing the model loss value through backpropagation until the model converges, resulting in a trained face quality model.
[0122] For example, the initial face quality model can use CNN (Convolutional Neural Network) or DNN (Deep Neural Network), etc.
[0123] Understandably, if the first score is lower than the set threshold (e.g., 0.9), it is considered that the image is blurry. In this case, the image frame to be detected can be rejected directly, meaning that the image frame to be detected will not participate in the subsequent selection of target image frames.
[0124] An expression recognition model is used to detect facial expressions in a face image region. A score is then assigned to the face image region of each image frame based on the degree of matching between the detected facial expression and a specified human posture, resulting in a second score. The expression recognition model can identify expressions such as happy, natural, smiling, disgusted, and aversion. Among facial expressions matching a specified human posture, positive expressions (happy, smiling) receive higher scores than "natural" expressions. If a negative expression (disgusted, aversion) is detected, a lower score is output, or the image frame is directly removed. The training process of the expression recognition model can be referred to the relevant description above, and will not be repeated here.
[0125] For example, in practical applications, the facial expression recognition model can include a binary classification network to identify two types of expressions: "natural" and "smiling". If the confidence scores of both types of expressions are lower than a preset confidence threshold (e.g., 0.1), the facial expression in the image frame can be considered poor, such as closed eyes or a tense mouth. In this case, the image frame can be directly removed.
[0126] For example, if a smiling expression is detected and the degree of matching between the smiling expression and the specified human posture is high, the confidence score of the smiling expression can be multiplied by a weight greater than 1 (e.g., 5) as the second score. If a natural expression is detected and the degree of matching between the natural expression and the specified human posture is high, the confidence score of the natural expression can be directly used as the second score. If a smiling expression or a natural expression with a low degree of matching with the specified human posture is detected, the image frame to be detected can be discarded to ensure that the facial expressions in the subsequently selected target image frames match the specified human posture.
[0127] For example, if a facial expression in any image frame to be detected is a natural expression that matches a specified human posture, the face image region can be input into an eye detection model to detect the degree of eye opening and closing. If the eyes in the image frame are detected as "closed," the image frame to be detected can be directly discarded. In this way, image frames with "natural expressions + closed eyes" can be discarded to remove image frames with out-of-control expressions, thus improving the content quality of the final selected target image frames. When detecting the degree of eye opening and closing, the face image region can be expanded proportionally to ensure that the image frame contains the complete eye area.
[0128] After obtaining the first and second scores, their average can be used as the facial score. Alternatively, the weighted sum of the two scores can be used as the facial score according to preset weights. The weights corresponding to the first and second scores can be set by relevant technical personnel based on experience. For example, the second score can be assigned a higher weight to primarily combine facial expressions and body posture when selecting target image frames.
[0129] In another implementation, before executing step S103, it can be determined whether the face image region in the image frame to be detected is smaller than a preset size. For example, it can be determined whether the width or height of the face image region in the image frame to be detected is less than 30 pixels. If the width or height of the face image region is too small, it is considered that the face information is insufficient to determine the face quality, and the image frame to be detected can be directly discarded. If the face image region is not smaller than the preset size, then step S103 is executed.
[0130] It should be noted that, in another embodiment, step S103 may be executed before step S102, or may be executed in parallel with step S102. Both are reasonable. The order of steps S103 and S102 is not limited in the embodiments of this application.
[0131] S104, Combine the face score and pose score of the image frame to be detected to generate the final score of the image frame to be detected;
[0132] For example, the face score and pose score of the image frame to be detected can be weighted and summed to obtain the final score of the image frame. The weights corresponding to the face score and pose score can be set by relevant technicians based on experience. In one implementation, the face score and pose score have the same weight to balance the impact of facial expression and human pose on the overall image quality.
[0133] S105, Select the target image frame from each image frame to be detected; wherein the final score of the target image frame is higher than the final score of the other image frames to be detected.
[0134] In this embodiment, after obtaining the final score of each image frame to be detected, it is reasonable to select image frames with a final score higher than a preset score as target image frames, or to select a preset number of image frames as target image frames according to the final scores from high to low. For example, the preset score could be 80% or 90% of the upper limit of the score range, etc. The preset number could be 3 or 5, etc.
[0135] After selecting the target image frame from each image frame to be detected, the selected target image frame can be directly used as the image frame to be sent to the user terminal, or the target image frame can be further optimized to use the optimized image frame as the image frame to be sent to the user terminal. Both are reasonable. For example, the optimization process may include centralized cropping, high-definition restoration, face whitening, wrinkle removal, etc. It should be noted that, for the sake of clarity in the scheme layout, the specific implementation method of the optimization process will be described in the following embodiments, and will not be repeated here.
[0136] Understandably, after selecting the target video frame, the target image frame can be sent to the user terminal, so that the user can directly obtain the target image frame through the user terminal. This allows the user to obtain the target image frame with a human posture close to the specified human posture and a high degree of matching between the facial expression and the specified human posture. This reduces the user's manual image selection and improves the efficiency of obtaining good images.
[0137] In another embodiment of this application, each second keypoint set further includes a sub-standard keypoint set. Each sub-standard keypoint set includes human body keypoints from the sub-standard image, and the similarity between the human body keypoints of the sub-standard image and the human body keypoints of the standard image is greater than a preset similarity threshold. For example, the preset similarity threshold could be 70% or 80%, etc.
[0138] For example, each set of second keypoints can be constructed through the following steps C1-C2:
[0139] C1, retrieves a pre-built library of beautiful poses and a library of ordinary poses;
[0140] C2, for each image in the beautiful pose image library and the ordinary pose image library, performs human key point recognition on the human image region of the image to obtain a second key point set.
[0141] In this embodiment, multiple standard images containing a specified human posture can be constructed into a beautiful posture image library, and multiple substandard images with a similarity to the specified human posture reaching a predetermined similarity threshold can be constructed into a normal posture image library.
[0142] For example, relevant technicians can pre-select multiple standard images containing specified human poses and construct a beautiful pose image library. Next, from the multiple candidate images, the similarity between the human keypoint set of each candidate image and the human keypoint set of the standard image can be calculated according to the method described above for calculating the matching degree between keypoint sets. Candidate images with a similarity greater than a preset similarity threshold with the human keypoint set of the standard image are then constructed into a general pose image library.
[0143] After constructing the graceful pose image library and the ordinary pose image library, human keypoints can be identified in the human image region of each image in each image library according to the above-mentioned human keypoint identification method, to obtain the sets of secondary keypoints. Among them, the set of human keypoints extracted from each image in the graceful pose image library is the standard keypoint set, and the set of human keypoints extracted from each image in the ordinary pose image library is the secondary standard keypoint set.
[0144] Accordingly, in this embodiment, the pose score of the image frame to be detected is: the weighted average of the highest number of matching degrees; the weight of the matching degree corresponding to the standard image is greater than the weight of the matching degree corresponding to the second-highest standard image.
[0145] For example, if the highest number of matching scores includes a matching score corresponding to the standard image, then the matching score is multiplied by a first weight to obtain a first score; if the highest number of matching scores includes a matching score corresponding to the second standard image, then the matching score is multiplied by a second weight to obtain a second score. The average of the first score and the second score is used as the pose score of the image frame to be detected. The first weight is greater than 1, and the second weight is less than 1 but greater than 0.
[0146] For example, the first weight could be 1.2 or 1.3, etc., and the second weight could be 0.7 or 0.8, etc.
[0147] Understandably, if the highest preset number of matches includes a match corresponding to the standard image, multiplying that match by a weight greater than 1 can improve the final pose score. Conversely, if the highest preset number of matches includes a match corresponding to the second-highest standard image, multiplying that match by a weight less than 1 can decrease the final pose score.
[0148] Understandably, by constructing a general pose image library, this solution ensures that even users who cannot accurately depict a specified human pose can still filter images that are relatively close to it. Furthermore, by assigning a greater weight to the matching degree of the standard image than to the matching degree of the next-lower standard image, image frames that closely match the human pose of the standard image receive higher pose scores. This allows for more accurate filtering of image frames that are closer to the specified human pose when combining pose scores for subsequent image frame selection.
[0149] Additionally, if the pose score is below a certain threshold (e.g., 0.2), it indicates poor pose, and the image frame to be detected can be directly filtered out to reduce computation. That is, image frames with excessively low pose scores are not included in the selection of target image frames. When selecting target image frames from each image frame to be detected, image frames with pose scores below the threshold have already been removed.
[0150] In another embodiment of this application, in Figure 1 Based on the illustrated embodiments, as Figure 2 As shown, the above image processing method may include steps S201-S204:
[0151] S201, For each selected target image frame, detect the area ratio of the human image region in the target image frame;
[0152] S202, if the detected area ratio is less than the preset ratio threshold, calculate the magnification factor required for the human image region in the target image frame to reach the preset ratio threshold.
[0153] S203, virtually magnify the target image frame according to the calculated magnification factor, and calculate the virtual cropping coordinates according to the predetermined cropping conditions; wherein, the predetermined cropping conditions are that the cropping meets the preset resolution, the preset aspect ratio, and the human image area in the cropped image area conforms to the golden ratio composition principle.
[0154] S204, the virtual cropping coordinates are mapped to the target image frame to crop the target image frame, and the resulting cropped image is magnified according to the magnification factor to obtain the image frame to be sent to the user terminal.
[0155] In this embodiment, the preset ratio threshold can be 50% or 30%, etc. If the area of the human image region in the target image frame is less than the preset ratio threshold, it indicates that the person is far away, and the target image frame is a distant view of the person. Therefore, to highlight the person in the target image frame, the magnification factor required to make the area of the human image region reach the preset ratio threshold can be calculated first. For example, if the area of the human image region in the target image frame is 25%, and the preset ratio threshold is 50%, then the magnification factor is 2.
[0156] It is understandable that virtually magnifying the target image frame means not storing the image after magnification, but only calculating the coordinates of each pixel after magnification. It is understandable that by virtually magnifying the target image frame, without performing actual scaling operations, image quality loss and storage resource waste can be reduced.
[0157] For example, the preset resolution could be 1920×1080 or 1280×720, etc. The preset aspect ratio could be a portrait aspect ratio of 9:16 or 3:4, or a landscape aspect ratio of 16:9 or 4:3, etc. The golden ratio composition principle is that the height of the cropped human image area and the height of the cropped image area conform to the golden ratio (i.e., 0.618), that is, the upper boundary of the cropped human image area is located at the height of the golden section line of the cropped image area.
[0158] Understandably, a unique cropping size can be determined based on the preset resolution and aspect ratio. Then, according to the golden ratio composition principle, the cropping center point can be determined. Based on this center point and the cropping size, the coordinates of each vertex of the area to be cropped can be determined, resulting in virtual cropping coordinates. Mapping these virtual cropping coordinates onto the original image coordinate system (i.e., the target image frame) allows for cropping of the target image frame. The resulting cropped image is then enlarged according to the determined magnification factor. This requires only one scaling operation, eliminating the need for multiple scaling and cropping operations, thus reducing resolution loss.
[0159] It is understandable that by adjusting the selected target image frames through this solution, the distant image of the person can be transformed into a close-up image of the person, thus making it suitable for scenarios that require highlighting the person. Furthermore, by adjusting the target image frames according to the preset aspect ratio of vertical format (9:16), it can be applied to scenarios such as vertical short video production.
[0160] Optionally, in one implementation, step S204, which involves enlarging the obtained cropped image by a magnification factor to obtain an image frame to be sent to the user terminal, may include steps D1-D3:
[0161] D1, enlarge the cropped image according to the magnification factor to obtain the intermediate image;
[0162] D2, perform portrait enhancement operation on the face image region of the intermediate image to obtain the enhanced image;
[0163] D3, performs a first number of high-definition restorations on the human image region in the obtained enhanced image, and performs a second number of high-definition restorations on the obtained enhanced image to obtain an image frame for sending to the user terminal; the first number is greater than the second number.
[0164] In this implementation, after enlarging the cropped image to obtain an intermediate image using the aforementioned magnification factor, portrait enhancement operations can be performed on the face image region of the intermediate image. These portrait enhancement operations include beautification operations such as whitening, acne removal, skin smoothing, fine line removal, and super-resolution restoration. For example, the intermediate image can be input into a pre-trained portrait enhancement model to perform portrait enhancement operations on the face image region of the intermediate image, obtaining the enhanced image output by the model. This portrait enhancement model can include: a face mask generation model for generating masks for the skin region of the face; a face super-resolution model for performing super-resolution high-definition restoration of the face; a face beautification model for performing operations such as whitening and acne removal; and a face smoothing tool for performing Gaussian kernel blur-based skin smoothing operations and removing fine lines.
[0165] Understandably, after performing portrait enhancement on the face image region, to avoid the face image region being too obviously separated from other regions in the obtained enhanced image (i.e., the face is too sharp while the background is too blurry), a first-order high-definition restoration can be performed on the human image region in the obtained enhanced image, and a second-order high-definition restoration can be performed on the obtained enhanced image. For example, the high-definition restoration can be performed using a super-resolution model (i.e., an image super-resolution model).
[0166] For example, the first quantity can be 3, and the second quantity can be 1. It is understandable that calling the super-resolution model for the human image region and performing 3 high-definition restorations, while performing 1 high-definition restoration for the entire enhanced image, can accelerate the restoration speed and avoid the computational pressure caused by the excessive time required for high-definition restoration of large images. Furthermore, the portrait enhancement operation and high-definition restoration involved in this implementation can also be used independently of the cropping process; that is, the above-mentioned portrait enhancement operation and high-definition restoration can be performed directly on the selected target image frame, which is also reasonable.
[0167] To better understand the image processing method provided in the embodiments of this application, a specific example will be used for illustration below.
[0168] The purpose of this example is to automatically select the best human images from a continuous sequence of motion videos or motion images to improve the efficiency of acquiring high-quality images and enhance the quality of the final product. This enables the device to stably output stunning images of the best moments with "good expressions, beautiful movements, and high clarity," which can be used in scenarios such as sports photography and AI sports scoring.
[0169] This example mainly consists of two parts: portrait image quality scoring and portrait image optimization.
[0170] like Figure 3 The diagram shows the process of scoring the quality of a person's image, including the following steps (1)-(5):
[0171] (1) The system (i.e. the device that performs the image processing method provided in the embodiments of this application) inputs the video to be detected into the human face detection model, so as to use the human face detection model to perform human face detection on the input video frame by frame, and obtain the human body region (corresponding to the human image region above) and face region (corresponding to the face image region above) of the person in each image frame.
[0172] (2) Human key points are identified in the detected human body area to obtain the coordinates and confidence of 17 human key points (corresponding to the first set of key points mentioned above).
[0173] (3) Input the face region into the face quality model for face quality detection.
[0174] If the height or width of the face region is too small (less than a set threshold, such as 30 pixels), the face information is considered insufficient to judge the face quality, and the frame image is directly rejected, i.e., the frame image is discarded.
[0175] Face quality detection includes face sharpness detection, eye detection, and expression detection. Specifically, the system inputs the detected face regions into the face quality model, which scores the face regions based on their sharpness and occlusion, resulting in a face score. If the score output by the face quality model is lower than a set threshold (e.g., 0.9), it is considered an image blur. In this case, the system can either output a lower face score (corresponding to the facial score mentioned above) or directly discard the frame.
[0176] If the score output by the face quality model is higher than a set threshold, the face region is determined to be a clear face. At this point, a facial expression recognition model can also be used to identify facial expressions within the face region. The expression recognition model can identify expressions such as "happy," "natural," "smiling," "disgusted," and "aversion." Specifically, the expression recognition model classifies expressions into two categories: "natural" and "smiling." If the scores for both categories are below a certain threshold (e.g., 0.1), it indicates a poor expression, such as closed eyes or a tight mouth, and the frame image is directly discarded. If the score for either expression is not lower than the threshold, the face region is considered a detected face. Furthermore, if a smiling expression is detected, it is given a higher weight (e.g., multiplied by 5). If it is a natural expression but has a high score, it is included in the expression score according to the natural expression score (without multiplying by any weight) (corresponding to the second score mentioned above).
[0177] If the facial expression recognition model identifies a natural expression, the eye opening / closing detection model (corresponding to the human eye detection model mentioned above) can be used to detect the degree of eye opening / closing. At this point, the outputs of the face quality model, the eye opening / closing detection model, and the facial expression recognition model can be combined to determine the final face score. During human eye detection, the face region in the image frame can be expanded outwards by a certain proportion to ensure that the detected image frame contains the complete eye area. If no smile is detected and the human eye detection model identifies closed eyes, the image frame is directly discarded. This avoids misidentifying "natural expression + closed eyes" as valid images. Therefore, images with "natural expression + closed eyes" will not be selected subsequently.
[0178] (4) Human posture comparison and posture scoring.
[0179] The system maintains two pose template libraries: Library 1 (corresponding to the standard keypoint set mentioned above) and Library 2 (corresponding to the secondary standard keypoint set mentioned above). These libraries correspond to graceful poses (corresponding to the specified human pose mentioned above) and ordinary poses (corresponding to human poses whose similarity to the specified human pose reaches a preset similarity threshold mentioned above), respectively. OKS similarity calculations are performed on the human keypoints in the current frame and on each human keypoint in both libraries to obtain two pose scores. In other words, the human keypoints in the current frame are compared with ordinary pose keypoints and graceful pose keypoints to obtain pose scores. These pose scores are then combined with the aforementioned face score to obtain the final photo score.
[0180] Library 1 and Library 2 are constructed as follows: Beautiful video clips in motion are manually selected as sample videos; a human face detection model is used to detect and crop the human body regions of the target person in the video; annotators can manually select the human poses based on the cropped images to obtain acceptable human poses (including beautiful and ordinary poses). These acceptable human poses are then input into a human keypoint recognition model to obtain the model's output human keypoints. Keypoints belonging to beautiful poses are stored as beautiful pose keypoints (i.e., Library 1, corresponding to the standard keypoint set mentioned above), and keypoints belonging to ordinary poses are stored as ordinary pose keypoints (i.e., Library 2, corresponding to the secondary standard keypoint set mentioned above). Library 1 and Library 2 can be dynamically updated as more data is used.
[0181] Pose alignment was performed using OKS similarity. The alignment algorithm specifically includes:
[0182] (a) Extraction and validity assessment of key data.
[0183] The input sequence of human keypoints is parsed into a keypoint matrix based on three values (x-coordinate, y-coordinate, and confidence score). The number of keypoints with a confidence score below a set threshold (e.g., 0.3) is counted. When the number of keypoints with low confidence scores exceeds a certain number (e.g., 10), the pose quality is considered poor, and a lower pose score is directly output to avoid misjudging low-quality poses. Alternatively, the video frame can be directly removed.
[0184] (b) Key point selection and weight design.
[0185] Considering that posture assessment in motion scenarios primarily focuses on limb movements, key points related to the human limbs are selected as the core key points for posture comparison. For example, the core key points include: shoulder (left, right), elbow (left, right), wrist (left, right), hip (left, right), knee (left, right), and ankle (left, right). Furthermore, this example assigns different weights to different key points, using different σ parameters to adjust the matching strictness. This results in greater weight being given to areas with large ranges of motion, such as the elbow, wrist, knee, and ankle, while less weight is given to stable points like the shoulder. In the OKS formula, a larger weight means a smaller σ value, and vice versa.
[0186] (c) For the current pose, compare it with each template in the template library as follows:
[0187] (i) Valid matching key point screening: Only limb key points that are valid coordinates (not occluded or missing) in both the current pose and the template pose are compared to ensure that the scoring is based on the same set of key points.
[0188] (ii) OKS Similarity Calculation: For each valid keypoint, calculate the Euclidean distance between the current keypoint and the template keypoint, and calculate the similarity according to the OKS formula. The OKS formula is as follows:
[0189] ;
[0190] in, σ is the square of the Euclidean distance between the current pose and the template pose corresponding to the key points (corresponding to the limb key points mentioned above). The smaller the Euclidean distance, the more similar the current pose is to the template pose. σ is the weight corresponding to the key point (the smaller the value, the stricter the matching, and the greater the weight of the key point).
[0191] Ultimately, the pose similarity corresponding to the template is the average of the OKS values of all valid keypoints. The larger the OKS, the more similar the current pose is to the template pose.
[0192] The system sorts the similarity of the current pose with the similarity of all templates in the template library, selects several pose samples with the highest similarity (e.g., the top three), and averages their OKS scores as the final pose score.
[0193] When the number of templates is less than a certain amount (e.g., less than three), the average score of all templates is taken directly. When the number of templates in the template library is large, the average of the top three highest-scoring templates is taken to improve scoring stability and reduce bias caused by accidental matching. The final output posture score ranges from 0 to 1, with a higher score indicating that the current posture is closer to a graceful posture.
[0194] (5) Perform multi-dimensional scoring weighted fusion on the human body image to obtain the final score.
[0195] The system first uses a full-body keypoint detector to acquire key points of the human body. If there are too many missing points in the detection results (for example, more than 10 key points are missing from the human body), the system determines that the image has no evaluation value and directly discards the image.
[0196] If the height or width of the face region is too small (less than a set threshold, such as 30 pixels), the face information is considered insufficient to judge the face quality, and the image is directly discarded.
[0197] If the current pose matches Library 2 better, the similarity score from Library 2 is used and penalized with a coefficient of 0.8 (corresponding to the second weight mentioned above). If the current pose matches Library 1 better, the similarity score from Library 1 is used and boosted with a coefficient of 1.2 (corresponding to the first weight mentioned above). If the final pose score is below a certain threshold (e.g., 0.2), it indicates a poor pose, and the image is directly removed. Finally, the system combines the expression score and pose score with equal weights. This allows for image selection by simultaneously considering both facial expression and body posture.
[0198] like Figure 4 The diagram shows the process of optimizing a person's image. Image optimization includes adaptive image cropping and scaling, which is used to center-crop images containing people. It is suitable for vertical (9:16 aspect ratio) video production and scenarios where people are highlighted. It can adjust the main subject of the person to the lower center of the screen without reducing image quality and automatically maintain the preset composition ratio. The optimization process includes steps (1)-(6):
[0199] (1) Localization of face region and body region.
[0200] Obtain the person detection bounding box, face detection bounding box, and facial key points from the input image, and map the face region to the global coordinate system of the original image for subsequent composition and quality optimization.
[0201] (2) Detect the proportion of the figures and perform virtual magnification.
[0202] Adaptive cropping and scaling of the input image. That is, it determines whether to enlarge the image based on the proportion of the person's area in the original image. When the person's area is less than a preset threshold (e.g., 1 / 5) of the image area, the image and the person's area are enlarged proportionally to simulate a closer portrait composition. This avoids actual scaling of the original image and thus prevents quality loss.
[0203] The virtual magnification ratio is calculated by the ratio of the target area to the current area and is limited to the maximum magnification range.
[0204] (3) Construct a cutting area that conforms to the golden ratio.
[0205] Based on the preset target resolution (e.g., 1080×1920) and target aspect ratio (9:16), the ideal width and height of the cropping area (corresponding to the cropping dimensions mentioned above) can be calculated. Combined with the golden ratio composition principle, the height ratio that the person should occupy in the target image (i.e., the cropped image) can be determined, thereby determining the center point of the cropping area.
[0206] (4) Optimize the location of the cropping area.
[0207] After determining the center point of the cropping area, this center point can be shifted downwards by a preset percentage for cropping. This ensures that the subject is located in the lower-middle area of the final cropped image, while avoiding excessive white space at the top, thus better conforming to the compositional habits of portrait photography. For example, the center point can be shifted downwards by 10% or 15% of the subject's height in the image, and so on.
[0208] The cropping area must also meet the following constraints: it must completely contain the character area; it must not exceed the boundaries of the virtual image; if the target aspect ratio is not met, the aspect ratio needs to be readjusted to maintain 9:16.
[0209] (5) Map back to the original image and perform a one-time scaling.
[0210] The determined cropping area is mapped back to the original image coordinate system, and the original image is cropped to obtain the original cropped image. Then, it is scaled up to the target size in one go. This one-time scaling replaces multiple scaling operations, reducing resolution loss. If the scaled image boundary exceeds the target size, a second center-cropping is performed to obtain the final image.
[0211] (6) Synchronous mapping of key points to face / body bounding boxes.
[0212] The person detection bounding box, face detection bounding box, and facial key points are synchronously mapped to the coordinate system of the final output image to ensure that post-processing (such as beautification and makeup) can be correctly positioned.
[0213] In addition, image optimization includes portrait enhancement processing, which further improves the clarity, visual quality, and facial detail of the human body area after cropping. This method can be used independently of the aforementioned cropping process or as a subsequent enhancement process.
[0214] Character enhancement processing includes the following steps (a)-(d):
[0215] (a) Region expansion based on person detection bounding box.
[0216] The expansion boundary is calculated based on the width and height of the figure area, forming an enhanced area that is slightly larger than the original figure frame, while ensuring that it does not exceed the original image boundary.
[0217] (b) Call the image super-resolution model.
[0218] For the human body portion, an image super-resolution model is used to perform two high-resolution restorations on the originally blurred person detection boxes. A single overall high-resolution restoration is then performed on the complete original image. This approach aims to accelerate the restoration process and avoid the computational burden caused by excessively long high-resolution restoration times for large images.
[0219] (c) Call the human image enhancement model.
[0220] For the face detection bounding box, after face parsing, a face restoration system performs high-resolution face restoration. The face restoration system includes: a face mask generation model to generate masks for the skin regions of the face; a super-resolution model specifically for faces to perform super-resolution high-resolution face restoration; a face beautification model to perform whitening, blemish removal, and other operations on the face; and a face smoothing tool to perform Gaussian kernel blur-based skin smoothing operations to remove fine lines, etc.
[0221] (d) Enhance regional posting and integration.
[0222] The enhanced image regions are mapped back to their corresponding positions in the original image, and region replacement is performed using interpolation or seamless blending to obtain a complete image output containing the enhanced person, resulting in the final image. The repaired face is also blended back to its original position at a certain ratio.
[0223] As can be seen, this solution constructs a comprehensive image quality evaluation system for motion capture scenarios, introducing multi-dimensional indicators such as pose scoring, face quality, expression quality, eye opening / closing status, and key point stability, enabling comprehensive quality assessment of moving images. Through OKS similarity, template library matching, and weighting mechanisms, the system can evaluate pose quality across actions, scenes, and angles, significantly improving ranking stability. Rejection strategies effectively filter low-quality samples, such as blurred faces, closed eyes, missing key points, and undetermined poses, automatically improving the quality of the final Top-N output. Through virtual magnification, golden ratio composition, face restoration, and body enhancement technologies, the final output images are more visually appealing and suitable for direct use in short videos, sports apps, and cover image generation. The pose template library, face quality model, and expression recognition model can all continuously learn and improve based on new data, ensuring continuous improvement over long-term use.
[0224] Corresponding to the above method embodiments, this application also provides an image processing apparatus, such as... Figure 5 As shown, it includes:
[0225] The recognition module 510 is used to identify human key points in the human image region of each image frame in the image sequence to be detected, and obtain a first set of key points.
[0226] The pose score determination module 520 is used to calculate the matching degree between the first key point set and the pre-stored second key point set to obtain the pose score of the image frame to be detected; wherein, the second key point set includes a standard key point set, and the standard key point set includes: human key points of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree;
[0227] The face score determination module 530 is used to determine the face score of the image frame to be detected according to the index value of the face image region in the image frame to be detected against the preset index; wherein, the preset index includes: image fidelity, and / or, the degree of matching between facial expression and the specified human posture; the face score of the image frame to be detected is positively correlated with the index value of the preset index;
[0228] The final score generation module 540 is used to combine the face score and pose score of the image frame to be detected to generate the final score of the image frame to be detected.
[0229] The selection module 550 is used to select a target image frame from each image frame to be detected; wherein the final score of the target image frame is higher than the final scores of other image frames to be detected.
[0230] Optionally, the attitude score determination module 520 includes:
[0231] The calculation submodule is used to calculate the matching degree between the first key point set and the second key point set based on the similarity between each limb key point in the first key point set and the corresponding limb key point in the pre-stored second key point set.
[0232] The determination submodule is used to obtain the pose score of the image frame to be detected based on the calculated matching degree.
[0233] Optionally, the similarity between any limb keypoint in the first keypoint set and the corresponding limb keypoint in the pre-stored second keypoint set is calculated using the following formula:
[0234] ;
[0235] Wherein, OKS represents similarity, d represents the Euclidean distance between the first coordinate of the limb keypoint in the first keypoint set in the image frame to be detected and the second coordinate of the corresponding limb keypoint in the second keypoint set in the standard image; σ represents the normalization factor corresponding to the limb keypoint; and the σ value of the preset active keypoint among the limb keypoints is less than the σ value of other limb keypoints.
[0236] Optionally, the second set of key points may be multiple; the attitude score is determined based on the highest preset number of matching degrees.
[0237] Optionally, each second key point set further includes a sub-standard key point set, and any sub-standard key point set includes: human body key points of the sub-standard image; the similarity between the human body key points of any sub-standard image and the human body key points of the standard image is greater than a preset similarity threshold.
[0238] The pose score is: the weighted average of the highest preset number of matching degrees; the weight of the matching degree corresponding to the standard image is greater than the weight of the matching degree corresponding to the next standard image.
[0239] Optionally, the device further includes:
[0240] The first detection module is used to detect the area ratio of the human image region in each selected target image frame.
[0241] The magnification factor determination module is used to calculate the magnification factor required for the human image region in the target image frame to reach the preset ratio threshold if the detected area ratio is less than the preset ratio threshold.
[0242] The virtual magnification module is used to virtually magnify the target image frame according to the calculated magnification factor, and calculate the virtual cropping coordinates according to the predetermined cropping conditions; wherein, the predetermined cropping conditions are that the cropping meets the preset resolution, the preset aspect ratio, and the human image area in the cropped image area conforms to the golden ratio composition principle.
[0243] The cropping module is used to map virtual cropping coordinates onto the target image frame to crop the target image frame, and then enlarge the cropped image according to the magnification factor to obtain an image frame for sending to the user terminal.
[0244] Optionally, the cropping module includes:
[0245] The magnification submodule is used to magnify the obtained cropped image according to the magnification factor to obtain an intermediate image;
[0246] The enhancement submodule is used to perform portrait enhancement operations on the face image region of the intermediate image to obtain an enhanced image;
[0247] The repair submodule is used to perform a first number of high-definition repairs on the human image region in the obtained enhanced image, and to perform a second number of high-definition repairs on the obtained enhanced image, to obtain an image frame for sending to the user terminal; the first number is greater than the second number.
[0248] Optionally, the preset index further includes image fidelity, and the face score determination module 530 includes:
[0249] The first score determination submodule is used to input the face image region of the image frame to be detected into the face quality model to obtain the first score; wherein, the face quality model is used to score the input image based on the image fidelity of the input image;
[0250] The second score determination submodule is used to input the face image region of the image frame to be detected into the expression recognition model to obtain a second score; wherein, the expression recognition model is used to score the input image based on the degree of matching between the facial expression of the input image and the specified human posture;
[0251] The face score generation submodule is used to combine the first score and the second score to generate the face score of the image frame to be detected.
[0252] This application also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0253] Memory 603 is used to store computer programs;
[0254] The processor 601, when executing the program stored in the memory 603, implements the steps of the above-described image processing method.
[0255] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0256] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0257] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0258] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0259] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described image processing methods.
[0260] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the image processing methods described in the above embodiments.
[0261] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0262] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0263] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, and computer-readable storage media are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0264] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. An image processing method, characterized in that, The method includes: For each image frame to be detected in the image sequence, human key points are identified in the human image region of the image frame to be detected to obtain the first set of key points. The matching degree between the first set of key points and the pre-stored second set of key points is calculated to obtain the pose score of the image frame to be detected; wherein, the second set of key points includes a standard set of key points, and the standard set of key points includes: human key points of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree; The facial score of the image frame to be detected is determined according to the index value of the facial image region in the image frame to be detected relative to the preset index; wherein, the preset index includes: the degree of matching between facial expression and the specified human posture; the facial score of the image frame to be detected is positively correlated with the index value of the preset index; The final score of the image frame to be detected is generated by combining the face score and pose score of the image frame to be detected. A target image frame is selected from each image frame to be detected; wherein the final score of the target image frame is higher than the final score of the other image frames to be detected.
2. The method according to claim 1, characterized in that, Calculate the matching degree between the first set of key points and the pre-stored second set of key points to obtain the pose score of the image frame to be detected, including: Based on the similarity between each limb key point in the first key point set and the corresponding limb key points in the pre-stored second key point set, the matching degree between the first key point set and the second key point set is calculated. Based on the calculated matching degree, the pose score of the image frame to be detected is obtained.
3. The method according to claim 2, characterized in that, The similarity between any limb keypoint in the first keypoint set and the corresponding limb keypoint in the pre-stored second keypoint set is calculated using the following formula: ; Wherein, OKS represents similarity, d represents the Euclidean distance between the first coordinate of the limb keypoint in the first keypoint set in the image frame to be detected and the second coordinate of the corresponding limb keypoint in the second keypoint set in the standard image; σ represents the normalization factor corresponding to the limb keypoint; and the σ value of the preset active keypoint among the limb keypoints is less than the σ value of other limb keypoints.
4. The method according to any one of claims 1-3, characterized in that, The second set of key points consists of multiple points; the attitude score is determined based on the highest preset number of matching degrees.
5. The method according to claim 4, characterized in that, Each set of second key points also includes a set of substandard key points. Each set of substandard key points includes: human body key points of the substandard image; the similarity between the human body key points of the substandard image and the human body key points of the standard image is greater than a preset similarity threshold. The pose score is: the weighted average of the highest preset number of matching degrees; the weight of the matching degree corresponding to the standard image is greater than the weight of the matching degree corresponding to the next standard image.
6. The method according to claim 1, characterized in that, The method further includes: For each selected target image frame, the area ratio of the human image region in that target image frame is detected; If the detected area ratio is less than a preset ratio threshold, then calculate the magnification factor required for the human image region in the target image frame to reach the preset ratio threshold. The target image frame is virtually magnified according to the calculated magnification factor, and the virtual cropping coordinates are calculated according to the predetermined cropping conditions; wherein, the predetermined cropping conditions are that the cropping meets the preset resolution, preset aspect ratio, and the human image area in the cropped image area conforms to the golden ratio composition principle. The virtual cropping coordinates are mapped onto the target image frame to crop the target image frame, and the resulting cropped image is enlarged according to the magnification factor to obtain an image frame for sending to the user terminal.
7. The method according to claim 6, characterized in that, The obtained cropped image is magnified according to the magnification factor to obtain an image frame for transmission to the user terminal, including: The cropped image is then magnified according to the stated magnification factor to obtain an intermediate image; A portrait enhancement operation is performed on the face image region of the intermediate image to obtain the enhanced image; The human image region in the obtained enhanced image is subjected to a first number of high-definition restorations, and the obtained enhanced image is subjected to a second number of high-definition restorations to obtain an image frame for sending to the user terminal; the first number is greater than the second number.
8. The method according to claim 1, characterized in that, The preset index also includes image fidelity. Determining the facial score of the image frame to be detected based on the index value of the facial image region in the image frame relative to the preset index includes: The face image region of the image frame to be detected is input into the face quality model to obtain the first score; wherein, the face quality model is used to score the input image based on the image fidelity of the input image; The facial image region of the image frame to be detected is input into the expression recognition model to obtain a second score; wherein, the expression recognition model is used to score the input image based on the degree of matching between the facial expression of the input image and the specified human posture; The facial score of the image frame to be detected is generated by combining the first score and the second score.
9. An image processing apparatus, characterized in that, The device includes: The recognition module is used to identify human key points in the human image region of each image frame in the image sequence to be detected, and obtain the first set of key points. The pose score determination module is used to calculate the matching degree between the first keypoint set and the pre-stored second keypoint set to obtain the pose score of the image frame to be detected; wherein, the second keypoint set includes a standard keypoint set, which includes human keypoints of a standard image containing a specified human pose; the pose score of the image frame to be detected is positively correlated with the matching degree; The face score determination module is used to determine the face score of the image frame to be detected according to the index value of the face image region in the image frame to be detected against the preset index; wherein, the preset index includes: image fidelity, and / or, the degree of matching between facial expression and the specified human posture; the face score of the image frame to be detected is positively correlated with the index value of the preset index; The final score generation module is used to combine the face score and pose score of the image frame to be detected to generate the final score of the image frame to be detected. The selection module is used to select a target image frame from each image frame to be detected; wherein the final score of the target image frame is higher than the final scores of the other image frames to be detected.
10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-8.