Video material image selection method and device, equipment and storage medium
By performing face detection and expression analysis on video frames, and selecting video frames with high-quality expressions as poster materials, the problem of low efficiency and poor quality in poster material selection in existing technologies is solved, and higher-quality poster generation is achieved.
Patent Information
- Application Number
- CN202110831155.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-22
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-09-16
AI Technical Summary
In existing technologies, poster material selection is based on the image parameters of video frames, which fails to accurately identify anomalies in the image content, resulting in low poster generation efficiency and poor quality.
By performing face detection on video frames, identifying facial regions, and analyzing facial expressions, video frames with high-quality expressions are selected as poster materials.
It improves the content quality and generation efficiency of poster materials, ensuring that the selected video frames can more accurately reflect the video content.
Smart Images

Figure CN113822136B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia, and in particular to a method, apparatus, device, and storage medium for selecting video material images. Background Technology
[0002] Intelligent poster material extraction refers to the process of using computer technology to extract video frames from a video stream, and then using computer programs to analyze them to select suitable video frames as poster materials.
[0003] In related technologies, the process of selecting poster materials includes analyzing and selecting video frames. First, video frames are obtained from the video stream. Then, the video frames are analyzed in terms of dimensions such as sharpness and color quality. The quality score of each video frame is obtained by combining the results of multi-dimensional analysis. The video frame with the highest quality score is selected as the poster material of the video stream.
[0004] However, in the above method, the selection of poster materials is based on the image parameters of the video frame itself (such as: clarity, contrast, brightness, etc.), and abnormalities in the image content cannot be accurately identified, resulting in poor quality of the determined video frame content and low poster generation efficiency. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for selecting video material images, which can improve the content quality of poster materials during the selection process of video poster materials. The technical solution is as follows:
[0006] On the one hand, a method for selecting video footage images is provided, the method comprising:
[0007] Acquire a target video stream, wherein the target video stream includes video frames;
[0008] Face detection is performed on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer;
[0009] Facial expression analysis is performed on the facial regions in the candidate video frames to obtain facial expression analysis results for the facial regions. These results are used to indicate the quality of facial expressions in the facial regions.
[0010] Based on the expression analysis results, a target video frame is determined from the n candidate video frames and used as the video material image of the target video stream. The video material image is used as a representative image of the target video stream.
[0011] On the other hand, a device for selecting video source images is provided, the device comprising:
[0012] The acquisition module is used to acquire a target video stream, wherein the target video stream includes video frames;
[0013] The first detection module is used to perform face detection on the video frame to obtain n candidate video frames containing face regions, where n≥2 and n is an integer;
[0014] The first analysis module is used to perform facial expression analysis on the face regions in the candidate video frames to obtain facial expression analysis results for the face regions. The facial expression analysis results are used to indicate the quality of facial expressions in the face regions.
[0015] The first determining module is used to determine a target video frame from the n candidate video frames based on the expression analysis results, as the video material image of the target video stream, and the video material image is used as a representative image of the target video stream.
[0016] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the video material image selection method as described in any of the above embodiments of this application.
[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the video material image selection method as described in any of the embodiments of this application above.
[0018] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video material image selection method described in any of the above embodiments.
[0019] The beneficial effects of the technical solutions provided in this application include at least the following:
[0020] By performing facial expression analysis on the facial regions in candidate video frames, the expression analysis results are obtained. Based on these results, the target video frame is determined from the candidate video frames and used as a representative image of the target video stream. This image is then used to generate the cover or poster image of the target video stream, improving the accuracy of video material image determination and the quality of image content in the video material images, thus increasing the efficiency of video material image generation. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;
[0023] Figure 2 This is a schematic diagram illustrating the process of selecting video material images provided in an exemplary embodiment of this application;
[0024] Figure 3 This is a flowchart of a method for selecting video material images provided in an exemplary embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the face eye analysis process provided in an exemplary embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the face and eye analysis process provided in another exemplary embodiment of this application;
[0027] Figure 6 This is a schematic diagram of the analysis process of a human face mouth provided in an exemplary embodiment of this application;
[0028] Figure 7 This is a schematic diagram of the analysis process of a human face and mouth provided in another exemplary embodiment of this application;
[0029] Figure 8 This is a structural diagram of an abnormal facial expression recognition model provided in an exemplary embodiment of this application;
[0030] Figure 9 This is a flowchart of a method for selecting video material images provided in another exemplary embodiment of this application;
[0031] Figure 10 This is a schematic diagram of human head posture analysis provided in an exemplary embodiment of this application;
[0032] Figure 11 This is a flowchart of a method for selecting video material images provided in another exemplary embodiment of this application;
[0033] Figure 12 This is a schematic diagram of the structure of a video material image generation framework provided in an exemplary embodiment of this application;
[0034] Figure 13 This is a structural block diagram of a video material image selection device provided in an exemplary embodiment of this application;
[0035] Figure 14 This is a structural block diagram of a video material image selection device provided in another exemplary embodiment of this application;
[0036] Figure 15 This is a structural block diagram of a video material image selection device provided in another exemplary embodiment of this application;
[0037] Figure 16 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0039] Currently, with the continuous development of Internet technology, the promotion of video content no longer relies on manual production of promotional images. Instead, computer technology is used to intelligently select video material images in batches as promotional images. However, the video material images selected by this technology cannot meet the promotional needs. The aesthetics of the images and the quality of the promotional content are relatively low, such as the tendency to produce distorted or exaggerated facial expressions.
[0040] This application provides a method for selecting video footage images. During implementation, this method can accurately identify the face region in the video frame and perform facial expression quality analysis to ensure the aesthetics and content quality of the video footage images.
[0041] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application, such as... Figure 1 As shown, the implementation environment includes a terminal 110 and a server 120, wherein the terminal 110 and the server 120 are connected through a communication network 130.
[0042] The terminal 110 may contain an application that provides video recommendation functionality. This application displays video clip images; that is, when a user runs the application on the terminal 110 and selects a video from the candidate videos provided in the application, they can see the corresponding video clip images for each candidate video, such as video cover images or video poster images. These video clip images are selected by the server 120 from the video frames of the candidate videos.
[0043] Server 120 is used to determine the target video frame from the middle of the video stream based on the face region corresponding to the video frame in the video stream of the candidate video. The target video frame is the video frame that will subsequently be used as video material. After determining the face region, server 120 determines the target video frame based on the facial expression quality obtained from the face region analysis.
[0044] For illustrative purposes, when terminal 110 needs to display the video cover image of a specified video, it sends a display request to server 120. The display request includes the video identifier of the specified video. Server 120 obtains the video cover image of the specified video based on the video identifier. The video cover image is selected based on the quality of the facial expression and is then fed back to terminal 110 for display.
[0045] The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart in-vehicle device, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0046] It is worth noting that the aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0047] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Based on the cloud computing business model, cloud technology encompasses network technology, information technology, integration technology, management platform technology, and application technology. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.
[0048] In some embodiments, the server described above can also be implemented as a node in a blockchain system. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and cryptographic algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0049] first Figure 2 This is a schematic diagram illustrating the process of selecting video material images provided in an exemplary embodiment of this application, such as... Figure 2 As shown, in this process, n candidate video frames 210 are first determined from the target video stream 200; facial expression analysis is performed on the facial regions in these n candidate video frames, and based on the obtained expression analysis results, video frames 220 with higher expression quality are determined from the candidate video frames 210. These video frames 220 are then used as video source images to obtain the cover image or poster image of the target video stream 200; for example... Figure 2 As shown, based on different video delivery requirements, a template for the image format corresponding to the video delivery requirements is determined, thereby obtaining the video cover image 221 or the poster image 222.
[0050] In light of the above implementation environment, the method for selecting video material images provided in this application embodiment will be described. This method can be executed by a terminal or a server, or it can be executed by both a terminal and a server. In this application embodiment, the method is executed by... Figure 1 The following explanation uses server 120 as an example. Figure 3 As shown, the method includes:
[0051] Step 310: Obtain the target video stream.
[0052] The target video stream includes video frames.
[0053] Optionally, the server obtains the target video stream in at least one of the following ways:
[0054] First, the server receives the target video stream uploaded by the terminal;
[0055] Second, the server receives a first video request sent by the terminal. The first video request includes a video identifier of the target video stream. The server retrieves the target video stream from the stored video library based on the video identifier.
[0056] Third, the server receives a second video request uploaded by the terminal. The second video request includes a video identifier and video segmentation conditions. The server retrieves the video stream from the stored video library based on the video identifier and extracts the target video stream from the video stream based on the video segmentation conditions.
[0057] In some embodiments, the target video stream is a video from which video image materials need to be selected. The target video stream is a complete film or television program, such as a TV series, movie, variety show, documentary, etc.; or, the target video stream is a video segment obtained by cutting and segmenting a film or television program.
[0058] In some embodiments, video frames from the target video stream are first acquired, and subsequent processing is performed based on these video frames. Illustratively, the video format of the target video stream is first corrected for compatibility, and then frame-level parallel multi-segment decoding is performed on the corrected target video stream to acquire video frames from the target video stream.
[0059] The acquired video frames are each video frame in the target video stream; or, the acquired video frames are a subset of video frames in the target video stream.
[0060] In some embodiments, a subset of video frames in the target video stream is selected through clustering. Illustratively, all video frames in the target video stream are clustered to obtain clustering results. These results include the clusters to which all video frames belong. A predetermined number of video frames are randomly selected from each cluster for subsequent processing. Alternatively, video frames meeting specified conditions (e.g., sharpness conditions, contrast conditions, etc.) are selected from each cluster for subsequent processing. In some embodiments, video frames are clustered by calculating their hash values in the target video stream.
[0061] In some embodiments, the target video stream is a film or television program featuring live actors; or, the target video stream is a film or television program featuring animated characters.
[0062] Specifically, when the target video stream contains live actors, there are video frames that include images of real people's faces; or, when the target video stream contains anime characters, there are video frames that include images of recognizable anime faces.
[0063] It is worth noting that the above-mentioned form of the target video stream is merely an illustrative example, and the specific form of the target video stream is not limited in the embodiments of this application.
[0064] Step 320: Perform face detection on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer.
[0065] In some embodiments, face detection is performed on video frames.
[0066] Optionally, face detection is performed on video frames using a pre-trained face detection model. Optionally, the face detection model is used to determine face regions by recognizing face detection points.
[0067] The process involves inputting video frames from the target video stream into a face detection model, which then outputs the face detection result for each video frame. The face detection result indicates the face region information contained in the current video frame. Based on the face detection result, n candidate video frames are determined from the video frames; these n candidate video frames are video frames containing face regions.
[0068] The face detection results also include region parameters of the face region, such as the region location and size. Optionally, face detection is performed on video frames to obtain video frames containing face regions. Based on the region parameters of the face regions in the video frames, the video frames are filtered to obtain n candidate video frames. The region parameters used for filtering the video frames include at least one of region size and region location.
[0069] This is an example of filtering out video frames where the face area is too large, too small, or located at the edge of the video frame.
[0070] Step 330: Perform facial expression analysis on the face regions in the candidate video frames to obtain the facial expression analysis results.
[0071] Optionally, the facial expression analysis results can be used to indicate the quality of facial expressions in the face region.
[0072] In some embodiments, during the face detection process, there may be a situation where multiple faces are included in the candidate video frames. When a video frame includes multiple faces, at least one of the following methods is used to determine the face that needs to be analyzed for expression from the multiple faces.
[0073] First, perform abnormal expression analysis on multiple faces in the candidate video frames.
[0074] Second, the face with the largest face area in the candidate video frame is selected as the face for expression analysis; that is, during the face detection process, the face region corresponding to each face is identified, and the face with the largest face region area is selected as the face for expression analysis based on the area of the face region.
[0075] Third, faces in candidate video frames that match face samples in a preset face database are selected as the faces for expression analysis. For example, the preset face database includes acquired and stored celebrity faces; when a candidate video frame includes a celebrity face, that celebrity face is selected as the face for expression analysis. In some embodiments, the celebrity faces in the preset face database are sorted in a preset order. When at least two celebrity faces are included, the celebrity face with the higher ranking in the preset face database is selected as the face for expression analysis.
[0076] Fourth, the faces in the candidate video frames that match the face samples in the preset character library are taken as the faces for which expression analysis is required; wherein, the preset character library is a face library set for the characters corresponding to the current target video stream. Optionally, the preset character library includes the faces corresponding to the main character in the current target video stream, so as to perform expression analysis on the main character faces in the candidate video frames.
[0077] Fifth, the sharpness of multiple faces in the candidate video frames is detected, and the face with the highest sharpness is selected as the face for expression analysis.
[0078] It is worth noting that the above-described method for determining the face for expression analysis is merely an illustrative example, and the embodiments of this application do not limit it.
[0079] In some embodiments, an abnormal expression recognition model is used to analyze facial expressions in candidate video frames. Specifically, the model inputs the facial expressions from candidate video frames into the abnormal expression recognition model and outputs the expression analysis results. This abnormal expression recognition model is a classification and regression model used to categorize facial expressions into different states, such as closed eyes, pouting, and half-open eyes.
[0080] In some embodiments, the facial region to be analyzed is segmented according to the distribution of facial features, and the expression analysis results are comprehensively evaluated by combining the states of each part of the facial features.
[0081] That is, the face region is segmented into sub-regions based on the distribution of facial features to obtain the face sub-regions corresponding to the facial features. Expression analysis is then performed on the face sub-regions corresponding to the facial features to obtain the expression analysis results of the face region.
[0082] In some embodiments, after facial key points are detected in a video frame using a face detection model, the positions of key points corresponding to each facial feature are determined based on the facial key points, thereby performing sub-region segmentation. For example, after facial key points are detected using a face detection model, the position of the eye key point is determined, and the eye sub-region is segmented from the face region; or, the position of the mouth key point is determined, and the mouth sub-region is segmented from the face region.
[0083] In this embodiment, as an illustration, the face sub-region includes a first sub-region corresponding to the eyes. Expression analysis is then performed on the first sub-region to obtain an eye state analysis result corresponding to the eyes. This eye state analysis result is used to indicate the degree of eye opening / closing in the face region. Optionally, the eye state analysis result can also be used to indicate the occlusion status of the eyes. For illustration purposes, please refer to... Figure 4 It illustrates a schematic diagram of the facial eye analysis process provided in an exemplary embodiment of this application, such as... Figure 4 As shown, after performing expression analysis on the face region 410, the analysis result of the eye state of the face region 410 is "closed eyes"; as Figure 5 As shown, after performing expression analysis on the face region 510, the analysis result of the eye state of the face region 510 is "eyes half open".
[0084] In some embodiments, the above-mentioned abnormal expression recognition model includes an eye state analysis model. After inputting the first sub-region into the eye state analysis model, the eye state analysis result can be output. The eye state analysis model is a model used for classification and regression of candidate eye states.
[0085] Alternatively, illustratively, in this embodiment of the application, the face sub-region includes a second sub-region corresponding to the mouth of the face. Expression analysis is then performed on the second sub-region to obtain a mouth state analysis result corresponding to the mouth of the face. This mouth state analysis result is used to indicate the appearance of the mouth in the face region. Optionally, the mouth state analysis result is also used to indicate the occlusion status of the mouth. For illustrative purposes, please refer to... Figure 6 It illustrates a schematic diagram of the analysis process of a human face mouth provided in an exemplary embodiment of this application, such as... Figure 6 As shown, after performing expression analysis on the face region 610, the analysis result of the mouth state of the face region 610 is "smiling"; as Figure 7 As shown, after performing expression analysis on the face region 710, the analysis result of the mouth state of the face region 710 is "pouting state".
[0086] In some embodiments, the above-mentioned abnormal expression recognition model includes a mouth state analysis model. After inputting the second sub-region into the mouth state analysis model, the mouth state analysis result can be output. The mouth state analysis model is a model used for classification and regression of candidate mouth states.
[0087] In some embodiments, the abnormal facial expression recognition model employs a coarse-grained classification + fine-grained facial feature classification method, and obtains a prediction of normal / abnormal through comprehensive decision-making. The model structure is as follows: Figure 8As shown, after inputting the facial image 810 into the feature extraction network 820 for feature extraction, global features 831, eye features 832, and mouth features 833 are obtained from the extracted features. Global expression analysis is performed using global features 831 to obtain global analysis results, which are determined from candidate states (normal, abnormal). Eye state analysis is performed using eye features 832 to obtain eye state analysis results, which are determined from candidate eye states (half-open, closed, open, squinting, uncertain, occluded). Mouth state analysis is performed using mouth features 833 to obtain mouth state analysis results, which are determined from candidate mouth states (pouting, baring teeth, showing teeth, wide open, slightly open, closed, occluded).
[0088] Optionally, during the training of the abnormal expression recognition model, approximately 40,000 face screenshots labeled with their states were collected for training, with a normal / abnormal image ratio of approximately 9:1.
[0089] Based on the above facial expression analysis results, the rules for judging facial expression states include the following:
[0090] 1. If any of the following rules are met: abnormal eyes (eyes closed or half-open), abnormal mouth (mouth baring teeth), abnormal overall expression, or abnormal combination (eyes squinting when the mouth is not wide open or smiling), then the expression is determined to be extremely poor.
[0091] 2. When the following rules are met: the eyes may be abnormal (eyes are covered or uncertain), the mouth may be abnormal (mouth is pouting, wide open or covered), or the combination is abnormal (eyes are not open and mouth is slightly open), then the facial expression is determined to be poor.
[0092] 3. When the following rule is met: the mouth is slightly open, the facial expression is considered acceptable;
[0093] 4. If any of the rules 1, 2, or 3 above are not met, then the expression state is determined to be normal.
[0094] Step 340: Based on the expression analysis results, determine the target video frame from the n candidate video frames and use it as the video material image of the target video stream.
[0095] Among them, video footage images are used as representative images of the target video stream, such as the cover image of the target video stream or the poster image of the target video stream.
[0096] In some embodiments, determining the target video frame based on facial expression analysis results includes at least one of the following methods:
[0097] 1. After removing candidate video frames whose expression analysis results do not meet the expression quality requirements, the target video frame is determined from the remaining candidate video frames based on the image parameters;
[0098] Indicatively, the first candidate video frames with normal or acceptable facial expressions are retained, while the second candidate video frames with extremely poor or poor facial expressions are removed. The target video frame is determined from the first candidate video frames based on parameters such as clarity, aesthetics, subtitle display, and pop-up ad display.
[0099] 2. The facial expression analysis result is used as a parallel parameter and weighted with other parameters to obtain the quality score of each candidate video frame. The target video frame is then determined based on the quality score.
[0100] It is worth noting that the above-described method for determining the target video frame is merely an illustrative example, and the embodiments of this application do not limit the method for determining the target video frame.
[0101] In summary, the video material image selection method provided in this application analyzes facial expressions in candidate video frames to obtain expression analysis results. Based on these results, a target video frame is determined from the candidate video frames as a representative image of the target video stream, which is then used to generate the cover or poster image of the target video stream. This improves the accuracy of video material image determination, the quality of image content in the video material images, and the efficiency of video material image generation.
[0102] In some embodiments, the quality of facial expressions in the aforementioned face region is affected by head posture. Figure 9 This is a flowchart illustrating a method for selecting video footage images according to an exemplary embodiment of this application. This method can be executed by a terminal or a server, or jointly by both. In this embodiment, the method is described using the server as an example. Figure 9 As shown, the method includes:
[0103] Step 901: Obtain the target video stream.
[0104] The target video stream includes video frames.
[0105] Step 902: Perform face detection on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer.
[0106] In some embodiments, face detection is performed on video frames.
[0107] Optionally, face detection is performed on video frames using a pre-trained face detection model. Optionally, the face detection model is used to determine face regions by recognizing face detection points.
[0108] Step 903: Perform head pose analysis on the face region to obtain the head pose results of the face region.
[0109] In some embodiments, a pre-trained head pose model is used to analyze the head pose of the face region in a video frame to obtain the head pose result of the face region. The head pose result is used to indicate the face rotation angle in the face region. In some embodiments, the face rotation angle includes at least one of the following: the face rotation angle in the pitch direction, the face rotation angle in the roll direction, and the face rotation angle in the yaw direction.
[0110] Figure 10 This is a schematic diagram of human head pose analysis provided by an exemplary embodiment of this application. This model can be used to determine the head pose result in the face region. Figure 10 As shown, by measuring the numerical value of the human head offset angle, the offset direction 1010 of the human head in the three-dimensional coordinate system is determined, including the head posture 1011 of pitch offset along the X-axis (such as looking up or down), the head posture 1012 of roll offset along the Y-axis (such as tilting the head to the left or right), and the head posture 1013 of yaw offset along the Z-axis (left side of the face or right side of the face). This is used to determine whether the face area in the video frame is frontal.
[0111] Step 904: Perform facial expression analysis on the face region based on the head pose results to obtain the facial expression analysis results.
[0112] Optionally, the facial expression analysis results can be used to indicate the quality of facial expressions in the face region.
[0113] In some embodiments, the abnormal expression recognition model uses the Laplacian operator to calculate the edge variance, and the resulting variance value is used as the clarity score of the face region in the video frame. The higher the score, the higher the clarity of the video frame.
[0114] Based on the above results of human head posture and abnormal expression recognition, the rules for recognizing complete abnormal expressions in the face region of the selected video frame include the following:
[0115] 1. When the following rules are met: the largest face in the image is a frontal, clear face with a normal expression; and the image does not contain any faces with poor or extremely poor expressions, then the expression is determined to be in excellent condition.
[0116] 2. When the following rules are met: the largest face in the image is a clear face with a normal expression; and there are no extremely poorly clear close-up faces, then the expression state is considered good.
[0117] 3. If the following rules are met: the face contains at least one clear face with a normal expression, and does not contain any extremely poor frontal, clear, close-up faces, then the expression state is determined to be normal.
[0118] 4. When the following rules are met: it contains at least one clear face with poor expression, and the largest face is not extremely poor or does not contain an extremely poor frontal, clear, close-up face, then the expression state is determined to be poor.
[0119] 5. If the following rule is met: the expression contains at least one clear face with poor expression, then the expression state is determined to be extremely poor.
[0120] 6. If the following rule is met: the selected video frame does not contain a clear face and the expression is not extremely poor, then the selected video frame is determined to be an invalid video frame.
[0121] Step 905: Based on the expression analysis results, determine the target video frame from the n candidate video frames and use it as the video material image of the target video stream.
[0122] Among them, video footage images are used as representative images of the target video stream, such as the cover image of the target video stream or the poster image of the target video stream.
[0123] In summary, the video material image selection method provided in this application analyzes facial expressions in candidate video frames to obtain expression analysis results. Based on these results, a target video frame is determined from the candidate video frames as a representative image of the target video stream, which is then used to generate the cover or poster image of the target video stream. This improves the accuracy of video material image determination, the quality of image content in the video material images, and the efficiency of video material image generation.
[0124] The method provided in this embodiment determines the offset direction of the human head in a three-dimensional coordinate system by analyzing the facial rotation angle in the face region of candidate video frames and measuring the magnitude of the head offset angle. Combining the head offset direction with the analysis data of facial expressions, the method obtains the facial expression quality results in the candidate video frames, thereby identifying the target video frame from n candidate video frames as the video material image for the target video stream. The method provided in this application improves the accuracy of facial expression analysis results in the target video frame and ensures the content quality of the target video frame.
[0125] In some embodiments, the target video frame is determined based on facial expression analysis results and diversity parameters. Figure 11 This is a flowchart illustrating a method for selecting video footage images according to an exemplary embodiment of this application. This method can be executed by a terminal or a server, or jointly by both. In this embodiment, the method is described using the server as an example. Figure 11 As shown, the method includes:
[0126] Step 1101: Obtain the target video stream.
[0127] The target video stream includes video frames.
[0128] In some embodiments, human detection is performed on video frames to obtain human key points (such as head, torso, and limbs), and it is determined whether the human head, torso, and limbs are truncated in the video frame. This determines the human state in each video frame. In response to the human state meeting the human integrity condition, video frames that meet the human integrity condition are retained in the target video stream, while the remaining video frames are not processed further.
[0129] Step 1102: Perform face detection on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer.
[0130] In some embodiments, face detection is performed on video frames.
[0131] Optionally, face detection is performed on video frames using a pre-trained face detection model. Optionally, the face detection model is used to determine face regions by recognizing face detection points.
[0132] Step 1103: Perform facial expression analysis on the face regions in the candidate video frames to obtain the facial expression analysis results.
[0133] Optionally, the facial expression analysis results can be used to indicate the quality of facial expressions in the face region.
[0134] Step 1104: Determine the quality parameters of the n candidate video frames based on the facial expression analysis results.
[0135] In some embodiments, by analyzing n candidate video frames, the clarity analysis score, aesthetics analysis score, face position analysis score, and expression score corresponding to the expression analysis result of the candidate video frames are obtained. The weighted sum of the clarity analysis score, aesthetics analysis score, face position analysis score, and expression score is determined as the quality parameter of the candidate video frame.
[0136] Step 1105: Determine the diversity parameters of the candidate video frames.
[0137] Optionally, before calculating the diversity parameters for candidate video frames, it is also necessary to perform clustering and filtering on the candidate video frames. The clustering and filtering methods include at least one of the following:
[0138] First, candidate video frames featuring the same person are clustered, and a predetermined number of candidate video frames are selected from the same cluster as candidate video frames for diversity analysis.
[0139] Second, candidate video frames with the same or similar scenes are clustered, and a preset number of candidate video frames are selected from the same cluster as candidate video frames for diversity analysis.
[0140] Third, cluster candidate video frames that contain the same combination of characters, and select a predetermined number of candidate video frames from the same cluster as candidate video frames for diversity analysis.
[0141] It is worth noting that the above-described clustering and filtering method for determining candidate video frames is merely an illustrative example, and the embodiments of this application do not limit the clustering and filtering method.
[0142] In some embodiments, the diversity parameter of a candidate video frame is calculated by summing the distances between the selected candidate video frame and other candidate video frames. The resulting distance sum is used as the diversity parameter of the selected candidate video frame. That is, for the i-th candidate video frame, the distance sum between the i-th candidate video frame and other candidate video frames in the n-th candidate video frame is determined, where 0 < i ≤ n. The diversity parameter of the i-th candidate video frame is determined based on the distance sum. When the value of the distance sum between the i-th candidate video frame and the n-th candidate video frame and other candidate video frames is larger, it indicates that the diversity quality of the i-th candidate video frame is higher.
[0143] Step 1106: Determine the target video frame from the n candidate video frames based on quality parameters and diversity parameters, and use it as the video material image of the target video stream.
[0144] Among them, video footage images are used as representative images of the target video stream, such as the cover image of the target video stream or the poster image of the target video stream.
[0145] In some embodiments, the target video frame among candidate video frames is determined by calculating a diversity quality score. The algorithm for calculating the diversity quality score is derived from the Maximum Marginal Relevance (MMR) algorithm. The purpose of this algorithm is to reduce redundancy in the ranking results while ensuring relevance. This allows for the recommendation of relevant products to users while maintaining diversity in the recommendation results; that is, the ranking results involve a trade-off between relevance and diversity. The specific formula is shown in Formula 1:
[0146] Formula 1:
[0147] Where Q represents the Query, S is the selected set, R represents the candidate set, and D...i D represents the current candidate result. j This means dividing D from the selected set S. i Other results include Sim1, which represents the relevance between candidate results and the query; Sim2, which represents the relevance between D; and λ, which is a weighting coefficient that modulates the relevance and diversity of recommendation results.
[0148] In this embodiment, in order to select target video frames from the target video stream while ensuring the quality and diversity of the target video frames, a formula for calculating the current score of candidate video frames, namely the diversity quality score, is proposed. The specific formula is shown in Formula 2:
[0149] Formula 2:
[0150] Analogous to Formula 1, where D i Let R be the current candidate video frame, S be the set of all candidate video frames, and D be the set of selected candidate video frames. j For the selected candidate video frames S excluding D i Other candidate video frames, λ is the weighting coefficient of the adjustment result and diversity, f(D) i ) represents the suitability score of candidate video frames as video poster advertisements, where dist(D) is the score. i D j ) represents the distance between candidate video frames. In the formula, f(D) i The `dist(D)` is a score formula related to image selection algorithms. It is a weighted average of analytical factors such as sharpness analysis score, aesthetics analysis score, face position analysis score, and expression score. The higher the score, the more suitable the current candidate video frame is as video poster material. i D j The distance formula is as follows: This scheme uses a weighted average of human body region image features, whole image features, and differences in face size and number of faces as the distance. Specifically, the quality parameters of the m-th candidate video frame are weighted and summed with the diversity parameters of the m-th video frame to obtain the material adaptation score of the m-th video frame, where 0 < m ≤ n. Based on the material adaptation scores corresponding to the n video frames, the target video frame is determined from the n video frames. In summary, the video material image selection method provided in this application provides expression analysis of the face regions in the candidate video frames to obtain expression analysis results. Based on these results, the target video frame is determined from the candidate video frames as a representative image of the target video stream, used to generate the cover or poster image of the target video stream. This improves the accuracy of video material image determination and the image content quality in the video material images, thus improving the efficiency of video material image generation.
[0151] Figure 12 This is a schematic diagram of the structure of a video material image generation framework 1200 provided in an exemplary embodiment of this application, as shown below. Figure 12 As shown, the framework includes the following parts:
[0152] The process includes decoding and clustering (1201), filtering and sorting (1202), defect detection (1203), image selection and processing (1204), design element processing (1205), and delivery feedback (1206). Each of these parts will be explained separately.
[0153] Decoding clustering 1201 refers to performing compatibility correction on the video format and then performing frame-level parallel multi-segment decoding on the corrected video stream to select a certain number of candidate frames. Specifically, it involves calculating the global image hash value of the candidate frames and clustering similar candidate frames to subsequently suppress the output of similar candidate frames.
[0154] Filtering and sorting 1202 refers to filtering and sorting candidate frames based on image content quality or parameters of the image itself, in order to confirm the images in subsequent video footage. This includes basic analysis 1211, clarity 1212, aesthetics 1213, OCR recognition 1214, key points 1215, face filtering (including face detection 1216), celebrity recognition 1217, abnormal expression recognition 1218, and scene recognition 1219.
[0155] First, basic analysis (1211) is used to calculate image brightness, contrast, and saturation values to initially filter out video frames that are too dark or too blurry. Then, based on clarity (1212), the video frames in the target video stream are scored to remove blurry scenes caused by factors such as camera shake and character movement, while selecting video frames with higher clarity quality. Next, based on aesthetics (1213), the video frames are scored for aesthetic appeal. The model used for this scoring process is trained using aesthetics, photography, and a proprietary cover frame dataset, and it has a good filtering effect on video frames with clever composition, clear lighting, and expressiveness.
[0156] Then, OCR recognition 1214 identifies text regions in the video frames, filtering out video frames where dialogue and face regions overlap, as well as commercial segments in movies and TV shows, and uses this information for subsequent cropping to remove subtitles. Next, keypoint 1215, i.e., human keypoint recognition technology, determines whether the video frame truncates the human body region, improving the pass rate of video poster materials, and can also be used to select video frames where the human body is not obscured. Face detection 1216 determines the position of faces and facial keypoints in the video frames, filtering out video frames where the face region is too large, too small, or located at the boundary, and uses this information for subsequent selection of target video frames. Next, abnormal expression 1218 identifies the specific state of facial features such as eyes and mouth, filtering out video frames containing abnormal expressions such as closed eyes and bared teeth, while selecting video frames with higher quality facial expressions from the target video stream; then, celebrity recognition 1217 identifies celebrities in the video frames, prioritizing video frames containing celebrities with high user attention or those with leading stars that are highly relevant to the candidate video frames to improve the relevance of the video frames and viewer attention; finally, scene recognition 1219 identifies the scenes in the video frames, prioritizing video frames that match the scenes in the target video stream.
[0157] Defect detection 1203 refers to detecting obvious defects in video frames, thereby filtering out video frames with obvious defects or eliminating defects in video frames. Specifically, defect detection 1203 includes black and white edge detection 1221, frosted glass detection 1222, and logo detection 1223. After detection, it can filter video frames containing black and white edges, or video frames with frosted glass effects, or remove logos from video frames.
[0158] Image selection processing 1204 refers to selecting video frames that meet the requirements based on diversity requirements and image size requirements, and then retouching the selected video frames. Image selection processing 1204 includes diverse image selection 1231, intelligent cropping 1232, and logo removal 1233. Among them, diverse image selection 1231 refers to selecting multiple different video frames according to diversity requirements 1231. Intelligent cropping 1232 refers to cropping the redundant parts, subtitle parts, or advertising parts of the candidate video frames according to different video frame templates set according to different advertising needs. Logo removal 1233 refers to intelligently removing the logo displayed on the video frame.
[0159] Design element processing 1205 refers to the process of generating a poster or cover image based on the target video frame after it has been determined. This includes template design 1241, element position selection 1242, color scheme selection 1243, and image enhancement 1244. Template design 1241 involves designing a template according to template requirements, inserting the edited target video frame into the template, and generating a poster or cover image. Element position selection 1242 involves determining the display position of elements in the poster or cover image. Color scheme selection 1243 involves designing the color scheme for the poster or cover image. Image enhancement 1244 involves purposefully emphasizing the overall or local characteristics of an image, making an originally unclear image clearer, emphasizing certain features of interest, and amplifying the differences between features of different objects in the image.
[0160] The "Platform Feedback 1206" refers to determining the effectiveness of a campaign based on the placement of the poster or cover image. This includes AI material tagging 1251, campaign effectiveness monitoring and iteration 1252, and a negative case feedback mechanism 1253. Specifically, AI material tagging 1251 uses artificial intelligence technology to assign corresponding film / television category tags to the poster or cover image. Campaign effectiveness monitoring and iteration 1252 involves updating campaign effectiveness data in real time based on feedback data after the poster or cover image is placed. The Badcase feedback mechanism 1253 receives feedback when the poster or cover image performs poorly after placement.
[0161] In this embodiment, using video frames from video streams such as movies and TV shows as video material images can effectively attract users to use the video app and retain and engage them. However, manually creating related materials is time-consuming, labor-intensive, and has limited output. It requires selecting suitable video frames from a large number of video streams before designing, which limits speed and output. By using technologies such as person recognition, scene understanding, and image analysis, high-quality video frames can be extracted from the video stream, and then intelligently cropped and graphic designed to automatically produce high-quality video material images. Under the demand for rapid deployment, AI poster production can respond faster and save time than manual material production; at the same time, AI's production capacity is not limited and can simultaneously and efficiently support the production of video material images from massive amounts of video.
[0162] The adoption rate is the proportion of video frames actually selected and used in a batch of video footage images out of all output target video frames. The main factors affecting the adoption rate are the quality and diversity of the video frames themselves. Defects in video frame quality, such as unnatural facial expressions or dark and blurry images, render the frames unusable as video footage images. Furthermore, when multiple video footage images have a high repetition rate, typically only the best image is selected for use, rendering the remaining duplicate images unusable. When directly using existing technologies for video footage image extraction, the usability rate is only 10%-20%, while the video footage image selection based on this solution is expected to achieve a usability rate of 60%-70%.
[0163] Figure 13 This is a structural block diagram of a video material image selection device provided in an exemplary embodiment of this application, such as... Figure 13 As shown, the device includes:
[0164] Acquisition module 1310 is used to acquire a target video stream, wherein the target video stream includes video frames;
[0165] The first detection module 1320 is used to perform face detection on the video frame to obtain n candidate video frames containing face regions, where n≥2 and n is an integer.
[0166] The first analysis module 1330 is used to perform expression analysis on the face region in the candidate video frame to obtain the expression analysis result of the face region, and the expression analysis result is used to indicate the quality of the facial expression in the face region;
[0167] The first determining module 1340 is used to determine a target video frame from the n candidate video frames based on the expression analysis results, as the video material image of the target video stream, and the video material image is used as a representative image of the target video stream.
[0168] In an optional embodiment, facial expression analysis is performed on the facial regions in the candidate video frames to obtain the facial expression analysis results of the facial regions;
[0169] The first analysis module 1330 is further configured to perform sub-region segmentation on the face region according to the distribution of facial features, to obtain face sub-regions corresponding to the facial features; and to perform expression analysis on the face sub-regions corresponding to the facial features respectively, to obtain the expression analysis result of the face region.
[0170] In an optional embodiment, the first analysis module 1330 is further configured to perform facial expression analysis on the first sub-region to obtain an eye state analysis result corresponding to the eyes of the face, wherein the eye state analysis result is used to indicate the degree of opening and closing of the eyes of the face in the face region.
[0171] In an optional embodiment, the first analysis module 1330 is further configured to perform expression analysis on the second sub-region to obtain a mouth state analysis result corresponding to the mouth of the face, wherein the mouth state analysis result is used to indicate the expression form of the mouth of the face in the face region.
[0172] In an optional embodiment, such as Figure 14 As shown, the device also includes:
[0173] The second analysis module 1350 is used to perform head posture analysis on the face region to obtain the head posture result of the face region, and the head posture result is used to indicate the face rotation angle in the face region; the first analysis module 1330 is also used to perform expression analysis on the face region based on the head posture result to obtain the expression analysis result.
[0174] In an optional embodiment, the first determining module 1340 is further configured to determine the quality parameters of the n candidate video frames based on the expression analysis results; determine the diversity parameters of the candidate video frames; and determine the target video frame from the n candidate video frames based on the quality parameters and the diversity parameters.
[0175] In an optional embodiment, the determination of the diversity parameters of the candidate video frames;
[0176] The first determining module 1340 is further configured to, for the i-th candidate video frame, determine the sum of distances between the i-th candidate video frame and other candidate video frames in the n-th candidate video frame, where 0 < i ≤ n; and determine the diversity parameters of the i-th candidate video frame based on the sum of distances.
[0177] In an optional embodiment, the quality parameters of the n candidate video frames are determined based on the facial expression analysis results;
[0178] The first determining module 1340 is further configured to obtain the clarity analysis score, aesthetics analysis score, face position analysis score, and expression score corresponding to the expression analysis result of the candidate video frame; and to determine the weighted sum of the clarity analysis score, the aesthetics analysis score, the face position analysis score, and the expression score as the quality parameter of the candidate video frame.
[0179] In an optional embodiment, the target video frame is determined from the n candidate video frames based on the quality parameter and the diversity parameter;
[0180] The first determining module 1340 is further configured to weight and sum the quality parameters of the m-th candidate video frame with the diversity parameters of the m-th video frame to obtain the material adaptation score of the m-th video frame, where 0 < m ≤ n; and to determine the target video frame from the n video frames based on the material adaptation scores corresponding to the n video frames respectively.
[0181] In an optional embodiment, face detection is performed on the video frames to obtain n candidate video frames containing face regions;
[0182] The first detection module 1320 is further configured to perform face detection on the video frame to obtain a face video frame containing a face region; and to filter the face video frame based on the region parameters of the face region in the face video frame to obtain the n candidate video frames, wherein the region parameters include at least one of region size and region position.
[0183] In an optional embodiment, in such a way Figure 13 Based on the acquisition module 1310, the first detection module 1320, the first analysis module 1330, and the first determination module 1340 shown, as follows: Figure 15 As shown, the device also includes:
[0184] The second detection module 1360 is used to perform human body detection on the video frame to obtain human body key points;
[0185] The second determining module 1370 is used to determine the human body state in the video frame based on the human body key points;
[0186] The retention module 1380 is configured to retain video frames in the target video stream that meet the human body integrity conditions in response to the human body state meeting the human body integrity conditions.
[0187] In summary, the video material image selection device provided in this application performs facial expression analysis on the facial regions in candidate video frames to obtain expression analysis results. Based on the expression analysis results, it determines the target video frame from the candidate video frames as a representative image of the target video stream, which is used to generate the cover or poster image of the target video stream. This improves the accuracy of video material image determination and the image content quality in the video material images, thereby improving the efficiency of video material image generation.
[0188] It should be noted that the video material image selection device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video material image selection device provided in the above embodiments belongs to the same concept as the video material image selection method embodiments, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0189] Figure 16 This illustration shows a structural block diagram of a computer device 1600 provided in an exemplary embodiment of this application. The computer device 1600 may be... Figure 1 The server or terminal shown.
[0190] Typically, computer device 1600 includes a processor 1601 and a memory 1602.
[0191] Processor 1601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0192] The memory 1602 may include one or more computer-readable storage media, which may be non-transitory. The memory 1602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1602 is used to store at least one instruction, which is executed by the processor 1601 to implement the video material image selection method provided in the method embodiments of this application.
[0193] In some embodiments, the computer device 1600 may optionally include a peripheral device interface 1603 and at least one peripheral device. The processor 1601, memory 1602, and peripheral device interface 1603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607, a positioning assembly 1608, and a power supply 1609.
[0194] Peripheral interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1601 and memory 1602. In some embodiments, processor 1601, memory 1602 and peripheral interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1601, memory 1602 and peripheral interface 1603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0195] The radio frequency (RF) circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1604 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1604 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0196] Display screen 1605 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1601 for processing. In this case, display screen 1605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1605, disposed on the front panel of computer device 1600; in other embodiments, there may be at least two display screens, disposed on different surfaces of computer device 1600 or in a folded design; in still other embodiments, display screen 1605 may be a flexible display screen, disposed on a curved or folded surface of computer device 1600. Furthermore, display screen 1605 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1605 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0197] The camera assembly 1606 is used to acquire images or videos. Optionally, the camera assembly 1606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0198] The audio circuit 1607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1601 for processing, or to the radio frequency circuit 1604 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the computer device 1600. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1601 or the radio frequency circuit 1604 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1607 may also include a headphone jack.
[0199] The positioning component 1608 is used to locate the current geographical location of the computer device 1600 in order to enable navigation or LBS (Location Based Service). The positioning component 1608 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0200] Power supply 1609 is used to supply power to the various components in computer device 1600. Power supply 1609 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1609 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0201] In some embodiments, the computer device 1600 further includes one or more sensors 1610. The one or more sensors 1610 include, but are not limited to: an accelerometer 1611, a gyroscope 1612, a pressure sensor 1613, a fingerprint sensor 1614, an optical sensor 1615, and a proximity sensor 1616.
[0202] Accelerometer 1611 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 1600. For example, accelerometer 1611 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1601 can control display screen 1605 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1611. Accelerometer 1611 can also be used for games or for acquiring user motion data.
[0203] The gyroscope sensor 1612 can detect the orientation and rotation angle of the computer device 1600. The gyroscope sensor 1612 can work in conjunction with the accelerometer sensor 1611 to acquire the user's 3D movements on the computer device 1600. Based on the data acquired by the gyroscope sensor 1612, the processor 1601 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0204] Pressure sensor 1613 can be disposed on the side bezel of computer device 1600 and / or on the lower layer of display screen 1605. When pressure sensor 1613 is disposed on the side bezel of computer device 1600, it can detect the user's grip signal on computer device 1600, and processor 1601 can perform left / right hand recognition or quick operation based on the grip signal collected by pressure sensor 1613. When pressure sensor 1613 is disposed on the lower layer of display screen 1605, processor 1601 can control operable controls on the UI interface based on the user's pressure operation on display screen 1605. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0205] The fingerprint sensor 1614 is used to collect a user's fingerprint. The processor 1601 identifies the user based on the fingerprint collected by the fingerprint sensor 1614, or vice versa. When the user's identity is identified as trusted, the processor 1601 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1614 can be located on the front, back, or side of the computer device 1600. When the computer device 1600 has physical buttons or a manufacturer's logo, the fingerprint sensor 1614 can be integrated with the physical buttons or the manufacturer's logo.
[0206] An optical sensor 1615 is used to collect ambient light intensity. In one embodiment, the processor 1601 can control the display brightness of the display screen 1605 based on the ambient light intensity collected by the optical sensor 1615. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1605 is increased; when the ambient light intensity is low, the display brightness of the display screen 1605 is decreased. In another embodiment, the processor 1601 can also dynamically adjust the shooting parameters of the camera assembly 1606 based on the ambient light intensity collected by the optical sensor 1615.
[0207] The proximity sensor 1616, also known as a distance sensor, is typically located on the front panel of the computer device 1600. The proximity sensor 1616 is used to detect the distance between the user and the front of the computer device 1600. In one embodiment, when the proximity sensor 1616 detects that the distance between the user and the front of the computer device 1600 is gradually decreasing, the processor 1601 controls the display screen 1605 to switch from a screen-on state to a screen-off state; when the proximity sensor 1616 detects that the distance between the user and the front of the computer device 1600 is gradually increasing, the processor 1601 controls the display screen 1605 to switch from a screen-off state to a screen-on state.
[0208] Those skilled in the art will understand that Figure 16 The structure shown does not constitute a limitation on the computer device 1600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0209] It should be noted that the video material image selection device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video material image selection device and the video material image selection method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0210] Embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the video material image selection method provided in the above-described method embodiments.
[0211] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the video material image selection method provided in the above-described method embodiments.
[0212] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the video material image selection methods described in the above embodiments.
[0213] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0214] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0215] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for selecting video source images, characterized in that, The method includes: Acquire a target video stream, wherein the target video stream includes video frames; Face detection is performed on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer; Facial expression analysis is performed on the facial regions in the candidate video frames to obtain facial expression analysis results for the facial regions. These results are used to indicate the quality of facial expressions in the facial regions. Based on the facial expression analysis results, determine the quality parameters of the n candidate video frames; For the i-th candidate video frame, determine the sum of distances between the i-th candidate video frame and other candidate video frames in the n-th candidate video frame, where 0 < i ≤ n; Based on the distance, the diversity parameters for determining the i-th candidate video frame are used; Based on the quality parameters and the diversity parameters, a target video frame is determined from the n candidate video frames and used as the video material image of the target video stream. The video material image is used as a representative image of the target video stream.
2. The method according to claim 1, characterized in that, The step of performing facial expression analysis on the facial regions in the candidate video frames to obtain the facial expression analysis results includes: The face region is segmented into sub-regions based on the distribution of facial features to obtain face sub-regions corresponding to the facial features. Expression analysis is performed on the facial sub-regions corresponding to the facial features to obtain the expression analysis results of the facial regions.
3. The method according to claim 2, characterized in that, The face sub-region includes a first sub-region corresponding to the eyes of the face; The step of performing expression analysis on the facial sub-regions corresponding to the facial features to obtain the expression analysis results for the facial regions includes: Expression analysis is performed on the first sub-region to obtain eye state analysis results corresponding to the eyes of the face. The eye state analysis results are used to indicate the degree of opening and closing of the eyes of the face in the face region.
4. The method according to claim 2, characterized in that, The face sub-region includes a second sub-region corresponding to the mouth of the face; The step of performing expression analysis on the facial sub-regions corresponding to the facial features to obtain the expression analysis results for the facial regions includes: Expression analysis is performed on the second sub-region to obtain the mouth state analysis result corresponding to the mouth of the face. The mouth state analysis result is used to indicate the expression form of the mouth of the face in the face region.
5. The method according to claim 1, characterized in that, The method further includes: A head pose analysis is performed on the face region to obtain the head pose result of the face region. The head pose result is used to indicate the face rotation angle in the face region. The step of performing facial expression analysis on the facial regions in the candidate video frames to obtain the facial expression analysis results includes: Based on the head pose results, facial expression analysis is performed on the facial region to obtain the expression analysis results.
6. The method according to any one of claims 1 to 5, characterized in that, The process of determining the quality parameters of the n candidate video frames based on the facial expression analysis results includes: Obtain the clarity analysis score, aesthetics analysis score, face position analysis score, and expression score corresponding to the expression analysis result of the candidate video frame; The weighted sum of the sharpness analysis score, the aesthetics analysis score, the face position analysis score, and the expression score is determined as the quality parameter of the candidate video frame.
7. The method according to any one of claims 1 to 5, characterized in that, Determining the target video frame from the n candidate video frames based on the quality parameter and the diversity parameter includes: The quality parameters of the m-th candidate video frame are weighted and summed with the diversity parameters of the m-th video frame to obtain the material adaptation score of the m-th video frame, where 0 < m ≤ n. The target video frame is determined from the n video frames based on the material adaptation scores corresponding to each of the n video frames.
8. The method according to any one of claims 1 to 5, characterized in that, The step of performing face detection on the video frames to obtain n candidate video frames containing face regions includes: Perform face detection on the video frames to obtain face video frames containing face regions; Based on the region parameters of the face region in the face video frame, the face video frame is filtered to obtain the n candidate video frames. The region parameters include at least one of region size and region location.
9. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Human body detection is performed on the video frames to obtain key human body points; The human body state in the video frame is determined based on the aforementioned human body key points; In response to the human body state meeting the human body integrity condition, video frames that meet the human body integrity condition are retained in the target video stream.
10. A device for selecting video material images, characterized in that, The device includes: The acquisition module is used to acquire a target video stream, wherein the target video stream includes video frames; The detection module is used to perform face detection on the video frames to obtain n candidate video frames containing face regions, where n≥2 and n is an integer; The first analysis module is used to perform facial expression analysis on the face regions in the candidate video frames to obtain facial expression analysis results for the face regions. The facial expression analysis results are used to indicate the quality of facial expressions in the face regions. The determination module is used to determine the quality parameters of the n candidate video frames based on the expression analysis results; for the i-th candidate video frame, determine the sum of distances between the i-th candidate video frame and other candidate video frames in the n candidate video frames, where 0 < i ≤ n; determine the diversity parameters of the i-th candidate video frame based on the sum of distances; and determine the target video frame from the n candidate video frames based on the quality parameters and the diversity parameters, as the video material image of the target video stream, wherein the video material image is used as a representative image of the target video stream.
11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the video material image selection method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The storage medium stores at least one program, which is loaded and executed by a processor to implement the video material image selection method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the method for selecting video material images as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Semi-supervised learning-based multi-gesture facial expression recognition method
CN103186774A
Method and device for generating poster of video
CN108833939A
Human face recognition method, device and equipment based on artificial intelligence, and medium
CN111242090A
Film and television poster automatic generation method and system based on deep learning
CN113010711A