Video cover determination method, device, equipment and storage medium
By analyzing the face position and angle information in the video frame, the target image of the video cover is solved, and the problem that the video cover cannot accurately represent the video content in the prior art is solved, and the accuracy of the cover determination is improved.
Patent Information
- Application Number
- CN202210210379.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-03-04
AI Technical Summary
The existing video cover scoring method only considers basic image information such as the blurring of the video frame, which results in the output video cover that cannot accurately represent the video content and has a low accuracy rate.
By acquiring candidate images from the image sequence of video to be processed, the face position analysis results are determined based on nose key points and size information, the face angle analysis results are determined based on eye and mouth key points, and the target image is determined from the candidate images for generating a video cover.
The accuracy of the determination of the target image of the video cover is improved, so that the generated video cover can more represent the video content.
Smart Images

Figure CN116719970B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a method, apparatus, device, and storage medium for determining a video cover. Background Art
[0002] In the video cover scoring method in the related art, usually, the video frames after sampling the video are analyzed in aspects such as illuminance, color vividness, and blurriness, and the scores of the video frames are comprehensively obtained. Finally, the video frames with higher scores are determined from all the sampled video frames as the output result, as Figure 1 shown.
[0003] However, the related art only considers relatively basic image information such as the blurriness of the video frames, and the output video cover cannot accurately represent the video content, and the accuracy of determining the video cover is relatively low. Summary of the Invention
[0004] To solve the above technical problems, this application provides a method, apparatus, device, and storage medium for determining a video cover.
[0005] On the one hand, this application proposes a method for determining a video cover, and the method includes:
[0006] Obtaining candidate images from the image sequence of the video to be processed;
[0007] Based on the nose key points in the candidate images and the first size information of the candidate images, determining the face position analysis result corresponding to the candidate images;
[0008] According to the eye key points and mouth key points in the candidate images, determining the face angle analysis result corresponding to the candidate images;
[0009] According to the comprehensive face analysis result, determining a target image from the candidate images; the target image is used to generate the cover of the video to be processed.
[0010] On the other hand, an embodiment of this application provides a device for determining a video cover, and the device includes:
[0011] A candidate image acquisition module, configured to obtain candidate images from the image sequence of the video to be processed;
[0012] A position analysis result determination module, configured to determine the face position analysis result corresponding to the candidate images based on the nose key points in the candidate images and the first size information of the candidate images;
[0013] An angle analysis result determination module, configured to determine the face angle analysis result corresponding to the candidate images according to the eye key points and mouth key points in the candidate images;
[0014] A comprehensive analysis result determination module, configured to determine a comprehensive face analysis result corresponding to the candidate image based on the face position analysis result and the face angle analysis result;
[0015] A target image determination module, configured to determine a target image from the candidate images according to the comprehensive face analysis result; the target image is used to generate a cover for the to-be-processed video.
[0016] On the other hand, the present application provides an electronic device for determining a video cover. The electronic device includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the video cover determination method as described above.
[0017] On the other hand, the present application provides a computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the video cover determination method as described above.
[0018] On the other hand, the present application provides a computer program product, and when the computer program is executed by a processor, it implements the video cover determination method as described above.
[0019] The video cover determination method, device, equipment and storage medium provided by the embodiments of the present application obtain candidate images from the image sequence of the to-be-processed video, determine the face position analysis result corresponding to the candidate image based on the nose key points in the candidate image and the first size information of the candidate image, determine the face angle analysis result corresponding to the candidate image based on the eye key points and mouth key points in the candidate image, determine the comprehensive face analysis result corresponding to the candidate image based on the face position analysis result and the face angle analysis result, and determine a target image for generating the cover of the to-be-processed video from the candidate images according to the comprehensive face analysis result, realizing the determination of the target image based on the position information and angle information of the face, making the target image for generating the cover more representative of the video content of the to-be-processed video, and improving the determination accuracy of the target image for generating the cover. Description of the Drawings
[0020] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is a flowchart of video cover scoring in the related art.
[0022] Figure 2 is a schematic diagram of an implementation environment of a video cover determination method shown according to an exemplary embodiment.
[0023] Figure 3 is a flowchart of a video cover determination method shown according to an exemplary embodiment Figure 1 .
[0024] Figure 4 is a flowchart of training a face detection network shown according to an exemplary embodiment.
[0025] Figure 5 is a flowchart of determining a face position analysis result shown according to an exemplary embodiment.
[0026] Figure 6 is a flowchart of determining a face angle analysis result shown according to an exemplary embodiment.
[0027] Figure 7 is a flowchart of another video cover determination method shown according to an exemplary embodiment.
[0028] Figure 8 is a process of a video cover determination method shown according to an exemplary embodiment Figure 2 .
[0029] Figure 9 is a process of a video cover determination method shown according to an exemplary embodiment Figure 3 .
[0030] Figure 10 is a comparison diagram of a cover image selected by using the method of the embodiment of the present application and a cover image selected by using the existing method shown according to an exemplary embodiment.
[0031] Figure 11 is a block diagram of a video cover determination device shown according to an exemplary embodiment.
[0032] Figure 12 is a hardware structure block diagram of a server for video cover determination provided by an embodiment of the present application. Detailed implementation manners
[0033] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0034] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0035] Specifically, the embodiments of the present application relate to image processing technology and face recognition technology in computer vision technology.
[0036] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present application.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0038] Figure 2 It is a schematic diagram of the implementation environment of a video cover determination method shown according to an exemplary embodiment. As Figure 2As shown, the implementation environment may at least include a terminal 01 and a server 02. The terminal 01 and the server 2 may be directly or indirectly connected through wired or wireless communication means, and this application does not limit this here.
[0039] Specifically, the terminal can be used to collect an image sequence of the video to be processed. Optionally, the terminal 01 may include, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.
[0040] Specifically, the server 02 can be used to obtain candidate images from the image sequence collected by the terminal; and to determine the face position analysis result corresponding to the candidate image based on the nose key points in the candidate image and the first size information of the candidate image; and to determine the face angle analysis result corresponding to the candidate image according to the eye key points and mouth key points in the candidate image; and to determine the comprehensive face analysis result corresponding to the candidate image based on the face position analysis result and the face angle analysis result; and to determine a target image for generating the cover of the above-mentioned video to be processed from the candidate images according to the comprehensive face analysis result. Optionally, the server 02 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0041] It should be noted that Figure 2 This is just an example. In other scenarios, other implementation environments may also be included. For example, the implementation environment may include a terminal that collects an image sequence of the video to be processed, determines the face position analysis result corresponding to the candidate image based on the nose key points in the candidate image and the first size information of the candidate image; determines the face angle analysis result corresponding to the candidate image according to the eye key points and mouth key points in the candidate image, and determines a target image for generating the cover of the above-mentioned video to be processed according to the face position analysis result and the face angle analysis result.
[0042] It should be noted that the embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0043] Figure 3 is a schematic flow chart of a video cover determination method shown according to an exemplary embodiment Figure 1 . This method can be used for Figure 2in the implementation environment. This specification provides the method operation steps as described in the embodiments or flowcharts, but may include more or fewer operation steps based on routine or non-creative labor. The step sequences listed in the embodiments are only one way among the execution sequences of numerous steps and do not represent the only execution sequence. When the actual system or server product is executed, it can be executed in the method sequence shown in the embodiments or the drawings or in parallel (for example, in an environment with parallel processors or multi-threaded processing). Specifically, as Figure 3 shown, the method may include:
[0044] S101. Obtain candidate images from the image sequence of the video to be processed.
[0045] The video cover determination method provided by the embodiments of this application is applied to an electronic device, which can be an intelligent device such as a mobile terminal or a server.
[0046] In an actual application scenario, after the electronic device for cover determination receives a processing request for cover determination of a certain video, it determines the video as the video to be processed, and based on the video to be processed, adopts the cover determination method provided by the embodiments of this application to perform corresponding processing.
[0047] Among them, the electronic device for cover determination receiving a processing request for cover determination of a certain video mainly includes at least one of the following situations:
[0048] Situation 1: When the terminal needs to process a certain video, it can input a service processing request to the intelligent device. After the intelligent device receives the service processing request for the video, it can send a processing request for cover determination of the video to the electronic device for cover determination.
[0049] Situation 2: When the intelligent device detects that a certain video has been recorded, it can also actively generate a processing request for cover determination of the video and send it to the electronic device for cover determination.
[0050] Situation 3: When the user needs to determine the cover of a certain video, it can input a cover processing request to the intelligent device. After the intelligent device receives the cover processing request for the video, it can send a processing request for cover determination of the video to the electronic device for cover determination.
[0051] It should be noted that the electronic device for cover determination can be the same as or different from the intelligent device.
[0052] In the specific implementation process, after the electronic device for determining the cover receives a processing request for determining the cover of a certain video, it determines the video as the video to be processed, obtains the pictures corresponding to the video frames in the video to be processed to obtain an image sequence, and obtains candidate images from the image sequence.
[0053] In an alternative implementation, each frame image in the video to be processed can be selected as a candidate image in the way of full-frame extraction.
[0054] In another alternative implementation, in order to reduce the data processing volume and be able to more comprehensively reflect the content covered by the video, some images can be extracted from the video as candidate covers. For example, multiple frames of images can be randomly extracted from the video as candidate images.
[0055] However, randomly extracting images from the video as candidate images may cause the distribution positions of multiple frames of candidate images in the video to be relatively concentrated, and it is easy to miss some wonderful content. In this way, the video cover selected from the candidate images may not be able to reflect the wonderful content that the client's account is interested in in the video, so that the selected candidate images cannot more comprehensively reflect the content shown in the video. Optionally, in order to make the selected candidate cover more comprehensively reflect the content shown in the video, the video can be first split into multiple consecutive video segments, and then at least one image is selected from each video segment as a candidate image. For example, the video can be evenly split into multiple video segments, and the number of image frames included in each video segment is the same; or considering the situation that multiple frames of images in the video cannot be evenly divided, the difference in the number of video frames between any two video segments can be less than a preset threshold; then, one frame of image is respectively selected from each video segment as a candidate image, so as to obtain multiple frames of images as candidate images.
[0056] S103. Based on the nose key points in the above candidate images and the first size information of the above candidate images, determine the face position analysis result corresponding to the above candidate images.
[0057] S105. Based on the eye key points and mouth key points in the above candidate images, determine the face angle analysis result corresponding to the above candidate images.
[0058] Optionally, the embodiments of the present application can determine the face position analysis result and face angle analysis result corresponding to the candidate images in various ways, which are not specifically limited herein.
[0059] In one implementation, a face image library can be established in advance. The face image library contains a large number of comparison images, and scores are given to the positions and angles of the faces in each comparison image to obtain the face position analysis result and face angle analysis result corresponding to each comparison image. The candidate image is matched with the comparison images in the face image library, and the face position analysis result and face angle analysis result of the comparison image that matches the candidate image are used as the face position analysis result and face angle analysis result corresponding to the candidate image.
[0060] In another implementation, the face position analysis result and face angle analysis result corresponding to the candidate image can be determined through a pre-trained face detection network.
[0061] Specifically, in the above step S103, based on the nose key points in the candidate image and the first size information of the candidate image, determining the face position analysis result corresponding to the candidate image:
[0062] Input the candidate image into the face detection network.
[0063] Based on the face detection network, process the nose key points in the candidate image and the first size information of the candidate image to obtain the face position analysis result.
[0064] Specifically, in the above step S105, according to the eye key points and mouth key points in the candidate image, determining the face angle analysis result corresponding to the candidate image may include:
[0065] Input the candidate image into the face detection network.
[0066] Based on the face detection network, process the eye key points and mouth key points in the candidate image to obtain the face angle analysis result.
[0067] Optionally, the candidate image can be input into a face detection network. The face detection network detects face bounding boxes for the candidate image. If the number of face bounding boxes is greater than 0, it indicates that there are faces in the candidate image. Since there may be at least one face in the candidate image, the face detection network can obtain the position information of the key points on each face, and based on the position information of the nose key points on each face and the first size information of the candidate image, obtain the face position analysis result for each face, and based on the position information of the eye key points and mouth key points on each face, obtain the face angle analysis result for each face. The sum of the face position analysis results of all faces is used as the face position analysis result corresponding to the candidate image, and the sum of the face angle analysis results of all faces is used as the face angle analysis result corresponding to the candidate image. Since the pre-trained face detection network can accurately detect the position information of the face detection bounding boxes and key points, and generate the face position analysis result and the above-mentioned face angle analysis result of the candidate image based on this, the determination accuracy of the face position analysis result and the above-mentioned face angle analysis result can be improved.
[0068] Figure 4 is a flowchart of training a face detection network shown according to an exemplary embodiment. As Figure 4 shown, in a feasible embodiment, the above method may further include the step of training a face detection network, and the step of training the face detection network may include:
[0069] S201. Obtain labeled sample candidate images from the sample video to be processed; the above labels represent the position annotation results and angle annotation results of the sample faces included in the sample candidate images.
[0070] S203. Process the sample nose key points in the sample candidate image and the size information of the sample candidate image based on a preset neural network to obtain the sample face position analysis result corresponding to the sample candidate image, and process the sample eye key points and sample mouth key points in the sample candidate image based on the above preset neural network to obtain the sample face angle analysis result corresponding to the sample candidate image.
[0071] S205. Determine the loss information of the sample candidate image based on the above sample face position analysis result, position annotation result, sample face angle analysis result, and angle annotation result.
[0072] S207. Train the above preset neural network according to the above loss information to obtain a face detection network.
[0073] Optionally, during the process of training the face detection network, sample data for training, i.e., the sample video to be processed, can be obtained, and sample candidate images can be obtained from the sample video to be processed; the position information of the sample human face in the sample candidate image is labeled to obtain a position labeling result, and the angle information of the sample human face is labeled to obtain an angle labeling result. The sample candidate image is input into a preset neural network, and the preset neural network performs sample key point detection on the sample candidate image to obtain the sample nose key points, sample eye key points, and sample mouth key points of the sample candidate image. According to the position information of the sample nose key points and the size information of the above sample candidate image, the sample human face position analysis result corresponding to the sample candidate image is obtained. According to the position information of the sample eye key points and the position information of the sample mouth key points, the sample human face angle analysis result corresponding to the sample candidate image is obtained. And according to the sample human face position analysis result and the position labeling result, position loss information is determined. According to the sample human face angle analysis result and the angle labeling result, angle loss information is determined. The position loss information and the angle loss information are weighted to obtain loss information. Finally, the preset neural network is trained according to the loss information until the sample human face position analysis result output by the preset neural network is close to the position labeling result, and the sample human face angle analysis result output by the preset neural network is close to the angle labeling result, and a face detection network is obtained.
[0074] Optionally, the determination process of the above sample candidate image is similar to the determination process of the candidate image, and will not be elaborated here.
[0075] It should be noted that the above sample human face position analysis result can represent the position score of the human face in the candidate image. For example, if the human face is located at the center of the candidate image, the position score is the highest; if the human face is located at the edge of the candidate image, the position score is lower.
[0076] It should be noted that the above sample human face angle analysis result can represent the score of the orientation of the human face in the candidate image (i.e., the front and side face score). For example, if the human face faces directly forward, the angle score is the highest; if the human face faces the side, the angle score is lower.
[0077] In the embodiment of the present application, the preset neural network is trained with the sample candidate image labeled with the position labeling result and the angle labeling result of the human face, so that the trained face detection network can accurately determine the sample human face position analysis result and the sample human face angle analysis result of the candidate image, improve the determination accuracy and efficiency of the sample human face position analysis result and the sample human face angle analysis result, and thus improve the generation accuracy and efficiency of the target image for generating the cover of the video to be processed.
[0078] In an optional embodiment, in the above step S103, based on the nose key points in the above candidate image and the first size information of the above candidate image, determining the face position analysis result corresponding to the above candidate image may include: obtaining the first position information of the above nose key points; based on the first position information and the first size information of the above candidate image, determining the face position analysis result. In the above step S105, according to the eye key points and mouth key points in the above candidate image, determining the face angle analysis result corresponding to the above candidate image may include: obtaining the second position information of the above eye key points and the third position information of the above mouth key points; according to the second position information and the third position information, determining the face angle analysis result.
[0079] In one implementation manner, the above step S103 and the above step S105 can also be implemented through a pre-established face image library. That is, according to the first position information of the nose key points of the candidate image and the first size information of the candidate image, obtain the first comparison image that matches the candidate image from the face image library, and use the face position score in the first comparison image as the face position analysis result, where matching may refer to that the position information of the nose key points in the comparison image matches the first position information of the nose key points in the candidate image, and the size information of the comparison image matches the first size information of the selected image. And according to the second position information of the eye key points and the third position information of the mouth key points of the candidate image, obtain the second comparison image that matches the candidate image from the face image library, and use the face angle score of the second comparison image as the face angle analysis result, where matching may refer to that the position information of the eye key points in the comparison image matches the second position information of the eye key points in the candidate image, and the position information of the mouth key points in the comparison image matches the third position information of the mouth key points in the candidate image.
[0080] In another implementation manner, the above step S103 and the above step S105 can also be implemented through a face detection network. That is, obtain the first position information of the nose key points, the second position information of the eye key points, and the third position information of the mouth key points through the face detection network, based on the first position information and the first size information, determine the face position analysis result, and according to the second position information and the third position information, determine the face angle analysis result.
[0081] In the embodiments of the present application, since the key points of the nose can represent the center of the human face, determining the face position analysis result based on the first position information of the key points of the nose and the first size information of the candidate image can accurately analyze the position of the human face in the candidate image and improve the determination accuracy of the face position analysis result. Since the orientations of the eyes and the mouth can represent the orientation of the human face, based on the second position information of the key points of the eyes and the third position information of the key points of the mouth, the angle of the human face in the candidate image can be accurately analyzed, and the determination accuracy of the face angle analysis result can be improved.
[0082] Figure 5 It is a flowchart showing a method for determining a face position analysis result according to an exemplary embodiment. As Figure 5 shown, in a specific embodiment, in the above step S103, the first size information of the candidate image includes the width information and the height information of the candidate image. Based on the key points of the nose in the candidate image and the first size information of the candidate image, determining the face position analysis result corresponding to the candidate image may include:
[0083] S1031. Obtain the first abscissa and the first ordinate included in the position information of the key points of the nose.
[0084] S1033. Based on the first abscissa of the key points of the nose and the width information of the candidate image, determine the first position analysis result corresponding to the candidate image.
[0085] S1035. According to the first ordinate of the key points of the nose and the height information of the candidate image, determine the second position analysis result corresponding to the candidate image.
[0086] S1037. Based on the first position analysis result and the second position analysis result, generate the face position analysis result.
[0087] In an alternative embodiment, the above steps S1031 - S1037 can be performed by a face detection network.
[0088] In one implementation, in the above steps S1031 - S1037, the face position analysis result can be calculated by the following formula:
[0089]
[0090] where Score pos refers to the face position analysis result, refers to the first position analysis result, refers to the second position analysis result, x nose refers to the first abscissa of the key points of the nose, y noseThe first ordinate of the nose key point is referred to as Y, the width information of the candidate image is referred to as W, and the height information of the candidate image is referred to as H.
[0091] In another embodiment, in the above steps S1031 - S1037, the sum of the absolute value of and the absolute value of can also be used as the face position analysis result.
[0092] In another embodiment, in the above steps S1031 - S1037, a first weight corresponding to the first position analysis result and a second weight corresponding to the second position analysis result can also be set. Calculate the first product of the first weight and and the second product of the second weight and The sum of the first product and the second product is used as the face position analysis result.
[0093] In the embodiments of the present application, by generating the face position analysis result in the above manner, not only the relationship between the abscissa of the nose key point and the width information of the candidate image is fully considered, but also the relationship between the ordinate of the nose key point and the height information of the candidate image is fully considered, improving the determination accuracy of the face position analysis result.
[0094] Figure 6 FIG. is a flowchart showing a method for determining a face angle analysis result according to an exemplary embodiment. As Figure 6 shown, in an exemplary embodiment, the above eye key points include the left eye center key point and the right eye center key point, and the above mouth key points include the left mouth corner key point and the right mouth corner key point. In the above step S105, determining the face angle analysis result corresponding to the above candidate image according to the above eye key points and mouth key points in the above candidate image may include:
[0095] S1051. Obtain the coordinate information included in the position information of the above left eye center key point, the coordinate information included in the position information of the above right eye center key point, the coordinate information included in the position information of the above left mouth corner key point, and the coordinate information included in the position information of the above right mouth corner key point.
[0096] S1053. Based on the coordinate information included in the position information of the above left eye center key point and the coordinate information included in the position information of the above right eye center key point, determine the distance between the above left eye center key point and the above right eye center key point, and the second abscissa and second ordinate of the eye center key point; the above eye center key point is the center point between the above left eye center key point and the above right eye center key point.
[0097] S1055. Determine the third abscissa and the third ordinate of the mouth center key point according to the coordinate information included in the position information of the left mouth corner key point and the coordinate information included in the position information of the right mouth corner key point; the mouth center key point is the center point between the left mouth corner key point and the right mouth corner key point.
[0098] S1057. Generate the face angle analysis result based on the distance between the left eye center key point and the right eye center key point, the difference between the second abscissa of the eye center key point and the third abscissa of the mouth center key point, and the difference between the second ordinate and the third ordinate.
[0099] In an alternative embodiment, the above steps S1051 - S1057 can be performed by a face detection network.
[0100] Exemplarily, the eye key points include a left eye center key point and a right eye center key point. In the above step S1051, the coordinate information of the left eye center key point includes the abscissa of the left eye center key point and the ordinate of the left eye center key point, and the coordinate information of the right eye center key point includes the abscissa of the right eye center key point and the ordinate of the right eye center key point. The average value of the abscissa of the left eye center key point and the abscissa of the right eye center key point can be calculated to obtain the second abscissa of the eye center key point, and the average value of the ordinate of the right eye center key point and the ordinate of the left eye center key point can be calculated to obtain the second ordinate of the eye center key point.
[0101] Exemplarily, in the above step S1051, the distance between the left eye center key point and the right eye center key point can be the Euclidean distance.
[0102] Exemplarily, the mouth key points include a left mouth corner key point and a right mouth corner key point. In the above step S1051, the coordinate information of the left mouth corner key point includes the abscissa of the left mouth corner key point and the ordinate of the left mouth corner key point, and the coordinate information of the right mouth corner key point includes the abscissa of the right mouth corner key point and the ordinate of the right mouth corner key point. The third abscissa of the mouth center key point can be determined according to the abscissa of the left mouth corner key point and the abscissa of the right mouth corner key point, and the third ordinate of the mouth center key point can be determined according to the ordinate of the left mouth corner key point and the ordinate of the right mouth corner key point.
[0103] In one implementation, in the above steps S1051 - S1057, the face angle analysis result can be calculated by the following formula:
[0104]
[0105] where Score angleRefers to the face angle analysis result, Dist eye Refers to the distance between the center key point of the left eye and the center key point of the right eye, x eye Refers to the second abscissa of the center key point of the eye, x mouth Refers to the third abscissa of the center key point of the mouth, y eye Refers to the second ordinate of the center key point of the eye, y mouth Refers to the third ordinate of the center key point of the mouth, (x eye -x mouth ) Refers to the difference between the third ordinate of the center key point of the eye and the third abscissa of the center key point of the mouth, (y eye -y mouth ) Refers to the difference between the second ordinate of the center key point of the eye and the third ordinate of the center key point of the mouth, σ refers to a constant, which can be set to 0.00001.
[0106] In another implementation manner, in the above steps S1051 - S1057, the face angle analysis result can be calculated through the following formula:
[0107]
[0108] Wherein, A refers to the weight corresponding to the abscissa of the center key point of the eye and the abscissa of the center key point of the mouth, and B refers to the weight corresponding to the ordinate of the center key point of the eye and the ordinate of the center key point of the mouth.
[0109] Exemplarily, the weight evaluation model can be obtained by pre-training a neural network through the coordinate information of the sample face key points marked with weight tags; the abscissa of the center key point of the eye and the abscissa of the center key point of the mouth are input into the weight evaluation model to obtain the weight corresponding to the abscissa of the center key point of the eye and the abscissa of the center key point of the mouth, and the ordinate of the center key point of the eye and the ordinate of the center key point of the mouth are input into the weight evaluation model to obtain the weight corresponding to the ordinate of the center key point of the eye and the ordinate of the center key point of the mouth.
[0110] In the embodiments of the present application, the face angle analysis result is generated through the distance between the center key point of the left eye and the center key point of the right eye, the difference between the second abscissa of the center key point of the eye and the third abscissa of the center key point of the mouth, and the difference between the second ordinate of the center key point of the eye and the third ordinate of the center key point of the mouth, so that the determination of the face angle analysis result fully considers the eye position, the mouth position, and the relative position relationship between the eye and the mouth, and improves the determination accuracy of the face angle analysis result.
[0111] S107. Based on the above-mentioned face position analysis result and the above-mentioned face angle analysis result, determine the comprehensive face analysis result corresponding to the above-mentioned candidate image.
[0112] In an optional embodiment, in the above S107, the determining the comprehensive face analysis result corresponding to the above-mentioned candidate image based on the above-mentioned face position analysis result and the above-mentioned face angle analysis result may include:
[0113] Perform weighted processing on the above-mentioned face position analysis result and the above-mentioned face angle analysis result based on a weight factor to obtain the above-mentioned comprehensive face analysis result.
[0114] It should be noted that the candidate image may include at least one face. In the case where the candidate image includes multiple faces, a corresponding face position analysis result and a face angle analysis result will be calculated for each face. The weighted sum of the face position analysis result and the face angle analysis result corresponding to each face is obtained to get the comprehensive face analysis result corresponding to each face. The calculation formula may be as follows:
[0115]
[0116] Where, refers to the comprehensive face analysis result corresponding to the i-th face, refers to the face position analysis result corresponding to the i-th face, refers to the face angle analysis result corresponding to the i-th face, α pos refers to the weight of the face position analysis result corresponding to the i-th face, α angle refers to the weight of the face angle analysis result corresponding to the i-th face.
[0117] After obtaining the comprehensive face analysis result corresponding to each face in the candidate image, the comprehensive face analysis results corresponding to each face can be added together to obtain the comprehensive face analysis result corresponding to the candidate image. The calculation formula may be as follows:
[0118]
[0119] Where, refers to the comprehensive face analysis result corresponding to the candidate image, refers to the comprehensive face analysis result corresponding to the i-th face.
[0120] In the embodiments of the present application, taking each face as a unit, based on the face position analysis result and the face angle analysis result corresponding to each face, a comprehensive face analysis result corresponding to the face is calculated, and based on the comprehensive face analysis results corresponding to all the faces in the candidate image, a final comprehensive face analysis result corresponding to the candidate image is calculated, so that the determination of the comprehensive face analysis result comprehensively considers each face in the candidate image, avoiding the defect of low quality of the target image caused by some faces having high scores and some faces having low scores in the candidate image, improving the determination accuracy of the comprehensive face analysis result corresponding to the candidate image, and further improving the determination accuracy of the target image for generating the video cover.
[0121] S109. Determine a target image from the above-mentioned candidate images according to the above-mentioned comprehensive face analysis result; the above-mentioned target image is used to generate the cover of the above-mentioned video to be processed.
[0122] Optionally, in the embodiments of the present application, the target image can be determined from the above-mentioned candidate images in various ways, which are not specifically limited herein.
[0123] In one implementation manner, when there is one candidate image, the candidate image can be used as the target image.
[0124] In another implementation manner, when there are at least two candidate images, the candidate image with the optimal comprehensive face analysis result (for example, the highest score) can be selected from the at least two candidate images according to the comprehensive face analysis results corresponding to the at least two candidate images, as the target image.
[0125] In a third implementation manner, when there are at least two candidate images, the at least two candidate images can be sorted in descending order according to the comprehensive face analysis results corresponding to the at least two candidate images to obtain a candidate image sequence, and the first preset number of candidate images are selected from the candidate image sequence as the target image. Since the number of target images is a preset number, the preset number of target images can be used as the cover of the video to be processed in turn.
[0126] Figure 7 It is a flowchart of another method for determining a video cover shown according to an exemplary embodiment. As Figure 7 shown, in a feasible embodiment, between the above-mentioned step S101 and step S103, the method may further include:
[0127] S102. Analyze the basic quality of the candidate image, and determine whether to execute the above-mentioned S103 according to the analysis result. Specifically, the S102 may include:
[0128] S1021. Perform an optical quality analysis on the above candidate images to obtain the optical quality analysis results of the above candidate images.
[0129] Exemplarily, the above optical quality may include but is not limited to: sharpness, blurriness, exposure, etc.
[0130] Optionally, embodiments of the present application may use various methods to determine the optical quality analysis results of candidate images, and no specific limitation is made here.
[0131] In one implementation, an optical quality analysis network may be pre-trained, and the candidate images are subjected to optical quality analysis through the optical quality analysis network to obtain the optical quality analysis results of the candidate images.
[0132] In another implementation, an image basic quality analysis library may be pre-established, and the optical quality analysis results are labeled for each image in the image basic quality analysis library. For example, the higher the sharpness, the better the optical quality analysis result. The candidate image is matched with the images in the image basic quality analysis library to obtain the target image in the image basic quality analysis library that matches the candidate image, and the optical quality analysis result labeled in the target image is used as the optical quality analysis result of the candidate image.
[0133] In a third method, taking the optical quality as sharpness as an example, the candidate image may first be converted into a grayscale image, the two gradients in the x and y directions of the grayscale image are calculated, then the sum of the squares of the two gradients in the x and y directions is calculated, and the sharpness value S is obtained by dividing the sum of the squares by the total number of pixels in the candidate image. The sharpness value S is used as the optical quality analysis result.
[0134] S1023. When the above optical quality analysis results meet the first preset condition, perform face detection on the above candidate images. When the detection results indicate that the candidate images include face images and the ratio between the second size information of the face images and the above first size information meets the second preset condition, execute the step of determining the face position analysis result corresponding to the above candidate images based on the nose key points in the above candidate images and the above first size information of the above candidate images in S103.
[0135] Optionally, the first preset condition may be a quality threshold characterizing the optical quality, which may be set according to the optical quality analysis results. For example, when the optical quality is sharpness, the quality threshold may be the sharpness threshold T s . When the sharpness value S is greater than the sharpness threshold T s , perform face detection on the above candidate images.
[0136] Optionally, the candidate image can be subjected to face detection through the above-mentioned face detection network. If the number of face bounding boxes obtained is greater than 0, it indicates the presence of a face, and subsequent steps can be performed.
[0137] Optionally, the second preset condition can be a size threshold representing size information, which can be set according to the size of the candidate image and the face. For example, if the width and height of the candidate image are W and H respectively, and the width and height of the i-th face in the candidate image are w i and h i , then calculate the ratio of the i-th face to the image (w i +h i ) / (W + H). If the ratio is greater than the threshold T f , it is considered that the i-th face is valid, and further calculate the face position analysis result and face angle analysis result of the i-th face. Otherwise, it is considered that the i-th face is invalid and continue to calculate the next face.
[0138] In an alternative embodiment, when the optical quality analysis result does not meet the first preset condition, the step of obtaining the face key points in the above candidate image is not performed, and the score of the candidate image is set to the lowest. For example, if the optical quality is sharpness, the quality threshold can be the sharpness threshold T s . When the sharpness value S is less than or equal to the sharpness threshold T s , face detection is not performed on the candidate image, and the score of the candidate picture is set to the lowest Score min .
[0139] In another alternative embodiment, when the optical quality analysis result meets the first preset condition, face detection is performed on the above candidate image. When the detection result indicates that the above candidate image does not include a face image, the step of judging the size information is not performed.
[0140] In another alternative embodiment, when the optical quality analysis result meets the first preset condition, face detection is performed on the above candidate image. When the detection result indicates that the above candidate image includes a face image, but the ratio between the second size information of the above face image and the above first size information does not meet the second preset condition, the step of determining the face position analysis result and face angle analysis result corresponding to the above candidate image based on the position information of the face key points in the above candidate image and the first size information of the above candidate image is not performed.
[0141] In the embodiments of the present application, before determining the face position analysis result and the face angle analysis result corresponding to the candidate image, a multi-layer conditional judgment step is set. The first layer of judgment is whether the optical quality analysis result meets the first preset condition to eliminate images with poor optical quality; and when the first preset condition is met, a second layer of judgment is made on whether the candidate image contains a face image to eliminate images that do not contain a face; and when it is met, a third layer of judgment is made on whether the ratio between the second size information of the face image and the above-mentioned first size information meets the second preset condition to eliminate invalid faces with size information not meeting the requirements. Only when all the above three layers of judgments are met, will the steps of determining the face position analysis result and the face angle analysis result corresponding to the candidate image be executed, thereby improving the accuracy of the face position analysis result and the face angle analysis result corresponding to the candidate image, and further improving the determination accuracy of the target image used to generate the cover of the video to be processed.
[0142] Figure 8 is the flow of the video cover determination method shown according to an exemplary embodiment Figure 2 As Figure 8 shown, in an optional embodiment, after determining the comprehensive face analysis result corresponding to the candidate image based on the above-mentioned face position analysis result and the above-mentioned face angle analysis result in step S107 and before step S109, the method may further include:
[0143] S108. Analyze the aesthetics of the candidate image to obtain the aesthetics analysis result of the candidate image.
[0144] Optionally, since images that meet the aesthetic degree are more suitable for use as video covers. For example, images with reasonable composition are more suitable for use as video covers. Therefore, a picture aesthetics network can be pre-trained, and the candidate image is input into the picture aesthetics network to obtain the aesthetics analysis result of the candidate image. This aesthetics analysis result can represent the aesthetic attribute score of the candidate image.
[0145] In an optional embodiment, the process of training the picture aesthetics network can be as follows: First, collect some training sample pictures, and then divide them into three grades of low, medium, and high according to the aesthetics of the training sample pictures. The pictures with aesthetics from low to high are scored 0, 1, and 2 respectively. Use the scored training sample pictures to train the picture aesthetics network, and a picture classification model, such as resnet, etc., can be used for score regression training. After the network training is completed, it can be used for subsequent processes to score the picture aesthetics.
[0146] Correspondingly, in step S109 above, determining the target image from the candidate images according to the above-mentioned comprehensive face analysis result; the target image is used to generate the cover of the above-mentioned video to be processed, and may include:
[0147] S1091. Based on the above aesthetic analysis result and the above comprehensive face analysis result, determine the target analysis result of the above candidate image.
[0148] S1093. Determine the above target image from the above candidate images according to the above target analysis result.
[0149] In one implementation, in the above S1091, the sum of the aesthetic analysis result and the comprehensive face analysis result can be used as the target analysis result of the candidate image.
[0150] In another implementation, in the above S1091, the aesthetic analysis result and the comprehensive face analysis result can also be weighted and added to obtain the target analysis result of the candidate image. The specific formula can be as follows:
[0151]
[0152] Among them, Score image refers to the target analysis result of the candidate image, Score face refers to the comprehensive face analysis result of the candidate image, refers to the aesthetic analysis result of the candidate image, β face refers to the weight of the comprehensive face analysis result, refers to the weight of the aesthetic analysis result.
[0153] Optionally, in the above step S1093, the process of determining the above target image from the above candidate images according to the above target analysis result is similar to the method of determining the target image from the above candidate images according to the above comprehensive face analysis result, and will not be elaborated here.
[0154] Figure 9 is the flowchart of a video cover determination method shown according to an exemplary embodiment Figure 3 . As Figure 9 shown, the embodiments of the present application are based on the aesthetic analysis result and the comprehensive face analysis result, and determine the target image for generating the cover of the to-be-processed video from the candidate images, which not only shows better character information, but also makes the cover image more aesthetic, makes the target image for generating the cover more representative, and improves the determination accuracy of the target image for generating the cover.
[0155] In an alternative embodiment, to ensure that the video cover meets the requirements of network civilization, based on the above embodiments, in the embodiments of the present application, before determining the face position analysis result and the face angle analysis result corresponding to the candidate image based on the position information of the face key points in the candidate image and the first size information of the candidate image, the method may further include:
[0156] Train an image violation detection network.
[0157] Based on the image violation detection network, determine whether the candidate image is in violation.
[0158] If it is determined that the candidate image is not in violation, then perform the step of determining the face position analysis result and the face angle analysis result corresponding to the candidate image based on the position information of the face key points in the candidate image and the first size information of the candidate image.
[0159] In the embodiments of the present application, to avoid the adverse effects brought by violation pictures, an image violation detection network may be pre-trained to determine whether the candidate image is in violation through the image violation detection network. Among them, the training process of the violation detection network may be as follows: Obtain a sample image, and the sample image is labeled with a violation label indicating whether the image is in violation; train a neural network with the sample image. During the training process, determine a loss value based on the output result of the neural network and the violation label, and adjust the network parameters of the neural network according to the loss value until the network parameters meet the preset conditions to obtain the violation detection network.
[0160] In a possible implementation manner, the probability value indicating whether the candidate image is in violation may be output through the image violation detection network. The probability value represents the possibility of the candidate image being in violation, and the larger the value, the greater the possibility of the picture being in violation.
[0161] In another possible implementation manner, the identification value indicating whether the candidate image is in violation may also be directly output through the image violation detection network. Based directly on the identification value output by the violation detection network, it can be determined whether the candidate image is in violation. For example, if the output identification value is "1", it indicates that the candidate image is in violation, and if the output identification value is "0", it indicates that the candidate image is not in violation.
[0162] In an alternative embodiment, the method may further include:
[0163] Train an image interest network.
[0164] Input the candidate image into the image interest network to obtain the interest analysis result of the candidate image (for example, an interest score); the interest analysis result is used to characterize the degree of interest of the client account in the candidate image.
[0165] Accordingly, based on the above-mentioned comprehensive face analysis results, a target image is determined from the above-mentioned candidate images; the above-mentioned target image is used to generate the cover of the above-mentioned video to be processed, and may include:
[0166] Based on the aesthetics analysis result, the comprehensive face analysis result, and the interest analysis result, determine the target analysis result of the above-mentioned candidate images; or based on the comprehensive face analysis result and the interest analysis result, determine the target analysis result of the above-mentioned candidate images.
[0167] According to the above-mentioned target analysis result, determine the above-mentioned target image from the above-mentioned candidate images.
[0168] In a feasible implementation manner, the training process of the above-mentioned image interest network may be as follows:
[0169] Input multiple sample images labeled with interest scores into the deep learning network to be trained, compare the interest scores of each sample image output by the deep learning network with the actually labeled interest scores of each sample image, and continuously adjust the network parameters of the deep learning network until the network parameters meet the preset conditions to obtain the image interest network.
[0170] In the implementation of this application, the target image for generating the video cover is determined based on the interest analysis result. Since the interest score of the image can be used to reflect the degree of interest of the client's account in the image, therefore, based on the interest scores of the multiple frames of images, selecting the target image for generating the video cover is beneficial to selecting an image from the video that can better reflect the content of interest to the client's account in the video as the video cover, so that the generated video cover has a higher degree of attraction to the client's account, and thus is beneficial to improving the click-through rate of the generated video cover. At the same time, since there is also a positive correlation between the degree of interest of the client's account in the image and the wonderfulness of the image, therefore, by the method of this application embodiment, while selecting the image of interest to the client's account from the candidate images as the video cover, it is actually also beneficial to select a more wonderful image in the video as the video cover, thereby being beneficial to improving the attraction of the video cover.
[0171] Figure 10 It is a comparison diagram of the cover image selected by using the method of this application embodiment shown in an exemplary embodiment and the cover image selected by using the existing method. As Figure 10 shown, the quality of the cover image selected by using the method of this application embodiment is higher, and it can better represent the content of the video to be processed.
[0172] In a feasible embodiment, for the video cover determination method disclosed in the present application, data such as the face position analysis result, face angle analysis result, face comprehensive analysis result, and target image can be stored on the blockchain.
[0173] Figure 11 It is a block diagram of a video cover determination device shown according to an exemplary embodiment. As Figure 12 shown, the device may at least include:
[0174] A candidate image acquisition module 301, configured to acquire candidate images from the image sequence of the video to be processed.
[0175] A position analysis result determination module 303, configured to determine the face position analysis result corresponding to the candidate image based on the nose key point in the candidate image and the first size information of the candidate image.
[0176] An angle analysis result determination module 305, configured to determine the face angle analysis result corresponding to the candidate image according to the eye key point and the mouth key point in the candidate image.
[0177] A comprehensive analysis result determination module 307, configured to determine the face comprehensive analysis result corresponding to the candidate image based on the face position analysis result and the face angle analysis result.
[0178] A target image determination module 309, configured to determine a target image from the candidate images according to the face comprehensive analysis result; the target image is used to generate the cover of the video to be processed.
[0179] In an optional embodiment, the position analysis result determination module 303 includes:
[0180] A nose coordinate determination unit, configured to acquire the first abscissa and the first ordinate included in the position information of the nose key point.
[0181] A first position analysis result determination unit, configured to determine the first position analysis result corresponding to the candidate image based on the first abscissa of the nose key point and the width information of the candidate image.
[0182] A second position analysis result determination unit, configured to determine the second position analysis result corresponding to the candidate image according to the first ordinate of the nose key point and the height information of the candidate image.
[0183] A face position analysis result generation unit, configured to generate the face position analysis result based on the first position analysis result and the second position analysis result.
[0184] In an optional embodiment, the above-mentioned eye key points include the left eye center key point and the right eye center key point, the above-mentioned mouth key points include the left mouth corner key point and the right mouth corner key point, and the above-mentioned angle analysis result determination module 305 includes:
[0185] An eye-mouth coordinate acquisition unit, configured to acquire the coordinate information included in the position information of the above-mentioned left eye center key point, the coordinate information included in the position information of the above-mentioned right eye center key point, the coordinate information included in the position information of the above-mentioned left mouth corner key point, and the coordinate information included in the position information of the above-mentioned right mouth corner key point.
[0186] A distance center point determination unit, configured to determine the distance between the above-mentioned left eye center key point and the above-mentioned right eye center key point, as well as the second abscissa and the second ordinate of the eye center key point, based on the coordinate information included in the position information of the above-mentioned left eye center key point and the coordinate information included in the position information of the above-mentioned right eye center key point; the above-mentioned eye center key point is the center point between the above-mentioned left eye center key point and the above-mentioned right eye center key point.
[0187] A mouth coordinate determination unit, configured to determine the third abscissa and the third ordinate of the mouth center key point according to the coordinate information included in the position information of the above-mentioned left mouth corner key point and the coordinate information included in the position information of the above-mentioned right mouth corner key point; the above-mentioned mouth center key point is the center point between the above-mentioned left mouth corner key point and the above-mentioned right mouth corner key point.
[0188] A face angle analysis result generation unit, configured to generate the above-mentioned face angle analysis result based on the distance between the above-mentioned left eye center key point and the above-mentioned right eye center key point, the difference between the second abscissa of the above-mentioned eye center key point and the third abscissa of the above-mentioned mouth center key point, and the difference between the second ordinate of the above-mentioned eye center key point and the third ordinate of the above-mentioned mouth center key point.
[0189] In an optional embodiment, the above-mentioned device further includes:
[0190] An optical analysis module, configured to perform optical quality analysis on the above-mentioned candidate image to obtain the optical quality analysis result of the above-mentioned candidate image.
[0191] An execution module, configured to perform face detection on the above-mentioned candidate image when the above-mentioned optical quality analysis result meets the first preset condition, and when the detection result indicates that the above-mentioned candidate image includes a face image and the ratio between the second size information of the above-mentioned face image and the above-mentioned first size information meets the second preset condition, execute the step of determining the above-mentioned face position analysis result corresponding to the above-mentioned candidate image based on the nose key point in the above-mentioned candidate image and the above-mentioned first size information of the above-mentioned candidate image.
[0192] In an alternative embodiment, the above-mentioned device further includes:
[0193] A beauty analysis module, configured to analyze the beauty of the above-mentioned candidate images to obtain the beauty analysis result of the above-mentioned candidate images.
[0194] Correspondingly, the target image determination module may include:
[0195] A target analysis result determination unit, configured to determine the target analysis result of the above-mentioned candidate images based on the above-mentioned beauty analysis result and the above-mentioned comprehensive face analysis result.
[0196] A target analysis result analysis unit, configured to determine the above-mentioned target image from the above-mentioned candidate images according to the above-mentioned target analysis result.
[0197] In an alternative embodiment, the above-mentioned position analysis result determination module includes:
[0198] An input unit, configured to input the above-mentioned candidate images into a face detection network.
[0199] A position and size processing unit, configured to process the above-mentioned nose key points and the first size information of the above-mentioned candidate images based on the above-mentioned face detection network to obtain the above-mentioned face position analysis result.
[0200] Correspondingly, the above-mentioned angle analysis result determination module includes:
[0201] An eye and nose key point processing unit, configured to process the above-mentioned eye key points and the above-mentioned mouth key points based on the above-mentioned face detection network to obtain the above-mentioned face angle analysis result.
[0202] In an alternative embodiment, the above-mentioned device further includes:
[0203] A sample candidate image acquisition module, configured to acquire labeled sample candidate images from a sample video to be processed; the above-mentioned label represents the position annotation result and the angle annotation result of the sample face included in the above-mentioned sample candidate images.
[0204] A sample position and angle determination module, configured to process the sample nose key points and the size information of the above-mentioned sample candidate images based on a preset neural network to obtain the sample face position analysis result and the sample face angle analysis result corresponding to the above-mentioned sample candidate images, and to process the sample eye key points and the sample mouth key points in the above-mentioned sample candidate images based on the above-mentioned preset neural network to obtain the sample face angle analysis result corresponding to the above-mentioned sample candidate images.
[0205] A loss information determination module, configured to determine the loss information of the sample candidate image based on the above-mentioned sample face position analysis result, the above-mentioned position annotation result, the above-mentioned sample face angle analysis result, and the above-mentioned angle annotation result.
[0206] A training module, configured to train the above-mentioned preset neural network according to the above-mentioned loss information to obtain the above-mentioned face detection network.
[0207] It should be noted that the device embodiment provided in the embodiment of the present application and the above-mentioned method embodiment are based on the same inventive concept.
[0208] The embodiment of the present application also provides an electronic device for determining a video cover. The electronic device includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and at least one instruction or at least one program segment is loaded and executed by the processor to implement the video cover determination method provided in any of the above embodiments.
[0209] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be set in a terminal to store at least one instruction or at least one program segment related to implementing a video cover determination method or a network training method for video cover determination in the method embodiment. At least one instruction or at least one program segment is loaded and executed by the processor to implement the video cover determination method provided in the above method embodiment.
[0210] Optionally, in the embodiment of the present specification, the storage medium may be located in at least one network server among multiple network servers of a computer network. Optionally, in this embodiment, the above-mentioned storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disc and other media that can store program codes.
[0211] The memory in the embodiment of the present specification can be used to store software programs and modules. The processor runs the software programs and modules stored in the memory to execute various functional application programs and data processing. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory may also include a memory controller to provide the processor with access to the memory.
[0212] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video cover determination method provided in the foregoing method embodiment.
[0213] The video cover determination method embodiment provided by the embodiment of the present application can be executed on a terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 12 is a hardware structure block diagram of a server for video cover determination provided by an embodiment of the present application. As Figure 12 shown, the server 400 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 410 (the central processing unit 410 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 430 for storing data, and one or more storage media 420 (such as one or more mass storage devices) for storing application programs 423 or data 422. Among them, the memory 430 and the storage media 420 may be transient storage or persistent storage. The program stored in the storage media 420 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processing unit 410 may be configured to communicate with the storage media 420 and execute a series of instruction operations in the storage media 420 on the server 400. The server 400 may also include one or more power supplies 460, one or more wired or wireless network interfaces 450, one or more input / output interfaces 440, and / or one or more operating systems 421, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0214] The input / output interface 440 may be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server 400. In one example, the input / output interface 440 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one example, the input / output interface 440 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0215] Those of ordinary skill in the art will understand that Figure 12 the structure shown is merely illustrative and does not limit the structure of the above-mentioned electronic device. For example, server 400 may also include more or fewer components than Figure 12 shown therein, or have a different configuration from Figure 12 that shown.
[0216] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of the present specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0217] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0218] Those of ordinary skill in the art will understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.
[0219] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining a video cover, characterized in that, the method includes: Obtaining candidate images from the image sequence of the video to be processed; When the optical quality analysis result of the candidate image meets the first preset condition, based on the nose key points in the candidate image and the first size information of the candidate image, determining the face position analysis result corresponding to the candidate image; Generating a face angle analysis result based on the square sum of the difference in the abscissa between the eye center key point and the mouth center key point, the difference in the ordinate between the eye center key point and the mouth center key point, and the distance between the left eye center key point and the right eye center key point; the eye center key point is the center point between the left eye center key point and the right eye center key point; the mouth center key point is the center point between the left mouth corner key point and the right mouth corner key point; Based on the face position analysis result and the face angle analysis result, determining the face comprehensive analysis result corresponding to the candidate image; Inputting the candidate image into an image interest network to obtain the interest analysis result of the candidate image; the interest analysis result is used to characterize the degree of interest of the client account in the candidate image; Determining a target image from the candidate images according to the face comprehensive analysis result and the interest analysis result; the target image is used to generate the cover of the video to be processed.
2. The method according to claim 1, characterized in that, the first size information of the candidate image includes the width information and height information of the candidate image; the determining the face position analysis result corresponding to the candidate image based on the nose key points in the candidate image and the first size information of the candidate image includes: Obtaining the first abscissa and the first ordinate included in the position information of the nose key point; Based on the first abscissa of the nose key point and the width information of the candidate image, determining the first position analysis result corresponding to the candidate image; According to the first ordinate of the nose key point and the height information of the candidate image, determining the second position analysis result corresponding to the candidate image; Based on the first position analysis result and the second position analysis result, generating the face position analysis result.
3. The method according to claim 1, characterized in that, the method further includes: Obtaining the coordinate information included in the position information of the left eye center key point, the coordinate information included in the position information of the right eye center key point, the coordinate information included in the position information of the left mouth corner key point, and the coordinate information included in the position information of the right mouth corner key point; Based on the coordinate information included in the position information of the left eye center key point and the coordinate information included in the position information of the right eye center key point, determining the distance between the left eye center key point and the right eye center key point, and the second abscissa and the second ordinate of the eye center key point; Determine the third abscissa and the third ordinate of the mouth center key point according to the coordinate information included in the position information of the left mouth corner key point and the coordinate information included in the position information of the right mouth corner key point; Generate the face angle analysis result based on the distance between the left eye center key point and the right eye center key point, the difference between the second abscissa of the eye center key point and the third abscissa of the mouth center key point, and the difference between the second ordinate of the eye center key point and the third ordinate of the mouth center key point.
4. The method according to claim 1, wherein, after obtaining the candidate image from the image sequence of the video to be processed, the method further includes: Performing optical quality analysis on the candidate image to obtain the optical quality analysis result of the candidate image; When the optical quality analysis result meets the first preset condition, performing face detection on the candidate image, and when the detection result indicates that the candidate image includes a face image and the ratio between the second size information of the face image and the first size information meets the second preset condition, performing the step of determining the face position analysis result corresponding to the candidate image based on the nose key point in the candidate image and the first size information of the candidate image.
5. The method according to claim 1, wherein, the method further includes: Analyzing the aesthetic degree of the candidate image to obtain the aesthetic degree analysis result of the candidate image; Determining the target analysis result of the candidate image based on the aesthetic degree analysis result and the comprehensive face analysis result; Determining the target image from the candidate images according to the target analysis result.
6. The method according to any one of claims 1 to 5, wherein, the determining the face position analysis result corresponding to the candidate image based on the nose key point in the candidate image and the first size information of the candidate image includes: Inputting the candidate image into a face detection network; Based on the face detection network, processing the nose key point and the first size information of the candidate image to obtain the face position analysis result; the method further includes: Based on the face detection network, processing the eye key points and the mouth key points to obtain the face angle analysis result.
7. The method according to claim 6, wherein, the method further includes: Obtaining a labeled sample candidate image from a sample video to be processed; the label represents the position annotation result and the angle annotation result of the sample face included in the sample candidate image; Based on a preset neural network, processing the sample nose key point in the sample candidate image and the size information of the sample candidate image to obtain the sample face position analysis result and the sample face angle analysis result corresponding to the sample candidate image, and based on the preset neural network, processing the sample eye key points and the sample mouth key points in the sample candidate image to obtain the sample face angle analysis result corresponding to the sample candidate image; Determine the loss information of the sample candidate image based on the sample face position analysis result, the position annotation result, the sample face angle analysis result, and the angle annotation result; Train the preset neural network according to the loss information to obtain the face detection network.
8. A video cover determination device Characterized in that The device includes: A candidate image acquisition module, configured to acquire candidate images from an image sequence of a video to be processed; A position analysis result determination module, configured to determine the face position analysis result corresponding to the candidate image based on the nose key point in the candidate image and the first size information of the candidate image when the optical quality analysis result of the candidate image meets a first preset condition; An angle analysis result determination module, configured to generate a face angle analysis result based on the sum of the squares of the difference in the abscissa between the eye center key point and the mouth center key point, the difference in the ordinate between the eye center key point and the mouth center key point, and the distance between the left eye center key point and the right eye center key point; the eye center key point is the center point between the left eye center key point and the right eye center key point; the mouth center key point is the center point between the left mouth corner key point and the right mouth corner key point; A comprehensive analysis result determination module, configured to determine the face comprehensive analysis result corresponding to the candidate image based on the face position analysis result and the face angle analysis result; input the candidate image into an image interest network to obtain the interest analysis result of the candidate image; the interest analysis result is used to characterize the degree of interest of the client account in the candidate image; A target image determination module, configured to determine a target image from the candidate images according to the face comprehensive analysis result and the interest analysis result; the target image is used to generate the cover of the video to be processed.
9. The device according to claim 8, Characterized in that The position analysis result determination module includes: A nose coordinate determination unit, configured to obtain the first abscissa and the first ordinate included in the position information of the nose key point; A first position analysis result determination unit, configured to determine the first position analysis result corresponding to the candidate image based on the first abscissa of the nose key point and the width information of the candidate image; A second position analysis result determination unit, configured to determine the second position analysis result corresponding to the candidate image according to the first ordinate of the nose key point and the height information of the candidate image; A face position analysis result generation unit, configured to generate the face position analysis result based on the first position analysis result and the second position analysis result.
10. The device according to claim 8, Characterized in that The angle analysis result determination module includes: An eye and mouth coordinate acquisition unit, configured to acquire the coordinate information included in the position information of the left eye center key point, the coordinate information included in the position information of the right eye center key point, the coordinate information included in the position information of the left mouth corner key point, and the coordinate information included in the position information of the right mouth corner key point; A distance center point determination unit, configured to determine the distance between the left eye center key point and the right eye center key point, as well as the second abscissa and the second ordinate of the eye center key point, based on the coordinate information included in the position information of the left eye center key point and the coordinate information included in the position information of the right eye center key point; A mouth coordinate determination unit, configured to determine the third abscissa and the third ordinate of the mouth center key point according to the coordinate information included in the position information of the left mouth corner key point and the coordinate information included in the position information of the right mouth corner key point; A face angle analysis result generation unit, configured to generate the face angle analysis result based on the distance between the left eye center key point and the right eye center key point, the difference between the second abscissa of the eye center key point and the third abscissa of the mouth center key point, and the difference between the second ordinate of the eye center key point and the third ordinate of the mouth center key point.
11. The apparatus according to claim 8, wherein, the apparatus further comprises: An optical analysis module, configured to perform optical quality analysis on the candidate image to obtain an optical quality analysis result of the candidate image; An execution module, configured to perform face detection on the candidate image when the optical quality analysis result meets the first preset condition, and execute the step of determining the face position analysis result corresponding to the candidate image based on the nose key point in the candidate image and the first size information of the candidate image when the detection result indicates that the candidate image includes a face image and the ratio between the second size information of the face image and the first size information meets the second preset condition.
12. The apparatus according to claim 8, wherein, the apparatus further comprises: An aesthetics analysis module, configured to analyze the aesthetics of the candidate image to obtain an aesthetics analysis result of the candidate image; A target analysis result determination unit, configured to determine the target analysis result of the candidate image based on the aesthetics analysis result and the face comprehensive analysis result; A target analysis result analysis unit, configured to determine the target image from the candidate image according to the target analysis result.
13. The apparatus according to claim 13, wherein, the position analysis result determination module comprises: An input unit, configured to input the candidate image into a face detection network; A position size processing unit, configured to process the nose key point and the first size information of the candidate image based on the face detection network to obtain the face position analysis result; the angle analysis result determination module comprises: An eye-nose key point processing unit, configured to process the eye key point and the mouth key point based on the face detection network to obtain the face angle analysis result.
14. The apparatus according to claim 13, wherein, the apparatus further comprises: A sample candidate image acquisition module, configured to acquire a labeled sample candidate image from a sample video to be processed; the label represents the position annotation result and the angle annotation result of the sample human face included in the sample candidate image; A sample position and angle determination module, configured to process the sample nose key points and the size information of the sample candidate image in the sample candidate image based on a preset neural network to obtain a sample human face position analysis result and a sample human face angle analysis result corresponding to the sample candidate image, and process the sample eye key points and the sample mouth key points in the sample candidate image based on the preset neural network to obtain a sample human face angle analysis result corresponding to the sample candidate image; A loss information determination module, configured to determine the loss information of the sample candidate image based on the sample human face position analysis result, the position annotation result, the sample human face angle analysis result, and the angle annotation result; A training module, configured to train the preset neural network according to the loss information to obtain the face detection network.
15. An electronic device for determining a video cover, characterized in that, the electronic device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the video cover determination method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, at least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the video cover determination method according to any one of claims 1 to 7.
17. A computer program product, including a computer program, characterized in that, the computer program, when executed by a processor, implements the video cover determination method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image processing method and device, electronic device and computer readable storage medium
CN110610171A
Video material image selection method and device, equipment and storage medium
CN113822136A
Method and device for generating driven image, computer equipment and storage medium
CN113903062A
Video cover recommendation method and device, video cover generation method and device, equipment and storage medium
CN113918763A