A video processing method and related devices
By performing expression detection and emotional analysis in long videos, emotional short video clips are extracted, which solves the problem of how to extract wonderful short videos from long videos, and improves the viewing rate and productivity of the video.
Patent Information
- Application Number
- CN202210588529.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-27
AI Technical Summary
How to extract wonderful short video clips from long videos to improve the viewing rate of long videos and the productivity of short videos.
By obtaining image frames of expressions to be detected in long videos, using a preset expression detection model for expression detection, determining the expression confidence in the video, and dividing and extracting short videos based on emotional confidence.
It realizes accurate identification of emotionally excited video clips in long videos, improving the production efficiency of short videos and the traffic promotion effect of long videos.
Smart Images

Figure CN114973366B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a video processing method and related devices. Background Art
[0002] In recent years, with the popularization of mobile terminals and the acceleration of network speed, short, flat, and fast large-flow propagation content has gradually gained people's favor. Among them, short videos, as video content that can be played on various new media platforms and is suitable for viewing in a mobile state and short leisure time, attract a large number of users in all age groups and educational levels.
[0003] For the video industry, by editing the wonderful content in long videos into short videos and using the short videos to drive the viewing of long videos, the viewing rate of long videos can be significantly increased. Therefore, how to obtain wonderful short videos from long videos has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a video processing method and related devices that overcome or at least partially solve the above problems. The technical solutions include:
[0005] A video processing method includes:
[0006] Obtaining at least one first image frame of a to-be-detected expression in a first video;
[0007] Inputting the first image frame into a preset expression detection model to obtain an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of a face image in the first image frame and a first expression confidence corresponding to the face expression result;
[0008] Determining a plurality of second videos in the first video by using at least a sliding window with a preset video time length;
[0009] For any one of the second videos: determining an emotion confidence corresponding to the second video by using the first expression confidences corresponding to the respective first image frames included in the second video;
[0010] Obtaining a preset number of third videos from the second videos according to the emotion confidence.
[0011] Optionally, the obtaining at least one first image frame of a to-be-detected expression in a first video includes:
[0012] Determining a plurality of second image frames to be face-detected in the first video at a preset image frame interval;
[0013] Input the second image frame into a preset multi-task face detection model to obtain the face detection result output by the preset multi-task face detection model, where the face detection result includes the image ratio of the face in the face image in the second image frame, the face confidence, and the face angle;
[0014] Perform convolution calculation on the second image frame using a harmonic operator to obtain the sharpness of the second image frame;
[0015] Use the image ratio of the face, the face confidence, the face angle, and the sharpness to obtain at least one first image frame of the expression to be detected in each of the second image frames.
[0016] Optionally, the step of using the image ratio of the face, the face confidence, the face angle, and the sharpness to obtain at least one first image frame of the expression to be detected in each of the second image frames includes:
[0017] Screen out at least one third image frame in each of the second image frames where the image ratio of the face is not less than a preset ratio threshold;
[0018] Screen out at least one fourth image frame in each of the third image frames where the face confidence is not less than a preset face confidence threshold;
[0019] Screen out at least one fifth image frame in each of the fourth image frames where the face angle meets a preset angle condition;
[0020] Screen out at least one first image frame of the expression to be detected in each of the fifth image frames where the sharpness is not less than a preset sharpness threshold.
[0021] Optionally, the step of inputting the first image frame into a preset expression detection model to obtain the expression detection result output by the preset expression detection model includes:
[0022] Input the first image frame into a preset expression detection model so that the preset expression detection model performs multi-expression type detection on the face image in the first image frame to obtain the second expression confidence corresponding to each expression of the face image under the preset expression types, determine the second expression confidence with the highest value as the first expression confidence, and determine the expression corresponding to the first expression confidence as the face expression result of the face image in the first image frame;
[0023] Obtain the expression detection result output by the preset expression detection model, including the first expression confidence and the face expression result.
[0024] Optionally, determining multiple second videos in the first video by using at least a sliding window with a preset video time length includes:
[0025] Determining multiple fourth videos in the first video by using a sliding window with a preset video time length;
[0026] Using a non-maximum suppression algorithm to remove overlapping videos in each of the fourth videos, and determining multiple second videos, where the overlap value between each of the second videos is less than a preset overlap threshold.
[0027] Optionally, obtaining a preset number of third videos in each of the second videos according to the emotion confidence includes:
[0028] Obtaining a preset number of third videos in each of the second videos in the order from largest to smallest of the emotion confidence.
[0029] Optionally, after obtaining a preset number of third videos in each of the second videos according to the emotion confidence, the method further includes:
[0030] Stitching each of the third videos to obtain a fifth video.
[0031] A video processing device includes: a first obtaining unit, a second obtaining unit, a first determining unit, a second determining unit, and a third obtaining unit.
[0032] The first obtaining unit is configured to obtain at least one first image frame of a to-be-detected expression in a first video.
[0033] The second obtaining unit is configured to input the first image frame into a preset expression detection model, and obtain an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of a face image in the first image frame and a first expression confidence corresponding to the face expression result.
[0034] The first determining unit is configured to determine multiple second videos in the first video by using at least a sliding window with a preset video time length.
[0035] The second determining unit is configured to, for any one of the second videos: determine an emotion confidence corresponding to the second video by using the first expression confidence corresponding to each of the first image frames included in the second video.
[0036] The third obtaining unit is configured to obtain a preset number of third videos in each of the second videos according to the emotion confidence.
[0037] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the video processing method described in any one of the above is implemented.
[0038] An electronic device, the electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory complete communication with each other through the bus; the processor is used to call program instructions in the memory to execute the video processing method described in any one of the above.
[0039] By means of the above technical solution, a video processing method and related devices provided by the present disclosure can obtain at least one first image frame of the expression to be detected in the first video; input the first image frame into a preset expression detection model to obtain an expression detection result output by the preset expression detection model, and the expression detection result includes the face expression result of the face image in the first image frame and the first expression confidence corresponding to the face expression result; at least use a sliding window of a preset video time length to determine multiple second videos in the first video; for any second video: use the first expression confidence corresponding to each first image frame included in the second video to determine the emotion confidence corresponding to the second video; according to the emotion confidence, obtain a preset number of third videos in each second video. The present disclosure can accurately identify video segments with excited emotions of characters in a long video through emotion confidence, which helps to improve the production efficiency of short videos and the drainage and promotion of long videos.
[0040] The above description is only an overview of the technical solution of the present disclosure. In order to be able to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically exemplified below. Description of the Drawings
[0041] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present disclosure. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0042] Figure 1 Shows a schematic flowchart of an implementation manner of the video processing method provided by an embodiment of the present disclosure;
[0043] Figure 2 Shows a schematic flowchart of an implementation manner of step S100 in the video processing method provided by an embodiment of the present disclosure;
[0044] Figure 3A schematic flowchart showing another implementation of step S100 in the video processing method provided by an embodiment of the present disclosure;
[0045] Figure 4 A schematic flowchart showing another implementation of the video processing method provided by an embodiment of the present disclosure;
[0046] Figure 5 A schematic flowchart showing another implementation of the video processing method provided by an embodiment of the present disclosure;
[0047] Figure 6 A schematic structural diagram of a video processing device provided by an embodiment of the present disclosure;
[0048] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0049] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0050] As Figure 1 shown, a schematic flowchart of an implementation manner of the video processing method provided by an embodiment of the present disclosure. The video processing method may include:
[0051] S100. Obtain at least one first image frame of the expression to be detected in the first video.
[0052] Wherein, the first video may be a video with a duration of more than half an hour, also known as a long video. The long video may be mainly a video mainly composed of movies and TV dramas.
[0053] Optionally, an embodiment of the present disclosure may split the first video into image frames, and then select and extract the first image frames from each image frame.
[0054] Optionally, an embodiment of the present disclosure may extract the first image frames from each image frame at a preset image frame interval. For example: an embodiment of the present disclosure may determine the image frames extracted every 5 frames in the first video as the first image frames.
[0055] Optionally, an embodiment of the present disclosure may perform video preprocessing on the first video to screen out the first image frames in the first video. Optionally, based on Figure 1 the method shown, as Figure 2As shown in the figure, it is a schematic flowchart of an implementation manner of step S100 in the video processing method provided by the embodiments of the present disclosure. Step S100 may include:
[0056] S110. Determine a plurality of second image frames to be face-detected in the first video according to a preset image frame interval.
[0057] Optionally, in the embodiments of the present disclosure, a plurality of second image frames to be face-detected may be determined in each image frame split from the first video according to a preset image frame interval.
[0058] Optionally, the preset image frame interval may be set according to the actual recognition efficiency and the risk of omission. Preferably, the preset image frame interval may be 5 frames.
[0059] Optionally, in the embodiments of the present disclosure, the original image size of the second image frame may be first adjusted to a preset image size, and then the adjusted second image frame may be input into a preset multi-task face detection model, where the preset image size is smaller than the original image size. By reducing the image size of the second image frame in the embodiments of the present disclosure, the efficiency of face detection and calculation can be improved.
[0060] S120. Input the second image frame into a preset multi-task face detection model to obtain a face detection result output by the preset multi-task face detection model, where the face detection result includes the ratio of the face image in the second image frame to the image, the face confidence, and the face angle.
[0061] Among them, the preset multi-task face detection model may be a MOS model. The MOS model is a low-latency and lightweight architecture for face detection, face landmark localization, and head pose estimation. After the second image frame is input into the preset multi-task face detection model in the embodiments of the present disclosure, the preset multi-task face detection model will predict the position, face confidence, and face angle of the face image in the second image frame. Among them, the position includes the upper left corner coordinates and the lower right corner coordinates of the face image in the second image frame, the face confidence is used to indicate the probability that the face image is a real face, and the face angle includes the pitch and yaw angles of the face in the face image. It can be understood that according to the position of the face image in the second image frame, the ratio of the face image to the second image frame can be calculated. In the preset multi-task face detection model, the face image determined in the second image frame is a rectangular area, and the ratio of the rectangular area to the second image frame can be used to determine the ratio of the face image to the second image frame.
[0062] S130. Perform convolution calculation on the second image frame by using a harmonic operator to obtain the clarity of the second image frame.
[0063] Among them, the harmonic operator is also called the Laplacian operator. In the embodiments of the present disclosure, each pixel block in the second image frame can be convolved with the harmonic operator to respectively obtain a convolution value corresponding to the pixel block, and the sum of squares of each convolution value is obtained to obtain the sharpness of the second image frame.
[0064] Specifically, the harmonic operator can be:
[0065]
[0066] In the embodiments of the present disclosure, according to the formula:
[0067]
[0068] the sharpness of the second image frame is obtained, where f is the sharpness of the second image frame, i and j represent the positions of the pixel block in the second image frame, and G(i, j) represents the convolution value corresponding to the pixel block at the position (i, j).
[0069] S140. Using the ratio of the face area in the image, face confidence, face angle, and sharpness, obtain at least one first image frame of the expression to be detected in each second image frame.
[0070] Optionally, in the embodiments of the present disclosure, at least one first image frame can be screened out in each second image frame by using the ratio of the face area in the image, face confidence, face angle, and sharpness.
[0071] Specifically, in the embodiments of the present disclosure, a second image frame with a face area ratio in the image not less than a preset ratio threshold, a face confidence not less than a preset face confidence threshold, a face angle satisfying a preset angle condition, and a sharpness not less than a preset sharpness threshold can be determined as the first image frame.
[0072] Optionally, the preset ratio threshold can be 0.1. Optionally, the preset face confidence threshold can be 0.55. Optionally, the preset angle condition can be that the pitch angle is not greater than 45 degrees and the yaw angle is not greater than 20 degrees. Optionally, the preset sharpness threshold can be 50. By screening the face detection results and sharpness predicted for the second image frame, the embodiments of the present disclosure can obtain a first image frame with a clear, large, and frontal face image.
[0073] Optionally, based on Figure 2 the method shown, as Figure 3 shown, a schematic flowchart of another implementation manner of step S100 in the video processing method provided by the embodiments of the present disclosure, step S140 may include:
[0074] S141. Screen out at least one third image frame with a face area ratio in the image not less than a preset ratio threshold in each second image frame.
[0075] S142. Screen out at least one fourth image frame with a face confidence not less than a preset face confidence threshold in each third image frame.
[0076] S143. Screen out at least one fifth image frame with a face angle satisfying a preset angle condition in each fourth image frame.
[0077] S144. Screen out at least one first image frame of the expression to be detected with a clarity not less than a preset clarity threshold in each fifth image frame.
[0078] S200. Input the first image frame into a preset expression detection model to obtain an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of the face image in the first image frame and a first expression confidence corresponding to the face expression result.
[0079] Among them, the preset expression detection model is a convolutional neural network model. Specifically, the preset expression detection model is resnet50. In the embodiments of the present disclosure, a plurality of face images with labeled face expression results can be collected in advance to train the expression detection model to obtain a trained expression detection model.
[0080] Optionally, in the embodiments of the present disclosure, the first image frame can be input into the preset expression detection model to enable the preset expression detection model to perform multi-expression type detection on the face image in the first image frame, obtain second expression confidences corresponding to each expression of the face image under the preset expression types, determine the second expression confidence with the highest value as the first expression confidence, and determine the expression corresponding to the first expression confidence as the face expression result of the face image in the first image frame; obtain an expression detection result including the first expression confidence and the face expression result output by the preset expression detection model.
[0081] Optionally, the preset expression types can be one or more expressions among neutral, happy, sad, angry, disgust, surprise, and fear. In the embodiments of the present disclosure, face images of one or more expressions can be selected to train the expression detection model.
[0082] Among them, the preset expression detection model can detect the values of the second expression confidence levels of each expression of the face image under the preset expression types, and determine the highest value of the second expression confidence level as the first expression confidence level corresponding to the first image frame of the face image. For example: Suppose the second expression confidence levels of each expression of the face image in the first image frame under the preset expression types are: neutral 10, happy 80, sad 15, angry 20, disgusted 18, surprised 35, and fearful 23. Then it is determined that the face expression result corresponding to the first image frame is happy and the first expression confidence level is 80.
[0083] S300. Determine multiple second videos in the first video by using at least a sliding window with a preset video time length.
[0084] Optionally, the preset video time length can be set according to actual needs. Optionally, the preset video time length can be 2 minutes. It can be understood that the sliding step length of the sliding window can be set according to actual needs, and the present disclosure does not further limit this here.
[0085] Optionally, the embodiments of the present disclosure can use a sliding window with a preset video time length to determine a video segment with a time length of the preset video time length sequentially selected in the first video as a second video.
[0086] Optionally, based on Figure 1 the method shown, as Figure 4 shown, a schematic flowchart of another implementation manner of the video processing method provided by the embodiments of the present disclosure, step S300 may include:
[0087] S310. Determine multiple fourth videos in the first video by using a sliding window with a preset video time length.
[0088] Specifically, the embodiments of the present disclosure can use a sliding window with a preset video time length to determine a video segment with a time length of the preset video time length sequentially selected in the first video as a fourth video.
[0089] S320. Use the non-maximum suppression algorithm to remove overlapping videos in each fourth video, and determine multiple second videos, where the overlap value between each second video is less than a preset overlap threshold.
[0090] Specifically, the embodiments of the present disclosure can use the Non-Maximum Suppression (NMS) algorithm to calculate the overlap value between any two fourth videos, and retain one fourth video between two fourth videos whose overlap value is not less than a preset overlap threshold, and determine the fourth video after removing the overlapping videos as the second video. By removing the overlapping videos, the embodiments of the present disclosure can avoid obtaining two third videos with high similarity, and improve the resource utilization rate of short video production.
[0091] S400. For any second video: use the first expression confidence corresponding to each first image frame included in the second video to determine the emotion confidence corresponding to the second video.
[0092] Optionally, the embodiments of the present disclosure can add the first expression confidences corresponding to each first image frame included in any second video, and determine the result after addition as the emotion confidence corresponding to the second video.
[0093] Optionally, when the preset expression types include neutral, for each first image frame included in any second video: add the first expression confidences corresponding to other first image frames except the first image frames with neutral face expression results, and determine the result after addition as the emotion confidence corresponding to the second video. It can be understood that the first image frames with neutral face expression results reflect small emotional fluctuations of the person, and the video content corresponding to the first image frames may be relatively dull. The embodiments of the present disclosure do not use the first expression confidence corresponding to the first image frames with dull video content as the basis for emotion confidence, so as to make the emotion confidence of the second video with more intense video content as high as possible.
[0094] Optionally, the embodiments of the present disclosure can calculate according to the formula:
[0095]
[0096] Calculate the emotion confidence corresponding to the second video, where Score_ij is the emotion confidence corresponding to the second video, Score_k is the first expression confidence of the kth frame, i is the number of the starting frame of the second video, and j is the number of the ending frame of the second video.
[0097] S500. Obtain a preset number of third videos from each second video according to the emotion confidence.
[0098] Optionally, the embodiments of the present disclosure can obtain a preset number of third videos from each second video in the order of decreasing emotion confidence.
[0099] Among them, the preset quantity can be set according to actual requirements. Optionally, the preset quantity can be 5. It can be understood that the third video is a short video with a video duration shorter than that of the first video.
[0100] Optionally, based on Figure 1 the method shown, as Figure 5 shown, a schematic flowchart of another implementation manner of the video processing method provided by the embodiments of the present disclosure. After step S500, the video processing method may further include:
[0101] S600. Stitch each third video to obtain a fifth video.
[0102] It can be understood that the fifth video is a short video with a video duration between that of the first video and the third video. Optionally, the embodiments of the present disclosure may stitch the third videos in the playing order in the first video to obtain the fifth video.
[0103] A video processing method provided by the present disclosure can obtain at least one first image frame of a to-be-detected expression in a first video; input the first image frame into a preset expression detection model to obtain an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of a face image in the first image frame and a first expression confidence corresponding to the face expression result; determine a plurality of second videos in the first video at least by using a sliding window of a preset video time length; for any second video: determine an emotion confidence corresponding to the second video by using the first expression confidences corresponding to the first image frames included in the second video; obtain a preset quantity of third videos from the second videos according to the emotion confidence. The present disclosure can accurately identify video segments with excited emotions of characters in a long video through the emotion confidence, which helps to improve the production efficiency of short videos and the drainage promotion of long videos.
[0104] Although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous.
[0105] It should be understood that the various steps recorded in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0106] Corresponding to the above method embodiments, the embodiments of the present disclosure further provide a video processing apparatus, the structure of which is as Figure 6 shown, and may include: a first obtaining unit 100, a second obtaining unit 200, a first determining unit 300, a second determining unit 400, and a third obtaining unit 500.
[0107] The first acquisition unit 100 is configured to acquire at least one first image frame of the expression to be detected in the first video.
[0108] The second acquisition unit 200 is configured to input the first image frame into a preset expression detection model, and acquire an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of a face image in the first image frame and a first expression confidence level corresponding to the face expression result.
[0109] The first determination unit 300 is configured to determine a plurality of second videos in the first video by using at least a sliding window of a preset video time length.
[0110] The second determination unit 400 is configured to, for any one of the second videos, determine an emotion confidence level corresponding to the second video by using the first expression confidence levels corresponding to the respective first image frames included in the second video.
[0111] The third acquisition unit 500 is configured to acquire a preset number of third videos from the respective second videos according to the emotion confidence level.
[0112] Optionally, the first acquisition unit 100 may include: a first determination subunit, a first acquisition subunit, a second acquisition subunit, and a third acquisition subunit.
[0113] The first determination subunit is configured to determine a plurality of second image frames to be face-detected in the first video at a preset image frame interval.
[0114] The first acquisition subunit is configured to input the second image frame into a preset multi-task face detection model, and acquire a face detection result output by the preset multi-task face detection model, where the face detection result includes a face image ratio, a face confidence level, and a face angle of a face image in the second image frame.
[0115] The second acquisition subunit is configured to perform convolution calculation on the second image frame by using a harmonic operator to acquire the sharpness of the second image frame.
[0116] The third acquisition subunit is configured to acquire at least one first image frame of the expression to be detected from the respective second image frames by using the face image ratio, the face confidence level, the face angle, and the sharpness.
[0117] Optionally, the third obtaining subunit may be specifically configured to screen out at least one third image frame in which the image ratio occupied by the face is not less than a preset ratio threshold from each second image frame; screen out at least one fourth image frame in which the face confidence is not less than a preset face confidence threshold from each third image frame; screen out at least one fifth image frame in which the face angle meets a preset angle condition from each fourth image frame; and screen out at least one first image frame of the expression to be detected with a clarity not less than a preset clarity threshold from each fifth image frame.
[0118] Optionally, the second obtaining unit 200 may be specifically configured to input the first image frame into a preset expression detection model, so that the preset expression detection model performs multi-expression type detection on the face image in the first image frame, obtains second expression confidences corresponding to each expression of the face image under the preset expression types, determines the second expression confidence with the highest value as the first expression confidence, and determines the expression corresponding to the first expression confidence as the face expression result of the face image in the first image frame; and obtain an expression detection result including the first expression confidence and the face expression result output by the preset expression detection model.
[0119] Optionally, the first determining unit 300 may include: a second determining subunit and a third determining subunit.
[0120] The second determining subunit is configured to determine a plurality of fourth videos in the first video by using a sliding window with a preset video time length.
[0121] The third determining subunit is configured to use a non-maximum suppression algorithm to remove overlapping videos in each fourth video and determine a plurality of second videos, where the overlap value between each second video is less than a preset overlap threshold.
[0122] Optionally, the third obtaining unit 500 is specifically configured to obtain a preset number of third videos in each second video in descending order of emotion confidence.
[0123] Optionally, the video processing device may further include: a video splicing unit.
[0124] The video splicing unit is configured to splice each third video to obtain a fifth video after the third obtaining unit 500 obtains a preset number of third videos in each second video according to the emotion confidence.
[0125] A video processing device provided by the present disclosure can obtain at least one first image frame of a to-be-detected expression in a first video; input the first image frame into a preset expression detection model to obtain an expression detection result output by the preset expression detection model, where the expression detection result includes a face expression result of a face image in the first image frame and a first expression confidence corresponding to the face expression result; determine a plurality of second videos in the first video by using at least a sliding window of a preset video time length; for any second video: determine an emotion confidence corresponding to the second video by using the first expression confidences corresponding to the respective first image frames included in the second video; and obtain a preset number of third videos from the respective second videos according to the emotion confidence. Through the emotion confidence, the present disclosure can accurately identify video segments in a long video where a person's emotion is excited, which helps to improve the production efficiency of short videos and the drainage and promotion of long videos.
[0126] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0127] The video processing device includes a processor and a memory. The above first obtaining unit 100, second obtaining unit 200, first determining unit 300, second determining unit 400, third obtaining unit 500, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0128] The processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the emotion confidence of video segments in the video can be calculated to accurately identify video segments in a long video where a person's emotion is excited, which helps to improve the production efficiency of short videos and the drainage and promotion of long videos.
[0129] An embodiment of the present disclosure provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the video processing method is implemented.
[0130] An embodiment of the present disclosure provides a processor, where the processor is used to run a program, and when the program runs, the video processing method is executed.
[0131] As Figure 7As shown, an embodiment of the present disclosure provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003. Among them, the processor 1001 and the memory 1002 communicate with each other through the bus 1003. The processor 1001 is used to call program instructions in the memory 1002 to execute the above video processing method. The electronic device herein may be a server, a PC, a PAD, a mobile phone, etc.
[0132] The present disclosure also provides a computer program product, which is adapted to execute a program initialized with the steps of the video processing method when executed on an electronic device.
[0133] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate means for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0134] In a typical configuration, an electronic device includes one or more processors (CPUs), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, etc.
[0135] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip. The memory is an example of a computer-readable medium.
[0136] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0137] In the description of the present disclosure, it should be understood that if terms such as "upper", "lower", "front", "rear", "left", and "right" are used to indicate the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present disclosure.
[0138] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.
[0139] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, system, or computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.
Claims
1. A video processing method, characterized in that, it includes: obtaining at least one first image frame of the expression to be detected in the first video; inputting the first image frame into a preset expression detection model, so that the preset expression detection model performs multi-expression type detection on the face image in the first image frame, obtaining second expression confidence levels corresponding to each expression of the face image under the preset expression types, determining the second expression confidence level with the highest value as the first expression confidence level, and determining the expression corresponding to the first expression confidence level as the face expression result of the face image in the first image frame; obtaining an expression detection result including the first expression confidence level and the face expression result output by the preset expression detection model, wherein the expression detection result includes the face expression result of the face image in the first image frame and the first expression confidence level corresponding to the face expression result; determining a plurality of second videos in the first video by using at least a sliding window of a preset video time length; for any one of the second videos: adding the first expression confidence levels corresponding to the first image frames other than the first image frames with a neutral face expression result, and determining the added result as the emotion confidence level corresponding to the second video; obtaining a preset number of third videos in each of the second videos according to the emotion confidence level.
2. The method according to claim 1, characterized in that, the obtaining at least one first image frame of the expression to be detected in the first video includes: determining a plurality of second image frames to be face-detected in the first video at a preset image frame interval; inputting the second image frame into a preset multi-task face detection model, and obtaining a face detection result output by the preset multi-task face detection model, wherein the face detection result includes the image ratio of the face in the face image in the second image frame, the face confidence level, and the face angle; performing convolution calculation on the second image frame by using a harmonic operator to obtain the clarity of the second image frame; obtaining at least one first image frame of the expression to be detected in each of the second image frames by using the image ratio of the face, the face confidence level, the face angle, and the clarity.
3. The method according to claim 2, characterized in that, the obtaining at least one first image frame of the expression to be detected in each of the second image frames by using the image ratio of the face, the face confidence level, the face angle, and the clarity includes: screening out at least one third image frame in each of the second image frames with the image ratio of the face not less than a preset ratio threshold; screening out at least one fourth image frame in each of the third image frames with the face confidence level not less than a preset face confidence level threshold; screening out at least one fifth image frame in each of the fourth image frames with the face angle satisfying a preset angle condition; screening out at least one first image frame of the expression to be detected in each of the fifth image frames with the clarity not less than a preset clarity threshold.
4. The method according to claim 1, characterized in that, Determining a plurality of second videos in the first video by using at least a sliding window with a preset video time length, including: Determining a plurality of fourth videos in the first video by using a sliding window with a preset video time length; Using a non-maximum suppression algorithm to remove overlapping videos in each of the fourth videos, and determining a plurality of second videos, wherein the overlap value between each of the second videos is less than a preset overlap threshold.
5. The method according to claim 1, wherein, Obtaining a preset number of third videos in each of the second videos according to the emotion confidence, including: Obtaining a preset number of third videos in each of the second videos in the order from the largest to the smallest of the emotion confidence.
6. The method according to claim 1, wherein, After obtaining a preset number of third videos in each of the second videos according to the emotion confidence, the method further includes: Stitching each of the third videos to obtain a fifth video.
7. A video processing device, wherein, including: A first obtaining unit, a second obtaining unit, a first determining unit, a second determining unit and a third obtaining unit, The first obtaining unit is configured to obtain at least one first image frame of the expression to be detected in the first video; The second obtaining unit is configured to input the first image frame into a preset expression detection model, and obtain an expression detection result output by the preset expression detection model, wherein the expression detection result includes a face expression result of the face image in the first image frame and a first expression confidence corresponding to the face expression result; The first determining unit is configured to determine a plurality of second videos in the first video by using at least a sliding window with a preset video time length; The second determining unit is configured to, for any one of the second videos: add the first expression confidences corresponding to the other first image frames except the first image frames with a neutral face expression result, and determine the added result as the emotion confidence corresponding to the second video; The third obtaining unit is configured to obtain a preset number of third videos in each of the second videos according to the emotion confidence; The second obtaining unit is specifically configured to input the first image frame into a preset expression detection model, so that the preset expression detection model performs multi-expression type detection on the face image in the first image frame, obtains second expression confidences corresponding to each expression of the face image under a preset expression type, determines the second expression confidence with the highest value as the first expression confidence, and determines the expression corresponding to the first expression confidence as the face expression result of the face image in the first image frame; obtain an expression detection result including the first expression confidence and the face expression result output by the preset expression detection model.
8. A computer-readable storage medium, on which a program is stored, wherein, The program, when executed by a processor, implements the video processing method according to any one of claims 1 to 6.
9. An electronic device, the electronic device includes at least one processor, and at least one memory and a bus connected to the processor; wherein, The processor and the memory complete communication with each other through the bus; The processor is used to call program instructions in the memory to execute the video processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Facial expression recognition method and system based on video stream
CN113111789A