Video editing processing method, device, electronic device and storage medium
By automatically generating expression tags and collections of video clips through user expression recognition technology, the problems of low efficiency and poor accuracy of video editing in existing technologies are solved, and video editing effects that better meet user needs are achieved.
Patent Information
- Application Number
- CN202110587602.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-05-27
AI Technical Summary
Existing video editing technology mainly relies on manual operation, resulting in low efficiency and poor accuracy, and cannot meet the real needs of users.
By obtaining facial data of users while watching videos, expression recognition processing is performed, expression tags of video clips are automatically generated, and video editing and clustering are performed based on expression tags to generate video collections corresponding to different expression tags.
The accuracy and efficiency of video editing have been improved, and the edited video clips are more in line with the real needs of users, improving the user's viewing experience.
Smart Images

Figure CN115484474B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a video editing processing method, device, electronic device and computer-readable storage medium. Background Art
[0002] Video editing technology is a technology that crops a video to obtain one or more video clips in the video. For example, taking a movie as an example, a movie with a total length of 60 minutes is edited to obtain the movie clips between the 5th minute and the 10th minute in the movie as the highlights of the movie.
[0003] However, in the video editing solutions provided by related technologies, editing operations are mainly completed manually, and it is necessary to rely on manual labor to judge the video content, and then manually identify the video clips that the user may be interested in and edit them. The entire process requires a lot of manpower and time costs, and it is also easy to miss or over-cut due to human negligence or subjective judgment of the editor. In other words, the video editing efficiency and accuracy of the solutions provided by related technologies are low, resulting in the edited video content being unable to meet the real needs of users. Summary of the Invention
[0004] The embodiments of the present application provide a video editing processing method, device, electronic device and computer-readable storage medium, which can achieve accurate video editing and automatically generate video collections corresponding to different expression tags.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present invention provides a video editing method, including:
[0007] Acquire facial data of at least one video, wherein the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user, and the facial image is collected from the user while the user is watching the video;
[0008] Performing expression recognition processing on each of the facial image sequences to obtain an expression label for at least one video clip in the video;
[0009] Edit the video according to the start time and end time corresponding to each video segment to obtain a file of each video segment;
[0010] Based on the expression tags of the video segments of the at least one video, clustering processing is performed on the files of the video segments of the at least one video to obtain a video collection corresponding to the at least one expression tag.
[0011] The present invention provides a video editing device, comprising:
[0012] an acquisition module, configured to acquire facial data of at least one video, wherein the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user, and the facial image is collected from the user while the user is watching the video;
[0013] An expression recognition module, configured to perform expression recognition processing on each of the facial image sequences to obtain an expression label for at least one video segment in the video;
[0014] An editing module, configured to edit the video according to the start time and end time corresponding to each video segment, to obtain a file of each video segment;
[0015] The clustering module is used to perform clustering processing on the files of the video segments of the at least one video based on the expression tags of the video segments of the at least one video to obtain a video collection corresponding to at least one expression tag.
[0016] In the above scheme, the expression recognition module is also used to perform the following processing on each frame of facial image in the facial image sequence: perform face detection processing on the facial image to obtain the facial area in the facial image; perform feature extraction on the facial area to obtain corresponding facial feature data; call the trained classifier based on the facial feature data to perform prediction processing to obtain the expression label corresponding to the facial image; determine the corresponding video segment in the video based on the collection time period corresponding to the consecutive facial images with the same expression label in the facial image sequence, and use the consecutive same expression labels as the expression label of the video segment.
[0017] In the above scheme, the expression recognition module is also used to extract features from the facial area to obtain a corresponding facial feature vector; wherein the dimension of the facial feature vector is smaller than the dimension of the facial area, and the facial feature vector includes at least one of the following: shape feature vector, motion feature vector, color feature vector, texture feature vector, and spatial structure feature vector.
[0018] In the above scheme, the expression recognition module is also used to detect key feature points in the facial area, and align and calibrate the facial image included in the facial area based on the key feature points; and edit the facial area including the aligned and calibrated facial image, wherein the editing process includes at least one of the following: normalization processing, cropping processing, and scaling processing.
[0019] In the above scheme, the device also includes a determination module, which is used to determine the number of each type of expression labels included in the video clip when the same expression labels of the video clip are determined through facial image sequences corresponding to multiple users; the determination module is also used to treat the expression labels whose number among the multiple expression labels is less than a quantity threshold as invalid labels; the device also includes a deletion module, which is used to delete the invalid labels.
[0020] In the above scheme, the determination module is also used to determine the number of each type of expression labels included in the video clip when multiple expression labels of the video clip are determined through facial image sequences corresponding to multiple users respectively; the device also includes a screening module for screening out expression labels whose number is greater than a quantity threshold from the multiple expression labels; the determination module is also used to determine the tendency ratio corresponding to each screened expression label; and is used to treat the expression labels whose tendency ratio is less than the ratio threshold among the multiple screened expression labels as invalid labels; the deletion module is also used to delete the invalid labels.
[0021] In the above scheme, the determination module is further used to perform the following processing on the video clip: when the same expression label of the video clip is determined by the facial image sequences corresponding to multiple users, the start time and end time corresponding to the video clip are determined in the following manner: based on the start time and end time of each user's expression label, a normal distribution curve is established; with the symmetry axis of the normal distribution curve as the center, an n% interval of the normal distribution curve is extracted, and the time corresponding to the start point of the interval is determined as the start time of the video clip, and the time corresponding to the end point of the interval is determined as the end time of the video clip; wherein n is a positive integer and satisfies 0 <n<100。
[0022] In the above scheme, the clustering module is also used to cluster the files of video clips with the same expression tags in the video into the same video collection when the number of the video is 1; and to cluster the files of video clips with the same expression tags in multiple videos into the same video collection when the number of the video is multiple, or, for videos of the same type in multiple videos, cluster the files of video clips with the same expression tags in the videos of the same type into the same video collection.
[0023] In the above scheme, the determination module is also used to determine the value of m according to the speed of change of the plot content of the video clip; determine the first time m seconds before the start time in the video; determine the second time m seconds after the end time in the video; the editing module is also used to edit the video based on the first time and the second time.
[0024] In the above scheme, the editing module is also used to obtain a first video segment in the video that is less than a duration threshold away from the first time, and a second video segment that is less than the duration threshold away from the second time; perform speech recognition processing on the first video segment to obtain a first text, perform integrity detection processing on the first text to obtain a first dialogue integrity detection result, and adjust the first time according to the first dialogue integrity detection result to obtain a third time; perform speech recognition processing on the second video segment to obtain a second text, perform integrity detection processing on the second text to obtain a second dialogue integrity detection result, and adjust the second time according to the second dialogue integrity detection result to obtain a fourth time; and edit a file including the video segment between the third time and the fourth time from the video.
[0025] In the above scheme, the editing module is also used to obtain a first video segment in the video that is less than a duration threshold away from the first time, and a second video segment that is less than the duration threshold away from the second time; perform frame extraction processing on the first video segment to obtain multiple first video image frames, perform comparison processing on the multiple first video frame images to obtain a first picture integrity detection result, and adjust the first time according to the first picture integrity detection result to obtain a fifth time; perform frame extraction processing on the second video segment to obtain multiple second video image frames, perform comparison processing on the multiple second video image frames to obtain a second picture integrity detection result, and adjust the second time according to the second picture integrity detection result to obtain a sixth time; and edit a file including the video segment between the fifth time and the sixth time from the video.
[0026] In the above scheme, the determination module is also used to perform the following processing for each of the video clips: when the number of users watching the video is 1, the start time and end time of the user's expression tag are used as the start time and end time corresponding to the video clip; when the number of users watching the video is multiple, the start time and end time corresponding to the video clip are determined based on the start time and end time of the expression tags of multiple users.
[0027] In the above scheme, the acquisition module is also used to perform the following processing for each of the videos: receiving at least one facial image sequence sent by the terminal of at least one user watching the video, wherein the facial image sequence is obtained by performing multiple facial captures on the user when the terminal is playing the video.
[0028] The present invention provides a video editing method, including:
[0029] Display a video interface, wherein the video interface is used to play a video or display a video list;
[0030] Displaying a viewing entry for a video collection, wherein the video collection is obtained through any of the above solutions;
[0031] In response to a triggering operation on a viewing entry for the video collection, the video collection is displayed.
[0032] The present invention provides a video editing device, comprising:
[0033] A display module is used to display a video interface, wherein the video interface is used to play videos or display a video list;
[0034] The display module is further configured to display a viewing entry for a video collection, wherein the video collection is obtained through any of the above-mentioned solutions;
[0035] The display module is also used to display the video collection in response to a triggering operation on the viewing entrance of the video collection.
[0036] In the above scheme, the display module is also used to receive input keywords through the viewing entrance; the device also includes an acquisition module, which is used to obtain a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; the display module is also used to play the matching video collection.
[0037] In the above scheme, the display module is also used to receive keywords input through the viewing entrance; the acquisition module is also used to obtain a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; and is used to obtain the user's historical behavior information; the device also includes a determination module for determining the type of video that the user is interested in based on the historical behavior information; the device also includes a screening module for screening out video clips of the same type from the matching video collection; the display module is also used to play a video collection consisting of the screened video clips.
[0038] An embodiment of the present application provides an electronic device, including:
[0039] a memory for storing executable instructions;
[0040] The processor is configured to implement the video editing processing method provided in the embodiment of the present application when executing the executable instructions stored in the memory.
[0041] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute and implement the video editing processing method provided in the embodiment of the present application.
[0042] An embodiment of the present application provides a computer program product, which includes computer-executable instructions for implementing the video editing processing method provided in the embodiment of the present application when executed by a processor.
[0043] The embodiments of the present application have the following beneficial effects:
[0044] By recognizing user expressions, the video content is fragmented and edited, and video collections corresponding to different expression tags are automatically generated. Since the changes in user expressions are the most realistic judgment of the video content, video editing through user expression recognition can make the judgment of editing timing (that is, the start time and end time corresponding to the video clip) more accurate, so that the edited video clips are more in line with the user's real needs and improve the user's viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 1 is a schematic diagram of the architecture of a video editing processing system 100 provided in an embodiment of the present application;
[0046] Figure 2A 2 is a schematic diagram of the structure of the server 200 provided in an embodiment of the present application;
[0047] Figure 2B 4 is a schematic diagram of the structure of the terminal 400 provided in an embodiment of the present application;
[0048] Figure 3 Schematic diagram of the video editing process provided by the embodiment of the present application;
[0049] Figure 4 Schematic diagram of the video editing process provided by the embodiment of the present application;
[0050] Figure 5A Schematic diagram of the video editing process provided by the embodiment of the present application;
[0051] Figure 5B Schematic diagram of the video editing process provided by the embodiment of the present application;
[0052] Figure 6 This is a schematic diagram of an application scenario of the video editing processing method provided in an embodiment of the present application;
[0053] Figure 7 This is a schematic diagram of an application scenario of the video editing processing method provided in an embodiment of the present application;
[0054] Figure 8 This is a schematic diagram of an application scenario of the video editing processing method provided in an embodiment of the present application;
[0055] Figure 9 Schematic diagram of the video editing process provided by the embodiment of the present application;
[0056] Figure 10 Schematic diagram of the process of facial expression recognition provided by the embodiment of the present application;
[0057] Figure 11 This is a schematic diagram of a process for preprocessing an input image provided by an embodiment of the present application;
[0058] Figure 12 Schematic diagram of the principle of performing expression recognition on an input image provided by an embodiment of the present application;
[0059] Figure 13 is a schematic diagram of optimizing video clips according to expressions of multiple users provided in an embodiment of the present application;
[0060] Figure 14 Schematic diagram of setting a single expression tag for a single video clip provided by an embodiment of the present application;
[0061] Figure 15 This is a schematic diagram of setting multiple expression tags for a single video clip provided by an embodiment of the present application;
[0062] Figure 16 The normal distribution curve established according to the expression generation time and disappearance time of multiple users provided in the embodiment of the present application;
[0063] Figure 17 is a schematic diagram of performing rough cutting and intelligent fine cutting on a video clip provided by an embodiment of the present application;
[0064] Figure 18 This is a schematic diagram of the process of generating different video collections for multiple video clips provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0066] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0067] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0069] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0070] 1) Video: A broad term for various technologies that capture, record, process, store, transmit, and reproduce a series of static images as electrical signals. When continuous image changes exceed 24 frames per second, the human eye cannot distinguish between individual static images due to the persistence of vision principle. The resulting image appears smooth and continuous, and this continuous image is called video.
[0071] 2) Facial expressions: also known as facial expressions, are part of human body language and are physiological and psychological reactions that are usually used to convey emotions. Facial expressions include basic expressions and compound expressions. Basic expressions include happiness, surprise, sadness, anger, disgust, and fear. In addition, human facial expressions also include 15 distinguishable compound expressions, such as surprise (happiness + surprise) and grief and indignation (sadness + anger).
[0072] 3) Expression label: a label used to represent the user's expression. For example, when the user's expression is determined to be happy, the corresponding expression label may be "makes people feel funny"; when the user's expression is determined to be sad, the corresponding expression label may be "makes people feel like crying".
[0073] 4) Expression recognition: Changes in key facial parts (such as the corners of the eyebrows, the tip of the nose, and the corners of the mouth) are captured through an image acquisition device (such as a mobile phone camera). Based on the captured facial images, a machine learning algorithm is called to predict the expression represented by the facial changes, such as happiness, anger, sadness, fear, etc.
[0074] 5) Video collection: A collection of multiple video clips that are merged according to certain thematic classifications.
[0075] 6) Client: An application (APP) running in a terminal to provide various services, such as an instant messaging client, a short video client, a live broadcast client, etc.
[0076] With the development of user demand and multimedia technology, the number of videos has exploded exponentially, and video editing has become a popular video processing method. Video editing technology is a video processing method that cuts the video to be edited into one or more video clips. It is often used in video editing scenarios such as short video production and video highlights.
[0077] At present, in the video editing solutions provided by related technologies, editing operations are mainly completed manually. It is necessary to rely on manual labor to judge the video content, and then manually identify the video clips that the user may be interested in and edit them. The whole process requires a lot of manpower and time costs, and it is also easy to miss or over-cut due to human negligence or subjective judgment of the editor. In other words, in the solutions provided by related technologies, the efficiency of video editing is low and the accuracy is poor, resulting in the edited video content being unable to meet the real needs of users.
[0078] In response to the above technical problems, the embodiments of the present application provide a video editing processing method, device, electronic device and computer-readable storage medium, which can achieve accurate video editing and automatically generate video collections corresponding to different emoticon tags. The following describes an exemplary application of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as a terminal, or as a server, or implemented in collaboration with a terminal and a server. The following is an example of the video editing processing method provided by the embodiment of the present application being implemented in collaboration with a terminal and a server.
[0079] See also Figure 1 , Figure 1 This is an architectural diagram of the video editing processing system 100 provided in an embodiment of the present application. In order to implement applications that support video editing and generate different types of video collections, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0080] A client 410 is running on the terminal 400. The client 410 can be an online video playback client, a short video client, a browser, etc. When the terminal 400 receives a face acquisition instruction triggered by a user (for example, user A) during the process of playing a video (for example, video A), the image acquisition device (for example, a camera built into the terminal) is called to perform multiple face acquisitions on user A, and a face image sequence corresponding to user A in the process of watching video A is obtained, wherein the face image sequence is arranged in the order of the acquisition time of user A's face image, and each face image in the face image sequence has an acquisition time based on the playback time axis of video A (that is, the playback time of video A), for example, the face image The acquisition time of the first face image in the face image sequence corresponds to the first second of video A (that is, when video A is played to the first second, the face of user A is acquired for the first time to obtain the first face image), the second face image in the face image sequence corresponds to the second second of video A (that is, when video A is played to the second second, the face of user A is acquired for the second time to obtain the second face image), and so on. The last face image in the face image sequence corresponds to the last second of video A (that is, when video A is played to the last second, the face of user A is acquired for the last time to obtain the last face image). In other words, the number of face images included in the face image sequence is positively correlated with the length of the video.
[0081] After obtaining the facial image sequence corresponding to user A when watching video A, terminal 400 can send the obtained facial image sequence to server 200 through network 300, so that server 200 performs expression recognition processing on the facial image sequence sent by terminal 400, and obtains at least one video segment in video A (the video segment here is only recorded according to the corresponding start time and end time, and no separate video segment file is edited. For example, when the start time and end time of a video segment are determined to be 15:00 and 15:30 respectively according to the expression recognition result, the video segment can be recorded at the corresponding position of the playback time axis of video A, so as to perform subsequent editing processing according to the recorded position). Then, the server 200 edits the video A according to the start time and end time corresponding to each video clip to obtain the file of each video clip (for example, assuming that for video A, a total of 10 video clip files are edited); then, the server 200 can cluster the files of the video clips of video A based on the expression tags of the video clips of video A (for example, for the files of the 10 video clips edited from video A, the files of the video clips with the same expression tags are clustered into the same video collection), and obtain a video collection corresponding to at least one expression tag (for example, a video collection that makes people want to cry, a video collection that makes people want to laugh).
[0082] After obtaining a video collection corresponding to at least one emoticon tag, the server 200 can send the obtained video collection to the terminal 400, so that the terminal 400 calls the human-computer interaction interface of the client 410 for presentation (for example, displaying the viewing entrance of the video collection in a browser or an online video client, and displaying the video collection when the terminal 400 receives a trigger operation from the user for the viewing entrance of the video collection). In this way, by editing the video content through the recognition of the user's expression, the judgment of the edited content can be made more accurate, so that the content of the edited video clips can be more in line with the user's real needs, thereby improving the user's content satisfaction and viewing time.
[0083] It should be noted that, in actual applications, the number of users watching video A may also be multiple, that is, the server 200 can receive facial image sequences sent by terminals of multiple users respectively (for example, including the facial image sequence corresponding to user B during the process of watching video A sent by the terminal of user B, the facial image sequence corresponding to user C during the process of watching video A sent by the terminal of user C, and the facial image sequence corresponding to user D during the process of watching video A sent by the terminal of user D, etc.). Then, for the start time and end time corresponding to each video segment in video A, as well as the expression label of the video segment, the server 200 can adjust the expression recognition results of the facial image sequences of multiple users (the adjustment process will be described in detail below). In this way, the editing timing (that is, the start time and end time corresponding to the video segment) and the type of video content are optimized based on the changes in the expressions of a large number of users, so that the accuracy of video editing is further improved.
[0084] In addition, it should be noted that in actual applications, the number of videos can also be multiple. For example, after editing multiple videos based on changes in user expressions, the server 200 can cluster the files of video clips with the same expression labels in multiple videos into the same video collection. For example, the files of video clips with the expression label "makes people want to cry" in video A (for example, a war film), the files of video clips with the expression label "makes people want to cry" in video B (for example, an emotional film), and the files of video clips with the expression label "makes people want to cry" in video C (for example, a documentary) are clustered into the same video collection, thereby obtaining a collection of all videos that make people want to cry; or, the server 200 can also make a fine-grained division of the video collection according to the video type. For example, for videos of the same type (for example, war films) in multiple videos, the files of video clips with the same expression label (for example, "makes people want to cry") in the war films are clustered into the same video collection, thereby obtaining a collection of videos in war films that make people want to cry.
[0085] In some embodiments, the embodiments of the present application can be implemented with the help of cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool that can be used on demand with flexibility and convenience. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources.
[0086] For example, Figure 1 The server 200 shown in the figure can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to these. The terminal 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0087] In other embodiments, the video editing processing method provided in the embodiments of the present application can also be implemented in combination with blockchain technology. For example, the terminal 400 and the server 200 can be node devices in the blockchain system.
[0088] Below Figure 1 The structure of the server 200 shown in FIG will be described. Figure 2A , Figure 2A is a schematic diagram of the structure of the server 200 provided in an embodiment of the present application, Figure 2A The server 200 shown includes: at least one processor 210, a memory 240, and at least one network interface 220. The various components in the server 200 are coupled together via a bus system 230. It is understood that the bus system 230 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 230 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 230 is not described in detail. Figure 2A Various buses are labeled as bus system 230 .
[0089] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0090] The memory 240 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 240 may optionally include one or more storage devices that are physically remote from the processor 210.
[0091] The memory 240 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 240 described in the embodiments of the present application is intended to include any suitable type of memory.
[0092] In some embodiments, memory 240 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0093] Operating system 241, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0094] The network communication module 242 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
[0095] In some embodiments, the video editing processing device provided in the embodiments of the present application can be implemented in a software manner. Figure 2A The video clip processing device 243 stored in the memory 240 is shown. It can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 2431, an expression recognition module 2432, a clipping module 2433, a clustering module 2434, a determination module 2435, a deletion module 2436, and a screening module 2437. These modules are logical and can be arbitrarily combined or further divided according to the functions implemented. It should be noted that in Figure 2A For the sake of convenience, all the above modules are shown at once, but it should not be considered that the video editing processing device 243 excludes the implementation of only including the acquisition module 2431, the expression recognition module 2432, the editing module 2433 and the clustering module 2434. The functions of each module will be explained below.
[0096] Next, continue Figure 1The structure of the terminal 400 shown in FIG will be described. Figure 2B , Figure 2B 4 is a schematic diagram of the structure of the terminal 400 provided in the embodiment of the present application. Figure 2B As shown, terminal 400 includes: a processor 420, a network interface 430, a user interface 440, a bus system 450, and a memory 460. The user interface 440 includes one or more output devices 441 that enable the presentation of media content, such as one or more speakers and / or one or more visual display screens. The user interface 440 also includes one or more input devices 442, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, and other input buttons and controls. The memory 460 includes: an operating system 461, a network communication module 462, a presentation module 463 for enabling the display of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 441 associated with the user interface 440 (e.g., a display screen, speakers, etc.), an input processing module 464 for detecting one or more user inputs or interactions from one of the one or more input devices 442 and translating the detected inputs or interactions, and a video clip processing device 465. The software modules in the video clip processing device 465 stored in the memory 460 include: a display module 4651, an acquisition module 4652, a determination module 4653 and a screening module 4654. These modules are logical and can be arbitrarily combined or further split according to the functions implemented. It should be noted that in Figure 2B For the sake of convenience, all the above modules are shown at once, but it should not be considered that the video clip processing device 465 excludes the implementation that only includes the display module 4651. The functions of each module will be explained below.
[0097] The video editing processing method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that the following description is based on the example of the server 200 described above as the execution subject of the video editing processing method.
[0098] See also Figure 3 , Figure 3 This is a flow chart of the video editing method provided by the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0099] In step S101, facial data of at least one video is obtained.
[0100] In some embodiments, the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user (that is, the facial image is collected from the user while the user is watching the video, rather than the facial image appearing in the video). The facial data of at least one video can be obtained in the following manner: the following processing is performed for each video: at least one facial image sequence is received from a terminal of at least one user watching the video, respectively, wherein the facial image sequence is obtained by performing multiple facial collections on the user watching the video when the terminal is playing the video.
[0101] For example, taking video A as an example, when the number of users watching video A is 1, for example, when there is only user A, during the process of playing video A on user A's terminal (for example, when receiving a face collection instruction triggered by user A), user A's face is collected multiple times to obtain a facial image sequence corresponding to user A in the process of watching video A, wherein the facial images in the facial image sequence are arranged in the order of the collection time of user A's face, that is, the first collected facial image is arranged in the front, and the last collected facial image is arranged in the back. In addition, each facial image has a collection time based on the playback timeline of video A (that is, the playback time of video A), that is, the collection time of the facial images in the facial image sequence is consistent with the playback time of video A. After obtaining the facial image sequence corresponding to user A in the process of watching video A, user A's terminal can send the obtained facial image sequence to the server.
[0102] For example, still taking video A as an example, when there are multiple users watching video A, for each user, a facial image sequence corresponding to each user in the process of watching video A can be obtained by adopting a method similar to the above-mentioned method of obtaining the facial image sequence corresponding to user A in the process of watching video A, such as the facial image sequence corresponding to user B in the process of watching video A (that is, during the process of playing video A, the face of user B is collected multiple times by the terminal of user B), the facial image sequence corresponding to user C in the process of watching video A (that is, during the process of playing video A, the face of user C is collected multiple times by the terminal of user C), and the facial image sequence corresponding to user D in the process of watching video A (that is, during the process of playing video A, the face of user D is collected multiple times by the terminal of user D). Subsequently, for video A, the server can receive the facial image sequence corresponding to user B in the process of watching video A sent by the terminal of user B, the facial image sequence corresponding to user C in the process of watching video A sent by the terminal of user C, and the facial image sequence corresponding to user D in the process of watching video A sent by the terminal of user D.
[0103] That is to say, the facial image sequence is related to the user and the video. For the same user, when watching different videos, the corresponding facial image sequence is different; for the same video, when different users watch it, the corresponding facial image sequence is also different.
[0104] In addition, it should be noted that in actual applications, the number of videos can also be multiple. For other videos (such as video B), a processing method similar to video A can be used to obtain the facial data of video B. The embodiments of this application will not be repeated here.
[0105] In step S102, expression recognition processing is performed on each facial image sequence to obtain an expression label of at least one video segment in the video.
[0106] In some embodiments, expression recognition processing can be performed on each facial image sequence in the following manner to obtain an expression label for at least one video segment in the video (the video segment here is only recorded according to the corresponding start time and end time, and no file including the video segment is edited to avoid unnecessary resource consumption): for each frame of facial image in the facial image sequence, the following processing is performed: face detection processing is performed on the facial image to obtain the facial area in the facial image; feature extraction is performed on the facial area to obtain corresponding facial feature data; based on the facial feature data, the trained classifier is called to perform prediction processing to obtain the expression label corresponding to the facial image; based on the collection period of the facial images corresponding to the consecutive identical expression labels in the facial image sequence, the corresponding video segment in the video is determined, and the consecutive identical expression labels are used as the expression labels of the video segment.
[0107] For example, taking the facial image sequence corresponding to user A when watching video A (hereinafter referred to as facial image sequence 1) as an example, for each frame of facial image in facial image sequence 1, face detection processing is performed on the facial image (for example, face detection processing is performed on the facial image using a convolutional neural network model) to obtain the facial region in the facial image; then, feature extraction is performed on the facial region to obtain corresponding facial feature data (for example, feature extraction can be performed on the facial image using a convolutional neural network model to obtain a corresponding facial feature vector; wherein the facial feature vector can be a shape feature vector, motion feature vector, color feature vector, texture feature vector, or spatial structure feature vector corresponding to the facial image, etc. In this way, by feature extraction of the facial region, Extraction can achieve data dimensionality reduction and improve the speed and accuracy of subsequent data operations), and then, based on the obtained facial feature data, the trained classifier (such as linear classifier, neural network classifier, support vector machine, hidden Markov model, etc.) can be called for prediction processing to obtain the expression label corresponding to the facial image (that is, according to the recognized expression, the corresponding expression label is set, for example, when the recognized expression is happy, the corresponding expression label can be "makes people want to laugh"; when the recognized expression is sad, the corresponding expression label can be "makes people want to cry"); finally, based on the collection period of consecutive facial images with the same expression label, the corresponding video segment in video A can be determined, and the consecutive same expression labels can be used as the expression labels of the corresponding video segments.
[0108] For example, assuming that the expression labels corresponding to the 10th to 20th facial images in facial image sequence 1 are all "makes people want to laugh", then the corresponding video segment 1 in video A can be determined based on the acquisition period of the 10th to 20th facial images (for example, the 10th to 20th seconds). For example, the 10th and 20th seconds on the playback timeline of video A can be marked as the identifier of the file for subsequently editing video segment 1 from video A, and the expression label "makes people want to laugh" can be used as the expression label of video segment 1. Assuming that the expression labels corresponding to the 50th to 70th facial images in the facial image sequence 1 are all "makes people want to cry", then the corresponding video clip 2 in video A can be determined according to the acquisition period of the 50th to 70th facial images (for example, the 50th second to the 70th second). For example, the 50th second and the 70th second positions on the playback timeline of video A can be marked as the identifier of the file for subsequently editing video clip 2 from video A, and the expression label "makes people want to cry" can be used as the expression label of video clip 2.
[0109] It should be noted that in actual applications, the number of consecutive identical expression tags can be flexibly adjusted according to actual conditions. For example, when the plot content of a video changes slowly, the user's expression changes slowly, and the number of consecutive identical expression tags can be set higher accordingly. For example, when the number of consecutive identical expression tags reaches 30, the corresponding video segment in the video will be determined. When the plot content of a video changes quickly, the user's expression changes quickly, and the number of consecutive identical expression tags can be set lower accordingly. For example, when the number of consecutive identical expression tags exceeds 10, the corresponding video segment in the video will be determined. In other words, the value of the number of consecutive identical expression tags is negatively correlated with the speed of change of the plot content of the video.
[0110] In other embodiments, continuing with the above embodiments, before extracting features from the face area, the following operations may be performed: detecting key feature points in the face area (such as the center of the eyes, the corners of the mouth, the tip of the nose, etc.), and aligning and calibrating the face image included in the face area based on the key feature points; editing the face area including the aligned and calibrated face image, wherein the editing process includes at least one of the following: normalization (i.e., a process of performing a series of standard processing transformations on the face image to transform it into a fixed standard form, for example, the pixel values of the face image may be normalized), cropping (i.e., cropping the size of the face area to obtain a face area of uniform size), and scaling (i.e., scaling the size of the face image included in the face area to make the size of the scaled face image uniform). In this way, the quality of the face image included in the face area can be improved, interference information can be eliminated, and the size, proportion, grayscale value and other information of the face image can be unified. In addition, by performing normalization on the face area, subsequent feature extraction and prediction classification processes are facilitated.
[0111] In some embodiments, after performing expression recognition processing on each facial image sequence to obtain the expression label of at least one video clip in the video, the following operations can also be performed: the following processing is performed on the expression label of each video clip: when the same expression label of a video clip is determined by facial image sequences corresponding to multiple users (that is, a video clip has only one expression label), the number of each type of expression label included in the video clip is counted; the expression labels whose number among the multiple expression labels is less than the quantity threshold are regarded as invalid labels, and the invalid labels are deleted.
[0112] For example, take video A as an example. When multiple users watch video A, each user's terminal will upload the facial image sequence obtained by the corresponding user in the process of watching video A to the server. Then, the server will perform expression recognition processing on each facial image sequence to obtain the expression label of the video segment in video A (for example, video segment 1). Among them, the corresponding expressions of different users when watching the same video segment may be different. For example, for user A's facial image sequence, the expression label of video segment 1 is "makes people want to cry", for user B's facial image sequence, the expression label of video segment 1 is "makes people want to laugh", and for user C's facial image sequence, the expression label of video segment 1 is "makes people want to cry". Subsequently, the number of each type of expression label included in video segment 1 is counted. For example, assuming that for video segment 1, a total of 3 different expression labels are obtained, among which the expression label is "makes people want to cry". The number of expression labels is 1000 (i.e., 1000 users have sad expressions when watching video clip 1), the number of expression labels with the label "makes you want to laugh" is 50 (i.e., 50 users have happy expressions when watching video clip 1), and the number of expression labels with the label "makes you want to laugh" is 30 (i.e., 30 users have fearful expressions when watching video clip 1); finally, the expression labels whose number among multiple expression labels is less than the number threshold (e.g., 800) are regarded as invalid labels (i.e., the expression labels of "makes you want to laugh" and "makes you afraid" are regarded as invalid labels), and the expression labels of "makes you want to laugh" and "makes you afraid" included in video clip 1 are deleted, and only the expression label of "makes you want to cry" is retained as the expression label of video clip 1 (this is the true judgment of most users for video clip 1). In this way, the expression labels of video clips are optimized based on the expression recognition results of massive users, so that the judgment of video content is more accurate.
[0113] It should be noted that the value of the above quantity threshold is related to the total number of users. For example, when the total number of users is 1,000, the corresponding quantity threshold can be set to 600; when the total number of users is 500, the corresponding quantity threshold can be set to 300.
[0114] In other embodiments, after performing expression recognition processing on each facial image sequence to obtain an expression label of at least one video segment in the video, the following operations may be performed: the following processing is performed on the expression label of each video segment: when multiple expression labels of a video segment are determined (i.e., a video segment has multiple expression labels) through facial image sequences corresponding to multiple users (i.e., facial image sequences collected when multiple users watch the same video (e.g., video A), for example, including a facial image sequence collected when user A watches video A, a facial image sequence collected when user B watches video A, etc.), , count the number of each type of expression tags included in the video clip; screen out expression tags whose number is greater than a threshold value from multiple expression tags, and determine the tendency ratio corresponding to each screened expression tag (that is, the proportion of the number of expression tags of a certain type to the total number of expression tags, for example, assuming that there are 1,000 expression tags in total, among which the number of expression tags of the type "makes people want to cry" is 500, then the tendency ratio corresponding to the expression tag "makes people want to cry" is 50%); treat the expression tags whose tendency ratio among the multiple screened expression tags is less than the threshold value as invalid tags, and delete the invalid tags.
[0115] For example, take video A as an example. When multiple users watch video A, each user's terminal will upload the facial image sequence obtained by the corresponding user in the process of watching video A to the server. Then, the server will perform expression recognition processing on each facial image sequence to obtain the expression label of the video segment in video A (such as video segment 1). Among them, the corresponding expressions of different users when watching the same video segment may be different. For example, for user A's facial image sequence, the expression label of video segment 1 is "makes people want to cry", for user B's facial image sequence, the expression label of video segment 1 is "makes people want to laugh", and for user C's facial image sequence, the expression label of video segment 1 is "makes people want to cry". Subsequently, the number of each type of expression label included in video segment 1 is counted. For example, assuming that for video segment 1, a total of 3 different expression labels are obtained, among which the number of expression labels with the expression label "makes people want to cry" is 1000, and the number of expression labels with the expression label "makes people want to cry" is 1000. The number of expressions labeled "makes you want to laugh" is 800, and the number of expressions labeled "makes you want to laugh" is 100. Subsequently, expression labels whose number is greater than the quantity threshold (assuming it is 500) are screened out from these three different types of expression labels, and the tendency ratio corresponding to each screened expression label is determined (that is, the tendency ratio corresponding to the expression label "makes you want to cry" and the tendency ratio of the expression label "makes you want to laugh" are determined, and for the expression label "makes you want to laugh", since its number is less than the quantity threshold, it is deleted as an invalid label). Finally, the expression labels whose tendency ratio is less than the proportion threshold (for example, 40%) among the multiple screened expression labels are deleted as invalid labels. Since the tendency ratio (52%) of the expression label "makes you want to cry" and the tendency ratio (42%) of the expression label "makes you want to laugh" are both greater than the proportion threshold, "makes you want to cry" and "makes you want to laugh" can be simultaneously used as expression labels for video clip 1.
[0116] In step S103, the video is edited according to the start time and end time corresponding to each video segment to obtain a file of each video segment.
[0117] In some embodiments, before editing the video according to the start time and end time corresponding to each video clip, the following operations may also be performed: the following processing is performed on the video clip: when the same expression label of the video clip is determined by the facial image sequences corresponding to multiple users, the start time and end time corresponding to the video clip are determined in the following manner: based on the start time and end time of each user's expression label, a normal distribution curve is established; with the symmetry axis of the normal distribution curve as the center, n% intervals of the normal distribution curve are extracted, and the time corresponding to the start point of the interval is determined as the start time of the video clip, and the time corresponding to the end point of the interval is determined as the end time of the video clip; wherein n is a positive integer and satisfies 0 <n<100。
[0118] Exemplarily, for the same video clip, since the generation time and the duration of each user's expression are different, but the generation and disappearance of multiple users' expressions will show a normal distribution on the video playback timeline. Therefore, the start time and end time corresponding to the video clip can be determined in the following way: Based on the start time and end time of each user's expression label (for example, for the same video clip, the start time and end time of user A's expression label are 15:01 and 15:40 respectively, the start time and end time of user B's expression label are 14:58 and 15:35 respectively, and the start time and end time of user C's expression label are 15:04 and 15:50 respectively), a normal distribution curve is established; then, with the symmetry axis of the normal distribution curve as the center, an n% (0 < n < 100, and the value of n can be adjusted according to the final editing effect) interval of the normal distribution curve is extracted, and the time corresponding to the start point of the interval (for example, 15:02) is used as the start time of the video clip, and the time corresponding to the end point of the interval (for example, 15:45) is used as the end time of the video clip. In this way, based on the expression recognition results of the face image sequences of a large number of users, the judgment of the editing timing (that is, the start time and end time corresponding to the video clip) can be made more accurate.
[0119] In some other embodiments, before performing video editing processing according to the start time and end time corresponding to each video clip, the following operations can also be performed: For each video clip, the following processing is performed: When the number of users watching the video is 1, the start time and end time of the user's expression label are used as the start time and end time corresponding to the video clip; when the number of users watching the video is multiple, the start time and end time corresponding to the video clip are determined based on the start time and end time of the expression labels of multiple users.
[0120] Exemplarily, taking video A as an example, when the number of users watching video A is 1, for example, only user A, after performing expression recognition processing on the face image sequence of user A and obtaining the expression labels corresponding to user A at different times, the start time and end time of the consecutive same expression labels of user A (assuming the start time and end time are 10:00 and 11:00 respectively) can be directly used as the start time and end time of the corresponding video clip (for example, video clip 1) (that is, the start time and end time corresponding to video clip 1 are 10:00 and 11:00 respectively).
[0121] For example, still taking video A as an example, when there are multiple users watching video A, such as user A, user B, user C, user D, etc., after performing expression recognition processing on the facial image sequence corresponding to each user and obtaining the expression labels corresponding to different users at different times, the final start time and end time of the corresponding video clip can be determined based on the start time and end time of the continuous expression labels of different users. For example, a normal distribution curve is established based on the start time and end time of the expression labels of multiple users, and the start time and end time of the corresponding video clip are determined based on the normal distribution curve. In this way, by adjusting the generation time and disappearance time of the expressions of a large number of users, the corresponding start time and end time of the video clip can be determined more accurately.
[0122] In some embodiments, Figure 3 Step S103 shown can be Figure 4 Steps S1031 to S1034 shown are implemented by combining Figure 4 The steps shown are explained.
[0123] In step S1031 , the value of m is determined according to the speed at which the plot content of the video clip changes.
[0124] In some embodiments, in order to avoid incomplete content of the edited video clip (for example, lack of the introduction of the picture story), after determining the start time and end time corresponding to the video clip based on the start time and end time of the user's expression tag, the time value that needs to be added to the video playback timeline before the start time and after the end time (that is, the value of m) can also be determined based on the speed at which the plot content of the current video clip changes.
[0125] For example, since the rhythm of different video contents is different, the duration of the user's expression is also different. For example, for fighting war films, the rhythm is faster, and accordingly, the user's expression changes faster, so the value of m can be set to a smaller value (for example, 3 seconds); and for emotional documentaries, the rhythm is slower, and accordingly, the user's expression changes slower, so the value of m can be set to a larger value (for example, 7 seconds).
[0126] In step S1032 , a first time m seconds before the start time in the video is determined.
[0127] In some embodiments, taking video A as an example, assuming that the start time corresponding to video segment 1 in video A is 10:00, and in step S1031, based on the speed of change in the plot content of video segment 1, the value of m is determined to be 3 seconds, then the position 09:57 on the playback timeline of video A can be marked as the first time.
[0128] In step S1033 , a second time m seconds after the end time in the video is determined.
[0129] In some embodiments, still taking video A as an example, assuming that the end time corresponding to the video segment 1 of video A is 11:00, and in step S1031, based on the speed of changes in the plot content of video segment 1, the value of m is determined to be 3 seconds, then the position 11:03 on the playback timeline of video A can be marked as the second time.
[0130] In step S1034, the video is edited based on the first time and the second time.
[0131] In some embodiments, continuing from the above, after determining the first time (09:57) and the second time (11:03), a file of the video segment between 09:57 and 11:03 can be edited from video A. In this way, by editing by adding m seconds to the start time and the end time, the edited video content can be more complete, thereby improving the user's viewing experience.
[0132] In other embodiments, Figure 4 Step S1034 shown can also be performed by Figure 5A Steps S10341A to S10344A shown are implemented by combining Figure 5A The steps shown are explained.
[0133] In step S10341A, a first video segment whose distance from the first time to the video is less than a duration threshold and a second video segment whose distance from the second time to the video is less than a duration threshold are obtained.
[0134] In some embodiments, taking video A as an example, after determining a first time m seconds before the start time corresponding to video segment 1 in video A and a second time m seconds after the end time corresponding to video segment 1, a first video segment in the video that is less than a duration threshold (for example, 2 seconds) from the first time and a second video segment that is less than a duration threshold from the second time can also be obtained. For example, when the first time is 10:00, the corresponding first video segment can be a video segment consisting of the video content from 09:58 to 10:02 in video A; when the second time is 11:00, the corresponding second video segment can be a video segment consisting of the video content from 10:58 to 11:02 in video A.
[0135] In step S10342A, speech recognition processing is performed on the first video clip to obtain a first text, and integrity detection processing is performed on the first text to obtain a first dialogue integrity detection result. The first time is adjusted according to the first dialogue integrity detection result to obtain a third time.
[0136] In some embodiments, in order to avoid incomplete dialogue in the edited video clip (for example, a sentence is cut off by 20%), after obtaining the first video clip, the first video clip can be subjected to speech recognition processing to convert the sound included in the first video clip into the corresponding first text. Then, the first text is subjected to completeness detection processing (for example, determining whether the first text lacks a subject, whether the narrative is complete, etc.) to obtain a first dialogue completeness detection result. Subsequently, the first time can be adjusted according to the first dialogue completeness detection result to obtain a third time. For example, when it is determined based on the first dialogue completeness detection result that the dialogue has not ended at the first time, the first time can be moved back a few seconds (the number of seconds moved corresponds to the dialogue completeness) to obtain the third time.
[0137] In step S10343A, speech recognition processing is performed on the second video clip to obtain a second text, and integrity detection processing is performed on the second text to obtain a second dialogue integrity detection result. The second time is adjusted according to the second dialogue integrity detection result to obtain a fourth time.
[0138] In some embodiments, after obtaining the second video clip, voice recognition processing can be performed on the second video clip to convert the sound included in the second video clip into a corresponding second text. Then, the second text can be subjected to integrity detection processing to obtain a second conversation integrity detection result. Subsequently, the second time can be adjusted according to the second conversation integrity detection result to obtain a fourth time. For example, when it is determined based on the second conversation integrity detection result that the conversation has ended at the second time, the second time can be moved forward a few seconds to obtain the fourth time.
[0139] In step S10344A, a file including the video segment between the third time and the fourth time is clipped from the video.
[0140] In some embodiments, taking video A as an example, for video segment 1 in video A, assuming that the start time corresponding to video segment 1 is 10:00 and the end time is 11:00, and the value of m determined according to the speed of change of the plot content of video segment 1 is 2 seconds, then the first time is 09:58 and the second time is 11:02. Then, assuming that the first time is adjusted according to the first dialogue completeness detection result, the third time obtained is 09:55, and the second time is adjusted according to the second dialogue completeness detection result, the fourth time obtained is 11:04, then a file of the video segment between 09:55 and 11:04 can be edited from video A. In this way, by adjusting the start time and end time of the video segment based on the dialogue completeness detection result, the accuracy of the video editing is improved, and the incomplete dialogue content of the edited video segment is avoided, thereby improving the user's viewing experience and viewing time.
[0141] In other embodiments, Figure 4 Step S1034 shown can be Figure 5B Steps S10341B to S10344B shown in FIG. 1 are implemented by combining Figure 5B The steps shown are explained.
[0142] In step S10341B, a first video segment whose distance from the first time to the video is less than a duration threshold and a second video segment whose distance from the second time to the video is less than a duration threshold are obtained.
[0143] In some embodiments, taking video A as an example, after determining a first time m seconds before the start time corresponding to video segment 1 in video A and a second time m seconds after the end time corresponding to video segment 1, a first video segment in the video that is less than a duration threshold (for example, 2 seconds) from the first time and a second video segment that is less than a duration threshold from the second time can also be obtained. For example, when the first time is 10:00, the corresponding first video segment can be a video segment consisting of the video content from 09:58 to 10:02 in video A; when the second time is 11:00, the corresponding second video segment can be a video segment consisting of the video content from 10:58 to 11:02 in video A.
[0144] In step S10342B, the first video clip is subjected to frame extraction processing to obtain multiple first video frame images, the multiple first video frame images are compared to obtain a first picture integrity detection result, and the first time is adjusted according to the first picture integrity detection result to obtain a fifth time.
[0145] In some embodiments, to avoid incomplete images in the edited video clip, after obtaining the first video clip, the first video clip may be subjected to frame extraction processing to obtain multiple first video image frames (e.g., five first video image frames, wherein the third first video image frame is the video image frame corresponding to the first time). Then, the third first video image frame is compared with the other first video image frames to obtain a first picture completeness detection result. For example, the similarity between the third first video image frame (i.e., the video image frame corresponding to the first time) and the other first video image frames may be compared using Peak Signal to Noise Ratio (PSNR) or Structural Similarity (SSIM) to determine whether the video image frame corresponding to the first time is complete. Subsequently, the first time is adjusted according to the first picture completeness detection result to obtain a fifth time. For example, when it is determined according to the first picture completeness detection result that the video frame corresponding to the first time is incomplete (e.g., lacks the preceding part of the picture content), the first time may be shifted forward by a few seconds (the number of seconds shifted corresponds to the picture completeness) to obtain the fifth time.
[0146] In step S10343B, the second video clip is subjected to frame extraction processing to obtain multiple second video frame images, the multiple second video frame images are compared to obtain a second picture integrity detection result, and the second time is adjusted according to the second picture integrity detection result to obtain a sixth time.
[0147] In some embodiments, after obtaining the second video clip, the second video clip can also be subjected to frame extraction processing to obtain multiple second video image frames (for example, 5 second video image frames, where the third second video image frame is the video image frame corresponding to the second time), and then, the third second video image frame is compared with other second video image frames to obtain a second picture integrity detection result (that is, to determine whether the picture of the video image frame corresponding to the second time is complete), and then, the second time is adjusted according to the second picture integrity detection result to obtain a sixth time. For example, when it is determined according to the second picture integrity detection result that the video frame picture corresponding to the second time is incomplete (for example, the subsequent part of the picture content is missing), the second time can be moved back a few seconds to obtain the sixth time.
[0148] In step S10344B, a file including the video segment between the fifth time and the sixth time is clipped from the video.
[0149] In some embodiments, taking video A as an example, for video segment 1 in video A, assuming that the start time corresponding to video segment 1 is 10:00 and the end time is 11:00, and the value of m determined according to the speed of change of the plot content of video segment 1 is 2 seconds, then the first time is 09:58 and the second time is 11:02. Then, assuming that the first time is adjusted according to the first picture completeness detection result, the fifth time obtained is 09:55, and the second time is adjusted according to the second picture completeness detection result, the sixth time obtained is 11:04, then a file of the video segment between 09:55 and 11:04 can be edited from video A. In this way, by adjusting the start time and end time of the video segment based on the picture completeness detection result, the accuracy of video editing is higher, and the incomplete picture content of the edited video segment is avoided, thereby improving the user's viewing experience and viewing time.
[0150] It should be noted that in actual applications, the start time and end time corresponding to the video clip can be adjusted in combination with the dialogue completeness detection results and the picture completeness detection results. In this way, by comprehensively considering the dialogue completeness and picture completeness, the editing accuracy can be further improved.
[0151] In step S104, based on the expression tags of the video segments of the at least one video, clustering processing is performed on the files of the video segments of the at least one video to obtain a video collection corresponding to the at least one expression tag.
[0152] In some embodiments, the above-mentioned expression tags based on video clips of at least one video can be implemented in the following manner, and the files of the video clips of at least one video are clustered to obtain a video collection corresponding to at least one expression tag: when the number of videos is 1, the files of video clips with the same expression tags in the video are clustered into the same video collection; when the number of videos is multiple, the files of video clips with the same expression tags in multiple videos are clustered into the same video collection, or, for videos of the same type in multiple videos, the files of video clips with the same expression tags in videos of the same type are clustered into the same video collection.
[0153] For example, when the number of videos is 1, such as only video A, after video A is edited to obtain files of multiple video clips, the files of video clips with the same expression labels in video A can be clustered into the same video collection. For example, the files of video clips with the expression label "makes people want to cry" in video A can be clustered into a video collection that makes people want to cry.
[0154] For example, when there are multiple videos, after editing each video to obtain files of multiple video clips corresponding to different videos, the files of video clips with the same expression label in multiple videos can be clustered into the same video collection. For example, the files of video clips with the expression label "makes people want to cry" in multiple videos can be clustered into the same video collection that makes people want to cry (that is, a video collection consisting of all video clips that make people want to cry), or, for videos of the same type in multiple videos (such as documentaries), the files of video clips in the documentaries with the same expression label can be clustered into the same video collection (for example, a video collection consisting of video clips in documentaries that make people want to cry).
[0155] The following describes in detail the video editing processing method provided in the embodiment of the present application from the terminal side.
[0156] In some embodiments, a client (such as a browser or an online video client) is running on the terminal (such as the terminal 400 described above), and a video interface is displayed on the human-computer interaction interface of the client, wherein the video interface is used to play videos or display a video list. In addition, a viewing entrance of a video collection can also be displayed on the human-computer interaction interface of the client, wherein the video collection can be a video collection that is created by the server through implementation. Figure 3 As shown in steps S101 to S104, after obtaining the video collection, the server can send the video collection to the terminal. When the terminal receives a trigger operation from the user for viewing the video collection displayed on the client's human-computer interaction interface, it responds by displaying the video collection on the client's human-computer interaction interface.
[0157] For example, when the terminal receives a keyword input by the user through the viewing entrance of a video collection, it can obtain a video collection matching the keyword from a video collection corresponding to at least one emoticon tag, and play the matched video collection. For example, when the keyword input by the user is "joy", it can obtain a video collection with an emoticon tag "makes people want to laugh" from a video collection corresponding to at least one emoticon tag, and play the video collection that makes people want to laugh.
[0158] For example, when the terminal receives a keyword input by the user through the viewing entrance of the video collection, a video collection matching the keyword can be obtained from the video collection corresponding to at least one emoticon tag. Then, the user's historical behavior information (such as the user's historical viewing records, search records, etc.) can be obtained, and based on the historical behavior information, the type of video that the user may be interested in is determined. Subsequently, video clips of the same type are filtered out from the matching video collection, and a video collection consisting of the filtered video clips is played. For example, when the keyword input by the user is "joy", a video collection with an emoticon tag of "makes people laugh" can be obtained from the video collection corresponding to at least one emoticon tag. Then, when it is determined that the user may be interested in war films based on the user's historical behavior information, war films can be further filtered out from all video collections that make people laugh, and video clips in war films that make people laugh can be played. In this way, by further refining the video collection based on the user's historical behavior information, it can better meet the user's real needs and enhance the user's viewing experience.
[0159] The video editing processing method provided in the embodiment of the present application fragments the video content by recognizing the user's expression, and automatically generates a video collection corresponding to different expression tags. Since the change of the user's expression is the most realistic judgment of the video content, the video editing through user expression recognition can make the judgment of the editing timing (that is, the start time and end time corresponding to the video clip) more accurate, so that the edited video clip is more in line with the user's real needs and improves the user's viewing experience.
[0160] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0161] Video editing technology is a video processing method that obtains one or more video clips in the video by editing the video to be edited. It is often used in video editing scenarios such as short video production and video highlights.
[0162] At present, the video editing solutions provided by related technologies are mainly based on manual methods to judge video content, editing content and merging content. Their efficiency is very low, and the judgment of video content is easily affected by the editor's personal subjective judgment, resulting in low accuracy of video editing.
[0163] In addition, related technologies also provide video editing methods based on picture content, such as using artificial intelligence (AI) to determine the content objects of the video picture and editing the video content based on the identification of the picture content objects. However, the video content edited by AI cannot reflect the user's true emotional perception of the video content, resulting in poor accuracy of video editing and failure to meet the real needs of users.
[0164] In response to the above technical problems, an embodiment of the present application provides a video editing and processing method. When a user watches a video, the user's face is captured through the camera built into the terminal (such as a mobile phone) to obtain a corresponding facial image sequence. Then, expression recognition processing is performed on the facial image sequence to obtain the user's expression, such as joy, anger, crying, happiness, fear, etc., and then the video content is fragmented and edited based on the recognized user expression. Finally, a video collection corresponding to different expressions is generated, such as a collection of scary videos, a collection of happy videos, etc. When the user is watching a video, the real-time expression changes are the most realistic judgment of the video content type, and the results can also be adjusted based on the changes in the expressions of a large number of users. In this way, the video editing method based on user expression recognition can make the judgment of the edited content more accurate, and make the edited video content more in line with the user's real needs, thereby improving the user's content satisfaction, viewing time, etc.
[0165] The video editing processing method provided in the embodiment of the present application is described in detail below.
[0166] For example, see Figure 6 , Figure 6 Schematic diagram of the application scenario of the video editing processing method provided in the embodiment of the present application. Figure 6 As shown, while the user is watching a video, a pop-up window 601 may be displayed on the video playback interface. When the terminal receives a click operation on the "Allow" button 602 displayed in the pop-up window 601 (i.e., the user authorizes the camera to enable the expression recognition function), the camera is called to capture the user's face multiple times. In other words, while the user is watching the video, the camera will capture the user's facial image in real time and perform expression recognition to determine the user's current expression type, such as happiness, fear, sadness, etc.
[0167] In some embodiments, the video editing processing method provided in the embodiments of the present application can label the video content watched by the user in real time according to the type of user expression recognized.
[0168] For example, see Figure 7 , Figure 7 Schematic diagram of the application scenario of the video editing processing method provided in the embodiment of the present application. Figure 7 As shown, when the video plays to 40:30, the user's expression is captured to become happy (for example, the user starts to laugh), and when the video plays to 40:40, the user's expression is captured to return to normal (for example, the user stops laughing). These two time points can be recorded, and a "makes you want to laugh" expression label can be added to the video clip consisting of the video content between these two time points (that is, the video clip from 40:30 to 40:40).
[0169] For example, see Figure 8 , Figure 8 Schematic diagram of the application scenario of the video editing processing method provided in the embodiment of the present application. Figure 8 As shown, when the video plays to 50:40, the user's expression is captured to become sad (for example, the user starts to cry), and when the video plays to 51:40, the user's expression is captured to return to normal (for example, the user stops crying). These two time points can be recorded, and a "makes you want to cry" expression label can be added to the video clip consisting of the video content between 50:40 and 51:40.
[0170] In other embodiments, the video editing processing method provided in the embodiments of the present application can also optimize the results based on big data of human faces, so as to make the optimal label judgment for the content of different video clips.
[0171] For example, for the same video clip, assuming that a total of 300 users' corresponding expressions when watching the video clip are obtained, among which 90% of the users have happy expressions when watching the video clip, 6% of the users have sad expressions when watching the video clip, and 4% of the users have terrifying expressions when watching the video clip, then the expression label corresponding to the video clip can be set to "makes people want to laugh".
[0172] In some embodiments, after obtaining the optimal labeling result corresponding to the video clip, the video can be edited according to the labeling result, wherein the editing process includes rough cutting and fine editing, thereby obtaining a video clip with an expression label. Subsequently, the video clips can be clustered into different video collections according to different dimensions and scene requirements. For example, all video clips with an expression label of "making people want to cry" can be clustered into the same video collection, thereby obtaining a collection of all videos that make people want to cry; all video clips with an expression label of "making people want to laugh" can be clustered into the same video collection, thereby obtaining a collection of all videos that make people want to laugh. Of course, users can also watch these video collections in different scenes, and the video clips in the video collection can be from one movie or multiple movies. In addition, the dimensions of the collection can also be of multiple types, such as a video collection of war movies that makes people want to cry, and a video collection of emotional movies that makes people want to cry.
[0173] For example, see Figure 9 , Figure 9 FIG. 1 is a flow chart of a video editing method according to an embodiment of the present invention. Figure 9 As shown, users need to authorize the camera to enable the expression recognition function while watching the video. For example, when the user clicks Figure 6 When the "Allow" button 602 in the pop-up window 601 is clicked, the client calls the camera of the terminal (such as a mobile phone) to collect the user's facial expressions in real time (that is, collect the user's facial image and perform expression recognition processing), and upload it to the server. After receiving the facial expressions sent by the terminal, the server matches the user's expressions with the corresponding video clips, and calculates the video clips of different expression types based on the effectiveness. Subsequently, the server can edit the video clips under the valid expressions according to certain editing rules to form video clips of different expression label types. Finally, the server can generate video collections of different dimensions according to different user needs and application scenarios, and send the generated video collections to the client to be presented in the client's human-computer interaction interface.
[0174] The following describes the process of facial expression recognition when a user is watching a video.
[0175] In some embodiments, for facial expression recognition, it is usually necessary to collect facial expression data for training using machine learning. Since the degree of expression of user facial expressions in different video contents is different, it is necessary to train user facial expression data in video viewing scenarios, and then use the trained classifier to predict user expressions, so as to obtain matching features of user expressions watching video content.
[0176] For example, see Figure 10 , Figure 10 FIG. 1 is a flow chart of the expression recognition process provided by the embodiment of the present application. Figure 10 As shown in the figure, the expression recognition processing process mainly includes image input, face detection, image preprocessing, feature extraction, pattern classification and recognition results, which are explained below.
[0177] Image input: When a user is watching a video, the user's image is captured through the mobile phone camera to obtain a static image or a dynamic image sequence.
[0178] Face detection: The core data requires facial expressions, but the input image may include non-face content, so a face detection algorithm is needed to determine the face area in the input image.
[0179] Image preprocessing: In order to facilitate subsequent feature extraction and classification, it is necessary to uniformly improve the quality of the input image, eliminate interference information, unify image size, ratio, grayscale value and other information, and normalize the input image.
[0180] For example, see Figure 11 , Figure 11 This is a flow chart of preprocessing an input image according to an embodiment of the present application. Figure 11 As shown in FIG, the preprocessing process for the input image includes detecting key feature points of the face image, scaling, rotating, denoising, and rendering the face image.
[0181] Feature extraction: In order for computers to understand different expressions through features, highly discriminative features need to be input into the computer. The core process of feature extraction is to convert image dot matrices into higher-level image representations, such as shape, motion, color, texture, and spatial structure, and perform dimensionality reduction processing on huge image data while ensuring stability and recognition rate as much as possible.
[0182] For example, feature extraction methods include geometric feature extraction, statistical feature extraction, frequency domain feature extraction and motion feature extraction. Among them, geometric feature extraction is mainly used to locate and measure the significant features of facial images, such as eyes, eyebrows, mouth, etc., to determine their size, distance, shape and mutual proportions, and perform expression recognition; the method based on overall statistical feature extraction mainly emphasizes retaining as much information as possible in the original facial image, and allows the classifier to discover relevant features in the facial image, and obtain features for identification by transforming the entire facial image; the feature extraction method based on the frequency domain is to convert the facial image from the spatial domain to the frequency domain to extract its features (i.e., lower-level features); the extraction method based on motion features is mainly to extract motion features of dynamic image sequences.
[0183] Pattern classification: The extracted feature data is trained through an algorithm to obtain an effective classifier. In the classifier design and selection stage of expression recognition, there are mainly the following methods: using linear classifiers, neural network classifiers, support vector machines, hidden Markov models and other classification and recognition methods.
[0184] Recognition results: The extracted facial expression features are input into the trained classifier, and the classifier is allowed to give the optimal prediction value, that is, to determine the final facial expression type.
[0185] For example, see Figure 12 , Figure 12 FIG. 1 is a schematic diagram of the principle of performing expression recognition on an input image provided by an embodiment of the present application, such as Figure 12As shown in the figure, by performing multiple convolution and downsampling processes on the input image, the feature data corresponding to the input image is obtained, and then the feature data is input into the trained classifier so that the classifier gives the probabilities corresponding to different expression types. Among them, the probability corresponding to the expression type of happiness is the largest, and the expression type corresponding to the input image is determined to be happy.
[0186] In some embodiments, in order to ensure the accuracy of expression tags, the expression data of all users watching the video can be integrated to calculate the tendency of a single expression tag or multiple expression tags for different video clips. For example, a video clip may have a tendency of multiple expression tags, and the tendency degree of different expression tags can be given through user data.
[0187] For example, see Figure 13 , Figure 13 FIG is a schematic diagram of optimizing a video clip according to expressions of multiple users provided in an embodiment of the present application, such as Figure 13 As shown, the expression labels corresponding to the video clips can be optimized according to the expression recognition results of multiple users (for example, the expression labels corresponding to the video clips can be adjusted according to the tendency ratios corresponding to different expressions), and the start time and end time corresponding to the video clips can be optimized (for example, a normal distribution curve can be established according to the generation time and end time of different users' expressions, and the start and end time corresponding to the video clips can be determined based on the normal distribution curve).
[0188] For example, see Figure 14 , Figure 14 Schematic diagram of setting a single expression tag for a single video clip provided by an embodiment of the present application, such as Figure 14 As shown, in the data reporting, each user's expression will be uploaded, but only the label expressions whose number of users reaches a certain level can be determined as valid labels. For example, the effective level U can be set. Only when the number of labels is greater than the effective level U can they be used as valid labels. For example, for a certain video clip, only when the number of expression labels of "frightening" is greater than the effective level U, the expression label "frightening" will be used as the expression label corresponding to the video clip, and other types of expression labels will be deleted.
[0189] For example, see Figure 15 , Figure 15 is a schematic diagram of setting multiple expression tags for a single video clip provided by an embodiment of the present application, such as Figure 15As shown, a video clip may have multiple different types of expression labels. For example, one user may see a happy expression, while another user may see a sad expression. When the number of different types of expression labels exceeds the effective level U, a tendency calculation can be performed. The tendency ratio is obtained by the number of labels of different expression types. For example, fear is m% (for example, m>80) and surprise is n% (for example, n>80). Then, this video clip can be used under both expression categories (that is, the expression labels of "fearful" and "surprising" are set for this video clip at the same time).
[0190] In some embodiments, for the same video clip, since the generation time of user expressions is uncertain (for example, the generation time of different users' expressions may be different), and the duration of different users' expressions is also different, the generation and disappearance of expressions will show a normal distribution on the video playback time axis. Therefore, in the calculation process, when the number of tags is greater than the effective magnitude U, the calculation begins, and the normal distribution curve composed of all tag data greater than U (for example, Figure 16 From the normal distribution curve shown, n% (n is a percentage, and the specific value can be continuously optimized according to the final effect) interval is extracted, and the start time and end time corresponding to the video segment are determined according to the extracted interval.
[0191] In some embodiments, based on the process of generating effective user expressions, the time when the label is generated and ended can be determined, but in specific editing applications, the time when the label is generated cannot be directly used for video slicing, because the video needs to be cut into segments with some picture stories before and after, so rough cutting and intelligent fine cutting are also required in the specific editing.
[0192] For example, see Figure 17 , Figure 17 Schematic diagram of rough cutting and intelligent fine cutting of video clips provided by an embodiment of the present application, such as Figure 17 As shown, the rough cut process involves adding n seconds forward or backward on the timeline of a specific tag. The value of n can be adjusted based on the video content and expression type. This is because different video content has different rhythms and different expressions require different pre-processing information. For example, for fighting videos, the tempo is fast and the user's expression changes quickly, so the value of n can be relatively small. On the other hand, for emotional documentaries, the tempo is slow and the user's expression changes slowly, so the value of n can be relatively large.
[0193] On the basis of the rough cut, fine intelligent adjustments can also be made based on the completeness of the dialogue and the completeness of the picture. Among them, the completeness of the dialogue is mainly based on the completeness of the sound at the beginning of the video content to avoid 10% of the sentence being cut off. For example, intelligent voice recognition can be used to convert the sound into text, and the completeness of the text can be used to determine whether it is a complete sentence; and the completeness of the picture mainly considers the continuity of the shot switching, presenting the complete picture as much as possible, that is, presenting the current shot content in full. For example, intelligent recognition of video pictures can be used to extract frames from the video and compare the extracted frame pictures to determine the degree of difference in the video pictures and determine whether the shot has been switched. For example, PSNR and SSIM can be used for similarity comparison.
[0194] For example, see Figure 18 , Figure 18 This is a schematic diagram of a process for generating different video collections for multiple video clips provided by an embodiment of the present application, such as Figure 18 As shown, video clips with different label tendencies can be stored in a database and synthesized for use according to different usage scenarios. For example, video clips with the same expression type can be searched from the database, and video collections can be automatically generated according to the needs of different scenarios. It can be a complete collection, such as all video clips that make people want to cry; or refined subsets can be generated according to different video types, such as video clips in war films that make people want to cry, and video clips in emotional films that make people want to cry. The splitting dimensions of the subsets can depend on the original classification information, time information, and user viewing volume of the video.
[0195] In some embodiments, the generated video collection can be presented in the client, where the presentation process can be divided into active presentation and passive presentation. Active presentation can present all video collections from a global perspective of the system, and can also allow users to search through different expression dimensions; passive presentation can present corresponding video collections based on the preferences of different users. For example, when it is determined through the user's past viewing history that the user's preference is war films, the recommended video clips that make people cry are video collections formed by video clips of war films.
[0196] The video editing processing method provided in the embodiments of the present application uses a camera to recognize the user's facial expressions, such as joy, anger, crying, happiness, fear, etc., in real time while the user is watching a video. The video content is then segmented and edited based on the recognition of the user's facial expressions, and a video collection is finally generated, such as a collection of scary videos or a collection of happy videos. When the user is watching a video, the real-time changes in facial expressions provide the most accurate judgment of the video content type, and the results can also be optimized based on the changes in the facial expressions of a large number of users. The video editing method based on user facial expression recognition can make the judgment of the edited content more accurate, making the edited content more in line with the user's actual needs, thereby improving user content satisfaction and viewing time.
[0197] The following continues to describe the exemplary structure of the video editing processing device 243 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the video clip processing device 243 of the memory 240 may include: an acquisition module 2431 , an expression recognition module 2432 , a clipping module 2433 and a clustering module 2434 .
[0198] An acquisition module 2431 is used to acquire facial data of at least one video, wherein the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user, and the facial image is collected from the user while the user is watching the video; an expression recognition module 2432 is used to perform expression recognition processing on each facial image sequence to obtain an expression label of at least one video segment in the video; a clipping module 2433 is used to clip the video according to the start time and end time corresponding to each video segment to obtain a file of each video segment; a clustering module 2434 is used to cluster the files of the video segments of at least one video based on the expression labels of the video segments of at least one video to obtain a video collection corresponding to at least one expression label.
[0199] In some embodiments, the expression recognition module 2432 is also used to perform the following processing on each frame of facial image in the facial image sequence: perform face detection processing on the facial image to obtain the facial area in the facial image; perform feature extraction on the facial area to obtain corresponding facial feature data; call the trained classifier to perform prediction processing based on the facial feature data to obtain the expression label corresponding to the facial image; determine the corresponding video segment in the video based on the collection time period corresponding to the consecutive facial images with the same expression label in the facial image sequence, and use the consecutive identical expression labels as the expression label of the video segment.
[0200] In some embodiments, the expression recognition module 2432 is also used to extract features from the facial area to obtain a corresponding facial feature vector; wherein the dimension of the facial feature vector is smaller than the dimension of the facial area, and the facial feature vector includes at least one of the following: shape feature vector, motion feature vector, color feature vector, texture feature vector, and spatial structure feature vector.
[0201] In some embodiments, the expression recognition module 2432 is also used to detect key feature points in the face area, and align and calibrate the face image included in the face area based on the key feature points; and edit the face area including the aligned and calibrated face image, wherein the editing process includes at least one of the following: normalization, cropping, and scaling.
[0202] In some embodiments, the video clip processing device 243 also includes a determination module 2435, which is used to determine the number of each type of expression labels included in the video clip when the same expression labels of the video clip are determined through facial image sequences corresponding to multiple users; the determination module 2435 is also used to treat the expression labels whose number among the multiple expression labels is less than a quantity threshold as invalid labels; the video clip processing device 243 also includes a deletion module 2436, which is used to delete invalid labels.
[0203] In some embodiments, the determination module 2435 is further used to determine the number of each type of expression labels included in the video clip when multiple expression labels of the video clip are determined by facial image sequences corresponding to multiple users; the video clip processing device 243 also includes a screening module 2437, which is used to screen out expression labels whose number is greater than a quantity threshold from multiple expression labels; the determination module 2435 is further used to determine the tendency ratio corresponding to each filtered expression label; and to treat the expression labels whose tendency ratio is less than the ratio threshold among the multiple filtered expression labels as invalid labels; the deletion module 2436 is further used to delete invalid labels.
[0204] In some embodiments, the determination module 2435 is further used to perform the following processing on the video clip: when the same expression label of the video clip is determined by the facial image sequences corresponding to multiple users, the start time and end time corresponding to the video clip are determined in the following manner: based on the start time and end time of each user's expression label, a normal distribution curve is established; with the symmetry axis of the normal distribution curve as the center, an n% interval of the normal distribution curve is extracted, and the time corresponding to the start point of the interval is determined as the start time of the video clip, and the time corresponding to the end point of the interval is determined as the end time of the video clip; wherein n is a positive integer and satisfies 0 <n<100。
[0205] In some embodiments, the clustering module 2434 is also used to cluster files of video segments with the same expression tag in the video into the same video collection when the number of videos is 1; and to cluster files of video segments with the same expression tag in multiple videos into the same video collection when the number of videos is multiple, or, for videos of the same type in multiple videos, cluster files of video segments with the same expression tag in videos of the same type into the same video collection.
[0206] In some embodiments, the determination module 2435 is also used to determine the value of m based on the speed at which the plot content of the video clip changes; determine a first time m seconds before the start time in the video; determine a second time m seconds after the end time in the video; and the editing module 2433 is also used to edit the video based on the first time and the second time.
[0207] In some embodiments, the editing module 2433 is further used to obtain a first video segment in the video that is less than a duration threshold away from a first time, and a second video segment that is less than a duration threshold away from a second time; perform speech recognition processing on the first video segment to obtain a first text, perform integrity detection processing on the first text to obtain a first conversation integrity detection result, adjust the first time according to the first conversation integrity detection result to obtain a third time; perform speech recognition processing on the second video segment to obtain a second text, perform integrity detection processing on the second text to obtain a second conversation integrity detection result, adjust the second time according to the second conversation integrity detection result to obtain a fourth time; and edit a file including the video segment between the third time and the fourth time from the video.
[0208] In some embodiments, the editing module 2433 is further used to obtain a first video segment in the video that is less than a duration threshold away from a first time, and a second video segment that is less than a duration threshold away from a second time; perform frame extraction processing on the first video segment to obtain multiple first video image frames, perform comparison processing on the multiple first video frame images to obtain a first picture integrity detection result, and adjust the first time according to the first picture integrity detection result to obtain a fifth time; perform frame extraction processing on the second video segment to obtain multiple second video image frames, perform comparison processing on the multiple second video image frames to obtain a second picture integrity detection result, and adjust the second time according to the second picture integrity detection result to obtain a sixth time; and edit a file including the video segment between the fifth time and the sixth time from the video.
[0209] In some embodiments, the determination module 2435 is also used to perform the following processing for each video clip: when the number of users watching the video is 1, the start time and end time of the user's expression tag are used as the start time and end time corresponding to the video clip; when the number of users watching the video is multiple, the start time and end time corresponding to the video clip are determined based on the start time and end time of the expression tags of multiple users.
[0210] In some embodiments, the acquisition module 2431 is also used to perform the following processing for each video: receiving at least one facial image sequence sent by the terminal of at least one user watching the video, wherein the facial image sequence is obtained by performing multiple facial captures on the user when the terminal is playing the video.
[0211] The following continues to describe the exemplary structure of the video clip processing device 465 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the video clip processing device 465 of the memory 460 may include: a display module 4651 .
[0212] Display module 4651 is used to display a video interface, wherein the video interface is used to play videos or display a video list; display module 4651 is also used to display a viewing entrance for a video collection, wherein the video collection is obtained through the video clip processing method provided by any of the above embodiments; display module 4651 is also used to display a video collection in response to a trigger operation on the viewing entrance of the video collection.
[0213] In some embodiments, the display module 4651 is also used to receive input keywords through the viewing entrance; the video clip processing device 465 also includes an acquisition module 4652, which is used to obtain a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; the display module 4651 is also used to play the matching video collection.
[0214] In some embodiments, the display module 4651 is also used to receive input keywords through the viewing entrance; the acquisition module 4652 is also used to obtain a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; and is used to obtain the user's historical behavior information; the video clip processing device 465 also includes a determination module 4653, which is used to determine the type of video that the user is interested in based on the historical behavior information; the video clip processing device 465 also includes a screening module 4654, which is used to screen out video clips of the same type as the determined type from the matching video collection; the display module 4651 is also used to play a video collection consisting of the screened video clips.
[0215] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated here. Figure 3 、 Figure 4 、 Figure 5A 、 Figure 5B ,or Figure 9 The present invention should be understood by referring to the description of any of the accompanying drawings.
[0216] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video editing method described above in the present invention.
[0217] The embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the embodiment of the present application, for example, Figure 3 、 Figure 4 、 Figure 5A 、 Figure 5B ,or Figure 9 The video clip processing method is shown.
[0218] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0219] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0220] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0221] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0222] To sum up, the embodiment of the present application fragments and edits the video content through the recognition of user expressions, and automatically generates video collections corresponding to different expression tags. Since the changes in user expressions are the most realistic judgment of the video content, video editing through user expression recognition can make the judgment of editing timing (that is, the start time and end time corresponding to the video clip) more accurate, so that the edited video clips are more in line with the user's real needs and improve the user's viewing experience.
[0223] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A video editing method, characterized in that: The method comprises: Acquire facial data of at least one video, wherein the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user, and the facial image is collected from the user while the user is watching the video; Performing prediction processing on each frame of the facial image in each facial image sequence to obtain an expression label corresponding to each frame of the facial image; Determining corresponding video segments in the video based on acquisition periods corresponding to consecutive facial images with the same expression label in the facial image sequence, and using the consecutive identical expression labels as expression labels for the video segments, wherein the number of the consecutive identical expression labels is negatively correlated with a speed of change of the plot content of the video; For each of the video clips, when identical expression tags are determined for the video clips through the facial image sequences corresponding to a plurality of users, the expression tags whose number among the plurality of expression tags included in the video clips is less than a threshold are treated as invalid tags, and the invalid tags are deleted; When determining a plurality of expression tags of the video clip through the facial image sequences corresponding to a plurality of users, screen out the expression tags whose number is greater than the number threshold from the plurality of expression tags, and define the expression tags whose tendency proportion is less than the proportion threshold among the plurality of screened expression tags as invalid tags, and delete the invalid tags; Edit the video according to the start time and end time corresponding to each video segment to obtain a file of each video segment; Based on the expression tags of the video segments of the at least one video, clustering processing is performed on the files of the video segments of the at least one video to obtain a video collection corresponding to the at least one expression tag.
2. The method according to claim 1, characterized in that The performing prediction processing on each frame of the facial image in each facial image sequence to obtain an expression label corresponding to each frame of the facial image includes: For each face image frame in the face image sequence, perform the following processing: Performing face detection processing on the face image to obtain a face area in the face image; Performing feature extraction on the face area to obtain corresponding face feature data; The trained classifier is called based on the facial feature data to perform prediction processing to obtain an expression label corresponding to the facial image.
3. The method according to claim 2, characterized in that The feature extraction of the face region to obtain corresponding face feature data includes: Performing feature extraction on the face area to obtain a corresponding face feature vector; The dimension of the facial feature vector is smaller than the dimension of the facial region, and the facial feature vector includes at least one of the following: a shape feature vector, a motion feature vector, a color feature vector, a texture feature vector, and a spatial structure feature vector.
4. The method according to claim 2, characterized in that Before extracting features from the face area, the method further includes: Detecting key feature points in the face area, and performing alignment and calibration on the face image included in the face area based on the key feature points; The face region including the aligned and calibrated face image is edited, wherein the editing process includes at least one of the following: normalization process, shearing process, and scaling process.
5. The method according to claim 1, wherein Before editing the video according to the start time and end time corresponding to each video segment, the method further includes: The following processing is performed on the video clip: When the same expression label of the video segment is determined by using facial image sequences corresponding to multiple users, the start time and end time corresponding to the video segment are determined by the following method: Based on the start time and end time of each user's expression tag, a normal distribution curve is established; Taking the axis of symmetry of the normal distribution curve as the center, extract n% intervals of the normal distribution curve, and Determining the time corresponding to the start point of the interval as the start time of the video segment, and determining the time corresponding to the end point of the interval as the end time of the video segment; Where n is a positive integer and satisfies 0 <n<100。 6. The method according to claim 1, characterized in that The step of clustering the files of the video segments of the at least one video based on the expression tags of the video segments of the at least one video to obtain a video collection corresponding to at least one expression tag includes: When the number of the videos is 1, clustering the files of the video clips with the same expression label in the videos into the same video collection; When there are multiple videos, the files of video clips with the same expression tags in the multiple videos are clustered into the same video collection; or, for videos of the same type in the multiple videos, the files of video clips with the same expression tags in the videos of the same type are clustered into the same video collection.
7. The method according to claim 1, characterized in that The video clipping process according to the start time and end time corresponding to each video segment includes: Determining the value of m according to the speed at which the plot content of the video clip changes; determining a first time m seconds before the start time in the video; determining a second time m seconds after the end time in the video; The video is edited based on the first time and the second time.
8. The method according to claim 7, characterized in that The editing of the video based on the first time and the second time includes: Acquire a first video segment of the video whose distance from the first time is less than a duration threshold, and a second video segment whose distance from the second time is less than the duration threshold; performing speech recognition processing on the first video clip to obtain a first text, performing integrity detection processing on the first text to obtain a first conversation integrity detection result, and adjusting the first time according to the first conversation integrity detection result to obtain a third time; performing speech recognition processing on the second video clip to obtain a second text, performing completeness detection processing on the second text to obtain a second conversation completeness detection result, and adjusting the second time according to the second conversation completeness detection result to obtain a fourth time; A file including a video segment between the third time and the fourth time is clipped from the video.
9. The method according to claim 7, characterized in that The editing of the video based on the first time and the second time includes: Acquire a first video segment of the video whose distance from the first time is less than a duration threshold, and a second video segment whose distance from the second time is less than the duration threshold; Performing frame extraction processing on the first video clip to obtain a plurality of first video image frames, performing comparison processing on the plurality of first video frame images to obtain a first picture integrity detection result, and adjusting the first time according to the first picture integrity detection result to obtain a fifth time; performing frame extraction processing on the second video clip to obtain a plurality of second video image frames, performing comparison processing on the plurality of second video image frames to obtain a second picture integrity detection result, and adjusting the second time according to the second picture integrity detection result to obtain a sixth time; A file including a video segment between the fifth time and the sixth time is clipped from the video.
10. The method according to claim 1, characterized in that Before editing the video according to the start time and end time corresponding to each video segment, the method further includes: The following processing is performed for each video segment: When the number of users watching the video is 1, the start time and end time of the user's expression tag are used as the start time and end time corresponding to the video segment; When there are multiple users watching the video, the start time and end time corresponding to the video segment are determined based on the start time and end time of the expression tags of the multiple users.
11. The method according to claim 1, wherein The acquiring of facial data of at least one video includes: The following processing is performed for each of the videos: Receive at least one facial image sequence respectively sent by a terminal of at least one user watching the video, wherein the facial image sequence is obtained by performing multiple facial captures on the user when the terminal is playing the video.
12. A video editing method, characterized in that: The method comprises: Display a video interface, wherein the video interface is used to play a video or display a video list; Displaying a viewing entrance for a video collection, wherein the video collection is obtained by the method according to any one of claims 1 to 11; In response to a triggering operation on a viewing entry for the video collection, the video collection is displayed.
13. The method according to claim 12, characterized in that The displaying of the video collection in response to a triggering operation on the viewing entrance of the video collection includes: receiving a keyword inputted through the viewing portal; Obtaining a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; Play the matching video collection.
14. The method according to claim 12, characterized in that The displaying of the video collection in response to a triggering operation on the viewing entrance of the video collection includes: receiving a keyword inputted through the viewing portal; Obtaining a video collection matching the keyword from a video collection corresponding to at least one emoticon tag; Obtain user historical behavior information; Determining the type of videos that the user is interested in based on the historical behavior information; Filtering video clips of the same type from the matching video collection; Play the video collection consisting of the filtered video clips.
15. A video editing device, characterized in that: The device comprises: an acquisition module, configured to acquire facial data of at least one video, wherein the facial data includes at least one facial image sequence, and each facial image sequence includes a facial image of a user, and the facial image is collected from the user while the user is watching the video; An expression recognition module is configured to perform prediction processing on each frame of the facial image in each facial image sequence to obtain an expression label corresponding to each frame of the facial image; determine the corresponding video segment in the video based on the acquisition time period corresponding to consecutive facial images with the same expression label in the facial image sequence, and use the consecutive identical expression labels as the expression label of the video segment; wherein the number of the consecutive identical expression labels is negatively correlated with the speed of change of the plot content of the video; for each video segment, when the same expression label of the video segment is determined through the facial image sequences corresponding to multiple users, treat the expression labels whose number among the multiple expression labels included in the video segment is less than a quantity threshold as invalid labels and delete the invalid labels; when the multiple expression labels of the video segment are determined through the facial image sequences corresponding to multiple users, screen out the expression labels whose number is greater than the quantity threshold from the multiple expression labels, treat the expression labels whose tendency proportion among the multiple screened expression labels is less than a proportion threshold as invalid labels, and delete the invalid labels; An editing module, configured to edit the video according to the start time and end time corresponding to each video segment, to obtain a file of each video segment; The clustering module is used to perform clustering processing on the files of the video segments of the at least one video based on the expression tags of the video segments of the at least one video to obtain a video collection corresponding to at least one expression tag.
16. The device according to claim 15, characterized in that The expression recognition module is also used to perform the following processing on each frame of the facial image in the facial image sequence: perform face detection processing on the facial image to obtain the facial area in the facial image; perform feature extraction on the facial area to obtain corresponding facial feature data; and call the trained classifier based on the facial feature data to perform prediction processing to obtain the expression label corresponding to the facial image.
17. The device according to claim 16, characterized in that The expression recognition module is further used to extract features from the facial area to obtain a corresponding facial feature vector; wherein the dimension of the facial feature vector is smaller than the dimension of the facial area, and the facial feature vector includes at least one of the following: a shape feature vector, a motion feature vector, a color feature vector, a texture feature vector, and a spatial structure feature vector.
18. A video editing and processing device, characterized in that: The device comprises: A display module is used to display a video interface, wherein the video interface is used to play videos or display a video list; The display module is further configured to display a viewing entrance for a video collection, wherein the video collection is obtained by the method according to any one of claims 1 to 11; The display module is also used to display the video collection in response to a triggering operation on the viewing entrance of the video collection.
19. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; The processor is configured to implement the video editing processing method according to any one of claims 1 to 11 or any one of claims 12 to 14 when executing the executable instructions stored in the memory.
20. A computer-readable storage medium, characterized in that Executable instructions are stored, which are used to implement the video editing processing method described in any one of claims 1-11 or any one of claims 12-14 when executed by a processor.
21. A computer program product comprising computer executable instructions, characterized in that When the computer executable instructions are executed by a processor, the video clipping processing method according to any one of claims 1 to 11 or any one of claims 12 to 14 is implemented.
Citation Information
Patent Citations
Method and device for clipping video based on feeling curve
CN107968961A
Hot video annotation processing method and device, computer equipment and storage medium
CN109819325A