Video processing method, device, electronic device and readable storage medium

By matching videos with storyboard and emotional labels, combined with background music synthesis technology, the problem of single and uncontrollable video processing results in the prior art is solved, and diversified video processing and target video generation with high emotional matching are achieved.

CN113572976BActive Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110164639.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2025-05-06
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

Existing deep learning technologies can only generate a single action video during video processing, and the output results are uncontrollable and cannot meet the diverse video processing needs.

Method used

By performing video storyboarding on the initial video, multiple sub-snippets are obtained, and the object identification and emotional label of each sub-snippet are determined, so as to select the target sub-snippet that meets the preset emotional label sequence and its corresponding background music to synthesize the target video.

Benefits of technology

The emotional matching degree of video images and background music in video processing is improved, and videos that meet the identification of different target objects and the emotions of targets can be generated according to needs. It is not limited to a single object, but also improves the accuracy of object recognition and video display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113572976B_ABST
    Figure CN113572976B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video processing technology for artificial intelligence, and discloses a video processing method, device, electronic device, and readable storage medium. The video processing method includes: obtaining an initial video to be processed, performing video storyboarding on the initial video to obtain multiple sub-segments; for each sub-segment, determining an object identifier and an emotion label corresponding to the sub-segment; determining at least one target sub-segment from the at least one sub-segment; obtaining target background music corresponding to the emotion label sequence, and synthesizing the target background music and the target sub-segment to obtain a target video. The video processing method provided in the present application can generate target videos that meet different target object identifiers and target emotions according to requirements, and is not limited to a single object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology. Specifically, the present application relates to a video processing method, device, electronic device and readable storage medium. Background Art

[0002] With the development of computer and network technology, the functions of electronic devices are becoming more and more diverse. For example, users can perform video editing through electronic devices.

[0003] Currently, deep learning can be used to process videos based on neural networks, for example, using video GAN (Generative Adversarial Networks).

[0004] However, the current deep learning method for video processing can only generate a video of a single person's action, and the output result is uncontrollable. Summary of the invention

[0005] The purpose of this application is to solve at least one of the above technical defects, and the following technical solutions are proposed:

[0006] In a first aspect, a video processing method is provided, comprising:

[0007] Obtaining an initial video to be processed, and performing video storyboarding on the initial video to obtain multiple sub-segments;

[0008] For each sub-segment, determining an object identifier and an emotion label corresponding to the sub-segment;

[0009] Determine at least one target sub-segment from at least one sub-segment; wherein the emotion label of the target sub-segment corresponds to a preset emotion label sequence; and the emotion label sequence includes at least one target emotion label;

[0010] The target background music corresponding to the emotion label sequence is obtained, and the target video is obtained by synthesizing the target background music and the target sub-segment.

[0011] In an optional embodiment of the first aspect, for each sub-segment, determining an object identifier corresponding to the sub-segment includes:

[0012] For each sub-segment, identifying at least one object present in the sub-segment;

[0013] Based on the at least one object that appears, an object identification corresponding to the sub-segment is determined.

[0014] In an optional embodiment of the first aspect, identifying at least one object appearing in the sub-segment includes:

[0015] Performing image detection on the video frame image of the sub-segment to determine the target vector of the sub-segment;

[0016] The target vectors are matched with standard target vectors of at least one object corresponding to the initial video to determine the object appearing in the sub-segment.

[0017] In an optional embodiment of the first aspect, determining an object identifier corresponding to the sub-segment based on at least one object that appears includes:

[0018] determining a valid object among the at least one object that appears;

[0019] The identity corresponding to the valid object is set to the object identity corresponding to the sub-fragment.

[0020] In an optional embodiment of the first aspect, determining a valid object among the at least one object that appears includes:

[0021] Determining the total number of video frame images in the sub-segment;

[0022] Determine a first number of video frame images in which the object appears in the sub-segment;

[0023] If the ratio of the first number to the total number is greater than a first preset ratio, the object is set as a valid object.

[0024] In an optional embodiment of the first aspect, determining the emotion label corresponding to the sub-segment includes:

[0025] Determine the emotion label of each video frame image in the sub-segment;

[0026] The emotion label with the highest frequency of occurrence in the video frame image of the sub-segment is set as the emotion label of the sub-segment.

[0027] In an optional embodiment of the first aspect, acquiring target background music corresponding to the emotion tag sequence includes:

[0028] Determine the target emotion tag with the highest frequency in the emotion tag sequence;

[0029] The target background music corresponding to the target emotion tag with the highest occurrence frequency is obtained from a preset music database; wherein a plurality of background music is set in the music database, and each background music is set with a corresponding emotion tag.

[0030] In an optional embodiment of the first aspect, before obtaining the target video based on the synthesis of the target background music and the target sub-segment, the method further includes:

[0031] For each target sub-segment, determining the emotion scale of each video frame image in the target sub-segment;

[0032] The target video is obtained by synthesizing the target background music and the target sub-segment, including:

[0033] For each target sub-segment, the target sub-segment is edited based on the target emotion scale, so that the emotion scale of the starting video frame image of the edited target sub-segment meets the target emotion scale;

[0034] The target video is synthesized based on the target background music and the edited target sub-segments.

[0035] In an optional embodiment of the first aspect, before obtaining the target video based on the synthesis of the target background music and the target sub-segment, the method further includes:

[0036] Generate at least one close-up video based on the target sub-segment; wherein the close-up video includes a close-up picture of the object in the target sub-segment;

[0037] The target video is obtained by synthesizing the target background music and the target sub-segment, including:

[0038] The target video is synthesized based on the target background music, the target sub-segment and the close-up video.

[0039] In an optional embodiment of the first aspect, obtaining a target video based on the synthesis of the target background music and the target sub-segment includes:

[0040] Determine the order of each target emotion label in the emotion label sequence;

[0041] synthesizing at least one target sub-segment to obtain a target segment based on the order of each target emotion label;

[0042] The target clip and the target background music are synthesized to obtain the target video.

[0043] In a second aspect, a video processing device is provided, comprising:

[0044] A storyboard module is used to obtain an initial video to be processed, and to storyboard the initial video to obtain multiple sub-segments;

[0045] A first determination module is used to determine, for each sub-segment, an object identifier and an emotion label corresponding to the sub-segment;

[0046] A second determination module is used to determine at least one target sub-segment from at least one sub-segment; wherein the emotion label of the target sub-segment corresponds to a preset emotion label sequence; and the emotion label sequence includes at least one target emotion label;

[0047] The synthesis module is used to obtain the target background music corresponding to the emotion label sequence, and obtain the target video based on the synthesis of the target background music and the target sub-segment.

[0048] In an optional embodiment of the second aspect, when the first determining module determines, for each sub-segment, the object identifier corresponding to the sub-segment, specifically:

[0049] For each sub-segment, identifying at least one object present in the sub-segment;

[0050] Based on the at least one object that appears, an object identification corresponding to the sub-segment is determined.

[0051] In an optional embodiment of the second aspect, when identifying at least one object appearing in the sub-segment, the first determining module is specifically configured to:

[0052] Performing image detection on the video frame image of the sub-segment to determine the target vector of the sub-segment;

[0053] The target vectors are matched with standard target vectors of at least one object corresponding to the initial video to determine the object appearing in the sub-segment.

[0054] In an optional embodiment of the second aspect, when the first determination module determines the object identifier corresponding to the sub-segment based on the at least one object that appears, it is specifically configured to:

[0055] determining a valid object among the at least one object that appears;

[0056] The identity corresponding to the valid object is set to the object identity corresponding to the sub-fragment.

[0057] In an optional embodiment of the second aspect, when determining a valid object among the at least one object that appears, the first determining module is specifically configured to:

[0058] Determining the total number of video frame images in the sub-segment;

[0059] Determine a first number of video frame images in which the object appears in the sub-segment;

[0060] If the ratio of the first number to the total number is greater than a first preset ratio, the object is set as a valid object.

[0061] In an optional embodiment of the second aspect, when determining the emotion label corresponding to the sub-segment, the first determination module is specifically configured to:

[0062] Determine the emotion label of each video frame image in the sub-segment;

[0063] The emotion label with the highest frequency of occurrence in the video frame image of the sub-segment is set as the emotion label of the sub-segment.

[0064] In an optional embodiment of the second aspect, when the synthesis module obtains the target background music corresponding to the emotion tag sequence, it is specifically used to:

[0065] Determine the target emotion tag with the highest frequency in the emotion tag sequence;

[0066] The target background music corresponding to the target emotion tag with the highest occurrence frequency is obtained from a preset music database; wherein a plurality of background music is set in the music database, and each background music is set with a corresponding emotion tag.

[0067] In an optional embodiment of the second aspect, a third determining module is further included, which is used to:

[0068] For each target sub-segment, determining the emotion scale of each video frame image in the target sub-segment;

[0069] When the synthesis module synthesizes the target video based on the target background music and the target sub-segment, it is specifically used to:

[0070] For each target sub-segment, the target sub-segment is edited based on the target emotion scale, so that the emotion scale of the starting video frame image of the edited target sub-segment meets the target emotion scale;

[0071] The target video is synthesized based on the target background music and the edited target sub-segments.

[0072] In an optional embodiment of the second aspect, a generating module is further included, which is used to:

[0073] Generate at least one close-up video based on the target sub-segment; wherein the close-up video includes a close-up picture of the object in the target sub-segment;

[0074] When the synthesis module synthesizes the target video based on the target background music and the target sub-segment, it is specifically used to:

[0075] The target video is synthesized based on the target background music, the target sub-segment and the close-up video.

[0076] In an optional embodiment of the second aspect, when the synthesis module synthesizes the target video based on the target background music and the target sub-segment, it is specifically used to:

[0077] Determine the order of each target emotion label in the emotion label sequence;

[0078] synthesizing at least one target sub-segment to obtain a target segment based on the order of each target emotion label;

[0079] The target clip and the target background music are synthesized to obtain the target video.

[0080] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video processing method shown in the first aspect of the present application is implemented.

[0081] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the video processing method shown in the first aspect of the present application is implemented.

[0082] The beneficial effects of the technical solution provided by this application are:

[0083] By dividing the video into sections, multiple sub-segments are obtained, the object identification and emotion label of each sub-segment are determined, and a target sub-segment that meets the target object identification and target emotion label is determined from the multiple sub-segments. The target sub-segment and the target background music are synthesized to obtain a target video. The emotion label sequence of the target background music corresponds to the emotion label of the target sub-segment, thereby improving the emotion matching degree between the video picture and the background music in the synthesized target video. In addition, target videos that meet different target object identifications and target emotions can be generated according to needs, not limited to a single object.

[0084] Furthermore, by performing image detection on the video frame images of the sub-segment, the target vector of the sub-segment is obtained, and by matching the target vector with the standard target vector of the object, the object appearing in the sub-segment is determined, which can improve the accuracy of object recognition. Furthermore, by performing close-up recognition on the target sub-segment to generate a close-up video, the target video is synthesized based on the target background music, the target sub-segment and the close-up video, so that the target video contains a close-up of the object, which can more clearly show the emotional changes of the object in the target video and improve the display effect of the target video.

[0085] Additional aspects and advantages of the present application will be partially given in the following description, which will become apparent from the following description, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0087] Figure 1 An application scenario diagram of a video processing method provided in an embodiment of the present application;

[0088] Figure 2 A schematic diagram of a video processing method provided in an embodiment of the present application;

[0089] Figure 3 A schematic diagram of a scheme for an initial video storyboard provided in an example of the present application;

[0090] Figure 4 A schematic diagram of a video processing method provided in an embodiment of the present application;

[0091] Figure 5 A schematic diagram of a solution for clipping a target segment provided in an example of the present application;

[0092] Figure 6 A schematic diagram of a scheme for synthesizing a target video provided in an example provided in this application;

[0093] Figure 7 A schematic diagram of a flow chart of a video processing method in an example provided in an embodiment of the present application;

[0094] Figure 8 A schematic diagram of the structure of a video processing device provided in an embodiment of the present application;

[0095] Fig. 9 A schematic diagram of the structure of an electronic device for video processing provided in an embodiment of the present application. DETAILED DESCRIPTION

[0096] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.

[0097] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. The terms "first" and "second" used herein are used to distinguish different features, and do not limit the order or quantity of features, and the number of features corresponding to "first" and "second" may be the same or different. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it may be directly connected or coupled to other elements, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0098] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0099] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0100] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0101] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0102] The solution provided in the embodiments of the present application involves artificial intelligence video processing technology, which is specifically described through the following embodiments.

[0103] The task of intelligent video production is to input a long video and generate a short video (video highlights) of related content through an algorithm. Highlights of characters (such as short videos related to a TV series on Bilibili or Weishi) generally include 10 to 30 seconds of interaction between the male and female protagonists, various cute or acting highlights of a star, etc. Traditional highlights are composed of steps such as manually selecting materials (image sequences or video clips), determining the order of materials, selecting soundtracks, and video synthesis. The time from collecting materials to the finished product to edit a video ranges from 3 hours to 10 hours. Generating video highlights is a very labor-intensive and brain-intensive task.

[0104] Currently, videos can be processed through manual processing or deep learning videoGAN.

[0105] The artificially generated model requires manual screening of materials, manual arrangement of materials, selection of music, etc. to synthesize the video.

[0106] VideoGAN can generate short videos of a single person in motion with the same background. The specific approach is: first, the camera's motion is not considered (the background does not move). Based on this assumption, the entire video is composed of a static background and a dynamic foreground. A two-stream architecture is designed to generate the background and foreground separately. The foreground f(z) and the background b(z) are fused as follows: a mask m(z) is used for linear fusion. The model inputs a random noise, and after two-stream generation and fusion, the output video is obtained.

[0107] The main problems of manual methods are: 1) manual analysis of basic materials is required: material analysis needs to be performed from the original video, such as video storyboards, which people are involved, and judgment of near and far shots, etc.; 2) manual design of the target video plot sequence is required: the final highlight collection style (such as expression collection, storyline collection, character mashup collection, etc.) needs to be designed in advance; 3) manual selection of the final materials and arrangement of the materials are required; 4) manual music arrangement: manual music arrangement is performed according to the materials and the scenario that needs to be presented.

[0108] The problems with videoGAN are: 1) The results are limited: it can only generate a video of a single person’s action, and does not support the generation of rich and colorful video styles; 2) It is difficult to migrate to new dramas: for each character in a TV series, it is necessary to re-label the data and train the two-way branch of GAN so that the generated characters and environments are consistent with the TV series; 3) The results are unpredictable: the input is only a random number perturbation, and the output result is not controllable; 4) Manual music is required: there is no music information, and manual music is still required.

[0109] This application is based on deep learning face recognition and face attribute recognition capabilities. By performing character recognition, emotion recognition, and music emotion recognition on the video, the target character's expression clips and background music are extracted to synthesize the video, realizing an automatic generation method for character highlights. It has the following effects:

[0110] 1) Reduce manual input, manual material selection and manual editing and synthesis steps;

[0111] 2) Materials can be screened through pre-trained deep learning models such as face recognition, music emotion recognition, and face emotion recognition;

[0112] 3) In the overall framework, different editing effects (such as different emotions, different faces alternating near and far, etc.) can be obtained by adding additional facial attributes to filter the material.

[0113] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0114] like Figure 1 As shown, the video processing method of the present application can be applied to Figure 1 In the scenario shown, specifically, the terminal 10 obtains the initial video to be processed, and sends the initial video to the server 20. The server 20 performs video storyboarding on the initial video to obtain multiple sub-segments. For each sub-segment, the object identifier and emotion label corresponding to the sub-segment are determined. The server 20 obtains the target background music corresponding to the emotion label sequence, obtains the target video based on the synthesis of the target background music and the target sub-segment, and the server 20 returns the target video to the terminal 10.

[0115] Figure 1 In the scenario shown, the above video processing method can be performed in the terminal and the server, or can be performed only in the server or the terminal.

[0116] Those skilled in the art can understand that the “terminal” used here can be a mobile phone, a tablet computer, a PDA (Personal Digital Assistant), a MID (Mobile Internet Device), a laptop computer (Laptop), a desktop computer (desktop computer), a tablet computer (Tablet Personal Computer), a smart TV (Smart TV), a smart watch, a smart car-mounted device, etc.; the “server” can be implemented with an independent server or a server cluster consisting of multiple servers, or a cloud server.

[0117] A possible implementation method is provided in the embodiment of the present application, such as Figure 2 As shown, a video processing method is provided, which can be executed by a terminal or a server, or by the terminal and the server together, and can include the following steps:

[0118] Step S201, obtaining an initial video to be processed, and performing video storyboarding on the initial video to obtain a plurality of sub-segments.

[0119] Among them, video storyboarding refers to dividing the initial video according to different lens conversions to obtain sub-segments under the storyboard of different scenes, which can also be called scene segmentation.

[0120] Specifically, pyscenedetect can be used to storyboard the initial video, wherein pyscenedetect is a code library that can detect and analyze scenes in a video and is used to segment the video into scenes.

[0121] like Figure 3 As shown, the total of 6 frames of the initial video can be divided into two storyboards 301 and 301. Two characters A and B appear in one storyboard 301, and another character C appears in the other storyboard 302.

[0122] Step S202: for each sub-segment, determine the object identifier and emotion label corresponding to the sub-segment.

[0123] Among them, the object can be a person, and the object identifier corresponding to the sub-segment can be a character identifier that mainly appears in the sub-segment. The character identifier can be used to indicate the identity of the person. For example, if the initial video is a TV series video, the character identifier can be the character's name, etc.

[0124] Among them, the emotion label can be a category used to represent the emotion of the object. For example, the emotion of a character can be happy, angry, sad, surprised, etc.

[0125] Specifically, the object identification can be determined by performing face recognition on the sub-segment; the emotion label can be determined by a neural network for recognizing facial emotions in images. The specific process of determining the object identification and emotion label will be described in detail below.

[0126] Step S203, determining at least one target sub-segment from at least one sub-segment; the emotion label of the target sub-segment corresponds to a preset emotion label sequence; and the emotion label sequence includes at least one target emotion label.

[0127] The preset emotion label sequence may be a randomly generated sequence or a user-defined emotion label sequence that needs to be generated. The emotion label sequence may include at least one target emotion label.

[0128] For example, the target object is identified as "character A", and the emotion label sequence may be "angry-happy-sad", then the three target sub-segments determined are segments of "character A", and the emotion labels are "angry", "happy" and "sad" respectively.

[0129] Specifically, the object identifier of each sub-segment may be matched with the target object identifier, and the emotion label of each sub-segment may be matched with each target emotion label in the emotion label sequence to obtain the target sub-segment.

[0130] Step S204, obtaining target background music corresponding to the emotion tag sequence, and synthesizing the target background music and the target sub-segment to obtain a target video.

[0131] Specifically, a music database may be pre-set, wherein a plurality of background music is set in the music database, each background music is set with a corresponding emotion label, and the target background music is determined from the background music database.

[0132] Specifically, the target sub-segments may be first synthesized into a target segment, and then the target segment and the target background music may be synthesized into a target video.

[0133] In the above embodiment, multiple sub-segments are obtained by storyboarding the video, the object identification and emotion label of each sub-segment are determined, and a target sub-segment that meets the target object identification and target emotion label is determined from the multiple sub-segments, and the target sub-segment and the target background music are synthesized to obtain a target video, and the emotion label sequence of the target background music corresponds to the emotion label of the target sub-segment, thereby improving the emotion matching degree between the video picture and the background music in the synthesized target video; in addition, target videos that meet different target object identifications and target emotions can be generated according to needs, not limited to a single object.

[0134] The process of identifying the object identifier of each sub-segment will be described below in conjunction with specific embodiments.

[0135] A possible implementation method is provided in the embodiment of the present application, such as Figure 4 As shown, step S202 of determining the object identifier corresponding to each sub-segment may include:

[0136] Step S210: for each sub-segment, identifying at least one object appearing in the sub-segment.

[0137] Specifically, the step S210 of identifying at least one object appearing in the sub-segment includes:

[0138] (1) performing image detection on the video frame image of the sub-segment to determine the target vector of the sub-segment;

[0139] (2) Matching the target vectors with the standard target vectors of at least one object corresponding to the initial video to determine the object appearing in the sub-segment.

[0140] The object may be a person, the image detection may be face recognition, and the target vector may be a face vector.

[0141] Specifically, a face recognition network can be used to perform face recognition on multiple video frame segments of a sub-segment to obtain face vectors that appear in the sub-segment. For each face vector, the face vector can be matched with the standard face vector of the initial video, that is, the above-mentioned standard target vector, to determine the object that appears in the sub-segment.

[0142] Specifically, a standard image of the object may be extracted from the initial video in advance to generate a corresponding standard target vector.

[0143] For example, if the initial video is a TV series, a frontal face image of a character can be extracted from the TV series, and a standard target vector of the character can be generated based on the frontal face image.

[0144] Specifically, a cosine similarity algorithm may be used. If the similarity between the target vector and the standard target vector is greater than a preset similarity threshold, the object corresponding to the standard target vector is the object of the target vector.

[0145] Step S220: determining an object identifier corresponding to the sub-segment based on the at least one object that appears.

[0146] Specifically, an object that appears in a sub-segment may have a small appearance time or probability and cannot be counted as a valid object in the sub-segment.

[0147] For example, if the initial video is a TV series, in a sub-segment of the TV series, a character may only appear in one frame, and the character may be just a passerby, and cannot be counted as a valid object in the sub-segment.

[0148] Specifically, determining the object identifier corresponding to the sub-segment based on the at least one object that appears in step S220 may include:

[0149] (1) determining a valid object among at least one object that appears;

[0150] (2) The identity identifier corresponding to the valid object is set as the object identifier corresponding to the sub-fragment.

[0151] The valid object is an object that can be used to represent the content of a sub-segment.

[0152] Specifically, determining a valid object among the at least one object that appears may include:

[0153] a. determining the total number of video frame images in the sub-segment;

[0154] b. determining a first number of video frame images in which the object appears in the sub-segment;

[0155] c. If the ratio of the first quantity to the total quantity is greater than a first preset ratio, the object is set as a valid object.

[0156] It is understandable that a sub-segment may have more than one object identifier. For example, if role A and role B in a sub-segment are both valid objects, the object identifier of the sub-segment is "role A+role B".

[0157] In the above embodiment, by performing image detection on the video frame image of the sub-segment to obtain the target vector of the sub-segment, and by matching the target vector with the standard target vector of the object to determine the object appearing in the sub-segment, the accuracy of object recognition can be improved.

[0158] The process of determining the object identifier corresponding to the sub-segment is described below with reference to an example.

[0159] In one example, the initial video is a TV series, and the process of identifying the object identifier of each sub-segment may include the following steps:

[0160] 1) Use the retinanet open source face detection model to perform face detection on the sub-segment (i.e. the above-mentioned image detection);

[0161] 2) Using the insightface open source face recognition model, the face frame obtained by face detection is used to extract the face embedding (i.e. the target vector mentioned above) through the face recognition model;

[0162] 3) For this drama, take a face picture of each of the target characters, such as the male lead, the second male lead, the female lead, and the second female lead, and extract the embedding of the face recognition model (i.e., the standard target vector of the object). Create a face seed library for the target characters and record the information of character ID-face embedding (i.e., establish the relationship between each object identifier and the standard target vector);

[0163] 4) Compare the face embeddings detected in all the storyboards with the embeddings in the seed library (using cosine similarity). The retrieval images with a value greater than the specified threshold thr are designated as the person ID (i.e., object identifier) ​​corresponding to the seed library.

[0164] 5) For each storyboard, take the face whose appearance ratio in the storyboard (the total number of times it appears in the video frame image of the sub-segment / the number of images in the sub-segment) is greater than the specified threshold thrFace (i.e., the first preset ratio) as the face ID (i.e., object identifier) ​​of the final storyboard.

[0165] 6) According to step 5), the face IDs of all video segments are obtained (i.e., the object identifiers of all sub-segments are determined).

[0166] The above example illustrates the process of determining the object identification of the sub-segment, and the process of determining the emotion label will be described below in conjunction with the embodiment.

[0167] A possible implementation method is provided in an embodiment of the present application. Determining the emotion label corresponding to the sub-segment in step S202 may include:

[0168] (1) Determine the emotion label of each video frame image in the sub-segment;

[0169] (2) The emotion label that appears most frequently in the video frame image of the sub-segment is set as the emotion label of the sub-segment.

[0170] Specifically, the representativeness of the emotion label can be measured by the frequency of occurrence, and the emotional expression with the highest frequency of occurrence can be used as the emotion label of the sub-segment.

[0171] For example, there are 10 video frames in a sub-segment, of which 6 frames have the emotion label "happy", 2 frames have the emotion label "angry", and 2 frames have the emotion label "sad", then the emotion label of the sub-segment can be set to "happy".

[0172] The present application provides a possible implementation method, in which step S204 of obtaining the target background music corresponding to the emotion tag sequence may include:

[0173] (1) Determine the target emotion tag with the highest frequency in the emotion tag sequence;

[0174] (2) Obtaining target background music corresponding to the target emotion tag with the highest occurrence frequency from a preset music database; wherein the music database is provided with a plurality of background music, and each background music is provided with a corresponding emotion tag.

[0175] Specifically, the target emotion label with the highest occurrence frequency may be used as the overall label of the target video to be synthesized, and then the target background music corresponding to the target emotion label with the highest occurrence frequency may be selected.

[0176] Specifically, a music database may be pre-built, a plurality of background music may be set, and a corresponding emotion label may be set for each background music.

[0177] A possible implementation method is provided in an embodiment of the present application. Before obtaining the target video based on the synthesis of the target background music and the target sub-segment in step S204, the method may further include: for each target sub-segment, determining the emotion scale of each video frame image in the target sub-segment.

[0178] Among them, the emotional scale can be the intensity of the target's emotions.

[0179] For example, the emotion scale may include levels such as calm, neutral, intense, etc.

[0180] Specifically, the emotion scale of each video frame image can be determined according to the emotion recognition network, and the video frame image is input into the emotion recognition network. If the probability of the output emotion label is larger, the corresponding emotion scale is larger.

[0181] Step S204 of synthesizing the target background music and the target sub-segment to obtain the target video may include:

[0182] (1) for each target sub-segment, the target sub-segment is edited based on the target emotion scale, so that the emotion scale of the starting video frame image of the edited target sub-segment meets the target emotion scale;

[0183] (2) The target video is obtained by synthesizing the target background music and the edited target sub-segments.

[0184] The target emotional scale may be an emotional scale that the target sub-segment needs to have.

[0185] In the specific implementation process, for each target sub-segment, the emotional scale of the starting video frame image before editing may not meet the target emotional scale, so the target sub-segment can be edited so that the emotional scale of the starting video frame image meets the target emotional scale.

[0186] like Figure 5 As shown, Figure 5 The change of the emotional scale of different frames of the target sub-segment is shown. If the emotional scale of a video frame image corresponding to 501 in the figure meets the target emotional scale, the video frame image before 501 shown in the figure can be removed, and the video frame image 501 is used as the starting video frame image. The shaded part in the figure can be the edited target sub-segment.

[0187] A possible implementation method is provided in an embodiment of the present application. Before obtaining the target video based on the synthesis of the target background music and the target sub-segment in step S204, the method may further include: generating at least one close-up video based on the target sub-segment.

[0188] The close-up video includes a close-up picture of the object in the target sub-segment, and the close-up video may be a video that specifically displays the details of the object.

[0189] If the object is a person, the close-up video may be a face close-up video, showing the details of the person's face.

[0190] In a specific implementation process, it may be first determined whether the target sub-segment includes a close-up video. If the target sub-segment does not include a close-up video, a close-up video may be generated based on the target sub-segment.

[0191] Specifically, a video frame image in the target sub-segment may be selected for close-up recognition to generate a close-up video.

[0192] Step S204 of synthesizing the target background music and the target sub-segment to obtain the target video may include:

[0193] The target video is synthesized based on the target background music, the target sub-segment and the close-up video.

[0194] Specifically, the close-up video is generated based on the target sub-segment, or the target sub-segment itself includes the close-up video, and the close-up video also conforms to the target object identification and the target object emotion.

[0195] In this embodiment, a close-up video is generated by performing close-up recognition on the target sub-segment, and the target video is synthesized based on the target background music, the target sub-segment and the close-up video. The target video includes a close-up of the object, which can more clearly show the emotional changes of the object in the target video and improve the display effect of the target video.

[0196] The present application provides a possible implementation method, in which step S204 of synthesizing the target background music and the target sub-segment to obtain the target video may include:

[0197] (1) Determine the order of each target emotion label in the emotion label sequence;

[0198] (2) synthesizing at least one target sub-segment to obtain a target segment based on the order of each target emotion label;

[0199] (3) The target clip and the target background music are synthesized to obtain the target video.

[0200] by Figure 6 As shown in the example, the emotion label sequence includes the emotion changes of "angry-happy-sad", then multiple target sub-segments can be sorted according to the emotion changes in the emotion label sequence, the emotion label of the target sub-segment 601 is "angry", the emotion label of the target sub-segment 602 is "happy", and the emotion label of the target sub-segment 603 is "sad", and the target sub-segments are sorted and synthesized to generate the target segment; then the target segment and the target background music 604 are synthesized to obtain the target video.

[0201] To better understand the above video processing methods, Figure 7 As shown, an example of a video processing method of the present invention is described in detail below:

[0202] In an example, taking the object in the video as a person as an example, the video processing method provided by the present application may include the following steps:

[0203] 1) Obtain the initial video to be processed;

[0204] 2) Storyboarding the initial video to obtain multiple sub-segments;

[0205] 3) Perform face recognition on each sub-segment and obtain close-up of the face;

[0206] 4) Identify the facial emotions of each sub-segment, i.e., emotion labels;

[0207] 5) Select the target person, that is, determine the target object identification;

[0208] 6) Select emotion, i.e. determine the target emotion label;

[0209] 7) Determine a target sub-segment from multiple sub-segments based on the target object identifier and the target emotion label. At this time, the target sub-segment contains a close-up of a face, i.e., the collection of people + emotion materials shown in the figure;

[0210] 8) Constructing a music database; namely, the emotion-music material library shown in the figure;

[0211] 9) Obtain the target background music corresponding to the emotion label sequence, that is, a piece of music under the emotion as shown in the figure;

[0212] 10) The target sub-segment and the target background music are synthesized to obtain the target video.

[0213] In order to better understand the above video processing method, the application scenario of this application will be explained below with examples.

[0214] The above-mentioned video processing method obtains multiple sub-segments by dividing the video into sections, determines the object identification and emotion label of each sub-segment, and determines the target sub-segment that meets the target object identification and target emotion label from the multiple sub-segments, synthesizes the target sub-segment and the target background music to obtain the target video, and the emotion label sequence of the target background music corresponds to the emotion label of the target sub-segment, thereby improving the emotion matching degree of the video picture and the background music in the synthesized target video; in addition, target videos that meet different target object identifications and target emotions can be generated according to needs, not limited to a single object.

[0215] Furthermore, by performing image detection on the video frame images of the sub-segment, the target vector of the sub-segment is obtained, and by matching the target vector with the standard target vector of the object, the object appearing in the sub-segment is determined, which can improve the accuracy of object recognition. Furthermore, by performing close-up recognition on the target sub-segment to generate a close-up video, the target video is synthesized based on the target background music, the target sub-segment and the close-up video, so that the target video contains a close-up of the object, which can more clearly show the emotional changes of the object in the target video and improve the display effect of the target video.

[0216] A possible implementation method is provided in the embodiment of the present application, such as Figure 8 As shown, a video processing device 80 is provided, and the video processing device 80 may include: a split mirror module 801, a first determination module 802, a second determination module 803 and a synthesis module 804, wherein:

[0217] The storyboard module 801 is used to obtain an initial video to be processed, and to storyboard the initial video to obtain a plurality of sub-segments;

[0218] A first determination module 802 is used to determine, for each sub-segment, an object identifier and an emotion label corresponding to the sub-segment;

[0219] The second determination module 803 is used to determine at least one target sub-segment from at least one sub-segment; wherein the emotion label of the target sub-segment corresponds to a preset emotion label sequence; and the emotion label sequence includes at least one target emotion label;

[0220] The synthesis module 804 is used to obtain the target background music corresponding to the emotion tag sequence, and synthesize the target background music and the target sub-segment to obtain the target video.

[0221] A possible implementation manner is provided in an embodiment of the present application. When the first determining module 802 determines the object identifier corresponding to each sub-segment, the first determining module 802 is specifically configured to:

[0222] For each sub-segment, identifying at least one object present in the sub-segment;

[0223] Based on the at least one object that appears, an object identification corresponding to the sub-segment is determined.

[0224] An embodiment of the present application provides a possible implementation manner, in which the first determining module 802 is specifically configured to:

[0225] Performing image detection on the video frame image of the sub-segment to determine the target vector of the sub-segment;

[0226] The target vectors are matched with standard target vectors of at least one object corresponding to the initial video to determine the object appearing in the sub-segment.

[0227] A possible implementation manner is provided in an embodiment of the present application. When the first determining module 802 determines the object identifier corresponding to the sub-segment based on at least one object that appears, it is specifically configured to:

[0228] determining a valid object among the at least one object that appears;

[0229] The identity corresponding to the valid object is set to the object identity corresponding to the sub-fragment.

[0230] A possible implementation manner is provided in an embodiment of the present application. When determining a valid object among the at least one object that appears, the first determining module 802 is specifically configured to:

[0231] Determining the total number of video frame images in the sub-segment;

[0232] Determine a first number of video frame images in which the object appears in the sub-segment;

[0233] If the ratio of the first number to the total number is greater than a first preset ratio, the object is set as a valid object.

[0234] A possible implementation manner is provided in an embodiment of the present application. When determining the emotion label corresponding to the sub-segment, the first determination module 802 is specifically configured to:

[0235] Determine the emotion label of each video frame image in the sub-segment;

[0236] The emotion label with the highest frequency of occurrence in the video frame image of the sub-segment is set as the emotion label of the sub-segment.

[0237] A possible implementation method is provided in an embodiment of the present application. When the synthesis module 804 obtains the target background music corresponding to the emotion tag sequence, it is specifically used to:

[0238] Determine the target emotion tag with the highest frequency in the emotion tag sequence;

[0239] The target background music corresponding to the target emotion tag with the highest occurrence frequency is obtained from a preset music database; wherein a plurality of background music is set in the music database, and each background music is set with a corresponding emotion tag.

[0240] A possible implementation manner is provided in an embodiment of the present application, further comprising a third determining module, configured to:

[0241] For each target sub-segment, determining the emotion scale of each video frame image in the target sub-segment;

[0242] When the synthesis module synthesizes the target video based on the target background music and the target sub-segment, it is specifically used to:

[0243] For each target sub-segment, the target sub-segment is edited based on the target emotion scale, so that the emotion scale of the starting video frame image of the edited target sub-segment meets the target emotion scale;

[0244] The target video is synthesized based on the target background music and the edited target sub-segments.

[0245] A possible implementation method is provided in an embodiment of the present application, further comprising a generating module for:

[0246] Generate at least one close-up video based on the target sub-segment; wherein the close-up video includes a close-up picture of the object in the target sub-segment;

[0247] When the synthesis module synthesizes the target video based on the target background music and the target sub-segment, it is specifically used to:

[0248] The target video is synthesized based on the target background music, the target sub-segment and the close-up video.

[0249] A possible implementation method is provided in an embodiment of the present application. When the synthesis module 804 synthesizes the target background music and the target sub-segment to obtain the target video, it is specifically used to:

[0250] Determine the order of each target emotion label in the emotion label sequence;

[0251] synthesizing at least one target sub-segment to obtain a target segment based on the order of each target emotion label;

[0252] The target clip and the target background music are synthesized to obtain the target video.

[0253] The above-mentioned video processing device obtains multiple sub-segments by splitting the video, determines the object identification and emotion label of each sub-segment, and determines the target sub-segment that meets the target object identification and target emotion label from the multiple sub-segments, synthesizes the target sub-segment and the target background music to obtain the target video, and the emotion label sequence of the target background music corresponds to the emotion label of the target sub-segment, thereby improving the emotion matching degree of the video picture and the background music in the synthesized target video; in addition, target videos that meet different target object identifications and target emotions can be generated according to needs, not limited to a single object.

[0254] Furthermore, by performing image detection on the video frame images of the sub-segment, the target vector of the sub-segment is obtained, and by matching the target vector with the standard target vector of the object, the object appearing in the sub-segment is determined, which can improve the accuracy of object recognition. Furthermore, by performing close-up recognition on the target sub-segment to generate a close-up video, the target video is synthesized based on the target background music, the target sub-segment and the close-up video, so that the target video contains a close-up of the object, which can more clearly show the emotional changes of the object in the target video and improve the display effect of the target video.

[0255] The video processing device for pictures of the embodiments of the present disclosure can execute a video processing method for pictures provided by the embodiments of the present disclosure, and the implementation principles are similar. The actions performed by each module in the video processing device for pictures in each embodiment of the present disclosure correspond to the steps in the video processing method for pictures in each embodiment of the present disclosure. For the detailed functional description of each module in the video processing device for pictures, please refer to the description in the corresponding video processing method for pictures shown in the previous text, which will not be repeated here.

[0256] Based on the same principle as the method shown in the embodiment of the present disclosure, an electronic device is also provided in the embodiment of the present disclosure, which may be a terminal or a server. The electronic device may include, but is not limited to: a processor and a memory; a memory for storing computer operation instructions; a processor for executing the video processing method shown in the embodiment by calling the computer operation instructions. Compared with the prior art, the video processing method in the present application can generate target videos that meet different target object identifications and target emotions according to requirements, and is not limited to a single object.

[0257] In an alternative embodiment, an electronic device is provided, such as Fig. 9 As shown, Fig. 9 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0258] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0259] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0260] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.

[0261] The memory 4003 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the application code stored in the memory 4003 to implement the contents shown in the above method embodiment.

[0262] Fig. 9 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0263] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run on a computer, the computer can execute the corresponding content in the above method embodiment. Compared with the prior art, the video processing method in the present application can generate target videos that meet different target object identifications and target emotions according to needs, and is not limited to a single object.

[0264] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0265] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0266] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0267] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0268] The embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that when the computer device executes the computer instructions, the following conditions are achieved:

[0269] Obtaining an initial video to be processed, and performing video storyboarding on the initial video to obtain multiple sub-segments;

[0270] For each sub-segment, determining an object identifier and an emotion label corresponding to the sub-segment;

[0271] At least one target sub-segment is determined from at least one sub-segment; the emotion label of the target sub-segment corresponds to a preset emotion label sequence; the emotion label sequence includes at least one target emotion label;

[0272] The target background music corresponding to the emotion label sequence is obtained, and the target video is obtained by synthesizing the target background music and the target sub-segment.

[0273] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0274] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0275] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a module does not limit the module itself in some cases. For example, a synthesis module may also be described as a "module for synthesizing a target video".

[0276] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

Claims

1. A video processing method, characterized in that: include: Acquire an initial video to be processed and a pre-built object seed library corresponding to the initial video, wherein the object seed library includes an object identifier of each object in at least one object appearing in the initial video and a standard target vector of the object; Performing video storyboarding on the initial video to obtain multiple sub-segments in different scenes; For each sub-segment, perform object recognition on the sub-segment to obtain a target vector of the object appearing in the sub-segment, match the target vector with each standard target vector in the object seed library, and determine the object identifier corresponding to the standard target vector matching the target vector as the object identifier corresponding to the sub-segment; For each sub-segment, determine the emotion label of each video frame image in the sub-segment; set the emotion label with the highest frequency of occurrence in the video frame image of the sub-segment as the emotion label of the sub-segment; according to the object identifier and emotion label corresponding to each sub-segment, determine at least one target sub-segment from the sub-segments in the multiple different scenes; wherein the at least one target sub-segment is a sub-segment including a target object identifier and an emotion label corresponding to a preset emotion label sequence; the emotion label sequence includes at least one target emotion label; The target background music corresponding to the emotion tag sequence is obtained, and the target video is obtained by synthesizing the target background music and the target sub-segment.

2. The method according to claim 1, characterized in that: The matching of the target vector with each standard target vector in the object seed library includes: Determining the similarity between the target vector and each standard target vector in the object seed library; The standard target vector whose corresponding similarity is greater than a preset threshold is determined as the standard target vector that matches the target vector.

3. The video processing method according to claim 1 or 2, characterized in that: For each sub-segment, determining the object identifier corresponding to the standard target vector that matches the target vector as the object identifier corresponding to the sub-segment includes: By performing object recognition on the sub-segment, a valid object among the at least one object that appears is determined, and an object identifier of a standard target vector that matches a target vector of the valid object is determined as an object identifier corresponding to the sub-segment.

4. The video processing method according to claim 3, characterized in that: The determining of a valid object among the at least one object that appears comprises: Determining the total number of video frame images in the sub-segment; Determine a first number of video frame images in which the object appears in the sub-segment; If the ratio of the first number to the total number is greater than a first preset ratio, the object is set as the valid object.

5. The video processing method according to claim 1, characterized in that: The step of acquiring target background music corresponding to the emotion tag sequence includes: Determine the target emotion tag with the highest occurrence frequency in the emotion tag sequence; The target background music corresponding to the target emotion tag with the highest occurrence frequency is obtained from a preset music database; wherein the music database is provided with a plurality of background music, and each background music is provided with a corresponding emotion tag.

6. The video processing method according to claim 1, characterized in that: Before obtaining the target video based on the target background music and the target sub-segment synthesis, the method further includes: For each target sub-segment, each video frame image in the target sub-segment is input into an emotion recognition network to obtain the probability of the emotion label output by the emotion recognition network, and the emotion scale of each video frame image in the target sub-segment is determined according to the probability of the emotion label output by the emotion recognition network, wherein the greater the probability of the emotion label, the greater the emotion scale represented; The step of synthesizing the target background music and the target sub-segment to obtain the target video includes: For each target sub-segment, a target emotion scale is obtained, a first video frame image whose emotion scale in each video frame image of the target sub-segment meets the target emotion scale is determined, and each video frame in the target sub-segment that is located before the first video frame image is deleted to obtain a clipped target sub-segment; The target video is obtained by synthesizing the target background music and the edited target sub-segments.

7. The video processing method according to claim 1, characterized in that: Before obtaining the target video based on the target background music and the target sub-segment synthesis, the method further includes: generating at least one close-up video based on the target sub-segment; wherein the close-up video includes a close-up picture of an object in the target sub-segment; The step of synthesizing the target background music and the target sub-segment to obtain the target video includes: A target video is synthesized based on the target background music, the target sub-segment and the close-up video.

8. The video processing method according to claim 1, characterized in that: The step of synthesizing the target background music and the target sub-segment to obtain the target video includes: Determining the order of each target emotion label in the emotion label sequence; synthesizing the at least one target sub-segment to obtain a target segment based on the order of each target emotion label; The target clip and the target background music are synthesized to obtain the target video.

9. A video processing device, characterized in that: include: A storyboard module is used to obtain an initial video to be processed and a pre-built object seed library corresponding to the initial video, wherein the object seed library includes an object identifier and a standard target vector of each object in at least one object appearing in the initial video; perform video storyboarding on the initial video to obtain sub-segments in multiple different scenes; A first determination module is used to perform object recognition on each sub-segment to obtain a target vector of the object appearing in the sub-segment, match the target vector with each standard target vector in the object seed library, and determine the object identifier corresponding to the standard target vector matching the target vector as the object identifier corresponding to the sub-segment; The second determination module is used to perform emotion recognition on each sub-segment to obtain an emotion label of the sub-segment; and determine at least one target sub-segment from the sub-segments in the multiple different scenes according to the object identifier and the emotion label corresponding to each sub-segment; wherein the at least one target sub-segment is a sub-segment including a target object identifier and an emotion label corresponding to a preset emotion label sequence, and the emotion label sequence includes at least one target emotion label; A synthesis module is used to obtain target background music corresponding to the emotion tag sequence, and synthesize the target background music and the target sub-segment to obtain a target video.

10. The device according to claim 9, characterized in that When the first determination module is used to match the target vector with each standard target vector in the object seed library, it is specifically used to: Determining the similarity between the target vector and each standard target vector in the object seed library; The standard target vector whose corresponding similarity is greater than a preset threshold is determined as the standard target vector that matches the target vector.

11. The device according to claim 9 or 10, characterized in that For each sub-segment, when the first determining module is used to determine the object identifier corresponding to the standard target vector matching the target vector as the object identifier corresponding to the sub-segment, it is specifically used to: By performing object recognition on the sub-segment, a valid object among the at least one object that appears is determined, and an object identifier of a standard target vector that matches a target vector of the valid object is determined as an object identifier corresponding to the sub-segment.

12. The device according to claim 11, characterized in that When the first determination module is used to determine a valid object among the at least one object that appears, it is specifically used to: Determining the total number of video frame images in the sub-segment; Determine a first number of video frame images in which the object appears in the sub-segment; If the ratio of the first number to the total number is greater than a first preset ratio, the object is set as the valid object.

13. The device according to claim 9, characterized in that When the synthesis module is used to obtain the target background music corresponding to the emotion tag sequence, it is specifically used to: Determine the target emotion tag with the highest occurrence frequency in the emotion tag sequence; The target background music corresponding to the target emotion tag with the highest occurrence frequency is obtained from a preset music database; wherein the music database is provided with a plurality of background music, and each background music is provided with a corresponding emotion tag.

14. The device according to claim 9, characterized in that The device also includes: A third determination module is used to input each video frame image in each target sub-segment into an emotion recognition network, obtain the probability of the emotion label output by the emotion recognition network, and determine the emotion scale of each video frame image in the target sub-segment according to the probability of the emotion label output by the emotion recognition network, wherein the greater the probability of the emotion label, the greater the emotion scale represented; When the synthesis module synthesizes the target video based on the target background music and the target sub-segment to obtain the target video, the synthesis module is specifically used to: For each target sub-segment, a target emotion scale is obtained, a first video frame image whose emotion scale in each video frame image of the target sub-segment meets the target emotion scale is determined, and each video frame in the target sub-segment that is located before the first video frame image is deleted to obtain a clipped target sub-segment; The target video is obtained by synthesizing the target background music and the edited target sub-segments.

15. The device according to claim 9, characterized in that The device also includes: A generating module, configured to generate at least one close-up video based on the target sub-segment; wherein the close-up video includes a close-up picture of an object in the target sub-segment; When the synthesis module synthesizes the target video based on the target background music and the target sub-segment to obtain the target video, the synthesis module is specifically used to: A target video is synthesized based on the target background music, the target sub-segment and the close-up video.

16. The device according to claim 9, characterized in that When the synthesis module synthesizes the target video based on the target background music and the target sub-segment to obtain the target video, the synthesis module is specifically used to: Determining the order of each target emotion label in the emotion label sequence; synthesizing the at least one target sub-segment to obtain a target segment based on the order of each target emotion label; The target clip and the target background music are synthesized to obtain the target video.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the video processing method according to any one of claims 1 to 8 is implemented.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the video processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video processing method and device and short video platform

    CN111460219A