Video processing method, device, equipment and medium
By extracting video content features and recommending materials that match these features, the method addresses inefficiencies in manual material selection, ensuring personalized and adaptive video beautification effects.
Patent Information
- Application Number
- JP2024550229
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-02-25
- Filing Date
- 2023-02-21
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-02-21
AI Technical Summary
Existing video clip software requires manual selection and addition of materials, which is time-consuming and reduces processing efficiency, and pre-made templates often fail to adapt intelligently to the user's video content, leading to suboptimal beautification effects.
A video processing method that extracts video content features, recommends materials matching these features, and processes the original video to generate a target video with enhanced personalization and content relevance, using techniques like voice recognition, semantic analysis, and deep learning to identify keywords and styles.
The method ensures that the added materials closely match the video content, providing personalized and adaptive beautification effects, improving processing efficiency and enhancing the video's atmosphere and style.
Smart Images

Figure 0007819337000003 
Figure 0007819337000004 
Figure 0007819337000005
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority from a Chinese patent application bearing application number 202210178794.5, filed on February 25, 2022, the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to the technical field of computer applications, and in particular to video processing methods, devices, equipment and media. [Background technology]
[0003] To enhance or enhance a captured video, users use video clip software to select and add appropriate materials to the video for decoration, but the process of selecting and adding each material increases time and reduces processing efficiency.
[0004] Currently, the related video clip software provides video templates or video decoration schemes for one-click video generation, incorporates the captured video or picture into the selected video template, and automatically clips the beautified video with template effects. Summary of the Invention
[0005] According to some embodiments of the present disclosure, a video processing method is provided, the method including: extracting video content features based on an analysis of an original video; obtaining at least one recommended material matching the video content features; processing the original video based on the recommended material; and generating a target video, which is a video generated after adding the recommended material to the original video.
[0006] According to some other embodiments of the present disclosure, a video processing device is further provided, the device including: an extraction module for extracting video content features based on analysis of an original video; an acquisition module for obtaining at least one recommended material matching the video content features; and a processing module for performing processing on the original video based on the recommended material and generating a target video, which is a video generated after adding the recommended material to the original video.
[0007] According to still further some embodiments of the present disclosure, there is further provided an electronic device, the electronic device including a processor and a memory for storing instructions executable by the processor, the processor being used to read the executable instructions from the memory and execute the instructions to realize a video processing method according to any embodiment of the present disclosure.
[0008] According to some further embodiments of the present disclosure, there is further provided a computer-readable storage medium having a computer program stored therein, the computer program being used to perform a video processing method according to any embodiment of the present disclosure.
[0009] According to some further embodiments of the present disclosure, a computer program is further provided, the computer program including instructions that, when executed by a processor, cause the processor to implement a video processing method according to any embodiment of the present disclosure. [Brief explanation of the drawings]
[0010] The features, advantages, and aspects of the above-described and other embodiments of the present disclosure will become more apparent when reference is made to the following detailed description in conjunction with the drawings. Throughout the drawings, identical or similar reference numerals represent identical or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0011] [Figure 1]1 is a flowchart of a video processing method according to an embodiment of the present disclosure. [Figure 2] 10 is a flowchart of another video processing method according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram of a video processing scenario according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 5] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 7] 10 is a flowchart of another video processing method according to an embodiment of the present disclosure. [Figure 8] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 9] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 10] 10 is a flowchart of another video processing method according to an embodiment of the present disclosure. [Figure 11] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 12] 10 is a flowchart of another video processing method according to an embodiment of the present disclosure. [Figure 13] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 14] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 15] 10 is a flowchart of another video processing method according to an embodiment of the present disclosure. [Figure 16] FIG. 2 is a schematic diagram of another video processing scenario according to an embodiment of the present disclosure. [Figure 17] 1 is a structural schematic diagram of a video processing device according to an embodiment of the present disclosure; [Figure 18] 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the drawings show several embodiments of the present disclosure, it should be understood that the present disclosure may be realized in many forms and should not be construed as being limited to the embodiments described herein, but rather that these embodiments are provided for a clearer and more complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are used for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0013] It should be understood that the steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel, and that method embodiments may include additional steps and / or omit performing steps as shown, and the scope of the present disclosure is not limited in this respect.
[0014] As used herein, the term "comprises" and variations thereof are open inclusions, including but not limited to. The term "based on" means "based at least in part on." The term "in one embodiment" means "at least one embodiment," the term "in another embodiment" means "at least one other embodiment," and the term "in some embodiments" means "at least some embodiments." Relevant definitions of other terms are provided in the following description. It should be noted that concepts such as "first," "second," etc., mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of functions performed by these devices, modules, or units. Those skilled in the art should note that the modifications of "one" and "multiple" mentioned in this disclosure are general, not limiting, and should be understood as "one or more" unless otherwise clearly indicated in the context. The names of messages or information transmitted between multiple devices in the embodiments of the present disclosure are used for descriptive purposes only and are not used to limit the scope of these messages or information.
[0015] The inventors discovered that because pre-made video templates are fixed, the template selected by the user over time may not be able to intelligently adapt to the original video introduced by the user and cannot be used directly. In addition, the number of effects of video templates is limited, and multiple videos may frequently use the same video template, making it impossible to perform effect processing such as adaptive beautification based on the specific video content.
[0016] In order to solve the problem mentioned above that when beautifying a video using conventional technologies, the beautification effect does not match the video content to a high degree, an embodiment of the present disclosure provides a video processing method, in which materials for effect processing are recommended based on the content of the video, so that the video after processing based on the recommended materials has a relatively high matching between the processing effect and the video content, and the processing effects of videos with different contents are obviously different, resulting in a processing effect of "different effects for different videos," thereby meeting the individual needs of video processing effects.
[0017] The method of the present disclosure will be introduced below along with specific examples.
[0018] 1 is a flowchart of a video processing method according to an embodiment of the present disclosure, which may be performed by a video processing device, which may be realized in software and / or hardware, and may generally be implemented in electronic equipment. As shown in FIG. 1, the method includes steps 101 to 103.
[0019] In step 101, video content features are extracted based on the analysis of the original video.
[0020] In some embodiments, in order to perform video effect processing in accordance with the unique characteristics of the video content, video content features are extracted based on analysis of the original video, where the original video is the uploaded video to be effect processed, and the video content features include, but are not limited to, one or more of the following: video audio features, video text features, video image features, video filter features, and subject features included in the video.
[0021] In step 102, at least one recommended material matching the video content characteristics is obtained.
[0022] In some embodiments, at least one recommended material matching the video content characteristics is obtained, including but not limited to one or more of audio material, sticker material, video material, filter material, etc. In practical applications, the manner of obtaining at least one recommended material matching the video content characteristics may vary depending on the scenario, and specific obtaining manners may be exemplarily described in subsequent embodiments and will not be further described here.
[0023] In step 103, the original video is processed based on the recommended material to generate a target video, where the target video is a video generated after adding the recommended material to the original video.
[0024] In this embodiment, a target video is generated by performing video processing on the original video based on the recommended material, i.e., the target video is generated after adding the corresponding recommended material to the original video. In the actual execution process, each material has a corresponding additional track, so that the corresponding material can be added based on the track of the corresponding material. For example, as shown in Table 1 below, each material track is defined by its corresponding field name, type, and description information, and Table 1 is an example of a material track.
[0025] [Table 1]
[0026] In addition, by including corresponding parameters for each recommended material, it becomes easier to make individual adjustments to the display effect when adding material, such as adjusting the size of the material after determining the material area in a subsequent embodiment. For example, the parameters of the text_template material shown in Table 2 below may include a scaling factor, a rotation angle, etc.
[0027] [Table 2]
[0028] Here, in the actual material adding process, different adding methods can be implemented according to different material types, and different adding methods can be distinguished by adding time, adding position, adding frequency, etc., thereby better responding to the corresponding recommended material and video content and showing a strong correlation between the recommended material and the presented video content. Specific adding methods will be exemplarily described in subsequent examples and will not be further described here.
[0029] To sum up, in the video processing method of the embodiment of the present disclosure, after extracting the video content features of the original video, at least one recommended material matching the video content features is obtained, and the recommended material is added to the original video to obtain a target video, thereby adding video material according to the video content of the video, improving the matching degree between the video content and the video material, and realizing personalized effect processing for the video.
[0030] As mentioned above, in the actual implementation process, the video content features will be different in different application scenarios, which will be exemplarily illustrated as follows.
[0031] In one embodiment of the present disclosure, video content features are extracted based on the text content of the original video to enhance the atmosphere of the video.
[0032] In this embodiment, as shown in FIG. 2, extracting video content features based on analyzing the original video includes steps 201-202.
[0033] In step 201, the target audio data of the original video is subjected to voice recognition processing to obtain the corresponding text content.
[0034] After obtaining the target audio data of the original video, in some embodiments, the preset video clip application can identify not only the video track of the original video but also each audio track included in the original video, where each audio track corresponds to a sound source. For example, in the case of original video A, the voices of user a and user b are included, and in this embodiment, the audio track corresponding to the voice of a and the audio track corresponding to the voice of b can be identified.
[0035] In some embodiments, to facilitate processing of each audio track, a video file of the original video retrieves all audio tracks displayed in the video clip application. As can be easily understood, the sound source corresponding to each audio track has a generation time, so in some embodiments, the audio tracks are displayed based on a time axis.
[0036] For example, as shown in FIG. 3, if the audio file of the original video is divided into a video track video and two audio tracks audio1 and audio2, the corresponding audio tracks can be displayed in the video clip application.
[0037] In some embodiments, still referring to Figure 3, all audio tracks are merged based on a time axis to generate total audio data, for example, audio1 and audio2 are merged based on a time axis to generate total audio data complex-audio, which includes all audio data in the original video.
[0038] Of course, as mentioned above, the total audio data is also related to time, so when the first time length of the total audio data is greater than the second time length of the second video of the original video, in order to ensure the consistency of the processing length and avoid some audio data having no corresponding video content, the first time length of the total audio data is trimmed to obtain the target audio data, where the time length of the target audio data is consistent with the second time length.
[0039] For example, still referring to FIG. 3, if the duration of complex-audio is longer than the duration of video, the portion of complex-audio that exceeds the duration of video on the time axis is trimmed to obtain target-audio, and the target-audio is aligned with the video on the time axis to facilitate subsequent video processing.
[0040] Of course, in actual implementation, the audio file corresponding to the original video may include not only audio data of interactions between subjects, but also background sounds, such as music playing in the environment or the sound of cars passing by on the road. Such background sounds are generally unrelated to the video content, which can improve the accuracy of subsequent video content feature extraction and prevent the background voice from interfering with the extraction of video content features (e.g., when extracting video text features, the text content in the background voice may be identified). In some embodiments, the background voice in the original video can be removed.
[0041] In some embodiments, an audio identifier for each audio track is detected, i.e., based on identifying voice features such as the voice spectrum of the audio corresponding to each audio track, the voice features of the audio corresponding to each audio track are matched with voice features corresponding to each predetermined audio identifier, an audio identifier for each audio track is determined based on the matching result, and once a target audio track representing the background music identifier is detected, an integration process can be performed on all audio tracks other than the target audio track based on a time axis to generate total audio data.
[0042] For example, as shown in FIG. 4, continuing with the scenario shown in FIG. 3, if it is determined that the audio identifiers of the audio tracks of the original video are audio1, audio2, and bg-audio, respectively, then bg-audio matches the background music identifier, and therefore, when generating the total audio data, only the audio tracks corresponding to audio1 and audio2 are integrated.
[0043] Of course, in the actual execution process, the target audio data may be obtained by integrating all audio tracks corresponding to the original video, or may be obtained by integrating only audio tracks that satisfy a certain type of pre-defined voice characteristic, etc., and can be specifically set according to the needs of the scenario, and no limitations are imposed here.
[0044] In some embodiments, after obtaining the original video, a voice recognition process is performed on the target audio data of the original video, and the corresponding text content is further obtained, and the text content may be obtained by identifying it using voice recognition technology.
[0045] In step 202, a semantic analysis process is performed on the text content to obtain a first keyword.
[0046] Here, the first keyword can match recommended materials for the video on a content dimension. For example, the first keyword may be an emotional keyword such as "haha, funny," and therefore, based on this first keyword, materials that render emotions can be recommended for the video, such as several laugh-out-loud sticker materials or several fireworks video materials. For example, the first keyword may be a professional field word such as "basin," and based on this first keyword, professional sticker materials introducing the corresponding field can be recommended for the video, making the corresponding professional field word more understandable.
[0047] In some embodiments, a semantic analysis is performed on the text content, and the analyzed semantic result is matched with a pre-defined keyword meaning to determine the first keyword that is successfully matched.
[0048] In some embodiments, to improve the efficiency and accuracy of identifying the first keywords, as shown in FIG. 5 , an Automatic Speech Recognition (ASR) technology can be used to identify the text content of the target audio data to obtain phrases, and a Natural Language Processing (NLP) technology can be used to understand the meaning of the corresponding text phrases and obtain the corresponding first keywords.
[0049] In some embodiments, the material recommended based on the first keyword can ensure the relevance of the recommended material to the video content in the content dimension, thereby better rendering the corresponding video content. For example, as shown in Figure 6 (in which the corresponding first keyword is displayed in the form of subtitles to make the solution easier for those skilled in the art to intuitively understand), after performing semantic analysis based on the text content of the original video, if the first keyword obtained is "haha", the material can be recommended as a sticker material of "applause". Therefore, in the processed video, the "applause" sticker is displayed for the audio of "haha", rendering a more joyful atmosphere, and the added recommended material has a higher degree of consistency with the video content, so that the addition of the recommended material does not seem abrupt.
[0050] In some embodiments of the present disclosure, as shown in FIG. 7, extracting video content features based on analyzing the original video includes steps 701-702.
[0051] In step 701, a voice detection process is performed on the target audio data of the original video to obtain the corresponding spectrum data.
[0052] In some embodiments, it is considered that in some scenarios, the target audio data can embody the content characteristics of the original video even if the target audio data cannot be converted into corresponding text content. For example, if the target audio data includes sounds such as "applause" or "explosion," adding recommended material based on such target audio data can further enhance the atmosphere of the original video in line with the corresponding audio.
[0053] Therefore, it is clear that the target audio data mentioned in the above embodiment can be subjected to voice detection processing, the corresponding spectrum data can be extracted, and some information can be extracted from the spectrum data that cannot be converted into text content but embodies the content characteristics of the video.
[0054] In step 702, an analysis process is performed on the spectral data to obtain a second keyword.
[0055] In some embodiments, an analytical process can be performed on the spectral data to obtain a second keyword, and then recommended materials corresponding to the corresponding spectral data can be obtained based on the second keyword.
[0056] In some embodiments, the spectral data can be input into a deep learning model that has been previously trained based on a large amount of sample data, and a second keyword output by the deep learning model can be obtained.
[0057] In some other embodiments, the acquired spectral data may be matched with spectral data of each preset keyword, and a second keyword corresponding to the spectral data may be determined based on the degree of matching. For example, if the degree of matching between the acquired spectral data and spectral data corresponding to the keyword "explosion" is greater than a preset threshold, the second keyword corresponding to the target audio data may be determined to be "explosion."
[0058] Referring to FIG. 8, when determining recommended materials, a first keyword and a second keyword can be jointly recommended, where the second keyword may be a corresponding second keyword identified based on audio event detection (AED) technology, and the corresponding recommended materials are determined based on the first keyword and the second keyword.
[0059] For example, as shown in FIG. 9 , after performing a voice detection process on the target audio data of the original video and obtaining the corresponding spectrum data, if the corresponding second keyword obtained based on the spectrum data is “explosion”, the matched recommended material is an “explosion” sticker, and therefore, the corresponding “explosion” sticker is displayed on the corresponding video frame, and the video content including the explosion audio is further rendered.
[0060] To summarize, the video processing method of the embodiment of the present disclosure takes any feature that reflects the video content as a video content feature, and the extracted video content feature has a strong correlation with the video content, ensuring the correlation between the recommended material and the video content based on the video content feature, and providing technical support for the personalized processing effect of the video.
[0061] According to the above embodiment, after obtaining the video content characteristics, the recommended material that matches the video content characteristics is further recommended, and the processing effect for the video is determined according to the recommended material. The following describes the determination of the recommended material in combination with a specific example.
[0062] In some embodiments of the present disclosure, as shown in FIG. 10, obtaining at least one recommended material matching a video content characteristic includes steps 1001-1002.
[0063] In step 1001, a video style feature is determined based on the video image of the original video.
[0064] As can be easily understood, even if the video content features are similar, the corresponding video styles are different, so adding similar recommended materials will affect the matching degree with the video content. For example, the first keyword obtained by semantic analysis based on the target audio data of original video S1 is "haha", and the first keyword obtained by semantic analysis based on the target audio data of original video S2 is also "haha". However, the person who utters "haha" in S1 is an animated character, and the person who utters "haha" in S2 is a real person. Therefore, if the recommended materials are adapted to these two styles, it will obviously affect the video processing effect.
[0065] In the embodiment of the present disclosure, in order to ensure the video processing effect, a video style feature is determined based on the video image of the original video, where the video style feature includes, but is not limited to, the image feature of the video content, the theme style feature of the video content, the subject feature included in the video, etc.
[0066] It should be noted that in different application scenarios, the manners of determining video style features based on video images of the original video are different, and an exemplary description is as follows:
[0067] In some embodiments, as shown in FIG. 11 , a convolutional network model is trained in advance based on a large amount of sample data, and a video image is input into the corresponding convolutional network model to obtain the video style features output by the convolutional network model.
[0068] In some embodiments, as shown in FIG. 12, determining video style characteristics based on video images of an original video includes steps 1201-1202.
[0069] In step 1201, an image identification process is performed on the video images of the original video, and at least one shooting target is determined based on the identification result.
[0070] Here, the subject may be the main body included in the video image, and includes, but is not limited to, people, animals, furniture, tableware, etc.
[0071] In step 1202, a weighting calculation is performed for at least one shooting object based on a preset object weight, and a video style feature corresponding to the original video is determined by matching the calculation result with a preset style classification.
[0072] In this embodiment, to determine the video style, the object type of each shooting object can be identified, a preset database can be queried, and an object weight for each shooting object can be obtained, where the database includes each shooting type and corresponding object weight obtained by training based on a large amount of sample data, and a weighting calculation can be performed for at least one shooting object based on the preset object weight, and a pre-set style classification can be matched based on the calculation result, and a video style feature corresponding to the style classification that has been successfully matched can be determined.
[0073] Here, as shown in Figure 13, by extracting multiple video frames from the original video and using the multiple video frames as video images of the original video, the efficiency of style identification can be further improved. In some embodiments, multiple video frames may be extracted from the original video based on a preset time interval (e.g., 1 second), or a corresponding video segment may be extracted at regular intervals based on a preset time length, and the multiple video frames included in the video segment may be used as video images of the original video.
[0074] In some embodiments, after obtaining the corresponding video image, the video image can be input into a pre-trained image smart identification model, and at least one shooting object can be determined based on the identification result. In the figure, the shooting objects include human faces, objects, environments, etc., and further, classification features t1, t2, t3 corresponding to each shooting object are identified, and the corresponding object weights are z1, z2, z3 respectively. The value of t1×z1+t2×z2+t3×z3 is calculated as the calculation result, and based on this calculation result, it is matched with a pre-set style classification to determine the video style feature corresponding to the original video.
[0075] In step 1002, at least one recommended material matching the video style characteristics and the video content characteristics is obtained.
[0076] In some embodiments, after obtaining the video style features, at least one recommended material that matches the video style features and the video content features is obtained, so that the recommended material matches the video content based on the video style features and the video content features, thereby further improving the video processing effect.
[0077] In some embodiments, as shown in FIG. 14, a material library matching the video style characteristics may be first obtained, and at least one recommended material matching the video content characteristics is obtained from the material library, thereby ensuring that the recommended material obtained not only matches the video content but also is consistent with the style of the video.
[0078] For example, when the video style feature is "girl animation," a material library consisting of various materials in girl style that match "girl animation" is obtained, and further, by matching video materials in the material library consisting of various materials in girl style based on the video content feature, it is guaranteed that all of the obtained recommended materials are in girl style.
[0079] In some embodiments of the present disclosure, as shown in FIG. 15, obtaining at least one recommended material matching a video content characteristic includes steps 1501-1504.
[0080] In step 1501, a playback time of a video frame corresponding to a video content feature in an original video is determined, where the video content feature is generated based on the video content of the video frame.
[0081] In some embodiments, since not all video images of each frame contain similar video content features, but the video content features are generated based on the video content of the video frames, determining the playback time of the video frames corresponding to the video content features in the original video facilitates recommending and adding corresponding material only for video frames that contain the corresponding video content features based on this playback time.
[0082] In step 1502, a time identifier is marked for a video content feature based on the playing time of the video frame.
[0083] In some embodiments, time identifiers are marked for video content features based on the playback duration of video frames to facilitate matching of recommended material over time.
[0084] In step 1503, if it is determined that there are multiple video content features corresponding to the time identifier, the multiple video content features are combined into a video feature set, and at least one recommended material matching the video feature set is obtained, where the multiple video content features include the video content feature.
[0085] In some embodiments, once it is determined that there are multiple corresponding video content features for a single time identifier, i.e., for a single video frame corresponding to a single time, the multiple video content features are combined into a video feature set to obtain at least one recommended material that matches the video feature set.
[0086] In some embodiments, a combination based on multiple video content features is generated to generate multiple video content feature combinations (video feature sets), a pre-defined correspondence is consulted to determine whether there is a corresponding enhancement material for each video content feature combination, and if there is no match with the enhancement material, the video content feature combination is divided into individual content features to match the recommended material, and if there is a match with the enhancement material, the enhancement material is used as the corresponding recommended material.
[0087] It should be understood that the enhanced material herein does not necessarily include a simple combination of recommended material corresponding to multiple video content features, but may be other recommended material with a stronger sense of atmosphere generated to further enhance the video atmosphere when there is a correlation between multiple video content features.
[0088] For example, among multiple video content features, if the first keyword corresponding to video content feature 1 is "haha" and the second keyword corresponding to video content feature 2 is "applause," the recommended material jointly determined based on the first keyword and the second keyword is a transition effect material, rather than the sticker material corresponding to "haha" and "applause" respectively mentioned above.
[0089] If step 1504 determines that a corresponding video content feature exists for a time identifier, then at least one recommended material matching the video content feature is obtained.
[0090] In some embodiments, if it is determined that a corresponding video content feature exists for a time identifier, at least one recommended material that matches the video content feature is obtained, i.e., if there is a single video content feature, at least one recommended material that matches alone is obtained.
[0091] Furthermore, after obtaining the corresponding recommended material, when performing video processing on the original video based on the recommended material to generate a target video, a material addition time of the recommended material matching the video content feature is set based on the time identifier of the video content feature, and this addition time is consistent with the display time of the video frame of the corresponding video content feature.
[0092] Furthermore, a target video is generated by clipping the original video based on the material addition time of the recommended material, so that the corresponding recommended material is added only when a video frame with the video content characteristics corresponding to the material is played, thereby avoiding the material addition being inconsistent with the video content.
[0093] In addition, in the actual execution process, some materials, such as music effect materials and transition effect materials, do not have size information, while some materials, such as sticker materials and text materials, have size information. When adding some materials with size information, it is necessary to determine the adding area of these materials with size information so as to avoid occluding important display contents in the video, such as avoiding occluding a person's face in the video frame.
[0094] In some embodiments, if the material type of the recommended material meets the predetermined target type, i.e., if the recommended material has an additional size information attribute, the corresponding material is deemed to meet the predetermined target type, and further, a target video frame corresponding to the material addition time of the recommended material is obtained from the original video, and image recognition is performed on the target video fragment frame to obtain the main area of the subject to be shot, where the main area may be displayed as any position information of the position where the subject to be shot is located, for example, it may be a central coordinate point, or for example, it may be a position range, etc.
[0095] For example, when the material to be added is added based on the first keyword "haha", the shooting target is the vocalization target corresponding to the "haha" audio, and after determining the main area of the shooting target, the material area for adding the recommended material to the target video frame is determined based on the main area of the shooting target.
[0096] Here, in some embodiments, a material type label of the recommended material is determined, and based on this material type label, a pre-defined correspondence is consulted to determine the region features of the material region (e.g., a background region on an image), and the region in the target video frame that matches this region feature is determined as the material region.
[0097] In some other embodiments, an object type label of the object to be photographed is determined, and based on this object type label, a pre-set correspondence is consulted to determine the region features of the material region (for example, if the object to be photographed is a human face type, the corresponding region feature corresponds to above the head, etc.), and the region in the target video frame that matches this region feature is determined as the material region.
[0098] After determining the material area, a clip process is performed on the original video based on the material addition time and material area of the recommended material to generate a target video, and the corresponding material is added to the material area in the video frame corresponding to the material addition time. Here, the material area may be represented as the coordinates of the center point where the material is added to the corresponding video frame, or as the coordinate range of the material where the material is added to the corresponding video frame. Here, the server that determines the body area, etc., does not have to be the same server as the server that performs the style feature identification. To improve identification efficiency, the server that performs the style feature identification may be a local server, and to reduce the computational load of the material addition time and material area analysis, the server identified based on the material addition time of the recommended material may be a remote server.
[0099] 16 , when the recommended materials include F1 and F2 with size attributes, the material addition times of the recommended materials matching the video content features are set as t1 and t2, respectively, based on the time identifiers of the video content features, and the original video is clipped based on the material addition times of the recommended materials to obtain video segments clip1 corresponding to F1 and video segments clip2 corresponding to F2. After clip1 and clip2 are sent to the corresponding server, the corresponding server optimizes the added materials, performs image recognition on the target video fragment frames to obtain the main area of the shooting target, and further determines the material area to add the recommended materials to the target video frames based on the main area of the shooting target. Based on the material addition times and material areas of the recommended materials, target video segments sticker1 corresponding to clip1 and target video segments sticker2 corresponding to clip2 are obtained. The original video is then edited based on sticker1 and sticker2 to obtain the corresponding target videos.
[0100] To summarize, the video processing method of the embodiment of the present disclosure determines video content features, and then determines at least one recommended material that matches the multi-dimensional video content features, ensures the positional and temporal correspondence between the material and the video frames, and further ensures that the video processing effect meets the unique characteristics of the video content.
[0101] To realize the above embodiment, the present disclosure further proposes a video processing device. Figure 17 is a structural schematic diagram of a video processing device according to an embodiment of the present disclosure, which may be realized by software and / or hardware, and may generally be implemented in electronic equipment. As shown in Figure 17, the device includes an extraction module 1710, an acquisition module 1720, and a processing module 1730.
[0102] The extraction module 1710 is used to extract video content features based on the analysis of the original video.
[0103] The acquisition module 1720 is used to acquire at least one recommended material that matches the video content characteristics.
[0104] The processing module 1730 is used to process the original video based on the recommended material and generate a target video, which is a video generated after adding the recommended material to the original video.
[0105] The video processing device according to the embodiments of the present disclosure can execute the video processing method according to any embodiment of the present disclosure, and includes functional modules corresponding to the execution of the method and beneficial effects, which will not be further described herein.
[0106] To realize the above embodiments, the present disclosure further proposes a computer program product including a computer program / instruction, which, when executed by a processor, realizes the video processing method in any of the above embodiments.
[0107] FIG. 18 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure.
[0108] 18 , a structural schematic diagram for implementing an electronic device 1800 according to an embodiment of the present disclosure is shown. The electronic device 1800 according to an embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG. 18 is merely an example and does not impose any limitations on the functionality and scope of use of the embodiment of the present disclosure.
[0109] 18, electronic device 1800 may include a processor (e.g., a central processor, a graphics processor, etc.) 1801, which can perform various appropriate operations and processes based on programs stored in read-only memory (ROM) 1802 or programs loaded from memory 1808 into random access memory (RAM) 1803. RAM 1803 stores various programs and data necessary for the operation of electronic device 1800. Processor 1801, ROM 1802, and RAM 1803 are connected to one another via bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.
[0110] Typically, input devices 1806 including, for example, a touch screen, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc., output devices 1807 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc., memory 1808 including, for example, a magnetic tape, hard disk, etc., and communication devices 1809 may be connected to the I / O interface 1805. The communication devices 1809 may enable the electronic device 1800 to exchange data with other devices by wireless or wired communication. While FIG. 18 shows the electronic device 1800 having various devices, it should be understood that it is not necessary to implement or include all of the devices shown. More or fewer devices may be implemented or included.
[0111] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the method illustrated in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 1809, or may be installed from the memory 1808 or the ROM 1802. When the computer program is executed by the processor 1801, it performs the functions defined in the video processing method of the embodiment of the present disclosure.
[0112] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of both. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of one or more thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in combination with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include, but is not limited to, a propagated data signal bearing computer-readable program code, either in baseband or as part of a carrier. Such a propagated data signal may take various forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which is capable of transmitting, propagating, or transporting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted over any suitable medium, including, but not limited to, electrical wire, optical cable, RF (radio frequency), etc., or any suitable combination thereof.
[0113] In some embodiments, clients and servers may communicate using any network protocol now known or later developed, such as HTTP (HyperText Transfer Protocol), and may be connected in communication (e.g., via a communications network) with digital data in any form or medium. Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and end-to-end networks (e.g., ad hoc end-to-end networks), and any networks now known or later developed.
[0114] The computer-readable medium may be included in the electronic device, or may exist separately from the electronic device.
[0115] The computer-readable medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to extract video content features of the original video, obtain at least one recommended material matching the video content features, and add the recommended material to the original video to obtain a target video, thereby adding video material according to the video content of the video, improving the matching degree between the video content and the video material, and realizing personalized effect processing for the video.
[0116] Computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and general procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When referring to a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0117] The flowcharts and block diagrams in the figures illustrate possible system architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from the order marked in the figures. For example, two blocks shown in succession may actually be executed essentially in parallel or in the reverse order, depending on the functionality involved. It should be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs the specified functions or operations, or in a combination of dedicated hardware and computer instructions.
[0118] The units described in the embodiments of the present disclosure may be implemented in a software or hardware manner, and the names of the units do not necessarily constitute limitations on the units themselves.
[0119] The functionality described herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific general purpose products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0120] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program used by or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of machine-readable storage media include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
[0121] According to some embodiments of the present disclosure, the present disclosure provides a video processing method, the video processing method comprising: Extracting video content features based on an analysis of the original video; obtaining at least one recommended material matching the video content characteristics; and processing the original video based on the recommended material to generate a target video, which is a video generated after adding the recommended material to the original video.
[0122] According to some embodiments of the present disclosure, in the video processing method according to the present disclosure, extracting video content features based on the analysis of the original video includes: performing a voice recognition process on the target audio data of the original video to obtain corresponding text content; and performing a semantic analysis process on the text content to obtain a first keyword.
[0123] According to some embodiments of the present disclosure, in the video processing method according to the present disclosure, extracting video content features based on the analysis of the original video includes: performing a voice detection process on the target audio data of the original video to obtain corresponding spectrum data; and performing an analysis process on the spectral data to obtain a second keyword.
[0124] According to some embodiments of the present disclosure, in a video processing method according to the present disclosure, the method for obtaining target audio data includes: Obtaining all audio tracks of the video file of the original video to be displayed in a video clip application; performing an integration process on all of the audio tracks based on a time axis to generate total audio data; comparing a first time length of the total audio data with a second time length of the original video, and if the first time length is greater than the second time length, trimming the first time length of the total audio data to obtain the target audio data whose time length matches the second time length.
[0125] According to some embodiments of the present disclosure, in the video processing method of the present disclosure, the step of performing integration processing on all the audio tracks based on a time axis to generate total audio data includes: detecting an audio identifier for each audio track; When a target audio track representing a background music identifier is detected, performing an integration process on all audio tracks other than the target audio track based on a time axis to generate total audio data.
[0126] According to some embodiments of the present disclosure, in the video processing method of the present disclosure, obtaining at least one recommended material that matches the video content characteristics includes: determining a video style characteristic based on video images of the original video; and obtaining at least one recommended material that matches the video style characteristics and the video content characteristics.
[0127] According to some embodiments of the present disclosure, in the video processing method according to the present disclosure, determining a video style feature based on a video image of the original video includes: performing an image identification process on the video images of the original video and determining at least one shooting target based on the identification result; Performing a weighting calculation on the at least one shooting object based on a preset object weight, and matching the calculation result with a preset style classification to determine a video style feature corresponding to the original video.
[0128] According to some embodiments of the present disclosure, in the video processing method of the present disclosure, obtaining at least one recommended material that matches the video content characteristics includes: determining, in the original video, a playback duration of a video frame corresponding to the video content feature generated based on video content of the video frame; marking a time identifier for the video content feature based on the playback time of the video frame; When determining that there are a plurality of video content features corresponding to the time identifier, combining the plurality of video content features into a video feature set and obtaining at least one recommended material matching the video feature set, wherein the plurality of video content features includes the video content feature; If it is determined that only video content characteristics correspond to the time identifier, obtaining at least one recommended material that matches the video content characteristics.
[0129] According to some embodiments of the present disclosure, in the video processing method according to the present disclosure, the step of performing video processing on the original video based on the recommended material to generate a target video includes: setting a material addition time of the recommended material matching the video content characteristics based on the time identifier of the video content characteristics; and generating a target video by performing clip processing on the original video based on the material addition time of the recommended material.
[0130] According to some embodiments of the present disclosure, in the video processing method of the present disclosure, generating a target video by performing clip processing on the original video based on the material addition time of the recommended material includes: If the material type of the recommended material satisfies a preset target type, obtaining a target video frame corresponding to the material addition time of the recommended material from the original video; performing image identification on the target video frame to obtain a main body region of the subject; determining a material area to add the recommended material to on the target video frame based on a main body area of the subject; and generating a target video by performing clip processing on the original video based on the material addition time of the recommended material and the material area.
[0131] According to some embodiments of the present disclosure, in the video processing method according to the present disclosure, the detecting of the audio identifier for each of the audio tracks comprises: identifying voice characteristics of the audio corresponding to each audio track; Matching voice characteristics of audio corresponding to each audio track with voice characteristics corresponding to each predetermined audio identifier; and determining an audio identifier for each audio track based on the matching results.
[0132] According to some embodiments of the present disclosure, in the video processing method of the present disclosure, obtaining at least one recommended material that matches the video feature set includes: Querying a predetermined correspondence relationship based on the video feature set to determine whether the video feature set corresponds to the enhancement material; If no match is found for the enhancement material, dividing the set of video features into recommended material that matches a single content feature; When matching with the reinforcement material, the reinforcement material is set as the corresponding recommended material.
[0133] According to some embodiments of the present disclosure, the present disclosure provides a video processing device, the video processing device comprising: an extraction module for extracting video content features based on an analysis of the original video; an acquisition module for acquiring at least one recommended material matching the video content characteristics; a processing module for performing video processing on the original video based on the recommended material to generate a target video, which is a video generated after adding the recommended material to the original video.
[0134] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the extraction module comprises: Performing a voice recognition process on the target audio data of the original video to obtain corresponding text content; The text content is subjected to a semantic analysis process to obtain a first keyword.
[0135] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the extraction module comprises: Perform a voice detection process on the target audio data of the original video to obtain corresponding spectrum data; The spectral data is subjected to an analytical process and used to obtain a second keyword.
[0136] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the extraction module obtains all audio tracks of the video file of the original video displayed in a video clip application; performing an integration process on all of the audio tracks based on a time axis to generate total audio data; The first time length of the total audio data is compared with the second time length of the original video, and if the first time length is greater than the second time length, the first time length of the total audio data is trimmed to obtain the target audio data whose time length matches the second time length.
[0137] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the extraction module comprises: Detecting the audio identifier for each audio track, When a target audio track representing a background music identifier is detected, the integration process is performed on all audio tracks other than the target audio track based on the time axis to generate total audio data.
[0138] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the acquisition module specifically comprises: determining a video style characteristic based on video images of the original video; The video style characteristics and the video content characteristics are used to obtain at least one recommended material that matches the video style characteristics and the video content characteristics.
[0139] According to some embodiments of the present disclosure, in the video processing device of the present disclosure, the acquisition module specifically performs an image identification process on the video images of the original video, and determines at least one shooting object based on the identification result; A weighting calculation is performed on the at least one shooting object based on a preset object weight, and the calculation result is matched with a preset style classification to determine a video style feature corresponding to the original video.
[0140] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the acquisition module comprises: determining, in the original video, a playback duration of a video frame corresponding to the video content feature generated based on the video content of the video frame; marking a time identifier for the video content feature based on the playback time of the video frame; If it is determined that there are multiple corresponding video content features for the same time identifier, combining the multiple video content features into a video feature set and obtaining at least one recommended material that matches the video feature set; If it is determined that a corresponding video content feature exists for the same time identifier, the video content feature is used to obtain at least one recommended material that matches the corresponding video content feature.
[0141] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the acquisition module comprises: determining a material addition time of the recommended material matching the video content characteristics based on the time identifier of the video content characteristics; The clip processing is used to generate a target video by clipping the original video based on the additional material time of the recommended material.
[0142] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the acquisition module comprises: If the material type of the recommended material satisfies a preset target type, a target video frame corresponding to the material addition time of the recommended material is obtained from the original video; Perform image identification on the target video fragment frame to obtain a main area of the shooting target; determining a material area to add the recommended material to on the target video frame based on a main body area of the subject; The clip processing is used to generate a target video by clipping the original video based on the material addition time of the recommended material and the material area.
[0143] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the extraction module comprises: Identifying the voice characteristics of the audio corresponding to each audio track; Matching voice characteristics of the audio corresponding to each audio track with voice characteristics corresponding to each predetermined audio identifier; It is used to determine an audio identifier for each audio track based on the matching results.
[0144] According to some embodiments of the present disclosure, in the video processing device according to the present disclosure, the acquisition module comprises: Querying a predetermined correspondence relationship based on the video feature set to determine whether the video feature set corresponds to the enhancement material; If no match is found for the reinforcement material, splitting the set of video features into recommended material that matches a single content feature; When matching with an enhancement material, it is used to make the enhancement material the corresponding recommended material.
[0145] According to some embodiments of the present disclosure, there is provided an electronic device, the electronic device comprising: a processor; a memory for storing instructions executable by the processor; The processor is adapted to read the executable instructions from the memory and execute the instructions to implement any one of the video processing methods according to the present disclosure.
[0146] According to some embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored therein, the computer program being used to execute any one of the video processing methods described herein.
[0147] The above description merely describes the preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the present disclosure is not limited to the technical solutions based on the specific combinations of the above technical features, but also encompasses other technical solutions formed by any combination of the above technical features or features equivalent thereto, without departing from the concept of the present disclosure, such as, for example, technical solutions formed by substituting the above features with technical features having similar functions disclosed in the present disclosure (but not limited to these).
[0148] It should be noted that although operations are illustrated using a particular sequence, this should not be understood as requiring these operations to be performed in the particular sequence or order shown. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes some specific implementation details, these should not be construed as limitations on the scope of the disclosure. Certain features that are described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may be implemented in multiple embodiments alone or in any suitable subcombination.
[0149] Although the present subject matter has been described using language specific to structural features and / or methodological logical operations, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are merely example forms of implementing the claims.
Claims
1. 1. A video processing method comprising: extracting video content features based on an analysis of the original video; obtaining at least one recommended material that matches the video content characteristics; processing the original video based on the recommended material to generate a target video, the target video being a video generated after adding the recommended material to the original video; The step of obtaining at least one recommended material that matches the video content characteristics includes: determining, in the original video, a playback duration of a video frame corresponding to the video content feature generated based on the video content of the video frame; marking a time identifier for the video content feature based on the playback time of the video frame; if it is determined that there are a plurality of video content features corresponding to the time identifier, combining the plurality of video content features into a video feature set and obtaining at least one recommended material matching the video feature set, wherein the plurality of video content features includes the video content feature; if it is determined that only video content characteristics correspond to the time identifier, obtaining at least one recommended material that matches the video content characteristics.
2. The step of extracting video content features based on the analysis of the original video includes: performing a voice recognition process on the target audio data of the original video to obtain corresponding text content; The video processing method according to claim 1 , further comprising: performing a semantic analysis process on the text content to obtain a first keyword.
3. The step of extracting video content features based on the analysis of the original video includes: performing a voice detection process on the target audio data of the original video to obtain corresponding spectrum data; The video processing method according to claim 1, further comprising the step of: performing an analysis process on the spectral data to obtain a second keyword.
4. The method for obtaining target audio data includes: obtaining all audio tracks of the video file of the original video to be displayed in a video clip application; performing an integration process on all the audio tracks based on a time axis to generate total audio data; comparing a first time length of the total audio data with a second time length of the original video, and if the first time length is greater than the second time length, trimming the first time length of the total audio data to obtain the target audio data whose time length matches the second time length.
5. The step of performing integration processing on all the audio tracks based on a time axis to generate total audio data includes: detecting an audio identifier for each audio track; 5. The video processing method of claim 4, further comprising: when a target audio track representing a background music identifier is detected, performing integration processing on all audio tracks other than the target audio track based on a time axis to generate total audio data.
6. The step of obtaining at least one recommended material that matches the video content characteristics includes: performing an image identification process on the video images of the original video, and determining at least one shooting object based on the identification result; performing a weighting calculation for the at least one photographed object based on a preset object weight, and matching the result of the weighting calculation with a preset style classification to determine a video style feature corresponding to the original video; and obtaining at least one recommended material that matches the video style characteristics and the video content characteristics.
7. The step of processing the original video based on the recommended material and generating a target video, which is a video generated after adding the recommended material to the original video, includes: setting a material addition time of the recommended material matching the video content characteristics based on the time identifier of the video content characteristics; 2. The video processing method according to claim 1, further comprising the step of: generating a target video by performing clip processing on the original video based on the material addition time of the recommended material.
8. The step of generating a target video by performing clip processing on the original video based on the material addition time of the recommended material includes: If the material type of the recommended material satisfies a preset target type, acquiring a target video frame corresponding to the material addition time of the recommended material from the original video; performing image identification on the target video frame to obtain a main body region of the subject; determining a material area to add the recommended material to on the target video frame based on a main body area of the subject; 8. The video processing method according to claim 7, further comprising the step of: generating a target video by performing clip processing on the original video based on the material addition time of the recommended material and the material area.
9. said step of detecting an audio identifier for each audio track comprising: identifying voice characteristics of the audio corresponding to each audio track; matching voice characteristics of the audio corresponding to each audio track with voice characteristics corresponding to each predetermined audio identifier; and determining an audio identifier for each audio track based on the matching results.
10. The step of obtaining at least one recommended material that matches the video feature set includes: Querying a predetermined correspondence relationship based on the video feature set to determine whether the video feature set corresponds to the enhancement material; If no match is found for the reinforcement material, dividing the set of video features into recommended material that matches a single content feature; and if the reinforcement material matches, setting the reinforcement material as the corresponding recommended material.
11. 1. A video processing device comprising: an extraction module for extracting video content features based on an analysis of the original video; an acquisition module for acquiring at least one recommended material matching the video content characteristics; a processing module for processing the original video based on the recommended material to generate a target video, the target video being a video generated after adding the recommended material to the original video; The acquisition module determines a playback time of a video frame in the original video corresponding to the video content feature generated based on the video content of the video frame, marks a time identifier for the video content feature based on the playback time of the video frame, and when it determines that there are multiple video content features corresponding to the time identifier, combines the multiple video content features into a video feature set and obtains at least one recommended material that matches the video feature set, and the multiple video content features include the video content feature, and when it determines that only video content features correspond to the time identifier, obtains at least one recommended material that matches the video content feature.
12. An electronic device, a processor; a memory for storing instructions executable by the processor; An electronic device, wherein the processor reads the executable instructions from the memory and executes the executable instructions to implement the video processing method of any one of claims 1 to 10.
13. 11. A computer-readable storage medium having a computer program stored thereon, the computer program being used to perform the video processing method of any one of claims 1 to 10.
14. A computer program comprising instructions which, when executed by a processor, implement the video processing method of any one of claims 1 to 10.
Citation Information
Patent Citations
Video and image processing method and device, electronic equipment and storage medium
CN111541936A
Viewing and listening content provision system and viewing and listening content provision method
JP2006099195A
Matrix decoder with constant output pairwise panning
JP2016529801A
Method and system of searching and collating video files, establishing semantic group, and program storage medium therefor
US20150169542A1