Video Processing Method, Apparatus, Device, and Storage Medium

By using video subtitle recognition technology between the terminal and the server, the problem of low efficiency in manually adding materials is solved, and the video editing efficiency and process automation is achieved.

CN118055292BActive Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410179053.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-06-27
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

In the prior art, manually adding materials to oral videos is inefficient and cumbersome. Especially for novice users, it is difficult to determine when to add materials and select suitable materials.

Method used

By uploading and identifying video subtitles between the terminal and the server, the material group is automatically recommended and determined based on the keywords in the subtitles, and adding the material to the video to generate a new video.

Benefits of technology

It improves video editing efficiency, reduces server load, enhances terminal autonomy, and makes the video material addition process more automated and efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118055292B_ABST
    Figure CN118055292B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method, apparatus, device, and storage medium. It relates to the field of artificial intelligence, especially fields such as material search, material recommendation, video editing, and intelligent interaction. The specific implementation solution is as follows: in response to receiving a material recommendation trigger operation for a first video, upload the subtitles of the first video; receive at least one material group returned by the server based on the subtitles of the first video; determine a target material group based on the at least one material group; add the materials in the target material group to the first video to generate a second video. According to the technical solution of the present disclosure, materials can be automatically added to the video, improving the video editing efficiency. In addition, compared with uploading the first video to the server, uploading the audio of the first video to the server first can make full use of the powerful computing resources of the server to quickly identify the subtitles of the first video, and can also improve the speed of obtaining subtitles, facilitating the quick acquisition of the material group from the server.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is a divisional application of a Chinese patent application with the application number 202210553116.2 and the title "Video Processing Method, Apparatus, Device, and Storage Medium", which was filed on May 20, 2022, and the entire content of which is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of artificial intelligence, and particularly to fields such as material search, material recommendation, video editing, and intelligent interaction. Background Art

[0004] With the rapid development of the video wave, some new types of videos such as oral broadcast videos have emerged. The main feature of this type of video is that a real person continuously outputs content facing the camera. Since the picture mainly shows people talking and lacks changes, it is often boring. To solve this problem, some materials are often added to this type of video, and the video is edited to make the edited video more interesting. However, manually adding materials is slow and the video editing efficiency is low. Summary of the Invention

[0005] The present disclosure provides a video processing method, apparatus, device, and storage medium.

[0006] According to a first aspect of the present disclosure, there is provided a video processing method applied to a terminal, including:

[0007] In response to receiving a material recommendation trigger operation for a first video, uploading the subtitles of the first video;

[0008] Receiving at least one material group returned by the server based on the subtitles of the first video;

[0009] Based on the at least one material group, determining a target material group;

[0010] Adding the materials in the target material group to the first video to generate a second video.

[0011] According to a second aspect of the present disclosure, there is provided a video processing method applied to a server, including:

[0012] Receiving the subtitles of a first video, where the subtitles of the first video are uploaded by a terminal when receiving a material recommendation trigger operation for the first video;

[0013] Identifying at least one keyword of the first video from the subtitles of the first video;

[0014] Based on the at least one keyword, determining at least one material group for the first video;

[0015] Send the at least one material group, where the at least one material group is used to indicate materials available for addition to the first video.

[0016] According to a third aspect of the present disclosure, there is provided a video processing apparatus, which is applied to a terminal and includes:

[0017] A first sending module, configured to upload the subtitles of the first video in response to receiving a material recommendation trigger operation for the first video;

[0018] A first receiving module, configured to receive at least one material group returned by the server based on the subtitles of the first video;

[0019] A first determining module, configured to determine a target material group based on the at least one material group;

[0020] A generating module, configured to add the materials in the target material group to the first video to generate a second video.

[0021] According to a fourth aspect of the present disclosure, there is provided a video processing apparatus, which is applied to a server and includes:

[0022] A second receiving module, configured to receive the subtitles of the first video, where the subtitles of the first video are uploaded by the terminal when receiving a material recommendation trigger operation for the first video;

[0023] A first recognition module, configured to recognize at least one keyword of the first video from the subtitles of the first video;

[0024] A second determining module, configured to determine at least one material group for the first video based on the at least one keyword;

[0025] A second sending module, configured to send the at least one material group, where the at least one material group is used to indicate materials available for addition to the first video.

[0026] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0027] At least one processor; and

[0028] A memory communicatively connected to the at least one processor; wherein,

[0029] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the methods provided in the first aspect and the second aspect above.

[0030] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the methods provided in the above first aspect and second aspect.

[0031] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the methods provided in the above first aspect and second aspect.

[0032] According to the technical solution of the present disclosure, materials can be automatically added to a video, improving the video editing efficiency.

[0033] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0035] Figure 1 is a flowchart of a video processing method according to an embodiment of the present disclosure; Figure 1 ;

[0036] Figure 2 is a schematic diagram showing materials displayed in a video image according to an embodiment of the present disclosure;

[0037] Figure 3 is a comparative schematic diagram of a video image before and after adding materials according to an embodiment of the present disclosure;

[0038] Figure 4 is a flowchart of a video processing method according to an embodiment of the present disclosure; Figure 2 ;

[0039] Figure 5 is a schematic diagram of the interaction process between a terminal and a server according to an embodiment of the present disclosure;

[0040] Figure 6 is a schematic structural diagram of a video processing apparatus according to an embodiment of the present disclosure; Figure 1 ;

[0041] Figure 7 is a schematic structural diagram of a video processing apparatus according to an embodiment of the present disclosure; Figure 2 ;

[0042] Figure 8 is a schematic diagram of a video processing scenario according to an embodiment of the present disclosure;

[0043] Figure 9It is a block diagram of an electronic device for implementing the video processing method of the embodiments of the present disclosure. Detailed implementation manners

[0044] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0045] The terms "first", "second", "third", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0046] In the related art, in a video mainly featuring people speaking, such as a voice-over video, in order to make the video more vivid, interesting and attractive, the following two processing methods can be adopted: First, add variety show captions and stickers, and match with sound effects to make the video more interesting; Second, add explanatory videos or pictures to change the picture content. However, generally, users need to perform a series of processes such as searching, screening, downloading, adding and adjusting by themselves, and the entire operation process is cumbersome and time-consuming. In addition, for novice users in the self-media industry, it is very difficult to judge at what time point to add materials and what kind of materials to add. Although some clients can provide some material packs, users still need to find the materials to be added in the video by themselves, and the operation process is cumbersome. Some clients can provide video templates that can be applied, but such video templates are highly homogenized and not suitable for videos that require different content for each period.

[0047] Embodiments of the present disclosure provide a video processing method, which can be applied to a terminal, specifically to a client installed on the terminal. The client has a video editing function and supports video import, material recommendation, video editing and video generation, etc. In practical applications, the terminal includes but is not limited to devices such as mobile phones, tablets, wearable devices or personal computers. As Figure 1 shown, the video processing method may include:

[0048] S101: In response to receiving a material recommendation trigger operation for a first video, upload the subtitles of the first video;

[0049] S102: Receive at least one material group returned by the server based on the subtitles of the first video;

[0050] S103: Determine a target material group based on the at least one material group;

[0051] S104: Add the materials in the target material group to the first video to generate a second video.

[0052] In the embodiments of the present disclosure, the first video is a video to be processed for adding materials. Here, the first video may be a video just recorded by the user, or a newly imported video previously recorded by the user, or a video stored locally.

[0053] Here, a material recommendation function may be set on the client. In some embodiments, when the terminal receives an instruction for triggering the material recommendation function input by the user through voice, it determines that a material recommendation trigger operation for the first video is received. In other embodiments, when the terminal detects an operation on the material recommendation trigger button, it determines that a material recommendation trigger operation for the first video is received. In still other embodiments, when the terminal detects that the recording completion time of the first video exceeds a preset time value, it automatically triggers and starts the material recommendation function, and determines that a material recommendation trigger operation for the first video is received. The present disclosure does not limit the triggering method of the material recommendation trigger operation.

[0054] In the embodiments of the present disclosure, the subtitles of the first video are subtitles corresponding to the audio in the first video. The subtitles may be subtitles recognized by the terminal according to the audio of the first video, or subtitles recognized by the server according to the audio of the first video. The present disclosure does not limit how to recognize the subtitles according to the audio of the first video, nor does it make a mandatory limitation on who is specifically responsible for the recognition.

[0055] In the embodiments of the present disclosure, the number of material groups is not limited. The number of material groups can be set or adjusted according to user needs. Exemplarily, the terminal instructs the server to return a certain number of material groups according to user needs, such as returning 3 material groups. In this way, when the server returns multiple material groups, the terminal displays multiple material groups for selection, which not only provides the diversity of material groups for the first video, but also enriches the various possibilities of editing.

[0056] In the embodiments of the present disclosure, the target material group is a material group specified by the user to be added to the first video. When multiple material groups are received, the target material group may be one of the multiple material groups. When only one material group is received, the target material group may be the material group, or a new material group obtained by adding or deleting materials in the material group.

[0057] In some embodiments, determining a target material group based on the at least one material group includes: receiving a first operation for indicating a selected material group; and determining the target material group from the at least one material group based on the first operation. In other embodiments, determining a target material group based on the at least one material group includes: determining the target material group from the at least one material group according to the receiving order or recommended ranking of each material group. For example, the material group with the earliest receiving time is used as the target material group. For another example, the material group with the topmost recommended ranking is used as the target material group. The above is only an exemplary illustration and does not limit all possible ways of determining the target material group, and only exhaustive listing is not done here.

[0058] In the embodiments of the present disclosure, the second video is a video with materials added. Compared with the first video, the second video has a more diverse, vivid and interesting display form.

[0059] In the embodiments of the present disclosure, the types of materials in the material group are not limited. Exemplarily, classified by the presence or absence of text, materials are divided into textless type materials and text-containing type materials. Another example is that classified by the presence or absence of sound effects, materials are divided into soundless type materials and sound-containing type materials.

[0060] In the technical solution of the embodiments of the present disclosure, the terminal uploads the subtitles of the first video in response to receiving a material recommendation trigger operation for the first video; receives at least one material group returned by the server based on the subtitles of the first video; determines a target material group based on the at least one material group; and adds the materials in the target material group to the first video to generate a second video. Compared with sending the first video to the server, the terminal sends the subtitles of the first video to the server, reducing the amount of data to be transmitted, improving the transmission speed, facilitating the server to quickly screen out a material group suitable for the first video based on the subtitles of the first video, and then enabling the terminal to quickly obtain the material group, thus helping to improve the video editing efficiency. In addition, compared with the server-side editing of the first video, the terminal determines the target material group according to at least one material group, which not only reduces the load on the server, but also enhances the autonomy of the terminal in editing videos, enabling the terminal to automatically add materials to the video and improving the video editing efficiency.

[0061] In some embodiments, before S101, the video processing method may further include: uploading the audio of the first video to the server when receiving the first video; receiving the subtitles of the first video returned by the server based on the audio of the first video; and saving the subtitles of the first video.

[0062] Here, the present disclosure does not limit how to obtain the audio from the first video. For example, if the terminal has a built-in audio and image separation function, after receiving the first video, the function is used to obtain the audio of the first video. For another example, the terminal extracts the audio from the first video through audio extraction technology.

[0063] Here, the present disclosure does not limit the storage location of the subtitles of the first video.

[0064] Here, the subtitles of the first video can be used both as a basis for requesting a material group from the server and as the subtitles of the generated second video.

[0065] In this way, after importing the first video, the audio of the first video is first uploaded to the server, which can make full use of the powerful computing resources of the server to quickly identify the subtitles of the first video, so as to prepare in advance for subsequent requests for material groups from the terminal. Compared with the method of uploading the first video to the server to obtain subtitles, the speed of obtaining subtitles is improved, and it also provides a subtitle basis for subsequent requests for material groups from the server, facilitating the quick acquisition of the material group from the server, thereby helping to improve the video editing efficiency.

[0066] In some embodiments, adding the materials in the target material group to the first video includes: determining the appearance time of each material in the target material group in the first video; adding the material corresponding to the corresponding appearance time at the corresponding appearance time on the time axis of the first video.

[0067] In some embodiments, determining the appearance time of each material in the target material group in the first video includes: according to the correspondence between the materials in the target material group and keywords, searching for the keywords included in the target material group from the subtitles of the first video, and taking the appearance time of the keyword in the first video as the appearance time of the material corresponding to the keyword.

[0068] Exemplarily, the material group includes m materials, denoted as m1, m2,..., mm respectively; material m1 corresponds to keyword c1, material m2 corresponds to keyword c2,..., material mm corresponds to keyword cm; if the time corresponding to keyword c1 is t1, the time corresponding to keyword c2 is t2,..., the time corresponding to keyword cm is tm, then the time corresponding to material m1 is t1, the time corresponding to material m2 is t2,..., the time corresponding to material mm is tm.

[0069] In this way, the appearance time of the materials in the first video can be automatically matched, improving the intelligence of video editing.

[0070] In some embodiments, adding the materials in the target material group to the first video to generate a second video includes: determining the display information of each material in the target material group; and adding each material to the image in the first video based on the display information of each material.

[0071] Among them, the display information includes but is not limited to: display position, display angle, and display size.

[0072] Here, the display position is the position of the material in the video image or picture, including display coordinates.

[0073] Here, the display angle is the angle of the material in the video image or picture relative to the horizontal line or the vertical line. The horizontal line and the vertical line here can be relative to the terminal display screen.

[0074] Here, the display size is the size of the material in the video image or picture.

[0075] Figure 2 A schematic diagram showing the display of materials in the video image is shown, as Figure 2 shown. In the current video image, when the host mentions "We should prepare a large wall-mounted bookshelf for the children", a picture material such as a background wall with a bookshelf as the background is displayed behind the host. At the same time, text materials such as "Whole wall!" and "Large bookshelf" are displayed in front of the host. From Figure 2 it can be seen that the display angle of "Whole wall!" is slightly tilted, having a certain angle with the horizontal line, making the entire image more vivid. The display angle of "Large bookshelf" is parallel to the horizontal line. The font sizes of "Whole wall!" and "Large bookshelf" are different. The background wall covers the background of the entire picture, making the content of the entire image more rich and the presentation form more diverse.

[0076] In this way, by adjusting the display information such as the display position, display size, and display angle of each material, the display form of the materials can be expanded, and the enhancement effect of the materials can be improved.

[0077] In some embodiments, adding the materials in the target material group to the first video to generate a second video includes: when there is a first material in the target material group, determining the volume of the first material according to a preset volume ratio; and adding the volume of the first material to the audio in the first video.

[0078] Among them, the preset volume ratio is equal to the ratio between the material volume and the video volume.

[0079] Here, the first material is a sound effect material, such as the sound of wind, rain, tsunami, bird calls, etc.

[0080] Here, the preset volume ratios corresponding to different first materials may be different. The preset volume ratio can be determined according to the content attributes of the first material. For example, the preset volume ratio for the content attribute of a tsunami can be greater than that for the content attribute of a bird call. Another example is that the preset volume ratio for the content attribute of rain sounds can be less than that for the content attribute of wind sounds.

[0081] Generally speaking, when the original volume of the sound effect material is relatively large, in order to prevent the sound effect from being too prominent, a certain volume ratio can be set.

[0082] In this way, the volume of the material can be adapted to the volume of the first video, and without affecting the volume of the first video, the material can play a better role in enhancing the atmosphere.

[0083] In some embodiments, adding the materials in the target material group to the first video to generate a second video includes: when there is a second material in the target material group, identifying the target image in the first video corresponding to the second material, and adding the second material to the corresponding preset position in the target image.

[0084] Here, the second material is a material used to decorate the target object in the target image. The target object includes but is not limited to people, animals, plants, etc.

[0085] Here, the second material mainly includes decorative or modifying materials. Exemplarily, the second material is a material related to the video anchor. For example, the second material includes but is not limited to materials such as blush, red eyes, eyeshadow, red lips, dimples, etc. Another example is that the second material includes but is not limited to materials such as bracelets, rings, necklaces, clothing, etc.

[0086] Taking the second material as blush as an example, the target image is the current video image, and the preset position is the position of the two cheeks of the anchor in the current video image.

[0087] Taking the second material as red eyes as an example, the target image is the current video image, and the preset position is the position of the two eyes of the anchor in the current video image.

[0088] Taking the second material as a ring as an example, the target image is the current video image, and the preset position is the position of the ring finger of the anchor in the current video image.

[0089] The embodiments of the present disclosure do not limit how to identify the preset position of the target object in the target image. For example, the positions of the facial features can be identified through face recognition technology. Another example is to identify the position of the human hand through hand detection technology.

[0090] In this way, by adding the second material to the corresponding preset position in the target image, so that the second material is added to the specified position, it can not only enrich the diversity of the video image, but also save the time cost of the user manually adjusting the image, such as image retouching.

[0091] Figure 3 shows a comparison schematic diagram of the video image before and after adding materials. As Figure 3 shown in the left figure, the anchor in the image says "First, we can open the teleprompter", and there is no material in this image. For Figure 3 the image shown in the left figure, after adding materials, the adding effect is as Figure 3 shown in the right figure. Specifically, the anchor in the image says "First, we can open the teleprompter", and there are 2 materials in this image, including: "First" and "Give it a try". Obviously, the display effect after adding materials is significantly better than that before adding materials.

[0092] An embodiment of the present disclosure provides a video processing method, which can be applied to a server. The server has a material recommendation function and supports material search, material screening, material group generation, etc. In practical applications, the server includes but is not limited to ordinary servers, cloud servers, etc. As Figure 4 shown, the video processing method may include:

[0093] S401: Receive the subtitles of the first video, where the subtitles of the first video are uploaded by the terminal when receiving a material recommendation trigger operation for the first video;

[0094] S402: Identify at least one keyword of the first video from the subtitles of the first video;

[0095] S403: Determine at least one material group for the first video based on the at least one keyword;

[0096] S404: Send the at least one material group, where the at least one material group is used to indicate the materials available for adding to the first video.

[0097] In an embodiment of the present disclosure, the material group includes at least one material. The present disclosure does not limit the number of materials included in the material group. The number of materials in the material group generally depends on the number of keywords identified from the subtitles. Taking a material group as an example, if one material is assigned to each keyword, the number of materials in the material group may be equal to the number of keywords in the subtitles. If multiple materials are assigned to some keywords, the number of materials in the material group will be greater than the number of keywords in the subtitles.

[0098] In an embodiment of the present disclosure, the materials in the material group are matched with the information represented by the keywords. For example, when the keyword is beach, the materials matched for the beach are all materials related to the beach, such as beach pictures. Another example is when the keyword is bookshelf, the materials matched for the bookshelf are all materials related to the bookshelf, such as the two characters "bookshelf", bookshelf pictures, etc.

[0099] The embodiments of the present disclosure do not limit the source of the material. For example, the material can be sourced from a material database, or from the material actively uploaded by the terminal, or from the material obtained from a third party such as a website.

[0100] In this way, the server can quickly screen out a material group suitable for the first video based on the subtitles of the first video, and then enable the terminal to quickly obtain at least one material group. The server provides the material group for the terminal, which not only increases the autonomy of video editing on the terminal side, but also helps to improve the video editing efficiency on the terminal side. In addition, compared with the server editing the first video, the load on the server is reduced, enabling the server to provide material recommendation services for more terminals at the same time.

[0101] In some embodiments, the video processing method may further include: when receiving the audio of the first video, identifying the audio of the first video to obtain the subtitles of the first video; and sending the subtitles of the first video. Here, the audio of the first video can be uploaded by the terminal when receiving the first video. For example, after the terminal imports the first video, it immediately uploads the audio of the first video to the server.

[0102] In some embodiments, identifying the audio of the first video to obtain the subtitles of the first video includes: the server uses audio recognition technology to convert the audio into text; and generates subtitles based on the text. In other embodiments, identifying the audio of the first video to obtain the subtitles of the first video includes: converting the audio into text through an audio converter; and generating subtitles based on the text. The present disclosure does not limit how to obtain subtitles according to audio recognition specifically.

[0103] In this way, by using the powerful computing resources of the server, subtitles can be quickly provided for the terminal, saving the time consumed by the terminal to generate subtitles and improving the speed of obtaining subtitles on the terminal side.

[0104] In some embodiments, identifying at least one keyword of the first video from the subtitles of the first video includes: splitting the subtitles of the first video to obtain a plurality of first target words of the first video; searching for the plurality of first target words of the first video in a preset word list library; and using at least one first target word that can be found in the preset word list library as at least one keyword of the first video.

[0105] Here, the first target word is the word after splitting the subtitles. Exemplarily, the first target words obtained by splitting the subtitle "Please open Baidu Map" include: "Please", "open", "Baidu Map".

[0106] Here, the preset word list library stores a large number of words.

[0107] Here, the sources of the words in the preset vocabulary library include, but are not limited to: (1) dictionaries, such as general dictionaries and common entries provided by Baidu Encyclopedia; (2) iconic words such as numerical words and blessing words, such as 1.3 billion, good luck in everything, etc.; (3) self-built emotion words: a batch of interactive words and chapter words determined by sampling from self-media videos and internal team evaluations, such as follow me, first, etc.

[0108] It should be noted that the words in the preset vocabulary library can be added or deleted. For example, new words can be added to the self-built emotion words. Another example is to delete some outdated blessing words.

[0109] In this way, determining the keywords in the subtitle based on the preset vocabulary library can improve the speed of keyword determination, which helps to quickly generate a material group for the first video.

[0110] In some embodiments, identifying at least one keyword of the first video from the subtitle of the first video includes: extracting a plurality of second target words in the subtitle of the first video through a semantic recognition algorithm; combining at least two of the plurality of second target words to obtain at least one combined word; using the at least one combined word as at least one keyword of the first video.

[0111] Here, the second target word is a word identified by the semantic recognition algorithm. Exemplarily, the second target word is a word with a certain amount of information content, such as Baidu Encyclopedia, birthday, happy, first, etc.

[0112] Here, the combined word refers to a word formed by combining at least two second target words. Exemplarily, "birthday" and "happy" are combined into a combined word "happy birthday". Another example is that "first" and "search purpose" are combined into a combined word "first, search purpose". Still another example is that "second" and "search result" are combined into a combined word "second, search result". In this way, by combining keywords, more complex keywords are formed, which are more suitable for clearly structured chapter titles.

[0113] In this way, by combining keywords, the determined keywords are more in line with semantics and context, improving the rationality of the selected keywords, which helps to improve the adaptability of the material group.

[0114] In some embodiments, based on at least one keyword, determining at least one material group for the first video includes: recalling multiple materials of at least one material type for each keyword in at least one keyword from the material database; determining a target number of materials for each keyword from the multiple materials of at least one material type corresponding to each keyword; and determining at least one material group for the first video according to the target number of materials corresponding to each keyword.

[0115] Here, a large number of materials are stored in the material database. The present disclosure does not limit the specific quantity of the material database. In practical applications, different types of materials can be stored in one material database. In practical applications, different types of materials can also be stored in different material databases, and each material database is used to store one type of material. For example, Material Database 1 is used to store text-type materials, Material Database 2 is used to store sticker-type materials, Material Database 3 is used to store sound-type materials, and Material Database 4 is used to store animation-type materials.

[0116] Here, the target quantity can be specified by the terminal side or determined by the server side. For example, the target quantity = 5, that is, 5 materials are selected for each keyword.

[0117] In this way, multiple material groups can be provided for the first video, which helps to provide rich material support for the terminal side and also supports the terminal side to perform video editing in different styles.

[0118] In some embodiments, determining the target quantity of materials for each keyword from multiple materials of at least one material type corresponding to each keyword includes: respectively sorting the priorities of multiple materials of the same type for each keyword; selecting different types of materials to be recommended for each keyword according to the priority sorting of multiple materials of the same type for each keyword; and determining the target quantity of materials for each keyword according to the different types of materials to be recommended for each keyword.

[0119] Here, selecting different types of materials to be recommended for each keyword may include: selecting materials with a high priority ranking and meeting the expected quantity as the materials to be recommended for each keyword.

[0120] It should be noted that the expected quantities corresponding to different types of materials can be different.

[0121] For example, keywords 1, 2, and 3 are identified from the subtitles of the first video; for keyword 1, materials S11 and S12 of the first type are selected; for keyword 1, material S21 of the second type is selected; for keyword 1, material S31 of the third type is selected; similarly, for keyword 2, materials S13 of the first type are selected; for keyword 2, material S22 of the second type is selected; for keyword 2, material S32 of the third type is selected; for keyword 3, materials S14, S15, and S16 of the first type are selected; for keyword 3, material S23 of the second type is selected; for keyword 3, material S33 of the third type is selected; then, multiple material groups can be generated, such as material group 1 = {S11, S13, S14}, material group 2 = {S12, S13, S15}, material group 3 = {S21, S22, S33}, material group 4 = {S31, S22, S33}, etc., and they will not be listed one by one here.

[0122] In this way, according to the materials to be recommended corresponding to different types of each keyword, the target number of materials for each keyword can be determined, which improves the diversity of the generated material groups and can meet the needs of the terminal side to frequently change the content and form of materials.

[0123] In some embodiments, after generating the material groups, it may further include: when there are two or more identical keywords in the material groups, performing deduplication processing on the materials corresponding to the identical keywords.

[0124] For example, the material group includes materials of keyword 1, materials of keyword 2, materials of keyword 3, and materials of keyword 4. If keyword 1 and keyword 3 are the same keyword, then perform deduplication processing on the materials of keyword 1 and the materials of keyword 3, and retain the materials of keyword 1 or the materials of keyword 3. It should be noted that keyword 1 and keyword 3 are the same keyword, but the materials of keyword 1 and the materials of keyword 2 can be the same or different.

[0125] In this way, deduplicating the materials of the same keyword in the same video can reduce the situation of a large number of repeated materials in the same video and improve the rendering effect of the materials.

[0126] In some embodiments, after generating the material groups, it may further include: when there are multiple materials of a target type in the material groups, performing deduplication processing on the recommended frequencies of the multiple materials of the target type.

[0127] Here, the target type may include one or more of the material types. For example, the target type may be materials that play a role in setting off the atmosphere, such as confetti and animated flying in and out.

[0128] Here, frequency deduplication processing includes: if there are multiple materials of the target type within a preset time period, one material within the preset time period is retained, such as retaining the first material.

[0129] Here, the preset time period can be set or adjusted according to the rendering effect. For example, the preset time period can be set to 1 second, 10 seconds, etc.

[0130] In this way, by deduplicating the recommendation density, the rendering effect of the materials can be improved.

[0131] Figure 5 The interaction schematic diagram between the terminal and the server is shown, such as Figure 5 As shown, the interaction process includes: the terminal client imports a video and uploads the audio of the video to the server; the server performs subtitle recognition on the audio and returns the recognized subtitles to the terminal client. When the terminal client receives a trigger operation for the material recommendation function, it uploads the subtitles to the server, so that the server extracts keywords according to the subtitles, filters the keywords, recalls materials, performs priority sorting, density filtering and other processing on the materials to obtain multiple material groups. The server returns the recommended material groups to the terminal client. Finally, the terminal client adds the materials in the material groups according to the material type strategy.

[0132] In this way, the cumbersome material addition process is shortened, and the video editing efficiency is improved; the threshold of video production is reduced, so that novice users can easily produce interesting videos; it can actively discover the interesting points ignored by users, create more possibilities, and thus increase the product popularity.

[0133] It should be understood that Figure 5 the interaction schematic diagram shown is only exemplary rather than restrictive, and it is extensible. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 5 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0134] The embodiments of the present disclosure provide a video processing device, which is applied to a terminal, such as Figure 6 As shown, the video processing device may include: a first sending module 601, configured to upload the subtitles of the first video in response to receiving a trigger operation for the material recommendation of the first video; a first receiving module 602, configured to receive at least one material group returned by the server based on the subtitles of the first video; a first determining module 603, configured to determine a target material group based on the at least one material group; a generating module 604, configured to add the materials in the target material group to the first video to generate a second video.

[0135] In some embodiments, the first sending module 601 is further configured to upload the audio of the first video to the server when receiving the first video; the first receiving module 602 is further configured to receive the subtitle of the first video returned by the server based on the audio of the first video.

[0136] In some embodiments, the video processing device may further include: a saving module 605 (not shown in the figure), configured to save the subtitle of the first video.

[0137] In some embodiments, the generating module 604 includes: a first determining sub-module, configured to determine the appearance time of each material in the target material group in the first video; a first adding sub-module, configured to add the material corresponding to the appearance time to the corresponding appearance time on the time axis of the first video.

[0138] In some embodiments, the generating module 604 includes: a second determining sub-module, configured to determine the display information of each material in the target material group; a second adding sub-module, configured to add each material in the target material group to the image of the first video based on the display information of each material in the target material group.

[0139] In some embodiments, the generating module 604 includes: a third determining sub-module, configured to determine the volume of the first material according to a preset volume ratio when the first material exists in the target material group; a third adding sub-module, configured to add the volume of the first material to the audio of the first video.

[0140] In some embodiments, the generating module 604 includes: a first identifying sub-module, configured to identify the target image in the first video corresponding to the second material when the second material exists in the target material group; a fourth adding sub-module, configured to add the second material to a preset position of the target image.

[0141] Those skilled in the art should understand that the functions of the processing modules in the video processing device according to the embodiments of the present disclosure can be understood with reference to the related descriptions of the video processing method applied to the terminal described above. Each processing module in the video processing device according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions described in the embodiments of the present disclosure, or can be implemented by software that executes the functions described in the embodiments of the present disclosure running on an electronic device.

[0142] The video processing device according to the embodiments of the present disclosure can automatically add materials to the video and improve the video editing efficiency.

[0143] The embodiments of the present disclosure provide a video processing device, which is applied to a server, such as Figure 7As shown, the video processing device may include: a second receiving module 701, configured to receive the subtitles of a first video, where the subtitles of the first video are uploaded by a terminal when receiving a material recommendation trigger operation for the first video; a first recognition module 702, configured to recognize at least one keyword of the first video from the subtitles of the first video; a second determination module 703, configured to determine at least one material group for the first video based on the at least one keyword; and a second sending module 704, configured to send the at least one material group, where the at least one material group is used to indicate the materials available for addition to the first video.

[0144] In some embodiments, the second receiving module 701 is further configured to receive the audio of the first video, where the audio of the first video is uploaded by the terminal when receiving the first video.

[0145] In some embodiments, the video processing device further includes: a second recognition module 705 (not shown in the figure), configured to recognize the audio of the first video to obtain the subtitles of the first video. Correspondingly, the second sending module 704 is further configured to send the subtitles of the first video.

[0146] In some embodiments, the first recognition module 702 includes: a splitting sub-module, configured to split the subtitles of the first video to obtain a plurality of first target words of the first video; and a second recognition sub-module, configured to use at least one first target word that can be found in a preset word list library as at least one keyword of the first video.

[0147] In some embodiments, the first recognition module 702 includes: an extraction sub-module, configured to extract a plurality of second target words from the subtitles of the first video through a semantic recognition algorithm; a combination sub-module, configured to combine at least two of the plurality of second target words to obtain at least one combined word; and a third recognition sub-module, configured to use the at least one combined word as at least one keyword of the first video.

[0148] In some embodiments, the second determination module 703 includes: a recall sub-module, configured to recall at least one type of a plurality of materials for each keyword from a material database; a fourth determination sub-module, configured to determine a target number of materials for each keyword based on the at least one type of a plurality of materials corresponding to each keyword; and a fifth determination sub-module, configured to determine at least one material group for the first video according to the target number of materials corresponding to each keyword.

[0149] In some embodiments, the fourth determination sub-module is configured to: respectively perform priority sorting on multiple materials of the same type for each keyword; select different types of materials to be recommended for each keyword according to the priority sorting of multiple materials of the same type for each keyword; and determine a target number of materials for each keyword according to the different types of materials to be recommended for each keyword.

[0150] In some embodiments, the video processing apparatus may further include: a first deduplication module 706 (not shown in the figure), configured to perform deduplication processing on materials corresponding to multiple identical keywords when a material group includes multiple identical keywords.

[0151] In some embodiments, the video processing apparatus may further include: a second deduplication module 707 (not shown in the figure), configured to perform deduplication processing on multiple materials of the target type when a material group includes multiple materials of the target type.

[0152] Those skilled in the art should understand that the functions of the processing modules in the video processing apparatus according to the embodiments of the present disclosure can be understood with reference to the related descriptions of the video processing method applied to the server described above. The processing modules in the video processing apparatus according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions described in the embodiments of the present disclosure, or can be implemented by software that executes the functions described in the embodiments of the present disclosure running on an electronic device.

[0153] The video processing apparatus according to the embodiments of the present disclosure can quickly screen out a material group suitable for the first video based on the subtitles of the first video, so that the terminal can quickly obtain at least one material group. The server provides the material group for the terminal, which can not only increase the autonomy of video editing on the terminal side, but also help improve the video editing efficiency on the terminal side.

[0154] Figure 8 shows a schematic diagram of a video processing scenario. From Figure 8 As can be seen, after an electronic device such as a cloud server receives the audio uploaded from each terminal, it generates and returns corresponding subtitles for each audio. After the electronic device receives the subtitles uploaded from each terminal, it identifies the subtitles based on a preset vocabulary library to obtain keywords; recalls materials from the material database based on the keywords to generate a material group, and returns the corresponding material group to the terminal. Then, the terminal edits the video based on the material group to generate a video with added materials.

[0155] The following are several video editing scenarios. For example, after a user records an oral video, it is imported into the editing client on the terminal, and the editing client automatically edits the oral video into a video with added materials. Another example is that after the user records a video through the editing client on the terminal, the material recommendation function is triggered, and one of the multiple material groups returned from the server is selected as the target material group, and the editing client generates a new video based on the target material group.

[0156] It should be understood that Figure 8 the shown scenario diagrams are merely illustrative rather than restrictive, and those skilled in the art can make various obvious changes and / or substitutions based on Figure 8 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0157] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0158] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0159] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0160] As Figure 9 shown, the device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 902 or the computer program loaded from the storage unit 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0161] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as a keyboard, a mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as a disk, an optical disc, etc.; and communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0162] Computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 901 include but are not limited to a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 901 executes the various methods and processes described above, such as the video processing method. For example, in some embodiments, the video processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the video processing method described above can be executed. Alternatively, in other embodiments, computing unit 901 can be configured to execute the video processing method in any other suitable manner (e.g., by means of firmware).

[0163] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or a general-purpose programmable processor that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0164] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0165] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0166] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0167] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0168] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0169] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps recited in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0170] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video processing method, comprising: obtaining, by a terminal, a first video, where the first video is a video that has been recorded or a video stored locally; uploading, from the terminal to a server, the audio of the first video; receiving, by the terminal from the server, the subtitles of the first video returned based on the audio; in response to receiving a material recommendation trigger operation for the first video, uploading, from the terminal to the server, the subtitles of the first video; receiving, by the terminal from the server, a material group returned based on the subtitles of the first video, where the material group is determined based on at least one keyword in the subtitles; and generating a second video by adding the materials in the material group to the first video.

2. The method according to claim 1, where generating the second video by adding the materials in the material group to the first video includes: determining the appearance time of each material in the material group in the first video; and adding, at the corresponding appearance time on the time axis of the first video, the material corresponding to the corresponding appearance time.

3. The method according to claim 2, where determining the appearance time of each material in the material group in the first video includes: searching, from the subtitles of the first video, for the keywords included in the material group according to the correspondence between the materials in the material group and the keywords; and determining the appearance time of the keyword in the first video as the appearance time of the material corresponding to the keyword.

4. The method according to claim 1, where generating the second video by adding the materials in the material group to the first video further includes: determining the display information of the materials in the material group; and adding the materials to the image in the first video based on the display information of the materials.

5. The method according to claim 4, where the display information includes a display position, a display angle, and a display size, the display position is the display coordinates of the material in the image of the first video, the display angle is the angle of the material in the image of the first video relative to the horizontal line or the vertical line, and the display size is the size of the material in the image of the first video.

6. The method according to claim 1, where generating the second video by adding the materials in the material group to the first video further includes: in the case where there is a first material belonging to sound effect materials in the material group, determining the volume of the first material according to a preset volume ratio; and adding the volume of the first material to the audio in the first video.

7. The method according to claim 1, where generating the second video by adding the materials in the material group to the first video further includes: in the case where there is a second material belonging to decorative materials in the material group, identifying the target image in the first video corresponding to the second material; and adding the second material to a corresponding preset position in the target image.

8. The method according to claim 1, wherein the material groups returned based on the subtitles of the first video received by the terminal from the server include: Receiving a plurality of material groups returned by the server based on the subtitles; And Determining the material groups based on the plurality of material groups.

9. The method according to claim 1, wherein the material groups do not include a plurality of materials of a target type within a preset time period.

10. The method according to claim 1, wherein the materials corresponding to the same keyword in the material groups are different.

11. The method according to claim 1, wherein the material groups include materials of text type, sticker type, sound type, and animation type.

12. A video processing method, comprising: The server receives the audio of the first video from the terminal, and the first video is a recorded video obtained by the terminal or a video stored locally; Obtaining subtitles of the first video based on the audio; Sending the subtitles of the first video from the server to the terminal; The server receives the subtitles of the first video from the terminal, and the subtitles are uploaded by the terminal when a material recommendation trigger operation for the first video is received; Based on the subtitles, determining a material group for the first video, wherein the material group is determined based on at least one keyword in the subtitles; And Sending the material group from the server to the terminal, and the material group is used to indicate materials that can be added to the first video.

13. The method according to claim 12, wherein determining at least one material group for the first video based on the subtitles includes: Identifying at least one keyword from the subtitles; Based on the at least one keyword, determining the material group for the first video, wherein identifying the at least one keyword from the subtitles includes: Splitting the subtitles of the first video to obtain a plurality of first target words of the first video; And Taking at least one first target word that can be found in a preset word list library as the at least one keyword of the first video.

14. The method according to claim 13, wherein identifying the at least one keyword from the subtitles further includes: Extracting a plurality of second target words from the subtitles through a semantic recognition algorithm; Combining at least two of the plurality of second target words to obtain at least one combined word; And Taking the at least one combined word as the at least one keyword.

15. The method according to claim 13, wherein determining the material group for the first video based on the at least one keyword includes: Recalling at least one type of a plurality of materials for each keyword from a material database; Based on the at least one type of a plurality of materials corresponding to each keyword, determining a target number of materials for each keyword; And According to the target number of materials corresponding to each keyword, determining at least one material group for the first video.

16. The method according to claim 15, wherein determining a target number of materials for each keyword based on the at least one type of multiple materials corresponding to each keyword includes: Sorting the priorities of multiple materials of the same type for each keyword respectively; Selecting materials to be recommended of different types for each keyword according to the priority sorting of multiple materials of the same type for each keyword; And Determining a target number of materials for each keyword according to the materials to be recommended of different types for each keyword.

17. The method according to claim 12, further comprising: When a material group includes multiple identical keywords, performing deduplication processing on the materials corresponding to the multiple identical keywords.

18. The method according to claim 12, further comprising: When a material group includes multiple materials of a target type, performing deduplication processing on the multiple materials of the target type.

19. The method according to claim 12, wherein the material group includes materials of text type, sticker type, sound type, and animation type.

20. A terminal for video processing, comprising: A module configured to obtain a first video, where the first video is a video that has been recorded or a video stored locally; An audio uploading module configured to upload the audio of the first video to a server; A subtitle receiving module configured to receive the subtitle of the first video returned by the server based on the audio; A first sending module configured to upload the subtitle of the first video to the server in response to receiving a material recommendation trigger operation for the first video; A first receiving module configured to receive a material group returned by the server based on the subtitle of the first video, where the material group is determined based on at least one keyword in the subtitle; And A generating module configured to generate a second video by adding the materials in the material group to the first video.

21. A server for video processing, comprising: An audio receiving module configured to receive the audio of a first video from a terminal, where the first video is a video that has been recorded and obtained by the terminal or a video stored locally; A second recognition module configured to obtain the subtitle of the first video based on the audio; A subtitle sending module configured to send the subtitle of the first video to the terminal; A second receiving module configured to receive the subtitle of the first video from the terminal, where the subtitle is uploaded by the terminal when receiving a material recommendation trigger operation for the first video; A second determination module configured to determine a material group for the first video based on the subtitle, where the material group is determined based on at least one keyword in the subtitle; And A second sending module configured to send the material group to the terminal, where the material group is used to indicate the materials available for adding to the first video.

22. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor, Wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-19.

23. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are for causing the computer to execute the method according to any one of claims 1-19.

24. A computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes computer instructions that, when executed, cause a machine to implement the method according to any one of claims 1-19.

Citation Information

Patent Citations

  • Video processing method and device, video playing method and device, storage medium and equipment

    CN112040263A

  • Automatically adding sound effects into audio files

    CN112041809A

  • Video editing method and device, computer equipment and storage medium

    CN114449310A