Method, apparatus, device and program product for matching video materials
By performing content understanding and multimodal analysis on audio and video materials, segmented text is generated and automatically matched, solving the problems of time-consuming, labor-intensive, and inaccurate material searching in video editing. This achieves efficient and accurate material matching, improving the quality of video works and user experience.
Patent Information
- Application Number
- CN202410578970.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, manually searching for video footage that matches the audio during video editing is time-consuming and labor-intensive, and the matching accuracy is difficult to guarantee, affecting the quality and viewing experience of the video.
By understanding the audio content, multiple segmented texts are generated, and multimodal content analysis is performed on the video footage to generate corresponding text descriptions, automatically matching audio and video footage.
It improves the efficiency and accuracy of video material matching, reduces production costs, enhances the quality and viewing experience of video works, and improves the user experience.
Smart Images

Figure CN120935402A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of computers, and more specifically to methods, apparatus, electronic devices, and products for matching video footage. Background Technology
[0002] Video editing refers to the post-processing of raw video footage during video production, including selection, trimming, and splicing, to optimize content, adjust pacing, and enhance expressiveness, ultimately creating a coherent and complete video work.
[0003] Selecting raw video footage is a crucial step in video editing, directly determining the quality and style of the final video. This process involves reviewing and analyzing a large amount of material to ensure that the selected footage best serves the video's theme and narrative needs. Adding audio to the selected video footage for matching is also a vital step in video editing. This step aims to ensure that the video and audio are consistent in content and rhythm, thereby enhancing the overall quality and viewing experience of the video. Summary of the Invention
[0004] Embodiments of this disclosure provide a method, apparatus, electronic device, and product for matching video footage.
[0005] According to a first aspect of this disclosure, a method for matching video footage is provided. The method includes generating multiple segmented texts of audio based on an understanding of the content of the audio. The method also includes generating multiple text descriptions of multiple video footage based on an understanding of the multimodal content of multiple video footage. Furthermore, the method includes determining video footage that matches each segmented text in the multiple segmented texts, based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video footage.
[0006] In a second aspect of this disclosure, an apparatus for matching video footage is provided. The apparatus includes a segmented text generation module configured to generate multiple segmented texts of audio based on an understanding of the content of the audio. The apparatus also includes a text description generation module configured to generate multiple text descriptions of the multiple video footage based on an understanding of the multimodal content of the multiple video footage. Furthermore, the apparatus includes a video footage determination module configured to determine video footage that matches each of the multiple segmented texts of the audio and the multiple text descriptions of the multiple video footage.
[0007] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.
[0008] In a fourth aspect of this disclosure, a computer program product is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the method of the first aspect.
[0009] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram of an example environment in which some embodiments of this disclosure may be implemented is shown;
[0012] Figure 2 Flowcharts of methods for matching video footage according to some embodiments of this disclosure are shown;
[0013] Figure 3 The diagram illustrates a process for matching audio and video materials, according to some embodiments of this disclosure.
[0014] Figure 4A Schematic diagrams illustrating some embodiments of this disclosure for importing audio are shown;
[0015] Figure 4B The illustration shows some embodiments of this disclosure for displaying imported audio on an audio track;
[0016] Figure 4C Schematic diagrams illustrating some embodiments of this disclosure for importing video footage are shown;
[0017] Figure 5 A schematic diagram illustrating the types of audio and video material matching in some embodiments of this disclosure is shown;
[0018] Figure 6A A schematic diagram illustrating the effect of matching audio narration content with video dialogue in some embodiments of this disclosure;
[0019] Figure 6B This diagram illustrates another example of the matching effect between audio narration and video dialogue in some embodiments of this disclosure;
[0020] Figure 6C A schematic diagram illustrating the effect of matching audio narration content with video footage in some embodiments of this disclosure;
[0021] Figure 6D This diagram illustrates the effect of matching audio narration with video footage in some embodiments of this disclosure;
[0022] Figure 6E The illustration shows a schematic diagram of the effect of matching audio narration content with video footage and dialogue in some embodiments of this disclosure;
[0023] Figure 7 Block diagrams of apparatus for matching video footage according to some embodiments of the present disclosure are shown; and
[0024] Figure 8 Block diagrams of electronic devices according to some embodiments of the present disclosure are shown.
[0025] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation
[0026] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0027] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0028] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0029] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0032] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.
[0033] Currently, when editing video works, users need to search for video materials related to imported audio one by one, which is inconvenient for users submitting video works. Searching for matching video materials one by one consumes a lot of time and energy, especially when the audio and video materials are complex and long. This not only prolongs the video production cycle but may also affect the user's creative enthusiasm and motivation. Secondly, matching accuracy is difficult to guarantee. Manually searching and matching video materials one by one is prone to inaccurate matching or omissions due to human factors. Such inaccurate matching may affect the overall quality and viewing experience of the video work, and may even change the original message it was intended to convey.
[0034] According to embodiments of this disclosure, by understanding the audio content, multiple segmented texts accurately reflecting the core content of corresponding parts of the audio can be generated. Simultaneously, by processing the multimodal content (e.g., visuals, audio, subtitles, etc.) of multiple video materials, key content can be extracted from the video materials, and corresponding text descriptions can be generated. This reduces the time video producers spend manually writing and organizing audio and materials, improving video production efficiency. After generating the audio segmented texts and video material text descriptions, the next step is to compare and analyze each audio segmented text with the video material text descriptions. This automatically determines the video material that best matches each audio segmented text, facilitating video producers to complete a large amount of material screening and matching work in a short time.
[0035] This method reduces the complex operations required for video editors in finding, placing, and editing materials, thereby improving the efficiency of material matching. This not only lowers production costs but also enhances the quality and visual appeal of video works, providing viewers with a superior visual experience and ultimately improving the experience for all users.
[0036] Figure 1 A schematic diagram of an example environment 100 in which some embodiments of this disclosure may be implemented is shown. For example... Figure 1 As shown, interface 102 provides a platform for users to import audio 120, video materials 122, and video materials 124. Interface 102 includes a menu bar 104 with options such as "Media," "Audio," "Text," and "Stickers," offering users a wealth of editing functions and operational possibilities. On the left side of interface 102, there is a material import area 106, containing multiple primary options such as "Local," "Cloud Materials," and "Material Library."
[0037] The "Cloud Media" option allows users to easily access and import media from cloud storage. During video editing, users can utilize these cloud-stored materials instead of importing them from their local devices each time. The "Media Library" option allows users to browse and select more media clips. The "Local" option in media import area 106 allows users to import video, audio, and image media from their computer's local storage devices (such as hard drives) into the editing software. This allows users to add various media files saved locally to their editing projects for subsequent editing, splicing, and effects addition.
[0038] The central part of interface 102 is the display area 108 for user-imported materials. Above the material display area 108 is a search box 110, where users can quickly locate the desired material by entering a file name, visual element, or dialogue. Below the search box 110 is a view editing area 112, which includes options such as "Layout," "Sort," and "Filter." Users can use the view editing area 112 to select and display imported material clips. If the user switches to the "Import" view 114, they can import the material files needed for the editing project using the "Import" button 116 in the material display area 108. For example, clicking "Import" 116 may bring up a file selection dialog box, allowing the user to browse and select local files. After selecting the desired files, these files will be imported into the editing software's material library, becoming available for editing.
[0039] In the material display area 108, imported audio 120, video material 122, and video material 124 are displayed. Each material clip has a corresponding thumbnail and playback duration marker. These material clips may have been selected from the user's imported material library, and users can select and add them to the editing sequence as needed. The file name and duration information are displayed below the thumbnail of each material clip for easy identification and management. When a user wants to quickly find a matching video material for the imported audio 120 among the imported video material 122 and imported video material 124, they can use "Automatic Material Matching" 118 to find suitable material for the audio. Interface 102 displays a comprehensive and user-friendly video editing software interface, which can be used by both beginners and professional editors to manage and edit materials.
[0040] The following will combine Figures 2 to 8 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0041] Figure 2 A flowchart of a method 200 for matching video footage, according to some embodiments of this disclosure, is shown. At block 202, multiple segmented texts of the audio are generated based on an understanding of the audio content. In some embodiments, the audio may be narration recorded by the user. In some embodiments, suppose there is an audio clip with the following content: “This video shows a busy city street scene. Many pedestrians and vehicles can be seen moving about the street. High-rise buildings line both sides of the street, creating a modern urban atmosphere. Meanwhile, the text on the billboards also attracts attention.” Based on this understanding of the audio content, multiple segmented texts can be generated: Segmented text 1: “This video shows a busy city street scene”; Segmented text 2: “Many pedestrians and vehicles can be seen moving about the street”; Segmented text 3: “High-rise buildings line both sides of the street, creating a modern urban atmosphere”; Segmented text 4: “The text on the billboards also attracts attention.” Each segmented text understands a key portion of the audio content and maintains the coherence and integrity of the content. Such segmented text not only helps to better understand and analyze audio content, but also facilitates subsequent audio-video matching.
[0042] In box 204, based on the understanding of the multimodal content of multiple video clips, multiple text descriptions for the multiple video clips are generated. Suppose there is a video clip about city life, which sequentially includes pedestrians, vehicles, buildings, billboards, and conversations between people on the street. This video can be edited into several video clip segments in sequence. In some embodiments, the conversations of pedestrians in the video can be converted into text, such as: "It's so lively here, people are coming and going." In some embodiments, the image content in the video can be analyzed to generate text descriptions such as: "There are many pedestrians and vehicles in the video, and tall buildings line both sides of the street." In some embodiments, text information in the video can be recognized. If there is text on billboards or shop signs in the video, it can be converted into editable text, such as: "The billboard says 'Welcome.'" In some embodiments, the above content can also be combined to generate a video content description like: "This video successfully captures the hustle and bustle and vibrancy of city life." This method helps to more comprehensively understand the content and characteristics of the video clips, providing strong support for subsequent video processing and applications.
[0043] In box 206, based on multiple segmented texts of audio and multiple text descriptions of video footage, video footage matching each segmented text is determined. In some embodiments, for segmented text three, "High-rise buildings line both sides of the street, creating a modern urban atmosphere," it can be matched with the text description generated in the video, "There are many pedestrians and vehicles in the video, and high-rise buildings line both sides of the street," thus making the corresponding video footage segment the appropriate video footage to match audio segmented text three. Similarly, other segmented texts, such as segmented text one, can be matched with the text descriptions of the generated video content to find the most suitable video footage.
[0044] In the embodiments of this disclosure, this method of generating multiple segmented texts from audio content enables the audio content to be structured, facilitating subsequent matching with video footage. Simultaneously, analyzing the multimodal content of the video footage allows for a more comprehensive extraction of its content, ensuring the accuracy and richness of the text description. This comprehensive extraction of video content also facilitates subsequent matching with the audio segmented text. Furthermore, by comparing the segmented audio text with the textual descriptions of the video footage, precise synchronization and correspondence between audio and video can be achieved. This method also improves video production efficiency, thereby enhancing the user experience.
[0045] Figure 3 A schematic diagram illustrating a process 300 for matching audio and video materials, representing some embodiments of this disclosure, is shown. (See also:) Figure 3The system allows importing audio, including background music, sound effects, dialogue, natural sounds, and narration. It supports various audio formats such as WAV, MP3, AAC, and WMA. This increases flexibility and compatibility, enabling users to more easily handle different types of materials.
[0046] The following will combine Figures 4A-4B The present disclosure provides schematic diagrams illustrating some embodiments of the present disclosure for importing audio and some embodiments of the present disclosure for displaying imported audio on an audio track. Figure 4A A schematic diagram of some embodiments of this disclosure for importing audio 400A is shown. Figure 4A It shows Figure 3 Imported audio as shown in page 310. (Reference) Figure 4A In interface 402A, users can import the audio materials they need 408A through the "Import" button 404A on the material display page 406A. Figure 4B Schematic diagrams illustrating some embodiments of this disclosure for displaying imported audio 400B on an audio track. For example... Figure 4B As shown, the audio imported by the user is located on audio track 404B at the bottom of the track editing page 402B.
[0047] Return to Figure 3 In section 320, import video footage. This footage can include actual shooting scenes, animations, special effects clips, or finished film clips. The following will combine... Figure 4C The following are schematic diagrams illustrating some embodiments of the present disclosure for importing video materials. Figure 4C A schematic diagram of some embodiments of this disclosure for importing video material 400C is shown. Figure 4C It shows Figure 3 The imported video footage shown in image 320. (Reference) Figure 4C In interface 402C, after importing the audio they need, users can continue to import the video footage they need using the "Import" button 404C 406C. There is no specific order for importing audio and video footage. When different types of files are imported into the video editing project, an "Auto-match Footage" option 408C will appear. Combined with... Figure 3 Then you can execute it through "Auto Match Material" 408C. Figure 3 The material shown matches 330.
[0048] By directly clicking a button to import materials, users no longer need to go through complicated menus or steps, simplifying the operation process, improving operational efficiency, and the user-friendly design makes material import more intuitive and simple, reducing the confusion and frustration that users may encounter during the operation process, and improving the overall user experience.
[0049] Continue to refer to Figure 3 In section 330, material matching is performed. For example, in a narration video editing project, once the user has imported audio and some video materials to match it, the system can automatically find the most suitable video materials for the narration audio. Specifically, in section 332, audio is understood and segmented text is generated. In some embodiments, a Big ASR (Big Audio Speech Model) can be used to deeply understand the content of the user-uploaded narration audio and generate multiple segmented texts. For example, for an audio clip like "Middle-aged, penniless, homeless, shop deserted, unable to pay rent, and all he does is smoke and gamble all day," the Big ASR can automatically generate: segment text 1 "Middle-aged"; segment text 2 "penniless, homeless"; segment text 3 "shop deserted"; segment text 4 "unable to pay rent"; segment text 5 "and all he does is smoke and gamble all day." This method helps users avoid manually segmenting the audio text into multiple segments, thereby improving user efficiency and enhancing the user experience.
[0050] In section 334, the video footage is understood and a text description of the video is generated. In some embodiments, a multimodal understanding model can be used to generate the text description of the video footage. In some embodiments, this multimodal understanding model has at least the following multimodal capabilities: large speech model, computer vision (CV), optical character recognition (OCR), highlight segmentation, etc. Highlight segmentation capability refers to the ability to automatically identify, extract, and present the most exciting and eye-catching segments in video, audio, or other multimedia content. These highlight segments are typically the most representative and attractive parts of the content. In some embodiments, the segmented text description of the video footage includes the dialogue of the video footage, the description of the video footage's visuals, and a combined description of the dialogue and visuals. This adaptive and flexible method of understanding video content through multimodal means enables the generation of more accurate text descriptions of the video footage.
[0051] The following will combine Figure 5 This describes the types of audio and video material matching in some embodiments of this disclosure. Figure 5 A schematic diagram of type 500 for matching audio and video materials according to some embodiments of this disclosure is shown. Figure 5 It shows Figure 3 The matching segment text and video text description types are shown in section 334. (Reference) Figure 5In 502, matching with dialogue. If the dialogue in the video footage is relatively dominant, then a text description of the video can be generated based on the dialogue in the video footage. The matching degree can then be calculated between this text description containing the video dialogue and the segmented text. In some embodiments, if the subtitle "Can't afford the rent" is dominant in the multimodal content of the video footage, the text description of the video footage generated by the multimodal understanding model is "Chen's shop can't afford the rent".
[0052] refer to Figure 5 In case of a 504 error, the text description matches the visuals. If visuals are relatively dominant in the video footage, a text description of the video can be generated based on those visuals. This text description containing video content can then be compared with segmented text to calculate the match. In some embodiments, if the visuals of an "old man with gray hair" dominate the multimodal content (i.e., only the visuals), a multimodal understanding model can generate a text description for the video footage as "father is sick."
[0053] refer to Figure 5 In section 506, matching is performed with both visuals and dialogue. If both dialogue and visuals are dominant in the video footage, a text description of the video can be generated based on the dialogue and visuals. This text description containing video content can then be compared with segmented text to calculate the matching degree. In some embodiments, if the dialogue "You warned me" and the visual "Chen is shouting and pushing others" are both dominant in the video footage, the text description of the video footage generated by the multimodal understanding model can be "Chen hits someone."
[0054] In some embodiments, a video clip depicts Chen standing in front of his modest health product shop, his brow furrowed and his expression grave. The shop window displays unsold health products, creating a desolate and bleak atmosphere. The shop is dimly lit, with few items on the shelves, and the entire space is filled with a sense of desolation and loneliness. Suddenly, Chen's phone rings; it's his landlord calling. He picks up the phone, listening to the landlord's unwavering demand for payment, his fingers fidgeting aimlessly, seemingly lost. His eyes reveal helplessness and anxiety, as if he's being suffocated by the burdens of life. He takes a deep breath, trying to calm himself, and then speaks to the landlord... He explained his predicament, but the landlord seemed unconcerned, coldly demanding he pay the rent immediately. Speechless, Chen could only silently hang up, his helplessness and despair deepening. He turned to look at everything in the shop; the goods that had once held his hopes and dreams had now become an inescapable burden. His gaze lingered on the shop, finally settling on the health supplements, as if pondering how to escape this predicament. (The text then mentions a multimodal understanding model that can identify the image and sound of Chen standing in front of the shop, hearing the landlord's phone call, and the dialogue of Chen trying to explain his predicament, generating the text description "Chen cannot pay the rent.")
[0055] In some embodiments, a multimodal understanding model can be used to identify the shop sign and the lack of customers in Chen's shop, generating the text description "Chen's shop is deserted." In some embodiments, the multimodal understanding model can also identify Chen's facial expressions, posture, and environmental features within the shop to generate the content description "Chen's brows are furrowed, his face is solemn, showing his inner anxiety and stress." In some embodiments, these content descriptions can be combined to ultimately generate the content description "Unable to pay rent."
[0056] return Figure 3 After multiple segmented texts of the audio and multiple text descriptions of the video clips are generated, step 336 involves matching the segmented texts with the text descriptions. In some embodiments, the matching degree between the segmented text of each audio clip and the multiple text descriptions of the user-uploaded video clips can be calculated. In some embodiments, the target video clip for adaptation is determined based on the matching degree.
[0057] In some embodiments, a segment of text in the audio is "unable to pay rent". Among the text descriptions of multiple video clips, the text description of video clip 1, "cannot pay, no money", is found to have a 95% match with the segment text "unable to pay rent". This match is the highest among all the video clips. Therefore, the target video clip that best matches the segment text "unable to pay rent" is video clip 1.
[0058] In some embodiments, the text description "Can't pay, no money" in video material 1 can be determined based on the dialogue in video material 1, or based on the visuals in video material 1, or based on both the visuals and the dialogue in video material 1.
[0059] In some embodiments, if both segmented text 1 and segmented text 2 are best matched with video clip 2, then the template video clip that matches which segmented text 2 can be determined based on the temporal order of segmented text 1 and segmented text 2 in the audio. For example, if segmented text 1 precedes segmented text 2, then segmented text 1 is the best match for video clip 2. In some embodiments, video clips that are second best matched with segmented text 2 can be selected for matching.
[0060] In some embodiments, if the duration of video clip 1 is longer than the duration of the segmented text "unable to pay rent," a multimodal understanding model can be used to determine the most suitable candidate video clip from video clip 1 based on the duration of the segmented text "unable to pay rent." For example, a highlight segment from video clip 1 with the same duration as the segmented text "unable to pay rent" can be determined as the most suitable candidate video clip. In some embodiments, if the duration of video clip 1 is shorter than the duration of the segmented text "unable to pay rent," video clips with a lower matching degree (e.g., video clip 2 with a matching degree of 90%) can be selected to fill the gap.
[0061] Continue to refer to Figure 3 In section 340, the results of the material matching are displayed to the user. For example, in a narration video editing project, the user can be shown video clips that match the uploaded audio, along with the corresponding video clip timestamps. The matching degree between the segmented text and the corresponding video clips can also be shown, or a brief summary of the video clip can be displayed.
[0062] By matching the audio segment titles with the content descriptions, the system automatically finds the most suitable video footage to fill each segment. This precise matching reduces the complexity of finding, placing, and editing materials, making the video production process simpler and more efficient, and lowering the production difficulty. This is especially user-friendly for beginners or non-professionals, thus improving the user experience. Simultaneously, the system displays the matching degree between the video footage and the audio segment text. Users can view the matching score and choose to keep, replace, or further optimize the audio footage to ensure the final video content and audio narration achieve optimal harmony. This flexible adjustment mechanism not only enhances the personalization of the production but also further ensures the quality of the video and the viewer experience.
[0063] The following will combine Figures 6A-6E This is to illustrate the display effect of matching audio and video materials. Figure 6A A schematic diagram of the effect 600A of matching audio segment text and video dialogue according to some embodiments of this disclosure is shown. Figure 6A It shows Figure 5 The audio content shown in section 502 matches the dialogue. Combined with... Figure 4C When a user imports the audio narration and corresponding video footage into a video editing project and clicks "Auto-match Footage" (408C), the following will appear: Figure 6A The interface shown is 602A. (Reference) Figure 6A The interface 602A includes a material import page 604A, a player page 606A, and a track editing page 608A. Figure 6A The material import page 604A shown is... Figure 4A Interface 402A or Figure 4C Interface 402C. The track editing page 608A includes editing areas such as the main video track area 610A, the subtitle area 612A, the matching result display area 614A, and the timeline pointer 616A. The main video track area 610A is used to place and edit video footage. The text track area 612A displays segmented text content of the audio explanations for the footage. The matching result display area 614A can display information such as a summary of the video footage corresponding to the audio content, the duration of the video footage, and the matching degree.
[0064] Continue to refer to Figure 6A , combined Figure 3 As shown in 332, the segmented text of one segment of the audio comprehension and explanation can be obtained as "Can't even afford the rent". Figure 3 The shown material matches 330. On video track 610A of track editing page 608A, video material 1 corresponding to the segmented text "Can't even afford the rent" in the audio content is displayed. When the timeline pointer 616A stops in the corresponding area, in the video screen of player page 606A, the dialogue 620A of video material 1 is displayed as "Can't pay, no money," and the corresponding segmented text subtitle 618A "Can't even afford the rent" is also displayed in the video. At this time, referring to the information displayed in the matching result display area 614A, it can be seen that the main content of 18-28S of video material 1 is "Chen's shop can't afford the rent," and the match degree with the segmented text "Can't even afford the rent" is 95%.
[0065] Figure 6B A schematic diagram of another audio segment text and video dialogue matching effect 600B, representing some embodiments of this disclosure, is shown. Figure 6B It shows Figure 5The audio content shown in section 502 matches the dialogue. Combined with... Figure 4C When a user imports the audio narration and corresponding video footage into a video editing project and clicks "Auto-match Footage" (408C), the following will appear: Figure 6B The interface shown is 602B. (Reference) Figure 6B , combined Figure 3 As shown in 332, the text content of the audio explanation can be obtained in the subtitle area 606B. For example, the segmented text of one audio segment is "Children are best with me." Figure 3 The material shown is matched 330. On video track 604B, video material 3, corresponding to the audio content segment text "It's best if the child stays with me," is displayed. When the timeline pointer stops at the corresponding area, in the video view on the player page, the dialogue 612B of video material 3 displays "It's best if the child stays with me," and the corresponding segment text subtitle 610B, "It's best if the child stays with me," is also displayed in the video. At this point, referring to the information displayed in the matching result display area 608B, it can be seen that the main content of segments 31-37S in video material 3 is "fighting for custody of the son," and the match rate with the segment text "It's best if the child stays with me" is 85%.
[0066] Figure 6C A schematic diagram of the effect 600C of matching audio segment text with video footage according to some embodiments of this disclosure is shown. Figure 600C illustrates... Figure 5 The audio content shown in error 504 matches the video. (Combined with...) Figure 4C When a user imports the audio narration and corresponding video footage into a video editing project and clicks "Auto-match Footage" (408C), the following will appear: Figure 6C The interface shown is 602C. (Reference) Figure 6C , combined Figure 3 As shown in 332, the text content of the audio narration can be obtained in the subtitle area 606C. For example, the segmented text of one audio segment is "Father is sick". Figure 3 The shown material matches 330, displaying video material 2 on video track 604C that corresponds to the audio content segment text "Father is sick". When the timeline pointer stops at the corresponding area, the video screen on the player page shows the scene of "Father is sick" in video material 2, and the corresponding segment text subtitle 610C "Father is sick" is also displayed in the video. At this time, referring to the information displayed in the matching result display area 608C, it can be seen that the main content of 28-31 seconds of video material 2 is "Father is sick and left in a nursing home", and the matching degree with the segment text "Father is sick" is 75%.
[0067] Figure 6D A schematic diagram of another audio segment text matching effect 600D with video footage, according to some embodiments of this disclosure, is shown. Figure 6D It shows Figure 5 The audio content shown in error 504 matches the video. (Combined with...) Figure 4C When a user imports the audio narration and corresponding video footage into a video editing project and clicks "Auto-match Footage" (408C), the following will appear: Figure 6D The interface shown is 602D. (Reference) Figure 6D , combined Figure 3 As shown in 332, the text content of the audio explanation can be obtained in the subtitle area 606D. For example, the segmented text of one audio segment is "The man takes off his thick three-layer mask". Figure 3 The material shown is matched 330. On video track 604D, video material 5, corresponding to the audio content segment text "The man takes off his thick three-layer mask," is displayed. When the timeline pointer stops at the corresponding area, in the video screen of the player page, the scene 612D of video material 5 shows "The man takes off his thick three-layer mask," and the corresponding segment text subtitle 610D, "The man takes off his thick three-layer mask," is also displayed in the video. At this time, referring to the information displayed in the matching result display area 608D, it can be seen that the main content of 48-53 seconds of video material 5 is "Old Wang, who has leukemia, meets Chen Chu," and the match rate with the segment text "The man takes off his thick three-layer mask" is 85%.
[0068] Figure 6E A schematic diagram of the effect 600E of matching audio segment text with video footage and dialogue according to some embodiments of this disclosure is shown. Figure 6E It shows Figure 5 The audio content shown in section 506 matches the visuals and dialogue. Combined with... Figure 4C When a user imports the audio narration and corresponding video footage into a video editing project and clicks "Auto-match Footage" (408C), the following will appear: Figure 6E The interface shown is 602E. (Reference) Figure 6E , combined Figure 3 As shown in 332, the text content of the audio explanation can be obtained in the subtitle area 606E. For example, the segmented text of one audio segment is "Only bullying the weak to bluff their way in." Figure 3The shown material matches 330, and video material 3, corresponding to the audio content segment text "Only bullying the weak to bluff," is displayed on video track 604E. When the timeline pointer stops at the corresponding area, in the video screen of the player page, the scene 612E of video material 3 shows "Chen shouting" and the line 610E "You warned me," while the corresponding segment text subtitle 614E "Only bullying the weak to bluff" is also displayed in the video. At this time, referring to the information displayed in the matching result display area 608E, it can be seen that the main content of 37-48 seconds of video material 3 is "Chen hitting people," and the match degree with the segment text "Only bullying the weak to bluff" is 88%.
[0069] This method of automatically generating and displaying video footage descriptions and segmented text on the track boundary page reduces the time video producers spend manually writing and organizing, allowing them to focus more on creating and optimizing video content, thereby improving overall production efficiency. At the same time, the generated descriptions and titles more accurately reflect the core content of the video footage and narration audio, and the displayed matching degree allows users to easily decide whether to adjust the video footage they are editing.
[0070] Figure 7 A block diagram of an apparatus 700 for matching video footage, according to some embodiments of the present disclosure, is shown. Figure 7 As shown, the device 700 includes a segmented text generation module 702, configured to generate multiple segmented texts of audio based on an understanding of the audio content. The device 700 also includes a text description generation module 704, configured to generate multiple text descriptions of multiple video clips based on an understanding of the multimodal content of multiple video clips. Furthermore, the device 700 includes a video clip determination module 706, configured to determine video clips that match each segmented text in the multiple segmented texts, based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video clips.
[0071] Figure 8 Block diagrams of electronic devices 800 according to some embodiments of the present disclosure are shown. Device 800 may be the device or apparatus described in the embodiments of the present disclosure. Figure 8As shown, device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RM) 803. RM 803 may also store various programs and data required for the operation of device 800. CPU / GPU 801, ROM 802, and RM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804. Although not shown in... Figure 8 As shown, device 800 may also include a coprocessor.
[0072] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0073] The various methods or processes described above can be executed by CPU / GPU 801. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into ROM 803 and executed by CPU / GPU 801, one or more steps or actions in the methods or processes described above can be performed.
[0074] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.
[0075] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0076] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0077] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (IS) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LN) or a wide area network (WN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGs), or programmable logic arrays (PLs), is personalized by utilizing the status information of the computer-readable program instructions to execute the computer-readable program instructions, thereby implementing various aspects of this disclosure.
[0078] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0079] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0081] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0082] The following are some example implementations of this disclosure.
[0083] Example 1. A method for matching video footage, comprising:
[0084] Based on the understanding of the audio content, multiple segmented texts of the audio are generated;
[0085] Based on the understanding of the multimodal content of multiple video materials, multiple text descriptions of the multiple video materials are generated;
[0086] Based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video clips, video clips that match each of the multiple segmented texts are determined.
[0087] Example 2. According to the method described in Example 1, generating multiple segmented texts of the audio based on an understanding of the audio content includes:
[0088] The speech model generates multiple segmented texts of the audio and multiple speech texts corresponding to the multiple segmented texts based on the audio content.
[0089] Example 3. According to any one of Examples 1-2, generating multiple text descriptions of the multiple video clips based on the understanding of the multimodal content of the multiple video clips includes:
[0090] Based on the multimodal content of the multiple video materials, the multimodal understanding model generates multiple text descriptions of the multiple video materials, wherein the multimodal content includes the audio content, visual content, and highlight video content of the video materials.
[0091] Example 4. The method according to any one of Examples 1-3, wherein determining the video material matching each of the plurality of segmented texts in the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials includes:
[0092] Determine multiple matching degrees between each segmented text and multiple text descriptions of the multiple video clips; and
[0093] Based on the multiple matching degrees, target video material that matches each segment of text is determined.
[0094] Example 5. The method according to any one of Examples 1-4, wherein determining the video material matching each segment of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further includes:
[0095] In response to the first segment text and the second segment text in the plurality of segment texts matching the same video material, and the first segment text being before the second segment text, the same video material is identified as the template video material that matches the first segment text in the time sequence.
[0096] Example 6. The method according to any one of Examples 1-5, wherein determining the video material matching each segment of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further includes:
[0097] The multimodal understanding model generates candidate video footage corresponding to the time range of each segmented text based on the time range of each segmented text.
[0098] Example 7. The method according to any one of Examples 1-6 further includes:
[0099] In response to the target video material's duration being shorter than the segmented text's duration, video material is added to the target video material, the added video material being determined based on the matching degree.
[0100] Example 8. The method according to any one of Examples 1-7 further includes:
[0101] In response to the detection of the audio upload, the audio is displayed on the audio track;
[0102] In response to detecting the upload of the plurality of video materials, the plurality of uploaded video materials are displayed on the material display page; and
[0103] A button for matching materials is displayed on the material display page.
[0104] Example 9. The method according to any one of Examples 1-8 further includes:
[0105] In response to detecting a user's click on the button, the plurality of segmented texts of the audio are displayed sequentially on the text track;
[0106] The candidate video materials corresponding to each segmented text are displayed sequentially on the main video track; and
[0107] The time range of each segmented text, the text description of the candidate video material, and the matching degree between each segmented text and the candidate video material are displayed.
[0108] Example 10. The method according to any one of Examples 1-9 further includes:
[0109] In response to detecting a user's click on the button, a subtitle corresponding to each segment of text is added to the candidate video material, the subtitle being the audio text corresponding to each segment of text.
[0110] Example 11. An apparatus for matching video footage, comprising:
[0111] The segmented text generation module is configured to generate multiple segmented texts of the audio based on an understanding of the audio content.
[0112] The text description generation module is configured to generate multiple text descriptions of the multiple video materials based on the understanding of the multimodal content of the multiple video materials;
[0113] The video material determination module is configured to determine video material that matches each segment of the multiple segmented texts based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video materials.
[0114] Example 12. The apparatus according to Example 11, wherein the segmented text generation module includes:
[0115] The first generation module is configured to generate, based on the audio content, multiple segmented texts of the audio and multiple speech texts corresponding to the multiple segmented texts using a speech model.
[0116] Example 13. The apparatus according to any one of Examples 11-12, wherein text description generation comprises:
[0117] The second generation module is configured to generate multiple text descriptions of the multiple video materials based on the multimodal content of the multiple video materials by the multimodal understanding model. The multimodal content includes the audio content, visual content, and highlight video content of the video materials.
[0118] Example 14. The apparatus according to any one of Examples 11-13, wherein the video material determining module comprises:
[0119] The matching degree determination module is configured to determine multiple matching degrees between each segmented text and multiple text descriptions of the multiple video materials; and
[0120] The target video material determination module is configured to determine the target video material that matches each segment of text based on the multiple matching degrees.
[0121] Example 15. The apparatus according to any one of Examples 11-14, wherein the video material determining module further comprises:
[0122] The template video material determination module is configured to determine the same video material as the template video material that matches the first segment text and the second segment text in the plurality of segment texts, and the first segment text precedes the second segment text.
[0123] Example 16. The apparatus according to any one of Examples 11-15, wherein the template video material further includes:
[0124] The candidate video material determination module is configured to generate candidate video materials corresponding to the time range of each segmented text based on the time range of each segmented text by the multimodal understanding model.
[0125] Example 17. The apparatus according to any one of Examples 11-16 further includes:
[0126] The fill module is configured to fill the target video material with video material in response to the target video material having a duration shorter than the duration of the segmented text, the fill video material being determined based on the matching degree.
[0127] Example 18. The apparatus according to any one of Examples 11-17 further includes:
[0128] A first display module is configured to display the audio on the audio track in response to detecting the upload of the audio;
[0129] The second display module is configured to display the uploaded video materials on a material display page in response to detecting the upload of the plurality of video materials; and
[0130] The third display module is configured to display a button for matching materials on the material display page.
[0131] Example 19. The apparatus according to any one of Examples 11-18 further includes:
[0132] The fourth display module is configured to, in response to detecting a user's click on the button, sequentially display the plurality of segmented texts of the audio on a text track;
[0133] The fifth display module is configured to sequentially display the candidate video materials corresponding to each segmented text on the main video track; and
[0134] The sixth display module is configured to display the time range of each segmented text, the text description of the candidate video material, and the matching degree between each segmented text and the candidate video material.
[0135] Example 20. The apparatus according to any one of Examples 11-19 further includes:
[0136] The seventh display module is configured to add subtitles corresponding to each segment of text to the candidate video material in response to detecting a click on the button by the user, wherein the subtitles are the audio text corresponding to each segment of text.
[0137] Example 21. An electronic device comprising:
[0138] Processor; and
[0139] A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform actions, the actions including:
[0140] Based on the understanding of the audio content, multiple segmented texts of the audio are generated;
[0141] Based on the understanding of the multimodal content of multiple video materials, multiple text descriptions of the multiple video materials are generated;
[0142] Based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video clips, video clips that match each of the multiple segmented texts are determined.
[0143] Example 22. The electronic device according to Example 21, wherein generating multiple segmented texts of the audio based on an understanding of the audio content includes:
[0144] The speech model generates multiple segmented texts of the audio and multiple speech texts corresponding to the multiple segmented texts based on the audio content.
[0145] Example 23. An electronic device according to any one of Examples 21-22, wherein generating multiple text descriptions of the multiple video materials based on an understanding of the multimodal content of the multiple video materials includes:
[0146] Based on the multimodal content of the multiple video materials, the multimodal understanding model generates multiple text descriptions of the multiple video materials, wherein the multimodal content includes the audio content, visual content, and highlight video content of the video materials.
[0147] Example 24. An electronic device according to any one of Examples 21-23, wherein determining the video material matching each of the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials includes:
[0148] Determine multiple matching degrees between each segmented text and multiple text descriptions of the multiple video clips; and
[0149] Based on the multiple matching degrees, target video material that matches each segment of text is determined.
[0150] Example 25. An electronic device according to any one of Examples 21-24, wherein determining the video material matching each of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further includes:
[0151] In response to the first segment text and the second segment text in the plurality of segment texts matching the same video material, and the first segment text being before the second segment text, the same video material is identified as the template video material that matches the first segment text in the time sequence.
[0152] Example 26. An electronic device according to any one of Examples 21-25, wherein determining the video material matching each of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further includes:
[0153] The multimodal understanding model generates candidate video footage corresponding to the time range of each segmented text based on the time range of each segmented text.
[0154] Example 27. The electronic device according to any one of Examples 21-26 further includes:
[0155] In response to the target video material's duration being shorter than the segmented text's duration, video material is added to the target video material, the added video material being determined based on the matching degree.
[0156] Example 28. The electronic device according to any one of Examples 21-27 further includes:
[0157] In response to the detection of the audio upload, the audio is displayed on the audio track;
[0158] In response to detecting the upload of the plurality of video materials, the plurality of uploaded video materials are displayed on the material display page; and
[0159] A button for matching materials is displayed on the material display page.
[0160] Example 29. The electronic device according to any one of Examples 21-28 further includes:
[0161] In response to detecting a user's click on the button, the plurality of segmented texts of the audio are displayed sequentially on the text track;
[0162] The candidate video materials corresponding to each segmented text are displayed sequentially on the main video track; and
[0163] The time range of each segmented text, the text description of the candidate video material, and the matching degree between each segmented text and the candidate video material are displayed.
[0164] Example 30. The electronic device according to any one of Examples 21-29 further includes:
[0165] In response to detecting a user's click on the button, a subtitle corresponding to each segment of text is added to the candidate video material, the subtitle being the audio text corresponding to each segment of text.
[0166] Example 31. A computer-readable storage medium having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 10.
[0167] Example 32. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of Examples 1 to 10.
[0168] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for matching video footage, comprising: Based on the understanding of the audio content, multiple segmented texts of the audio are generated; Based on the understanding of the multimodal content of multiple video materials, multiple text descriptions of the multiple video materials are generated; Based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video clips, video clips that match each of the multiple segmented texts are determined.
2. The method according to claim 1, wherein generating multiple segmented texts of the audio based on an understanding of the audio content includes: The speech model generates multiple segmented texts of the audio and multiple speech texts corresponding to the multiple segmented texts based on the audio content.
3. The method according to claim 1, wherein generating multiple text descriptions of the multiple video materials based on the understanding of the multimodal content of the multiple video materials includes: Based on the multimodal content of the multiple video materials, the multimodal understanding model generates multiple text descriptions of the multiple video materials, wherein the multimodal content includes the audio content, visual content, and highlight video content of the video materials.
4. The method of claim 1, wherein determining the video material matching each segment of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials comprises: Determine multiple matching degrees between each segment of text and multiple text descriptions of the multiple video materials; as well as Based on the multiple matching degrees, target video material that matches each segment of text is determined.
5. The method of claim 4, wherein determining the video material matching each segment of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further comprises: In response to the first segment text and the second segment text in the plurality of segment texts matching the same video material, and the first segment text precedes the second segment text, the same video material is identified as the template video material that matches the first segment text in the time sequence.
6. The method of claim 5, wherein determining the video material matching each segment of the plurality of segmented texts based on the plurality of segmented texts of the audio and the plurality of text descriptions of the plurality of video materials further comprises: The multimodal understanding model generates candidate video footage corresponding to the time range of each segmented text based on the time range of each segmented text.
7. The method according to claim 6, further comprising: In response to the target video material's duration being shorter than the segmented text's duration, video material is added to the target video material, the added video material being determined based on the matching degree.
8. The method according to claim 6, further comprising: In response to the detection of the audio upload, the audio is displayed on the audio track; In response to the detection of the upload of the multiple video materials, the multiple uploaded video materials are displayed on the material display page; as well as A button for matching materials is displayed on the material display page.
9. The method according to claim 8, further comprising: In response to detecting a user's click on the button, the plurality of segmented texts of the audio are displayed sequentially on the text track; The candidate video materials corresponding to each segmented text are displayed sequentially on the main video track; as well as The time range of each segmented text, the text description of the candidate video material, and the matching degree between each segmented text and the candidate video material are displayed.
10. The method according to claim 9, further comprising: In response to detecting a user's click on the button, subtitles corresponding to each segment of text are added to the candidate video material, wherein the subtitles are audio text corresponding to each segment of text.
11. An apparatus for matching video footage, the apparatus comprising: The segmented text generation module is configured to generate multiple segmented texts of the audio based on an understanding of the audio content. The text description generation module is configured to generate multiple text descriptions of the multiple video materials based on the understanding of the multimodal content of the multiple video materials; The video material determination module is configured to determine video material that matches each segment of the multiple segmented texts based on the multiple segmented texts of the audio and the multiple text descriptions of the multiple video materials.
12. An electronic device, comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 10.