Video retrieval method and device

By configuring a video search table of a large language model locally, receiving storyboard description information and searching video slices, the problem of inefficient search accuracy and efficiency in video clips in the prior art is solved, and efficient and accurate video slice retrieval and addition in local video materials is realized, and video clipping efficiency is improved.

CN120523992APending Publication Date: 2025-08-22SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510555292.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When searching for video clips during video editing, the search accuracy and efficiency are inefficient and accurate, and the required video slices cannot be retrieved from the Internet cloud efficiently and accurately.

Method used

By configuring the video search table output from the large language model locally, receiving target storyboard description information, searching target video slices from the local material search table, and displaying and adding these slices in the video editing interface, supporting preview and cloud-based slicing videos to generate structured indexes.

Benefits of technology

It realizes efficient and accurate retrieval and add required video slices in local video materials, improving the efficiency and user experience of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523992A_ABST
    Figure CN120523992A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video retrieval method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of video processing. The video retrieval method comprises the following steps: displaying a video editing interface, wherein a search entry is configured on the video editing interface; receiving target sub-mirror description information through the search entry; according to the sub-mirror description information, retrieving a corresponding target video slice from a preset local material retrieval table; wherein the local material retrieval table comprises video retrieval information of each video slice output based on the large language model. According to the technical scheme provided by the embodiment of the invention, the required video slice can be accurately retrieved from the local video material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video processing technology, and in particular to a video retrieval method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] From creating personal creative works to disseminating various media, video editing is an indispensable application. During the video editing process, it is often necessary to search for video clips and edit them into the creative work. However, the current method of finding the required video clips relies on a large-scale search on the internet cloud, which is inaccurate and inefficient.

[0003] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the Invention

[0004] The embodiments of the present application provide a video retrieval method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0005] One aspect of an embodiment of the present application provides a video retrieval method, the method comprising: Displaying a video editing interface, wherein the video editing interface is provided with a search entry; Receive target storyboard description information through the search entry; According to the storyboard description information, the corresponding target video slice is retrieved from a preset local material retrieval table; The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

[0006] Optionally, the video editing interface is further configured with a material frame and a video track; and the method further comprises: Displaying the retrieved target video slice in the material frame; In response to an add operation on the target video slice, the target video slice is added to the video track.

[0007] Optionally, the video editing interface is further configured with a preview window; and the method further comprises: In response to the material frame being selected, the target video slice is played in the preview window.

[0008] Optionally, the method further includes: In response to the search entry being selected, a plurality of character avatars are displayed.

[0009] Optionally, the method further includes: Creating a storyboard search page for a target video through a background configuration page, wherein the background configuration page is configured with an upload interface; Uploading the target video to the cloud through the upload interface, so that the cloud divides the target video into multiple video slices using the large language model, and outputs and returns video slice retrieval information for each video slice, wherein the video slice retrieval information includes slice description information and video understanding information of the corresponding video slice; The local material retrieval table with a structured index is generated according to the video slice retrieval information of each video slice.

[0010] Optionally, the cloud is also used to: Extracting frames from the target video to obtain multiple video frames; Comparing the multiple video frames with corresponding frames of an existing source video in the cloud, where the existing source video in the cloud is an existing video corresponding to the target video; When the comparison result meets the preset requirements, it is determined to segment the target video and perform video understanding.

[0011] Optionally, the video editing interface is further configured with a text generation video control; the method further includes: In response to the text-generated video control being selected, displaying a Vincent video pop-up window; the Vincent video pop-up window is configured with a create control; Receive text content through the Wensheng video pop-up window and receive text description information through the text-generated video entry; In response to the creation control being selected, a target video corresponding to the text content is obtained and generated based on a Wensheng video model.

[0012] Another aspect of the embodiments of the present application provides a video retrieval device, the device comprising: A display module is used to display a video editing interface, wherein the video editing interface is provided with a search entry; A receiving module, configured to receive target storyboard description information through the search entry; A retrieval module is used to retrieve the corresponding target video slice from a preset local material retrieval table according to the storyboard description information; The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

[0013] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0014] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0015] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the above-mentioned method when executed by a processor.

[0016] The above technical solution employed in the embodiments of the present application may provide the following advantages: Because the local material retrieval table includes video retrieval information for each video slice based on the output of the large language model, the local material retrieval table contains video content understanding information for each video slice. Based on this, by entering the target storyboard description information in the search entry on the video editing page, the corresponding target video slice can be retrieved from the pre-configured local material retrieval table, thereby achieving efficient and accurate retrieval of the desired video slice from the local video material. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0018] Figure 1 The following schematically shows an operating environment diagram of the video retrieval method according to the first embodiment of the present application; Figure 2 The flowchart of the video retrieval method according to the first embodiment of the present application is schematically shown; Figure 3 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 4 The following schematically shows a new flow chart of the video retrieval method according to the first embodiment of the present application; Figure 5 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 6 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 7The following schematically shows a new flow chart of the video retrieval method according to the first embodiment of the present application; Figure 8 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 9 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 10 The following schematically shows a new flow chart of the video retrieval method according to the first embodiment of the present application; Figure 11 The following schematically shows a new flow chart of the video retrieval method according to the first embodiment of the present application; Figure 12 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 13 A schematic diagram illustrating an example of the application of the video retrieval method according to the first embodiment of the present application is shown; Figure 14 A block diagram schematically shows a video retrieval device according to the second embodiment of the present application; and Figure 15 The following schematically shows a hardware architecture diagram of a computer device according to the third embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of this application more clear, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] It should be noted that the descriptions of "first", "second", etc. in the embodiments of the present application are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0021] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed. They are only used to facilitate the description of this application and to distinguish each step. Therefore, they cannot be understood as limitations on this application.

[0022] First, an explanation of the terms involved in this application is provided: The Large Language Model (LLM) is an AI model trained on large-scale datasets, boasting powerful language understanding and generation capabilities. By learning patterns and regularities from massive amounts of text data, it can accurately analyze input text and generate responses. In terms of video understanding, the LLM can be combined with a visual encoder to convert video frames into textual information, understanding the content, events, and semantics of the video. It can also infer, summarize, and answer questions based on the video content, achieving in-depth understanding and analysis of the video.

[0023] This embodiment of the present application provides a video retrieval technology solution. In this technology solution, a local (PC) video editor can accurately retrieve the required video slices from the local video material. See below for details.

[0024] For ease of understanding, an exemplary operating environment is provided below.

[0025] like Figure 1 As shown, the operating environment diagram includes: a service platform 2, one or more clients 4, and a network 6.

[0026] The service platform 2 and one or more clients 4 may be coupled via a network 6 to enable information transmission interaction.

[0027] The service platform 2, the client 4, and the network 6 are respectively described in detail below.

[0028] The service platform 2 can be comprised of a single or multiple computing devices. The one or more computing devices may include virtualized computing instances. Virtualized computing instances may include virtual machines, such as emulations of computer systems, operating systems, servers, and the like. A computing device may load a virtual machine based on a virtual image and / or other data defining the specific software (e.g., operating system, specialized application, server) used for the emulation. As demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device. The service platform 2 may run one or more services or software applications that enable execution of the methods described herein. The service platform 2 may also provide other services or software applications, which may include both non-virtualized and virtualized environments. In some embodiments, these services may be provided as web-based or cloud services, for example, provided to users of client 4 under a software-as-a-service (SaaS) model.

[0029] The service platform 2 may include one or more components that implement the functions performed by the service platform 2. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. In some embodiments, the service platform 2 may provide services such as video storage and video publishing, such as providing a video publishing service to the client. In other embodiments, users operating the client 4 may sequentially utilize one or more client applications to interact with the service platform 2 to utilize the services provided by these components.

[0030] Clients 4 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service kiosks, service robots, gaming systems, thin clients, various messaging devices, sensors, or other electronic devices. These computer devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as Google ChromeOS); or various mobile operating systems, such as Microsoft Windows, Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants, etc. Wearable devices may include head-mounted displays (such as smart glasses), etc. Gaming systems may include various handheld gaming devices and internet-enabled gaming devices. Client devices are capable of executing a variety of different applications, such as various internet-related applications, communication applications (such as email applications), and short message service (SMS) applications, and may utilize various communication protocols.

[0031] Based on the operating system described above, the client 4 may also be installed with one or more application programs, such as a video editing tool. The client 4 may use the video editing tool to obtain local video materials and then edit the video materials. After editing is completed, the client 4 may upload the edited video to the service platform 2 so that the service platform 2 can store and publish the edited video.

[0032] Client 4 may include an input / output interface. The input interface may include a touchpad, touch screen, mouse, keyboard, or other sensor elements. The input interface may be configured to receive user commands, which may cause client 4 to perform various operations, such as editing a video. The output interface is used to output information to the user, such as display information.

[0033] A network 6 may be used as a transmission medium between the service platform 2 and the client 4. The network 6 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network may include physical links, such as coaxial cable links, twisted-pair cable links, optical fiber links, or combinations thereof, or wireless links, such as cellular links, satellite links, and Wi-Fi links.

[0034] It should be noted that the above devices are exemplary, and the number and type of devices can be adjusted in different scenarios or according to different needs.

[0035] The following describes the technical solutions of the present application through multiple embodiments, taking the client as the execution subject. It should be noted that these embodiments can be implemented in many different forms and should not be construed as being limited to the embodiments described here.

[0036] Example 1 Figure 2 The flowchart of the video retrieval method according to the first embodiment of the present application is schematically shown.

[0037] like Figure 2 As shown, the video retrieval method may include steps S200 to S204, wherein: Step S200: displaying a video editing interface, wherein a search entry is configured on the video editing and decoding interface.

[0038] Step S202: Receive target storyboard description information through the search entry.

[0039] Step S204: According to the storyboard description information, the corresponding target video slice is retrieved from a preset local material retrieval table.

[0040] The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

[0041] In this embodiment, the video retrieval information for each video slice is output through a large language model and stored in a local material retrieval table. This allows the video editor to receive the target shot description (descriptive information of the desired video slice, such as "Help me find a shot of fire") input by the user through the search portal displayed on the video editing page. Based on this target shot description (or the keywords within it or keywords derived from semantic understanding), the corresponding target video slice is retrieved from the local material retrieval table pre-configured on the client, thereby efficiently and accurately retrieving the desired video slice from the local video material.

[0042] The following combination Figure 2, each step in steps S200~S204 and other optional steps are described in detail.

[0043] Step S200 , displaying the video editing interface, wherein the video editing interface is configured with a search entry.

[0044] The video editing interface is a visual operation area for performing various editing operations on the video. The search portal is a functional module set in the video editing interface. According to actual needs, it can be presented in the form of an input box, button, etc. Through this search portal, the user can enter relevant information to initiate a search operation. In some embodiments, the local material in the client is searched for video slices through the search portal. In some embodiments, Figure 3 In the video editing interface shown, the search entry is presented in the form of an input box.

[0045] Step S202 , receiving target storyboard description information through the search entry.

[0046] The target storyboard description information may be a descriptive content input by the user in order to find a video clip that meets a specific requirement. Figure 3 As shown in the example, you can enter the target storyboard description information "Help me find a shot of a volcano" in the input box, and then search based on this description to display the retrieved relevant content. In addition, you can also select the content to be displayed according to actual needs, such as all, video, audio, and pictures.

[0047] Step S204 According to the storyboard description information, the corresponding target video slice is retrieved from the preset local material retrieval table.

[0048] The target storyboard description information can be analyzed and processed, such as natural language understanding, keyword extraction and other operations, to increase the retrieval accuracy. The local material retrieval table is used to store relevant information of local materials, such as video retrieval information of each video slice. The video retrieval information can be stored in a structured manner, for example, it can include the physical storage location of each video slice (such as file path and database identifier, etc.), character portrait, scene number, end time, character name, time, scene, action, emotion, lens, key frame, subtitle field. Figure 3 As shown, according to the target storyboard description information of "help me find a shot of a volcano", the relevant pictures are retrieved and the corresponding timestamps are displayed, and the lines corresponding to the pictures are also displayed.

[0049] In an optional embodiment, the video editing interface is further configured with a material frame and a video track; Figure 4 As shown, the method further includes: Step S400: Display the retrieved target video slice in the material frame.

[0050] Step S402: In response to an adding operation on the target video slice, the target video slice is added to the video track.

[0051] In the video editing interface, the material box is an area for displaying and managing search results (one or more video slices). The material box can be an independent window, panel, or a specific display area. The adding operation can take many forms, such as selecting the target video slice in the material box, clicking the plus button, or directly dragging the selected target video slice from the material list to the video track area. The video track is a virtual track used to organize and arrange video clips. The video track can be regarded as a timeline, on which users can arrange, edit, and combine different video clips as needed. Each video track can independently carry one or more video clips, and video editing and special effects production can be achieved by adjusting parameters such as the position, duration, and sequence of the video clips on the track. In some embodiments, such as Figure 5 As shown, the clip frame is displayed in three dimensions: picture, character, and line. After adding the target video slice to the video track, a thumbnail of the corresponding time can be displayed on the video track, and the corresponding audio chart can be displayed below the thumbnail.

[0052] In this embodiment, the retrieved target video slice is displayed in the material frame so that the user can intuitively see the screen content of the target video slice. By adding the target video slice to the video track, the user can quickly integrate the selected target video slice to edit the video, thereby improving the efficiency of video editing.

[0053] In an optional embodiment, the video editing interface is further configured with a preview window. The method further includes: In response to the material frame being selected, the target video slice is played in the preview window.

[0054] The preview window is used to provide a display area for users to view video content in real time. The selection of a material frame indicates that the target video slice corresponding to the material frame can be further viewed or operated, and then the target slice video can be played in the preview window. In some embodiments, Figure 3 As shown in the following figure, when a clip frame is selected, the corresponding target video slice will be displayed in the preview window. You can control the playback of the target video slice in the preview window, such as rewind, fast forward, pause, and play. You can also select a video track to preview the video in the track.

[0055] In this embodiment, the target slice video is played in the preview window so that the searched video slice can be quickly confirmed to meet one's needs during video editing. If it does not meet the needs, it can be retrieved again in time.

[0056] Regarding the search entry, in an optional embodiment, the method further includes: in response to the search entry being selected, displaying multiple character avatars.

[0057] The multiple character portraits can be identified by a large language model. By displaying multiple character portraits, the target video slice can be retrieved from the local material according to the character portraits. In some embodiments, a search record can also be displayed so that when retrieving the same or similar video slices, the search record can be directly used for retrieval. In some embodiments, Figure 6 As shown, the selected search entry displays the character portrait. If no character portrait is identified, no portrait is displayed or a preset default portrait is displayed. In this embodiment, by selecting multiple character portraits to be displayed to retrieve corresponding video slices, the search time is reduced and the user experience is improved.

[0058] Regarding the local material retrieval table, in an optional embodiment, as Figure 7 As shown, the method further includes: Step S700: creating a storyboard search page for the target video through a background configuration page, wherein the background configuration page is configured with an upload interface.

[0059] Step S702: Upload the target video to the cloud through the upload interface so that the cloud divides the target video into multiple video slices through the large language model, and outputs and returns video slice retrieval information of each video slice, wherein the video slice retrieval information includes slice description information and video understanding information of the corresponding video slice.

[0060] Step S704: Generate the local material retrieval table with a structured index according to the video slice retrieval information of each video slice.

[0061] The background configuration page can be used to manage the video slice retrieval information of each local material. The storyboard retrieval page can be used to manage the video slice retrieval information of the specified local material. The cloud refers to a remote service platform based on cloud computing technology that provides huge computing resources, storage resources, and software services to users through the network. Since the cloud has more computing resources, video segmentation through the cloud can improve the efficiency and quality of video segmentation. Based on the large language model, the target video can be segmented according to the video speech semantics or scene changes. Slice description information and video understanding information can be generated according to each video slice. Video understanding information can be obtained through understanding the video language semantics or analyzing the video picture content. Its essence is the understanding of the video content. Video understanding information can be generated by a large language model. In some embodiments, the slice description information may include the slice number and the timestamp range corresponding to the video slice in the target video. In some embodiments, such as Figure 8 The background configuration page shown in the figure shows the auto-increment ID, unique ID, file name, file type, operation and creation time. Select the relevant name under the file name list to enter the corresponding storyboard search page. Figure 8 In the file name list, enter the specified name Figure 9 The storyboard search page shown here displays storyboard information including video number, shot number, start, end, keyframe, character, time, scene, action, emotion, shot, and dialogue. You can upload the corresponding episode video by clicking the Upload Video button for content verification.

[0062] In this embodiment, a local material retrieval table with a structured index is generated according to the video slice retrieval information of each video slice, so that the user can more accurately retrieve the required video slice through the local material retrieval table.

[0063] Regarding video content verification, in an optional embodiment, as Figure 10 As shown, the cloud is also used for: Step 1000: extract frames from the target video to obtain multiple video frames.

[0064] Step 1002: Compare the multiple video frames with corresponding frames of an existing source video stored in the cloud, where the existing source video in the cloud is an existing video corresponding to the target video. Step 1004 : If the comparison result meets the preset requirements, determine to segment the target video and perform video understanding.

[0065] Frame extraction refers to selecting some frames from a continuously played video according to certain time intervals or rules. The corresponding frames can be determined based on the correspondence between the target video and the existing material video on the cloud on the timeline. If the comparison result does not meet the preset requirements, a pop-up window may be displayed indicating that the video content verification has failed. In some embodiments, the compliance with the standards can be determined by comparing the timeline matching of multiple video frames with the corresponding frames of the existing material video on the cloud, and identifying whether the matching rate reaches 98%. In some embodiments, the target video can be partially cut according to the video content of the existing material video on the cloud to improve the efficiency of video content verification, such as cutting the beginning and end of the video accordingly.

[0066] In this embodiment, by comparing multiple video frames with corresponding frames of existing material videos in the cloud, content verification is performed on the target video and the existing material videos in the cloud, thereby effectively verifying whether the video content sequence of the target video is consistent with that of the existing material videos in the cloud.

[0067] It should be noted that the above video slices can be obtained based on local material segmentation, or can be generated in real time.

[0068] In an optional embodiment, the video editing interface is further configured with a text generation video control. Figure 11 As shown, the method further includes: Step 1100: In response to the text-generated video control being selected, a Vincent video pop-up window is displayed; the Vincent video pop-up window is configured with a creation control.

[0069] Step 1102: receiving text content through the text-generated video pop-up window and receiving text description information through the text-generated video entry.

[0070] Step 1104 : In response to the creation control being selected, a target video corresponding to the text content and generated based on a Wensheng video model is obtained.

[0071] The text-generated video control is an interactive element on the video editing interface and can be presented in the form of buttons, icons, etc. The text-generated video model is an artificial intelligence model that can automatically generate videos based on the input text content. Based on the text description information, the cloud-based materials can be associated, and then the associated materials can be added to the video track to generate the target video corresponding to the text content. In some embodiments, the selected Figure 5 The Vincent video control in the display is as follows Figure 12 The Wensheng Video pop-up window shown above appears. You can enter text content in this pop-up window and select the Create control to generate a target video corresponding to the text content based on the cloud-based assets. You can also select local assets to generate a target video corresponding to the text content based on the local assets.

[0072] In this embodiment, the demand for quickly generating videos from text content is met, thereby improving user experience.

[0073] In order to make this application easier to understand, an exemplary application is provided below.

[0074] Taking the retrieval of local video files as local materials as an example, the operation process is as follows: (1) Local video file preprocessing stage: S1. The client displays the video editing interface and imports local materials (such as Figure 13 to upload local materials to the cloud.

[0075] S2: The cloud checks whether there is a local search table. If not, it proceeds to S3; if so, it proceeds to S4.

[0076] S3,The cloud uses a large language model to perform multi-level segmentation of the local video material and returns the segment retrieval information of each segment to the client.

[0077] For example, you can split the video into multiple outlines (first-level video slices), which can be further divided into multiple sub-outlines (second-level video slices), which can then be further divided into multiple strips (third-level video slices). Typically, third-level video slices are shot-level.

[0078] Slice search information can be structured and stored in a local search table on the client. This information can also be synchronized to the cloud and shared with other users. The local search table can be configured with multiple dimensions, such as slice number, timestamp, and character.

[0079] S4. The cloud sends the local search table to the client.

[0080] (2) Retrieval stage: S5. Receive target storyboard description information input by the user through the search entry configured in the video editing interface.

[0081] like Figure 3 As shown, enter "Help me find a shot of a volcano" into the search portal.

[0082] S6. Understand the target storyboard description information through the large language model and output search keywords.

[0083] S7. According to the search keyword, the corresponding target video slice is retrieved from the preset local material search table.

[0084] like Figure 3 As shown, the retrieved video slice is displayed, and the lines associated with the target storyboard description information are displayed.

[0085] S8. In response to the adding operation on the target video slice, add the target video slice to the video track.

[0086] like Figure 5 As shown in the figure, add the target video slice to the video track.

[0087] In this exemplary application, it is achieved to accurately retrieve the required video slices from the local video material.

[0088] Example 2 Figure 14 The block diagram of the video retrieval device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Figure 14 As shown, the apparatus 1400 may include: a display module 1410, a receiving module 1420, and a retrieval module 1430, wherein: Display module 1410, used to display a video editing interface, wherein the video editing interface is configured with a search entry; Receiving module 1420, configured to receive target storyboard description information through the search entry; A retrieval module 1430 is configured to retrieve a corresponding target video slice from a preset local material retrieval table according to the storyboard description information; The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

[0089] As an optional embodiment, the video editing interface is further configured with a material frame and a video track; the retrieval module 1430 is further configured to: Displaying the retrieved target video slice in the material frame; In response to an add operation on the target video slice, the target video slice is added to the video track.

[0090] As an optional embodiment, the video editing interface is further configured with a preview window; the retrieval module 1330 is further configured to: In response to the material frame being selected, the target video slice is played in the preview window.

[0091] As an optional embodiment, the retrieval module 1430 is further configured to: In response to the search entry being selected, a plurality of character avatars are displayed.

[0092] As an optional embodiment, the apparatus 1400 further includes a local search management module, and the local search management module further includes: Creating a storyboard search page for a target video through a background configuration page, wherein the background configuration page is configured with an upload interface; Uploading the target video to the cloud through the upload interface, so that the cloud divides the target video into multiple video slices using the large language model, and outputs and returns video slice retrieval information for each video slice, wherein the video slice retrieval information includes slice description information and video understanding information of the corresponding video slice; The local material retrieval table with a structured index is generated according to the video slice retrieval information of each video slice.

[0093] As an optional embodiment, the cloud is further used for: Extracting frames from the target video to obtain multiple video frames; Comparing the multiple video frames with corresponding frames of an existing source video in the cloud, where the existing source video in the cloud is an existing video corresponding to the target video; When the comparison result meets the preset requirements, it is determined to segment the target video and perform video understanding.

[0094] As an optional embodiment, the apparatus 1400 further includes a Vincent video module, wherein the Vincent video module is configured to: In response to the text-generated video control being selected, displaying a Vincent video pop-up window; the Vincent video pop-up window is configured with a create control; Receive text content through the Wensheng video pop-up window and receive text description information through the text-generated video entry; In response to the creation control being selected, a target video corresponding to the text content is obtained and generated based on a Wensheng video model.

[0095] Example 3 Figure 15 The following schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing a video retrieval method according to the third embodiment of the present application. In some embodiments, the computer device 10000 can be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. Figure 15 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other via a system bus. Memory 10010 includes at least one type of computer-readable storage medium, including flash memory, a hard disk, a multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, and the like. In some embodiments, memory 10010 may be an internal storage module of computer device 10000, such as a hard disk or memory of computer device 10000. In other embodiments, memory 10010 may also be an external storage device of computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, and the like equipped on computer device 10000. Of course, memory 10010 may also include both internal storage modules and external storage devices of computer device 10000. In this embodiment, the memory 10010 is generally used to store an operating system and various application software installed on the computer device 10000, such as program codes of a video retrieval method, etc. In addition, the memory 10010 can also be used to temporarily store various data that has been output or is to be output.

[0096] In some embodiments, processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. Processor 10020 is typically used to control the overall operation of computer device 10000, such as performing control and processing related to data exchange or communication with computer device 10000. In this embodiment, processor 10020 is used to execute program code stored in memory 10010 or process data.

[0097] Network interface 10030 may include a wireless network interface or a wired network interface. Network interface 10030 is typically used to establish a communication link between computer device 10000 and other computer devices. For example, network interface 10030 is used to connect computer device 10000 to an external terminal via a network, establishing a data transmission channel and a communication link between computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, or Wi-Fi.

[0098] It should be pointed out that Figure 15 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementing all of the shown components is not a requirement, and more or fewer components may alternatively be implemented.

[0099] In this embodiment, the video retrieval method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0100] Example 4 An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the video retrieval method in the embodiment are implemented.

[0101] In this embodiment, computer-readable storage media include flash memory, hard disks, multimedia cards, card-type memories (e.g., SD or DX memories), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disks, optical disks, and the like. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the computer device's hard disk or memory. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, and the like. Of course, the computer-readable storage medium may also include both the internal storage unit and external storage devices of the computer device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the computer device, such as the program code of the video retrieval method described in the embodiment. Furthermore, the computer-readable storage medium may also be used to temporarily store various types of data that has been output or is about to be output.

[0102] Example 5 An embodiment of the present application further provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0103] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented using general-purpose computer devices. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Alternatively, they can be implemented using program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0104] It should be noted that the above are only preferred embodiments of the present application and do not limit the scope of patent protection of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of the present application.

Claims

1. A video retrieval method, characterized in that: The method comprises: Displaying a video editing interface, wherein the video editing interface is provided with a search entry; Receive target storyboard description information through the search entry; According to the storyboard description information, the corresponding target video slice is retrieved from a preset local material retrieval table; The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

2. The method according to claim 1, characterized in that The video editing interface is further configured with a material frame and a video track; the method further comprises: Displaying the retrieved target video slice in the material frame; In response to an add operation on the target video slice, the target video slice is added to the video track.

3. The method according to claim 2, characterized in that The video editing interface is further configured with a preview window; the method further comprises: In response to the material frame being selected, the target video slice is played in the preview window.

4. The method according to claim 1, wherein The method further comprises: In response to the search entry being selected, a plurality of character avatars are displayed.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Creating a storyboard search page for a target video through a background configuration page, wherein the background configuration page is configured with an upload interface; Uploading the target video to the cloud through the upload interface, so that the cloud divides the target video into multiple video slices using the large language model, and outputs and returns video slice retrieval information for each video slice, wherein the video slice retrieval information includes slice description information and video understanding information of the corresponding video slice; The local material retrieval table with a structured index is generated according to the video slice retrieval information of each video slice.

6. The method according to claim 5, characterized in that The cloud is also used to: Extracting frames from the target video to obtain multiple video frames; Comparing the multiple video frames with corresponding frames of an existing material video in the cloud, where the existing material video in the cloud is an existing video corresponding to the target video; When the comparison result meets the preset requirements, it is determined to segment the target video and perform video understanding.

7. The method according to claim 5, characterized in that The video editing interface is further configured with a text generation video control; the method further comprises: In response to the text-generated video control being selected, displaying a Vincent video pop-up window; the Vincent video pop-up window is configured with a create control; Receive text content through the Wensheng video pop-up window and receive text description information through the text-generated video entry; In response to the creation control being selected, a target video corresponding to the text content is obtained and generated based on a Wensheng video model.

8. A video retrieval device, characterized in that: The device comprises: A display module is used to display a video editing interface, wherein the video editing interface is provided with a search entry; A receiving module, configured to receive target storyboard description information through the search entry; A retrieval module is used to retrieve the corresponding target video slice from a preset local material retrieval table according to the storyboard description information; The local material retrieval table includes video retrieval information of each video slice output based on the large language model.

9. A computer device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.