Generative video retrieval method and device oriented to abstract text, equipment and medium
By searching for videos matching abstract text from the video material library, and determining the target video using vector encoding and semantic similarity calculation, the problem of low quality of abstract text generation video is solved, and video generation with higher quality and rich content is achieved.
Patent Information
- Application Number
- CN202510414162.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-08
AI Technical Summary
When facing abstract descriptive text, the generated video picture quality is not high, the content is not rich, and it is difficult to show rich details and vivid scenes.
By obtaining abstract text, several first videos are generated, a second video matching the first video is searched from the preset video material library, and the second video is matched with the abstract text, and the target video is determined using vector encoding and semantic similarity calculation.
Improve the quality of video images generated based on abstract text, and enhance the richness and vividness of video content.
Smart Images

Figure CN120448583A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross-modal technical field of the intersection of natural language processing and computer vision, and in particular relates to a generative video retrieval method, apparatus, device and storage medium for abstract text. Background Art
[0002] Currently, mainstream text-to-video generation models have demonstrated amazing generation effects when dealing with specific descriptive texts, such as Figure 1 Figure (a) in the figure is a video image generated based on the specific text "A fashionable woman walks on the streets of Tokyo, which are full of warm neon lights and animated city signs. She wears a black leather jacket, a red long skirt and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is wet and reflective, forming a mirror effect with the colorful lights. Many pedestrians walk around." and Figure 1 Figure (b) is a video image generated based on the specific text "Inside the gym, a woman in sportswear is running on a treadmill. Side angle. Realistic, indoor lighting, professional". Figure 1 As can be seen from (a) and (b), the generated video image quality is high, with rich details and excellent dynamic performance.
[0003] However, when faced with abstract descriptive texts, the generation effects of these models are not satisfactory and are still far from the acceptable standards. For example, the video images generated based on the abstract texts "Embrace sensuality and confidence through intimate expression", "Natural beauty in motion", and "Explore a relaxed and comfortable modern digital lifestyle" are as follows: Figure 1 As shown in (c), (d), and (e) in the figure, from the perspective of video semantics and scenes, these models still have some highlights in generating abstract texts, which are worth affirming. For example, Figure 1 In Figure (c), the model uses a female image; Figure 1 In the figure (d), reed elements are used; Figure 1 Figure (e) shows a laptop computer. These are all consistent with the abstract text to a certain extent, but these highlights cannot cover up its shortcomings in the overall generation effect, which are specifically manifested in two aspects: first, the picture quality is significantly reduced, with problems such as blur and noise; second, the details and dynamic presentation in the video are greatly reduced, making it difficult to show rich content and vivid scenes.
[0004] Abstract descriptive text is ubiquitous in everyday life. This is because humans have the ability to think abstractly and often use abstract language to express their thoughts and feelings. Therefore, video generation based on abstract text should occupy a key position in the field of video generation. Without a good solution for video generation based on abstract text, the challenge of generating videos from text cannot be truly overcome. Summary of the Invention
[0005] The purpose of the present invention is to provide a generative video retrieval method, device, equipment and medium for abstract text, aiming to solve the problem that videos generated based on abstract text are not vivid, have poor content and low picture quality due to the existing technology.
[0006] In one aspect, the present invention provides a generative video retrieval method for abstract text, the method comprising the following steps:
[0007] Obtain abstract text for video generation;
[0008] generating a plurality of first videos according to the abstract text;
[0009] Searching for second videos that match the first videos respectively from a preset video material library;
[0010] Each of the second videos is matched with the abstract text, and a target video generated from the abstract text is determined according to the matching result.
[0011] Preferably, the step of searching for second videos that match each of the first videos from a preset video material library includes:
[0012] Performing vector encoding on all material videos in the video material library to obtain a corresponding vector database;
[0013] Performing vector encoding on all the first videos to obtain corresponding first video vectors;
[0014] Each first video vector is compared with all material vectors in the vector database for similarity, and the second video is determined based on the comparison result.
[0015] Preferably, the step of determining the second video according to the comparison result includes:
[0016] According to the comparison result, several material vectors having the highest similarity to the first video vector are selected from the vector database, and the material videos corresponding to the material vectors are determined as the second video.
[0017] Preferably, the step of matching each of the second videos with the abstract text and determining a target video generated from the abstract text according to the matching result includes:
[0018] Generating corresponding descriptive text for each second video;
[0019] Calculating the matching degree between each of the descriptive texts and the abstract text according to a preset matching degree calculation formula;
[0020] The target video is determined according to the calculated matching degree.
[0021] Preferably, the step of generating corresponding descriptive text for each second video includes:
[0022] Generating corresponding image description text for each frame of image in the second video;
[0023] The descriptive text corresponding to the second video is generated according to the preset prompt text and the image description text.
[0024] Preferably, the step of calculating the matching degree between each of the descriptive texts and the abstract text according to a preset matching degree calculation formula includes:
[0025] Calculating the semantic similarity between the descriptive text and the abstract text to obtain a first similarity;
[0026] Calculating semantic similarity between the abstract text and a second video corresponding to the descriptive text to obtain a second similarity;
[0027] The matching degree is calculated using the matching degree calculation formula according to the first similarity, the second similarity, and the video quality score of the second video corresponding to the descriptive text.
[0028] In another aspect, the present invention provides a generative video retrieval device for abstract text, the device comprising:
[0029] An abstract text acquisition unit, used to acquire abstract text for video generation;
[0030] A first video generating unit, configured to generate a plurality of first videos according to the abstract text;
[0031] A second video search unit, configured to search a preset video material library for second videos that respectively match the first videos;
[0032] The target video determining unit is configured to match each of the second videos with the abstract text, and determine a target video generated from the abstract text according to the matching result.
[0033] Preferably, the second video search unit includes:
[0034] A first vector encoding unit is used to perform vector encoding on all material videos in the video material library to obtain a corresponding vector database;
[0035] A second vector encoding unit, configured to perform vector encoding on all the first videos to obtain corresponding first video vectors;
[0036] The second video determining unit is configured to compare the similarity between each of the first video vectors and all the material vectors in the vector database, and determine the second video according to the comparison result.
[0037] On the other hand, the present invention also provides a computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned generative video retrieval method for abstract text are implemented.
[0038] On the other hand, the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the above-mentioned generative video retrieval method for abstract text.
[0039] The present invention obtains abstract text for video generation, generates several first videos based on the abstract text, searches for second videos that match each first video from a preset video material library, matches each second video with the abstract text, and determines the target video generated by the abstract text based on the matching result, thereby improving the picture quality of the video generated based on the abstract text and improving the richness and vividness of the video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of a video image generated by a traditional text-to-video generation model based on specific description text or abstract description text;
[0041] Figure 2 This is a flowchart of an implementation of a generative video retrieval method for abstract text provided in the first embodiment of the present invention;
[0042] Figure 3 Schematic diagram of the structure of a generative video retrieval device for abstract text provided in the second embodiment of the present invention;
[0043] Figure 4 This is a schematic diagram of the preferred structure of a generative video retrieval device for abstract text provided in the second embodiment of the present invention;
[0044] Figure 5 It is a structural diagram of the computing device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0046] It should be understood that the terms "include" and "have" and any variations thereof in the present invention are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components is not necessarily limited to those components explicitly listed, but may include other components not explicitly listed or inherent to these products or devices.
[0047] In this disclosure, the term "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0048] Unless otherwise specified, the term "plurality" in the present invention refers to two or more, and other quantifiers are similar to them.
[0049] The terms "first," "second," "third," etc., used in the present invention are used to distinguish between similar or similar objects or entities and are not necessarily intended to limit a particular order or precedence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable where appropriate, e.g., the embodiments of the present disclosure can be implemented in an order other than that shown or described in the drawings or descriptions.
[0050] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0051] The following describes the specific implementation of the present invention in detail with reference to specific embodiments:
[0052] Example 1:
[0053] Figure 2 The following illustrates an implementation process of a generative video retrieval method for abstract text provided in the first embodiment of the present invention. For ease of illustration, only the portion related to the embodiment of the present invention is shown, which is described in detail as follows:
[0054] In step S201 , abstract text for video generation is obtained.
[0055] Embodiments of the present invention are applicable to computing devices, such as personal computers and servers. In embodiments of the present invention, abstract text used for video generation refers to text that does not describe specific, perceptible objects or phenomena, but instead focuses on expressing abstract content such as concepts, ideas, emotions, principles, and theories. The source of abstract text can be a sentence or multiple paragraphs input by a user through a human-computer interaction interface (such as a text input box or voice input), or it can be generated by artificial intelligence (AI). The source of abstract text is not specifically limited herein.
[0056] In a feasible embodiment, after obtaining the abstract text used for video generation, the obtained abstract text is detected for sensitive words. Specifically, first, a list of sensitive words covering various types of sensitive words is prepared in advance as a basis for detection. Then, a vocabulary scan is performed on the abstract text description to determine whether the abstract text contains sensitive words in the sensitive word list. If sensitive words are detected, the user is prompted that the abstract text is illegal, and all sensitive words in the abstract text are displayed to guide the user to modify the text description. If no sensitive words are detected, the relevant steps of subsequent video generation can be continued, thereby improving the legality of the subsequent video generated based on the abstract text.
[0057] In step S202, a plurality of first videos are generated according to the abstract text.
[0058] In an embodiment of the present invention, a plurality of currently popular and open-source text-to-video generation models are used to generate corresponding videos based on abstract texts. Specifically, a single or multiple text-to-video generation models (such as HunyuanVideo, CogVideoX, OpenSora, ModelScope, etc.) are used, and the abstract texts are respectively input into each text-to-video generation model. Each text-to-video generation model generates a corresponding video based on the abstract text received. For example, when there are N text-to-video generation models, each text-to-video generation model will generate a corresponding video, and finally N videos are obtained, which are respectively recorded as V1, V2, ..., V N ,For the convenience of description and distinction, the video generated by each ,text-to-video generation model is referred to as the first video.
[0059] In step S203, a second video that matches each first video is searched from a preset video material library.
[0060] In an embodiment of the present invention, a second video that matches each first video is searched from a preset video material library (such as Visual China, Pexels Videos, etc.).
[0061] In a feasible embodiment, searching for second videos that match each first video from a preset video material library is achieved through the following steps:
[0062] (S203.1) Perform vector encoding on all material videos in the video material library to obtain a corresponding vector database;
[0063] In an embodiment of the present invention, all material videos in a video material library are encoded into vectors to form a vector database of the video. Specifically, a video encoding tool is used to perform vector encoding on all material videos in the video material library. The video encoding tool can be derived from an open source tool library (such as OpenCV, PyTorchVideo, etc.) or from a published related paper (such as "VVS: Video-to-Video Retrieval with Irrelevant Frame Suppression"). The video encoding tool is not specifically limited here.
[0064] (S203.2) Perform vector encoding on all first videos to obtain corresponding first video vectors;
[0065] In the embodiment of the present invention, the same video encoding tool as above is used to encode all first videos into vectors to obtain corresponding first video vectors, which are recorded as Vec1, Vec2, ..., Vec N .
[0066] (S203.3) Compare the similarity of each first video vector with all material vectors in the vector database, and determine the second video based on the comparison results.
[0067] In the embodiment of the present invention, each first video vector is compared with all material vectors in the vector database for similarity. Specifically, each first video vector Vec is calculated. i (1≤i≤N) and the cosine similarity between each material vector in the vector database, and determining the comparison result of the similarity comparison based on the calculated cosine similarities. Then, the second video is determined according to the comparison result.
[0068] When determining the second video based on the comparison result, preferably, based on the comparison result, several material vectors with the highest similarity to the first video vector are selected from the vector database, and the material videos corresponding to the material vectors are determined as the second video.
[0069] In this embodiment of the present invention, for each first video vector Vec i(1≤i≤N), according to the comparison results, select several (denoted by M) material vectors with the highest similarity to the first video vector from the vector database, and use the material videos corresponding to these M material vectors as the videos matching the first video corresponding to the first video vector, that is, each first video will get M matching material videos. Finally, N first videos correspond to M·N material videos. For the convenience of description and distinction, the material video searched from the video material library that matches the first video is called the second video, denoted by V′ 1j , V′ 2j ,...,V′ Nj , where 1≤j≤M.
[0070] In a feasible embodiment, M is set to 3, that is, each first video V i (1≤i≤N) will get 3 matching material videos V′ i1 , V′ i2 , V′ i3 .
[0071] The second video is obtained through the above steps S203.1 to S203.3, thereby improving the generation effect of subsequent videos.
[0072] In step S204, each second video is matched with the abstract text, and a target video generated from the abstract text is determined according to the matching result.
[0073] In the embodiment of the present invention, a video that best matches the abstract text is selected from all second videos, and the video is used as the target video generated from the abstract text.
[0074] In a feasible embodiment, the target video is determined by the following steps:
[0075] (S204.1) Generate corresponding descriptive text for each second video;
[0076] In an embodiment of the present invention, a corresponding descriptive text is generated for each second video. The descriptive text is a text description of the main content, characteristics, theme, key elements or actions included in the video and other core information.
[0077] In a feasible embodiment, the descriptive text is generated by the following steps:
[0078] (S204.1.1) Generate corresponding image description text for each frame of the second video;
[0079] In the embodiment of the present invention, each second video V ij′(1≤i≤N, 1≤j≤M) is divided into a series of images according to the frame, and then an image description model is used to generate a text describing the image content for each image, namely, the image description text. Among them, the image description model can use Tag2Text or other image description models.
[0080] (S204.1.2) Generate descriptive text corresponding to the second video based on the preset prompt text and image description text.
[0081] In the embodiment of the present invention, the preset prompt text and the second video V ij All the corresponding image description texts are input into the Large Language Model (LLM), and the LLM summarizes all the image description texts into a paragraph (i.e., descriptive text) according to the prompt text. This paragraph is the specific description corresponding to the video. The large language model can use QWen or other large language models, which are not specifically limited here.
[0082] For example, the prompt text used is "Please summarize the key events in the video based on the frame descriptions provided below. Your summary should convey the overall storyline or message of the video through a coherent narrative, while condensing and synthesizing the most important details of each frame. Focus on describing the main characters, scenes, actions, plot points, and any other elements in all frames that advance the central storyline. You should not directly copy or repeat the description of the frame, but should reorganize and integrate key information to summarize the entire video in a concise and descriptive manner." This prompt text can be optimized according to user habits.
[0083] The descriptive text is generated through the above steps S204.1.1 to S204.1.2, so that the descriptive text can express the content of the second video more accurately and richly.
[0084] (S204.2) Calculating the matching degree between each descriptive text and the abstract text according to a preset matching degree calculation formula;
[0085] In this embodiment of the present invention, the calculation of the matching degree is achieved through the following steps:
[0086] (S204.2.1) calculating the semantic similarity between the descriptive text and the abstract text to obtain a first similarity;
[0087] In an embodiment of the present invention, the semantic similarity between the descriptive text and the abstract text is calculated. Specifically, the descriptive text and the abstract text are encoded using the SimCSE model, and then the cosine similarity of the encoded texts is calculated to obtain a first similarity, which is recorded as SentenceSim.
[0088] (S204.2.2) Calculating the semantic similarity of the second video corresponding to the abstract text and the descriptive text to obtain a second similarity;
[0089] In an embodiment of the present invention, the semantic similarity of the abstract text and the second video corresponding to the descriptive text is calculated. Specifically, the abstract text and the second video are encoded using the CLIP model, and then the cosine similarity of the encoded text and the second video is calculated to obtain a second similarity, which is recorded as CLIPSIM.
[0090] (S204.2.3) Calculate the matching degree using a matching degree calculation formula based on the first similarity, the second similarity, and the video quality score of the second video corresponding to the descriptive text.
[0091] In an embodiment of the present invention, the DOVER (Disentangled Objective Video Quality Evaluator) model is used to calculate the video quality score of the second video. DOVER is a model trained on a specific dataset that can comprehensively evaluate video quality from both aesthetic and technical perspectives. The resulting score is recorded as VQA (Video Quality Assessment), i.e., the video quality score. Here, SentenceSim and CLIPSIM together constitute the semantic similarity portion of the matching index, and VQA constitutes the video quality portion of the matching index. The weights of each component are controlled by weight factors α and β, respectively. In actual use, users can adjust the values of these weight factors according to their own needs. Finally, based on SentenceSim, CLIPSIM, and VQA, the matching degree Score of the second video corresponding to each descriptive text and the abstract text is calculated using the matching degree calculation formula Score = α(SentenceSim+CLIPSIM)+β*VQA.
[0092] Through the above steps S204.2.1 to S204.2.3, the matching degree is obtained by integrating multi-dimensional factors, thereby improving the accuracy of the semantic similarity judgment between the descriptive text and the abstract text, and further improving the accuracy of the subsequent screening of the second video that is highly consistent with the abstract text.
[0093] (S204.3) Determine the target video based on the calculated matching degree.
[0094] In the embodiment of the present invention, a larger value of the matching degree Score indicates that the second video matches the abstract text more closely. Based on this, the second video corresponding to the maximum value Score is determined as the target video.
[0095] The target video is determined through the above steps S204.1 to S204.3, thereby improving the degree of matching between the target video and the abstract text.
[0096] In an embodiment of the present invention, an abstract text for video generation is obtained, and a plurality of first videos are generated based on the abstract text. Second videos that match each of the first videos are searched from a preset video material library, and each second video is matched with the abstract text. Based on the matching result, a target video generated by the abstract text is determined, thereby improving the picture quality of the video generated based on the abstract text and improving the richness and vividness of the video content.
[0097] Example 2:
[0098] Figure 3 The structure of a generative video retrieval device for abstract text provided in the second embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, including:
[0099] An abstract text acquisition unit 31 is used to acquire abstract text for video generation;
[0100] A first video generating unit 32 is configured to generate a plurality of first videos according to the abstract text;
[0101] The second video search unit 33 is used to search for second videos that match the first videos from a preset video material library;
[0102] The target video determining unit 34 is configured to match each second video with the abstract text, and determine a target video generated from the abstract text according to the matching result.
[0103] Preferably, if Figure 4 As shown, the second video search unit 33 includes:
[0104] The first vector encoding unit 331 is used to perform vector encoding on all material videos in the video material library to obtain a corresponding vector database;
[0105] A second vector encoding unit 332 is configured to perform vector encoding on all first videos to obtain corresponding first video vectors;
[0106] The second video determining unit 333 is configured to compare the similarity between each first video vector and all material vectors in the vector database, and determine the second video according to the comparison result.
[0107] The target video determination unit 34 includes:
[0108] A text generation unit 341 is configured to generate corresponding descriptive text for each second video;
[0109] A matching degree calculation unit 342 is used to calculate the matching degree between each descriptive text and the abstract text according to a preset matching degree calculation formula;
[0110] The video determination subunit 343 is configured to determine a target video according to the calculated matching degree.
[0111] Still preferably, the second video determination unit 333 includes:
[0112] The video screening unit is used to select several material vectors with the highest similarity to the first video vector from the vector database according to the comparison result, and determine the material video corresponding to the material vector as the second video.
[0113] The text generation unit 341 includes:
[0114] An image text generating unit, configured to generate corresponding image description text for each frame of image in the second video;
[0115] The text generation subunit is used to generate a descriptive text corresponding to the second video according to the preset prompt text and image description text.
[0116] The matching degree calculation unit 342 includes:
[0117] A first similarity calculation unit is used to calculate the semantic similarity between the descriptive text and the abstract text to obtain a first similarity;
[0118] A second similarity calculation unit is used to calculate the semantic similarity of the second video corresponding to the abstract text and the descriptive text to obtain a second similarity;
[0119] The matching degree calculation subunit is used to calculate the matching degree using a matching degree calculation formula according to the first similarity, the second similarity and the video quality score of the second video corresponding to the descriptive text.
[0120] In the embodiment of the present invention, each unit of the generative video retrieval device for abstract text can be implemented by corresponding hardware or software units. Each unit can be an independent hardware or software unit, or can be integrated into a single hardware or software unit, without limiting the present invention. Specifically, the implementation of each unit can refer to the description of the aforementioned embodiment 1 and will not be repeated here.
[0121] Example 3:
[0122] Figure 5 The structure of a computing device provided by the third embodiment of the present invention is shown. For ease of description, only the parts related to the embodiment of the present invention are shown.
[0123] The computing device 5 of the embodiment of the present invention includes a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer program 52, the steps of the above-mentioned generative video retrieval method for abstract text are implemented, such as Figure 2 Alternatively, when the processor 50 executes the computer program 52, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 3 Function of the unit shown.
[0124] In an embodiment of the present invention, an abstract text for video generation is obtained, and a plurality of first videos are generated based on the abstract text. Second videos that match each of the first videos are searched from a preset video material library, and each second video is matched with the abstract text. Based on the matching result, a target video generated by the abstract text is determined, thereby improving the picture quality of the video generated based on the abstract text and improving the richness and vividness of the video content.
[0125] The computing device of the embodiment of the present invention can be a personal computer or a server. The steps implemented when the processor 50 in the computing device 5 executes the computer program 52 to implement the generative video retrieval method for abstract text can be referred to the description of the aforementioned method embodiment and will not be repeated here.
[0126] Example 4:
[0127] In an embodiment of the present invention, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned generative video retrieval method for abstract text are implemented. For example, Figure 2 Alternatively, when the computer program is executed by a processor, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 3 Function of the unit shown.
[0128] In an embodiment of the present invention, an abstract text for video generation is obtained, and a plurality of first videos are generated based on the abstract text. Second videos that match each of the first videos are searched from a preset video material library, and each second video is matched with the abstract text. Based on the matching result, a target video generated by the abstract text is determined, thereby improving the picture quality of the video generated based on the abstract text and improving the richness and vividness of the video content.
[0129] The computer-readable storage medium of the embodiment of the present invention may include any entity, device, or recording medium capable of carrying computer program code, for example, ROM / RAM, magnetic disk, optical disk, flash memory, or other memory.
[0130] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A generative video retrieval method for abstract text, characterized by: The method comprises the following steps: Obtain abstract text for video generation; generating a plurality of first videos according to the abstract text; Searching for second videos that match the first videos respectively from a preset video material library; Each of the second videos is matched with the abstract text, and a target video generated from the abstract text is determined according to the matching result.
2. The method according to claim 1, wherein The step of searching for second videos that match the first videos respectively from a preset video material library includes: Performing vector encoding on all material videos in the video material library to obtain a corresponding vector database; Performing vector encoding on all the first videos to obtain corresponding first video vectors; Each first video vector is compared with all material vectors in the vector database for similarity, and the second video is determined based on the comparison result.
3. The method according to claim 2, wherein The step of determining the second video according to the comparison result includes: According to the comparison result, several material vectors having the highest similarity to the first video vector are selected from the vector database, and the material videos corresponding to the material vectors are determined as the second video.
4. The method according to claim 1, wherein The step of matching each of the second videos with the abstract text and determining a target video generated from the abstract text according to the matching result includes: Generating corresponding descriptive text for each second video; Calculating the matching degree between each of the descriptive texts and the abstract text according to a preset matching degree calculation formula; The target video is determined according to the calculated matching degree.
5. The method according to claim 4, wherein The step of generating corresponding descriptive text for each second video includes: Generating corresponding image description text for each frame of image in the second video; The descriptive text corresponding to the second video is generated according to the preset prompt text and the image description text.
6. The method according to claim 4, wherein The step of calculating the matching degree between each of the descriptive texts and the abstract text according to a preset matching degree calculation formula includes: Calculating the semantic similarity between the descriptive text and the abstract text to obtain a first similarity; Calculating semantic similarity between the abstract text and a second video corresponding to the descriptive text to obtain a second similarity; The matching degree is calculated using the matching degree calculation formula according to the first similarity, the second similarity, and the video quality score of the second video corresponding to the descriptive text.
7. A generative video retrieval device for abstract text, characterized in that: The device comprises: An abstract text acquisition unit, used to acquire abstract text for video generation; A first video generating unit, configured to generate a plurality of first videos according to the abstract text; A second video search unit, configured to search a preset video material library for second videos that respectively match the first videos; The target video determining unit is configured to match each of the second videos with the abstract text, and determine a target video generated from the abstract text according to the matching result.
8. The device according to claim 7, characterized in that The second video search unit includes: A first vector encoding unit is used to perform vector encoding on all material videos in the video material library to obtain a corresponding vector database; A second vector encoding unit, configured to perform vector encoding on all the first videos to obtain corresponding first video vectors; The second video determining unit is configured to compare the similarity between each of the first video vectors and all the material vectors in the vector database, and determine the second video according to the comparison result.
9. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.