Video generation method, device, equipment, computer readable medium and program product
By generating a candidate event localization video set through a pre-trained multimodal large language model, and using target hit accuracy and text modal feature information to filter out accurate event localization videos, the problem of high computational complexity and insufficient accuracy in existing technologies is solved, and efficient and accurate video localization is achieved.
Patent Information
- Application Number
- CN202511416733.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Existing technologies suffer from high computational complexity and insufficient accuracy in video localization, making it difficult to effectively utilize multimodal large language models for accurate video localization.
A candidate event location video set is generated by a pre-trained multimodal large language model, and a subset of candidate event location videos is selected by using the target hit accuracy. The accurate event location videos are further selected by combining the video description text and the event description text.
It improves the accuracy and efficiency of video localization, reduces the computational load of subsequent models, and achieves higher accuracy in video localization without further adjustments to the multimodal large language model.
Smart Images

Figure CN121334458A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of artificial intelligence technology, and more specifically to video generation methods, apparatus, devices, computer-readable media, and program products. Background Technology
[0002] Currently, with the continuous development of the intelligent era, multimodal large language models are being applied more and more widely in various scenarios. Specifically, text-driven video event localization is a multimodal task. Regarding how to query the video corresponding to the descriptive text from the video, existing technologies often instruct multimodal large language models to extract more accurate video modal feature information and text modal feature information corresponding to the descriptive text. This approach may require complex adjustments to the model structure, training method, or prompt information of the multimodal large language model, often resulting in high computational complexity, and the accuracy of video localization cannot be effectively guaranteed. Summary of the Invention
[0003] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0004] Some embodiments of this disclosure provide video generation methods, apparatuses, devices, computer-readable media, and program products to address the technical problems mentioned in the background section above.
[0005] In a first aspect, some embodiments of this disclosure provide a video generation method, comprising: generating a candidate event location video set using a pre-trained first multimodal large language model based on an acquired target video and event description text; selecting a target number of candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, thereby obtaining a subset of candidate event location videos, wherein the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos; generating a video description text set corresponding to the subset of candidate event location videos using a pre-trained second multimodal large language model; and selecting event location videos corresponding to the event description text from the subset of candidate event location videos based on the video description text set and the event description text using a pre-trained third multimodal large language model.
[0006] Optionally, the above-mentioned selection of candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, to obtain a subset of candidate event location videos, includes: selecting candidate event location videos from the candidate event location video set whose video accuracy information ranks among the top target numbers, to obtain a subset of candidate event location videos.
[0007] Optionally, the above-mentioned method of selecting event location videos corresponding to the event description texts from the candidate event location video subset using a pre-trained third multimodal large language model, based on the above-mentioned video description text set and the above-mentioned event description texts, includes: using the above-mentioned third multimodal large language model to generate a question-answer pair set corresponding to the above-mentioned event description texts, wherein the question-answer pair includes: question and answer content; for each candidate event location video, performing a first generation step; for each question-answer pair in the above-mentioned question-answer pair set, performing a second generation step; inputting the video description text corresponding to the above-mentioned candidate event location video and the question in the above-mentioned question-answer pair into the above-mentioned third multimodal large language model to obtain the description answer content; determining the first content difference information between the answer content in the above-mentioned question-answer pair and the description answer content; generating video accuracy information corresponding to the above-mentioned candidate event location video based on the obtained first content difference information set; and selecting candidate event location videos from the above-mentioned candidate event location video subset whose video accuracy information meets the target accuracy condition as event location videos.
[0008] Optionally, generating the precise video information corresponding to the candidate event location video based on the obtained first content difference information set includes: for each question-answer pair in the question-answer pair set, performing a third generation step: inputting the candidate event location video and the questions in the question-answer pair into the third multimodal large language model to obtain the video description response content; determining the second content difference information between the response content in the question-answer pair and the video description response content; and generating the precise video information based on the first content difference information set and the obtained second content difference information set.
[0009] Optionally, the above-mentioned generation of a candidate event location video set based on the acquired target video and event description text using a pre-trained first multimodal large language model includes: inputting the target video and the event description text into the first multimodal large language model to obtain a candidate location time information set; determining the candidate event location video corresponding to each candidate location time information in the candidate location time information set to obtain the candidate event location video set.
[0010] Optionally, before selecting candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, and obtaining a subset of candidate event location videos, the method further includes: selecting hit accuracy rates from the hit accuracy rate sequence that satisfy the difference condition and / or whose values in the sequence are initially higher than a predetermined threshold, as the target hit probability, wherein the hit accuracy rate sequence corresponds to the first multimodal large language model, the hit accuracy rate sequence is determined based on the candidate event location video sequence, and the candidate event location video sequence is obtained by sorting the candidate event location video set according to the video accuracy information.
[0011] Secondly, some embodiments of this disclosure provide a video generation apparatus, comprising: a first generation unit configured to generate a candidate event location video set based on an acquired target video and event description text, using a pre-trained first multimodal large language model; a first filtering unit configured to filter a target number of candidate event location videos for video content verification from the candidate event location video set according to the target hit accuracy corresponding to the first multimodal large language model, thereby obtaining a subset of candidate event location videos, wherein the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos; a second generation unit configured to generate a video description text set corresponding to the subset of candidate event location videos using a pre-trained second multimodal large language model; and a second filtering unit configured to filter event location videos corresponding to the event description text from the subset of candidate event location videos based on the video description text set and the event description text, using a pre-trained third multimodal large language model.
[0012] Optionally, the first filtering unit can be configured to: filter out candidate event location videos whose video accuracy information is among the top target number from the above candidate event location video set, thereby obtaining a subset of candidate event location videos.
[0013] Optionally, the second filtering unit can be configured to: generate a set of question-and-answer pairs corresponding to the event description text using the aforementioned third multimodal large language model, wherein the question-and-answer pairs include: questions and answers; for each candidate event location video, perform a first generation step; for each question-and-answer pair in the aforementioned question-and-answer pair set, perform a second generation step; input the video description text corresponding to the aforementioned candidate event location video and the questions in the aforementioned question-and-answer pairs into the aforementioned third multimodal large language model to obtain the description and answer content; determine the first content difference information between the answer content in the aforementioned question-and-answer pairs and the description and answer content; generate the video accuracy information corresponding to the aforementioned candidate event location video based on the obtained first content difference information set; and filter out candidate event location videos whose video accuracy information meets the target accuracy condition from the aforementioned candidate event location video subset as event location videos.
[0014] Optionally, the second filtering unit can be configured to: for each question-answer pair in the above question-answer pair set, perform a third generation step: input the above candidate event location video and the questions in the above question-answer pair into the above third multimodal large language model to obtain the video description response content; determine the second content difference information between the response content in the above question-answer pair and the above video description response content; and generate the above video precision information based on the above first content difference information set and the obtained second content difference information set.
[0015] Optionally, the first generation unit can be configured to: input the target video and the event description text into the first multimodal large language model to obtain a candidate location time information set; determine the candidate event location video corresponding to each candidate location time information in the candidate location time information set to obtain a candidate event location video set.
[0016] Optionally, the steps further include: selecting from the hit accuracy sequence the hit accuracy rates where the difference in accuracy changes satisfies the difference condition and / or the values in the sequence are initially higher than a predetermined threshold, as the target hit probability, wherein the hit accuracy sequence corresponds to the first multimodal large language model, the hit accuracy sequence is determined based on the candidate event localization video sequence, and the candidate event localization video sequence is obtained by sorting the candidate event localization video set according to the video accuracy information.
[0017] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0018] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0019] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0020] The above embodiments of this disclosure have the following beneficial effects: The video generation model method of some embodiments of this disclosure can more accurately and efficiently select the event location video that best matches the event description text from multiple candidate event location videos output by the multimodal model. Specifically, the reason why the related event location videos are not accurate and efficient is that complex adjustments are needed to the model structure, training method, or prompt information of the multimodal large language model, often resulting in high computational complexity, and the accuracy of video location cannot be effectively guaranteed. Based on this, the video generation method of some embodiments of this disclosure firstly generates a candidate event location video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model. Here, by generating the candidate event location video set, each location video can be used as a candidate to further select the most accurate event location video from among them. Then, based on the target hit accuracy corresponding to the first multimodal large language model, a target number of candidate event location videos are selected from the candidate event location video set for video content verification, resulting in a subset of candidate event location videos. Here, the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos. By using the target hit accuracy, a subset of candidate event location videos with high hit accuracy can be selected, ensuring that the hit accuracy meets the corresponding accuracy requirements (e.g., the hit accuracy is higher than the target accuracy) within a suitable number of candidate event location videos. Accurate selection of the candidate event location video subset based on the target hit accuracy not only accurately identifies event location videos but also reduces subsequent model computation. Next, using the pre-trained second multimodal large language model, a video description text set corresponding to the aforementioned candidate event location video subset can be accurately generated. Here, further selection of event location videos from the candidate event location video subset from a text modality perspective can further improve the accuracy of event location video determination. Finally, based on the aforementioned video description text set and event description text, the pre-trained third multimodal large language model can accurately filter out the event location videos corresponding to the aforementioned event description texts from the aforementioned candidate event location video subset. In summary, after determining the candidate event location video subset with the target hit accuracy, by using the feature information in the text modality, the multimodal large language model is used again to re-predict event location videos from the coarse output result (i.e., the output candidate event location video subset). This can improve the accuracy of generating event location videos without further adjusting the multimodal large language model for further feature extraction in both video and image modalities. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation method according to some embodiments of the present disclosure;
[0023] Figure 2 This is a flowchart of some embodiments of the video generation method according to the present disclosure;
[0024] Figure 3 This is a flowchart of some other embodiments of the video generation method according to the present disclosure;
[0025] Figure 4 These are schematic diagrams illustrating the structure of some embodiments of the video generation apparatus according to this disclosure;
[0026] Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0030] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0031] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0032] Before performing any of the operations involving the collection, storage, or use of user personal information (such as user profiles and user historical behavior) disclosed in this disclosure, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, informing the personal information subjects, and obtaining prior authorization and consent from the personal information subjects.
[0033] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0034] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation method according to some embodiments of the present disclosure.
[0035] exist Figure 1 In the application scenario, firstly, the electronic device 101 can generate a candidate event location video set 106 based on the acquired target video 103 and event description text 102, using a pre-trained first multimodal large language model 104. In this application scenario, the event description text 102 could be "a person walked to the window and looked out." Then, the electronic device 101 can select a target number of candidate event location videos for video content verification from the candidate event location video set 106 based on the target hit accuracy 105 corresponding to the first multimodal large language model 104, obtaining a candidate event location video subset 107. The target hit accuracy 105 is the probability of finding an event location video among the target number of candidate event location videos. In this application scenario, the target hit accuracy 105 could be "95%". Next, the electronic device 101 can use a pre-trained second multimodal large language model 108 to generate a video description text set 109 corresponding to the candidate event location video subset 107. Finally, the electronic device 101 can use the pre-trained third multimodal large language model 110 to select the event location video 111 corresponding to the event description text 102 from the candidate event location video subset 107 based on the video description text set 109 and the event description text 102.
[0036] It should be noted that the aforementioned electronic device 101 can be either hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the electronic device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.
[0037] It should be understood that Figure 1 The number of electronic devices shown is merely illustrative. Any number of electronic devices can be used depending on the implementation requirements.
[0038] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of a video generation method according to the present disclosure. The video generation method includes the following steps:
[0039] Step 201: Based on the acquired target video and event description text, generate a candidate event localization video set using a pre-trained first multimodal large language model.
[0040] In some embodiments, the execution entity of the above video generation method (e.g. Figure 1 The electronic device 101 shown can generate a candidate event location video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model. The target video can be the original video for at least one event. That is, the target video corresponds to a video content description containing at least one event. In practice, in the logistics field, the target video can be a warehouse video or a platform video. The event description text can be a descriptive text of the event video viewing content. The event video can be a video with the event content as its main body. Here, the event video can be a segment of the target video. The event video corresponds to one of at least one events. For example, for a warehouse video as the target video, the corresponding event description text could be "determining the process of employee A packing package A". In specific practice, the event description text can be the text describing the event content entered by the target user on the target page according to the video viewing requirements. The first multimodal large language model can be a large language model that supports various processing of data under multiple modalities. In practice, multiple modalities can include: image modality, video modality, and speech modality. The first multimodal large language model can be a commercially available large language model pre-trained on a task of extracting video content corresponding to event description text. For example, the first multimodal large language model can be the MS-2D-TAN model. Candidate event localization videos can be candidate event localization videos whose video content, to be further determined, corresponds to the event description text. Event localization videos can be event videos located from the target video that correspond to the event description text. That is, each candidate event localization video in the candidate event localization video set is a video segment from the target video.
[0041] As an example, firstly, the aforementioned execution entity can generate a first generated prompt message based on the acquired target video and event description text, generating multiple candidate event location videos corresponding to the event description text. Then, the first generated prompt message is input into a first multimodal large language model to obtain a candidate event location video set.
[0042] In some optional implementations of certain embodiments, the aforementioned execution entity can generate a candidate event localization video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model, including the following steps:
[0043] The first step involves inputting the target video and the event description text into the first multimodal large language model to obtain a set of candidate localization time information. This candidate localization time information can be the time segment of the candidate event localization video within the target video; that is, the time segment corresponding to the fragment in the target video. The candidate localization time information may include the start and end times of the candidate event localization video.
[0044] As an example, firstly, the aforementioned executing entity can generate a third generated prompt message based on the target video and event description text to generate a candidate location time information set. Then, the third generated prompt message is input into the first multimodal large language model to obtain the candidate location time information set.
[0045] The second step is to determine the candidate event location video corresponding to each candidate location time information in the above candidate location time information set, thus obtaining the candidate event location video set.
[0046] As an example, for each candidate location time information, a video segment corresponding to the time period of the above candidate location time information is extracted from the target video as a candidate event location video, thus obtaining a candidate event location video set.
[0047] In some optional implementations of certain embodiments, before step 201, the steps further include:
[0048] The aforementioned execution entity can filter out hit accuracy rates from the hit accuracy rate sequence that meet the difference condition and / or whose numerical values in the sequence initially exceed a predetermined threshold, as the target hit probability. The hit accuracy rate sequence corresponds to the aforementioned first multimodal large language model. The hit accuracy rate sequence is determined based on the candidate event localization video sequence. The candidate event localization video sequence is obtained by sorting the candidate event localization video set according to the video accuracy information. The hit accuracy rate sequence can be sorted in ascending order of the number of corresponding candidate event localization videos. The hit accuracy rate sequence can represent the accuracy of each candidate event localization video pair output by the first multimodal large language model. The accuracy change difference can be the numerical difference in the hit accuracy rate within the hit accuracy rate sequence. In practice, the accuracy change difference can be obtained by subtracting every two adjacent hit accuracy rates in the hit accuracy rate sequence. For example, subtracting the hit accuracy rate of an earlier sequence position from the hit accuracy rate of a later sequence position. The corresponding accuracy change difference can also be information in sequence form. That is, the accuracy variation difference can be an accuracy difference sequence. The accuracy difference can be the difference between the hit accuracy at a later sequence position and the hit accuracy at an earlier sequence position. The difference condition can be that the maximum hit accuracy corresponding to the largest accuracy difference is used as the target hit accuracy. The largest accuracy difference can be the largest accuracy difference in the accuracy difference sequence. The maximum hit accuracy corresponding to the largest accuracy difference can be the hit accuracy at the later sequence position corresponding to the largest accuracy difference. The predetermined threshold can be a pre-set accuracy threshold.
[0049] In practice, the hit accuracy sequence can be generated through the following steps:
[0050] The first step is to set a number sequence corresponding to the hit accuracy sequence based on each candidate event localization video in the candidate event localization video sequence. The numbers in the number sequence are arranged in ascending order.
[0051] The second step involves performing the following determination steps for each number in the sequence:
[0052] Sub-step 1 involves sequentially selecting a number of candidate event location videos from the candidate event location video sequence, from left to right. The candidate event location videos are arranged in descending order of their accuracy information.
[0053] Sub-step 2: Determine the probability that an event location video exists among the aforementioned number of candidate event location videos, and use this probability as the hit accuracy corresponding to the number.
[0054] Step 202: Based on the target hit accuracy corresponding to the first multimodal large language model, select the target number of candidate event location videos for video content verification from the candidate event location video set to obtain a subset of candidate event location videos.
[0055] In some embodiments, the execution entity can, based on the target hit accuracy corresponding to the first multimodal large language model, filter out a target number of candidate event location videos for video content verification from the candidate event location video set, thus obtaining a subset of candidate event location videos. The target hit accuracy is the probability that an event location video exists among the target number of candidate event location videos. In practice, the target hit accuracy represents a very high probability that the target number of candidate event location videos includes the event location video. The target hit accuracy can be a hit accuracy selected from various hit accuracy rates. In practice, each hit accuracy rate has a corresponding number. This number can be the number of candidate event location videos. The corresponding hit accuracy rate can represent the probability that the corresponding number of candidate event location videos includes the event location video. The hit accuracy rate is a value between 0 and 1. The higher the value, the higher the probability that the corresponding number of candidate event location videos includes the event location video. Each hit accuracy rate can be obtained statistically based on the accuracy of the model's output. For example, the candidate event localization video set includes: candidate event localization video A, candidate event localization video B, candidate event localization video C, and candidate event localization video D. Each hit accuracy rate can include: a first hit accuracy rate, a second hit accuracy rate, and a third hit accuracy rate. The first hit accuracy rate corresponds to a number of 2, so the candidate event localization videos corresponding to the first hit accuracy rate can include: candidate event localization video A and candidate event localization video B. Here, the candidate event localization videos corresponding to the first hit accuracy rate can be selected from the candidate event localization video set based on a predetermined selection rule. The predetermined selection rule can be a pre-set selection rule for candidate event localization videos. For example, the predetermined selection rule can be to select candidate event localization videos according to the order in which the first multimodal large language model outputs the videos. When the candidate event localization videos corresponding to the first hit accuracy rate include candidate event localization video A and candidate event localization video B, the first hit accuracy rate can represent the probability that an event localization video exists in candidate event localization videos A and B. Similarly, if the number corresponding to the second hit accuracy is 3, then the candidate event localization videos corresponding to the second hit accuracy can include: candidate event localization video A, candidate event localization video B, and candidate event localization video C. When the candidate event localization videos corresponding to the second hit accuracy include candidate event localization video A, candidate event localization video B, and candidate event localization video C, the second hit accuracy can represent the probability of hitting an event localization video among candidate event localization videos A, B, and C. If the number corresponding to the third hit accuracy is 4, then the candidate event localization videos corresponding to the third hit accuracy can include: candidate event localization video A, candidate event localization video B, candidate event localization video C, and candidate event localization video D.The third hit accuracy rate characterizes the probability of finding a video that locates an event among candidate event location videos A, B, C, and D. This corresponds to a first hit accuracy rate of 50%, a second hit accuracy rate of 60%, a third hit accuracy rate of 80%, and a fourth hit accuracy rate of 90%. The target hit accuracy rate can be greater than or equal to 80% and has the fewest number of videos. That is, the target hit accuracy rate is the third hit accuracy rate.
[0056] In some optional implementations of certain embodiments, step 202 above may include the following steps:
[0057] The aforementioned executing entity can filter from the candidate event location video set to select the top-ranking number of candidate event location videos with accurate video information, thus obtaining a subset of candidate event location videos. The accurate video information can be the precise information confirming that a candidate event location video is indeed an event location video. For example, the accurate video information could be the video accuracy rate. A higher video accuracy rate indicates a higher probability that a candidate event location video is indeed an event location video.
[0058] As an example, firstly, based on the accurate video information corresponding to each candidate event location video, the candidate event location videos in the candidate event location video set are sorted in descending order to obtain a candidate event location video sequence. Then, a target number of candidate event location videos are selected from left to right to obtain a subset of candidate event location videos.
[0059] Step 203: Using the pre-trained second multimodal large language model, generate video description text sets corresponding to the above candidate event localization video subsets.
[0060] In some embodiments, the aforementioned execution entity can utilize a pre-trained second multimodal large language model to generate a video description text set corresponding to the aforementioned candidate event location video subset. The second multimodal large language model can also be a large language model that supports various processing of data across multiple modalities. The second multimodal large language model can be the same as or different from the first multimodal large language model. For example, the second multimodal large language model can be the shareGPT4video[3] model. There is a one-to-one correspondence between the candidate event location videos in the candidate event location video subset and the video description texts in the video description text set. The video description text can be a description text describing the video content corresponding to the candidate event location video.
[0061] As an example, firstly, second generated prompts are generated for the video description text corresponding to each candidate event location video in the candidate event location video subset. The second generated prompts include relevant video content from the candidate event location video subset. Then, the second generated prompts are input into a second multimodal large language model to obtain the video description text set.
[0062] Step 204: Based on the above video description text set and the above event description text, use the pre-trained third multimodal large language model to select the event location video corresponding to the above event description text from the above candidate event location video subset.
[0063] In some embodiments, the execution entity can, based on the video description text set and the event description text, utilize a pre-trained third multimodal large language model to filter out the event location videos corresponding to the event description texts from the candidate event location video subset. The third multimodal large language model can be a large language model that supports various processing of data across multiple modalities. The second multimodal large language model can be the same as or different from the first or second multimodal large language model. The event location video can be a video segment from the target video that is related to the description content corresponding to the event description text.
[0064] As an example, firstly, the aforementioned execution entity can generate a third-generation prompt message that filters out the video description text from the video description text and finds the one with the closest semantic content to the corresponding event description text. Then, the third-generation prompt message is input into a third multimodal large language model to obtain the target video description text. Finally, candidate event location videos corresponding to the target video description text are determined and used as the event location videos corresponding to the event description text.
[0065] The above embodiments of this disclosure have the following beneficial effects: The video generation model method of some embodiments of this disclosure can more accurately and efficiently select the event location video that best matches the event description text from multiple candidate event location videos output by the multimodal model. Specifically, the reason why the related event location videos are not accurate and efficient is that complex adjustments are needed to the model structure, training method, or prompt information of the multimodal large language model, often resulting in high computational complexity, and the accuracy of video location cannot be effectively guaranteed. Based on this, the video generation method of some embodiments of this disclosure firstly generates a candidate event location video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model. Here, by generating the candidate event location video set, each location video can be used as a candidate to further select the most accurate event location video from among them. Then, based on the target hit accuracy corresponding to the first multimodal large language model, a target number of candidate event location videos are selected from the candidate event location video set for video content verification, resulting in a subset of candidate event location videos. Here, the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos. By using the target hit accuracy, a subset of candidate event location videos with high hit accuracy can be selected, ensuring that the hit accuracy meets the corresponding accuracy requirements (e.g., the hit accuracy is higher than the target accuracy) within a suitable number of candidate event location videos. Accurate selection of the candidate event location video subset based on the target hit accuracy not only accurately identifies event location videos but also reduces subsequent model computation. Next, using the pre-trained second multimodal large language model, a video description text set corresponding to the aforementioned candidate event location video subset can be accurately generated. Here, further selection of event location videos from the candidate event location video subset from a text modality perspective can further improve the accuracy of event location video determination. Finally, based on the aforementioned video description text set and event description text, the pre-trained third multimodal large language model can accurately filter out the event location videos corresponding to the aforementioned event description texts from the aforementioned candidate event location video subset. In summary, after determining the candidate event location video subset with the target hit accuracy, by using the feature information in the text modality, the multimodal large language model is used again to re-predict event location videos from the coarse output result (i.e., the output candidate event location video subset). This can improve the accuracy of generating event location videos without further adjusting the multimodal large language model for further feature extraction in both video and image modalities.
[0066] Further reference Figure 3The diagram illustrates a flow 300 of another embodiment of the video generation method according to the present disclosure. This video generation method includes the following steps:
[0067] Step 301: Based on the acquired target video and event description text, generate a candidate event localization video set using a pre-trained first multimodal large language model.
[0068] Step 302: Based on the target hit accuracy corresponding to the first multimodal large language model, select the target number of candidate event location videos for video content verification from the candidate event location video set to obtain a subset of candidate event location videos.
[0069] Step 303: Using the pre-trained second multimodal large language model, generate video description text sets corresponding to the above candidate event localization video subsets.
[0070] In some embodiments, the specific implementation of steps 301-303 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-203 in the corresponding embodiments will not be repeated here.
[0071] Step 304: Using the aforementioned third multimodal large language model, generate a set of question-answer pairs corresponding to the event description text.
[0072] In some embodiments, the executing entity (e.g. Figure 1 The electronic device 101 shown can utilize the aforementioned third multimodal large language model to generate a set of question-answer pairs corresponding to the event description text. Each question-answer pair includes a question and a response. Specifically, a question-answer pair can be a question and its corresponding response constructed based on the text content corresponding to the event description text. For example, the event description text could be "determining the process of employee A packing package A," with the question in question-answer pair A being "who is the object processing package A," and the response being "employee A." Similarly, the question in question-answer pair B is "what kind of processing does employee A perform on package A," and the response is "packaging processing."
[0073] As an example, firstly, a fourth generated prompt is generated for the question-answer pair set corresponding to the event description text. Then, the fourth generated prompt is input into the third multimodal large language model mentioned above to obtain the question-answer pair set.
[0074] Step 305: For each candidate event location video, perform the first generation step:
[0075] Step 3051: For each question-answer pair in the above question-answer pair set, perform the second generation step:
[0076] Step 30511: Input the video description text corresponding to the above candidate event location video and the questions in the above question-answer pair into the above third multimodal large language model to obtain the description response content.
[0077] In some embodiments, the executing entity may input the video description text corresponding to the candidate event location video and the question in the question-answer pair into the third multimodal large language model to obtain the descriptive response content. The descriptive response content is the response content related to the question in the video description text. That is, the third multimodal large language model extracts the text content from the video description text, answers the question, and obtains the response content (i.e., the descriptive response content).
[0078] As an example, firstly, the aforementioned executing entity can generate a first response prompt message to extract text content from the video description text to answer the questions in the question-answer pair. Then, the first response prompt message is input into the third multimodal large language model to obtain the descriptive response content.
[0079] Step 30512: Determine the first content difference information between the response content in the above question-and-answer pair and the above-described response content.
[0080] In some embodiments, the executing entity may determine first content difference information between the response content and the descriptive response content in the question-and-answer pair. The first content difference information may be a description of the content differences between the response content and the descriptive response content. In practice, the first content difference information may be in the form of a label or in the form of a numerical value. For example, the first content difference information may be one of the following: a difference exists, or no difference exists. The first content difference information may also be a numerical value between 0 and 100. The higher the value, the more severe the content difference.
[0081] As an example, the aforementioned implementing entity can utilize a third multimodal large language model to determine the first content difference information between the response content in the aforementioned question-answer pair and the aforementioned descriptive response content.
[0082] Step 3052: Based on the obtained first content difference information set, generate the precise video information corresponding to the above candidate event location video.
[0083] In some embodiments, the executing entity can generate precise video information corresponding to the candidate event location video based on the obtained first content difference information set. The precise video information can be the precise information that the candidate event location video is an event location video. The precise video information can be in the form of tags or in the form of values. For tag-based information, the precise video information can be one of the following: the candidate event location video provides a precise question-and-answer pair, or the candidate event location video provides an inaccurate question-and-answer pair. For numerical information, the larger the value, the more precise the question-and-answer pair of the candidate event location video.
[0084] As an example, for the first content difference information set in numerical form, the aforementioned execution entity can perform weighted summation of each first content difference information in the first content difference information set to obtain accurate video information.
[0085] As another example, for a first set of content difference information in the form of tags, in response to determining that all content difference information in the first set of content difference information contains no difference, candidate event location videos are generated to provide accurate video information for question-and-answer matching. In response to determining that at least one content difference information in the first set of content difference information contains a difference, candidate event location videos are generated to provide accurate video information for question-and-answer matching.
[0086] In some optional implementations of certain embodiments, the execution entity may generate precise video information corresponding to the candidate event location video based on the obtained first content difference information set, including the following steps:
[0087] First, for each question-answer pair in the above question-answer pair set, perform the third generation step:
[0088] Sub-step 1 involves inputting the aforementioned candidate event location videos and the questions from the aforementioned question-answer pairs into the aforementioned third multimodal large language model to obtain video description response content. The video description response content can be the question-related response content from the candidate event location videos. That is, the third multimodal large language model extracts video content from the candidate event location videos to answer the questions, thereby obtaining response content (i.e., description response content).
[0089] As an example, firstly, the aforementioned executing entity can generate a second response prompt message to extract video content within the video of the candidate event location to answer the question in the question-answer pair. Then, the second response prompt message is input into a third multimodal large language model to obtain the video description response content.
[0090] Sub-step 2 involves determining the second content difference information between the response content in the question-and-answer pair and the response content in the video description. This second content difference information can be a description of the content differences between the response content and the response content in the video description. In practice, the second content difference information can be in the form of a label or a numerical value. For example, the second content difference information can be one of the following: there is a difference, or there is no difference. The second content difference information can also be a numerical value between 0 and 100. The higher the value, the more severe the content difference.
[0091] As an example, the aforementioned implementing entity can utilize a third multimodal large language model to determine the second content difference information between the response content in the aforementioned question-and-answer pair and the response content in the aforementioned video description.
[0092] The second step is to generate the aforementioned precise video information based on the first content difference information set and the obtained second content difference information set.
[0093] As an example, for the first content difference information set and the second content difference information set in numerical form, the aforementioned executing entity can perform weighted summation on each first content difference information set in the first content difference information set and each second content difference information set in the second content difference information set to obtain accurate video information.
[0094] As another example, for a first set of content difference information in the form of tags, in response to determining that all first content difference information in the first set of content difference information contains no difference and all second content difference information in the second set of content difference information contains no difference, candidate event location videos are generated to provide accurate video information for question-and-answer matching. In response to determining that at least one first content difference information in the first set of content difference information contains a difference and / or at least one second content difference information in the second set of content difference information contains a difference, candidate event location videos are generated to provide accurate video information for question-and-answer matching that does not provide accurate video information.
[0095] 306. From the above subset of candidate event location videos, select candidate event location videos whose video accuracy information meets the target accuracy condition, and use them as event location videos.
[0096] In some embodiments, the aforementioned executing entity may select candidate event location videos from the aforementioned subset of candidate event location videos whose video accuracy information meets the target accuracy condition, and use them as event location videos. The target accuracy condition may be one of the following: the candidate event location video with the highest video accuracy information, or the candidate event location video whose video accuracy information corresponds to an accurate question-and-answer session with the candidate event location video.
[0097] from Figure 3 It can be seen from this that, with Figure 2Compared to the description of some corresponding embodiments, Figure 3 In some corresponding embodiments, the video generation method process 300 involves constructing question-and-answer pairs to perform semantic content difference matching between the video description text set and the event description text in a text modal. Therefore, by using the semantic content differences in the text modal, the most accurate candidate event location video can be selected from the candidate event location video subset and used as the event location video.
[0098] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a video generation apparatus, which are similar to... Figure 2 Corresponding to the method embodiments shown, this video generation apparatus can be specifically applied to various electronic devices.
[0099] like Figure 4 As shown, a video generation device 400 includes: a first generation unit 401, a first filtering unit 402, a second generation unit 403, and a second filtering unit 404. The first generation unit 401 is configured to generate a candidate event location video set based on the acquired target video and event description text using a pre-trained first multimodal large language model; the first filtering unit 402 is configured to filter out a target number of candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, thereby obtaining a subset of candidate event location videos, wherein the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos; the second generation unit 403 is configured to generate a video description text set corresponding to the subset of candidate event location videos using a pre-trained second multimodal large language model; and the second filtering unit 404 is configured to filter out the event location videos corresponding to the event description text from the subset of candidate event location videos based on the video description text set and the event description text using a pre-trained third multimodal large language model.
[0100] In some optional implementations of some embodiments, the first filtering unit 402 may be further configured to: filter out candidate event location videos whose video accuracy information is among the top target number from the candidate event location video set, thereby obtaining a subset of candidate event location videos.
[0101] In some optional implementations of certain embodiments, the second filtering unit 404 may be further configured to: generate a set of question-and-answer pairs corresponding to the event description text using the third multimodal large language model, wherein the question-and-answer pairs include: questions and answers; for each candidate event location video, perform a first generation step: for each question-and-answer pair in the question-and-answer pair set, perform a second generation step: input the video description text corresponding to the candidate event location video and the questions in the question-and-answer pairs into the third multimodal large language model to obtain the description and answer content; determine the first content difference information between the answer content in the question-and-answer pairs and the description and answer content; generate the video accuracy information corresponding to the candidate event location video based on the obtained first content difference information set; and select candidate event location videos whose video accuracy information meets the target accuracy condition from the subset of candidate event location videos as event location videos.
[0102] In some optional implementations of some embodiments, the second filtering unit 404 may be further configured to: for each question-answer pair in the question-answer pair set, perform a third generation step: input the candidate event location video and the questions in the question-answer pair into the third multimodal large language model to obtain video description response content; determine the second content difference information between the response content in the question-answer pair and the video description response content; and generate the video accuracy information based on the first content difference information set and the obtained second content difference information set.
[0103] In some optional implementations of some embodiments, the first generation unit 401 may be further configured to: input the target video and the event description text into the first multimodal large language model to obtain a candidate location time information set; determine the candidate event location video corresponding to each candidate location time information in the candidate location time information set to obtain a candidate event location video set.
[0104] In some optional implementations of certain embodiments, the apparatus 400 further includes a third filtering unit (not shown in the figure). This third filtering unit can be configured to: filter from the hit accuracy sequence those hit accuracy rates where the difference in accuracy changes satisfies a difference condition and / or the value in the sequence initially exceeds a predetermined threshold, as target hit probabilities. The hit accuracy sequence corresponds to the first multimodal large language model, and the hit accuracy sequence is determined based on candidate event location video sequences, which are obtained by sorting the candidate event location video set according to the video accuracy information.
[0105] It is understandable that the units described in the video generation device 400 are related to the reference. Figure 2The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the video generation apparatus 400 and the units contained therein, and will not be repeated here.
[0106] The following is for reference. Figure 5 It illustrates electronic devices suitable for implementing some embodiments of this disclosure (e.g., Figure 1 A schematic diagram of the structure of electronic device 101)500. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0107] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory 502 or a program loaded from a storage device 508 into a random access memory 503. The random access memory 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the read-only memory 502, and the random access memory 503 are interconnected via a bus 504. An input / output interface 505 is also connected to the bus 504.
[0108] Typically, the following devices can be connected to the input / output interface 505: input devices 506 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 508 including, for example, magnetic tape, hard disk, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 5 Each box shown can represent a device or multiple devices as needed.
[0109] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a read-only memory 502. When the computer program is executed by the processing device 501, it performs the functions defined above in the methods of some embodiments of this disclosure.
[0110] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0111] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0112] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: generate a candidate event location video set using a pre-trained first multimodal large language model based on the acquired target video and event description text; select a target number of candidate event location videos from the candidate event location video set for video content verification based on the target hit accuracy corresponding to the first multimodal large language model, thereby obtaining a subset of candidate event location videos, wherein the target hit accuracy is the probability of finding an event location video among the target number of candidate event location videos; generate a video description text set corresponding to the subset of candidate event location videos using a pre-trained second multimodal large language model; and select the event location video corresponding to the event description text from the subset of candidate event location videos using a pre-trained third multimodal large language model based on the video description text set and the event description text.
[0113] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0115] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first generation unit, a first filtering unit, a second generation unit, and a second filtering unit. The names of these units do not necessarily limit the specific unit itself; for example, the first generation unit may also be described as "a unit that generates a candidate event location video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model."
[0116] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0117] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the video generation methods described above.
[0118] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A video generation method, comprising: Based on the acquired target video and event description text, a candidate event localization video set is generated using a pre-trained first multimodal large language model; Based on the target hit accuracy corresponding to the first multimodal large language model, a target number of candidate event location videos are selected from the candidate event location video set for video content verification to obtain a subset of candidate event location videos. The target hit accuracy is the probability that an event location video exists among the target number of candidate event location videos. Using a pre-trained second multimodal large language model, a video description text set corresponding to the candidate event location video subset is generated; Based on the video description text set and the event description text, a pre-trained third multimodal large language model is used to filter out the event location videos corresponding to the event description text from the candidate event location video subset.
2. The method according to claim 1, wherein, The step of selecting candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, resulting in a subset of candidate event location videos, includes: From the candidate event location video set, select the candidate event location videos whose accurate video information is in the number of the preceding targets to obtain a subset of candidate event location videos.
3. The method according to claim 1, wherein, The step of selecting the event location video corresponding to the event description text from the candidate event location video subset using a pre-trained third multimodal large language model based on the video description text set and the event description text includes: Using the third multimodal large language model, a question-answer pair set corresponding to the event description text is generated, wherein the question-answer pair includes: question and answer content; For each candidate event location video, perform the first generation step: For each question-answer pair in the question-answer pair set, perform the second generation step: The video description text corresponding to the candidate event location video and the question in the question-answer pair are input into the third multimodal large language model to obtain the description response content; Determine the first content difference information between the response content and the descriptive response content in the question-answer pair; Based on the obtained first content difference information set, generate the precise video information corresponding to the candidate event location video; Candidate event location videos that meet the target accuracy criteria are selected from the subset of candidate event location videos and used as event location videos.
4. The method according to claim 3, wherein, The step of generating precise video information corresponding to the candidate event location video based on the obtained first content difference information set includes: For each question-answer pair in the question-answer pair set, perform the third generation step: The candidate event location video and the questions in the question-answer pair are input into the third multimodal large language model to obtain the video description response content; Determine the second content difference information between the response content in the question-and-answer pair and the response content in the video description; The precise video information is generated based on the first content difference information set and the obtained second content difference information set.
5. The method according to claim 1, wherein, The step of generating a candidate event localization video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model, includes: The target video and the event description text are input into the first multimodal large language model to obtain a candidate location time information set; The candidate event location video corresponding to each candidate location time information in the candidate location time information set is determined to obtain the candidate event location video set.
6. The method according to claim 1, wherein, Before the step of filtering candidate event location videos for video content verification from the candidate event location video set based on the target hit accuracy corresponding to the first multimodal large language model, and obtaining a subset of candidate event location videos, the method further includes: The hit accuracy rate sequence is selected from the hit accuracy rate sequence if the difference in accuracy rate meets the difference condition and / or the value in the sequence is higher than a predetermined threshold for the first time, and is used as the target hit probability. The hit accuracy rate sequence corresponds to the first multimodal large language model. The hit accuracy rate sequence is determined based on the candidate event location video sequence. The candidate event location video sequence is obtained by sorting the candidate event location video set according to the video accuracy information.
7. A video generation apparatus, comprising: The first generation unit is configured to generate a candidate event localization video set based on the acquired target video and event description text, using a pre-trained first multimodal large language model. The first filtering unit is configured to filter out a target number of candidate event positioning videos for video content verification from the candidate event positioning video set based on the target hit accuracy corresponding to the first multimodal large language model, thereby obtaining a subset of candidate event positioning videos, wherein the target hit accuracy is the probability of finding an event positioning video among the target number of candidate event positioning videos. The second generation unit is configured to use a pre-trained second multimodal large language model to generate a set of video description texts corresponding to the candidate event location video subset. The second filtering unit is configured to filter out the event location video corresponding to the event description text from the candidate event location video subset based on the video description text set and the event description text, using a pre-trained third multimodal large language model.
8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Event detection method and device, equipment and storage medium
CN116824455A
Search playing method and device, equipment and medium
CN116886947A
Visual model-based large language model video time sequence positioning method and product
CN117851638A
Video text understanding model training method and system based on contrast learning
CN120689793A
Efficient and fine-grained video retrieval
US20200302294A1