Hierarchical video search ranking optimization method and device
Through the combination of multimodal large language model and visual language model, the easiest and most difficult semantics of video clips are screened, and the video retrieval method is optimized, which solves the problem of insufficient cross-modal alignment and long-term modeling, and achieves more accurate video retrieval results.
Patent Information
- Application Number
- CN202510590203.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing video retrieval methods have shortcomings in cross-modal alignment and long-term modeling, resulting in poor accuracy of video retrieval results. The traditional feedback mechanism relies on manual labeling inefficient efficiency and fails to effectively utilize semantic supervision signals in negative samples.
The multimodal large language model is used to semantically annotate candidate video clips, and the easiest and most difficult semantics are filtered through the basic semantics of the query statement, the similarity score is updated to strengthen difficult semantic recognition and suppress easy confusing semantics, and the multimodal features are extracted in combination with the visual language model for initial similarity ranking optimization.
It significantly improves the accuracy of video clip ranking, improves the distinction between positive and negative samples, solves the problem of long-tail semantic matching, and provides fine-grained supervised signals to improve the accuracy of search results.
Smart Images

Figure CN120448584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information retrieval technology, and in particular to a hierarchical video search ranking optimization method and device. Background Art
[0002] With the explosive growth of multimedia content in recent years, video understanding and retrieval technologies are facing increasing demands and challenges. Understanding feature films requires in-depth analysis of character relationships, scene dynamics, and long-term narrative logic, while real-time video retrieval must process complex semantic queries input by users. Traditional methods face core challenges in both areas: insufficient fusion of multimodal information, difficulty modeling long-term relationships, and inefficient utilization of user feedback.
[0003] Specifically, early video understanding techniques focused on single-modal feature extraction, such as actor recognition through face matching or action timing modeling using 3D convolutional networks. However, film narratives have strong contextual dependencies, and single modalities such as vision or text have difficulty capturing complex interactions.
[0004] Existing methods are particularly vulnerable to multimodal alignment and long-term temporal modeling. While two-stream networks and Transformer architectures have improved short-term action recognition, the evolution of cross-scene character relationships still lacks effective modeling.
[0005] Therefore, existing video retrieval methods lack the dynamics of cross-modal alignment and the understanding and utilization of the semantics of the retrieved text, resulting in poor accuracy of existing video retrieval results. Summary of the Invention
[0006] The present invention provides a hierarchical video search ranking optimization method and device to address the defects in the existing technology of insufficient semantic utilization of cross-modal information and retrieval text, and realizes a hierarchical video search ranking optimization method and system based on a multimodal large language model.
[0007] The present invention provides a hierarchical video search ranking optimization method, comprising: Generate an initial similarity ranking based on the similarity scores between the query and the candidate video clips from large to small; Using a multimodal large language model to semantically annotate a preset number of candidate video clips ranked before the initial similarity, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing result of the query statement; Verifying the accuracy of basic semantics for each candidate video segment in the negative sample set, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; The initial similarity score is updated using the most difficult semantics and the easiest semantics to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and a search result of the video clip is generated based on the updated initial similarity score.
[0008] According to a hierarchical video search ranking optimization method provided by the present invention, before the step of generating an initial similarity ranking based on the descending order of similarity scores between the query statement and the candidate video clips, the method further includes: Extracting multimodal features of the candidate video clips using multiple visual language models; Generate a query image of the query statement using an image generation model, and extract and aggregate features of the query image and query text to obtain multimodal features of the query statement; Based on the similarity between the multimodal features of each candidate video segment and the multimodal features of the query sentence, a similarity score between the query sentence and the candidate video segment is determined.
[0009] According to a hierarchical video search ranking optimization method provided by the present invention, the step of generating an initial similarity ranking based on the descending order of similarity scores between the query statement and the candidate video clips specifically includes: The average fusion method is used to calculate the similarity scores of several visual language models for each candidate video segment as the candidate similarity score of each candidate video segment; An initial similarity ranking is generated based on the descending order of the candidate similarity scores.
[0010] According to a hierarchical video search ranking optimization method provided by the present invention, the step of verifying the accuracy of basic semantics of candidate video clips in the negative sample set one by one to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics specifically includes: For each basic semantics, calculate the ratio between the number of candidate video segments in the negative sample set that accurately matches its semantics and the number of all candidate video segments in the negative sample set as the accuracy rate of each basic semantics; The basic semantics with the lowest accuracy is regarded as the most difficult semantics, and the basic semantics with the highest accuracy is regarded as the easiest semantics.
[0011] According to a hierarchical video search ranking optimization method provided by the present invention, the step of generating search results for video clips based on the updated initial similarity scores specifically includes: Based on the semantic matching accuracy of each candidate video segment, the similarity score of each visual language model for each candidate video segment is modified, so as to determine the model weight of each visual language model based on the modified similarity score; For each candidate video clip, the updated initial similarity scores determined based on each visual language model are weightedly fused using the model weight of each visual language model to obtain the re-aggregated similarity score of each candidate video clip; The search results of the video clips are generated based on the similarity scores of each candidate video clip after re-aggregation.
[0012] According to a hierarchical video search ranking optimization method provided by the present invention, the step of determining the model weight of each visual language model based on the semantic matching accuracy of each candidate video segment specifically includes: For each candidate video clip, calculate the product of the initial similarity score corresponding to each visual language model and the semantic matching accuracy of the candidate video clip as the modified similarity score of each candidate video clip in each visual language model; The inverse of the standard deviation of the corrected similarity scores of all candidate video clips in each visual language model is used as the model weight of each visual language model.
[0013] The present invention also provides a hierarchical video search ranking optimization device, comprising: An initial ranking generation module is used to generate an initial similarity ranking based on the similarity scores between the query statement and the candidate video clips in descending order; A coarse-grained screening module is configured to use a multimodal large language model to semantically annotate a preset number of candidate video clips ranked before the initial similarity, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing results of the query statement; A fine-grained screening module is used to verify the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; The initial ranking update module is used to update the initial similarity score using the most difficult semantics and the easiest semantics to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and generate search results for the video clip based on the updated initial similarity score.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-described hierarchical video search ranking optimization methods is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for optimizing hierarchical video search ranking as described above is implemented.
[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned hierarchical video search ranking optimization methods.
[0017] The hierarchical video search ranking optimization method and device provided by the present invention semantically annotates candidate video clips through a multimodal large language model, and obtains the easiest and most difficult semantics by screening the accuracy of each basic semantic of the query statement, thereby effectively mining the correct semantic information implicit in the negative samples and providing fine-grained supervision signals for video ranking optimization; on this basis, the most difficult and easiest semantics are used as pseudo-queries to enhance the discrimination ability of difficult semantics and suppress the interference of easily confused semantics, thereby improving the discrimination between positive and negative samples, significantly improving the long-tail semantic matching problem, and ultimately obtaining more accurate video clip ranking results. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is one of the flow charts of the hierarchical video search ranking optimization method provided by the present invention; Figure 2 This is the second flow chart of the hierarchical video search ranking optimization method provided by the present invention; Figure 3 It is a structural diagram of the hierarchical video search ranking optimization device provided by the present invention; Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0021] First, let’s introduce the following contents: Video retrieval technology involves video understanding on the one hand and retrieval technology on the other.
[0022] Early video understanding technologies focused on unimodal feature extraction. However, film narratives are highly context-dependent, and a single modality, such as visual or text, struggles to capture complex interactions. For example, character relationship recognition relies on cross-modal cues such as dialogue text, character age, and scene objects, while traditional graph matching methods model affinity solely through face-name co-occurrence frequency, ignoring the synergy between conversational context and scene semantics. Furthermore, while location recognition fuses global and local features through a multi-resolution network, it fails to incorporate character behavior dynamics, such as the differences between scenes like "conversation in an office" and "chasing outdoors."
[0023] Existing methods are particularly deficient in multimodal alignment and long-term temporal modeling. While two-stream networks and Transformer architectures have improved short-term action recognition, they still lack effective modeling of the evolution of character relationships across scenarios. For example, while single-slice multimodal feature concatenation combined with visual and character tracking can capture instantaneous interactions, it struggles to correlate behavioral patterns across multiple scenarios. Furthermore, the generation of knowledge graphs relies on manual rules and cannot automatically construct fine-grained semantic relationships, such as "emotional conflict."
[0024] In terms of retrieval technology, live video retrieval tasks require processing queries containing complex semantics such as people, actions, and locations. However, traditional retrieval models face the problem of semantic combination explosion. While methods based on pre-trained visual language models can calculate text-video similarity, the heterogeneity of different pre-trained visual language models limits their fusion effectiveness. For example, the BLIP model excels at character semantics, while CLIP excels at scene matching. Simple average weighting strategies cannot dynamically adapt to query requirements.
[0025] A more fundamental challenge lies in the inefficiency of the feedback mechanism. Traditional relevance feedback relies on manual labeling of positive and negative samples, which is time-consuming and overlooks some correct semantics in negative samples. Research has shown that a large number of negative samples contain only one incorrect underlying semantic meaning. For example, when searching for "children climbing," a negative sample may correctly contain "children" but incorrectly identify "climbing action." Existing methods only optimize weights using positive samples and fail to exploit the implicit semantic supervision signals in negative samples.
[0026] In recent years, the rise of multimodal large language models has provided a new approach to solving the above problems. Based on the multimodal large language model's ability to automatically parse complex semantics, the hierarchical video search ranking optimization method of the present invention is proposed.
[0027] The following combination Figure 1 and Figure 2 The hierarchical video search ranking optimization method of the present invention is introduced. Figure 1As shown, including: Step 101: Generate an initial similarity ranking based on the similarity scores between the query and the candidate video clips in descending order; The query statement is the input search statement, which represents the video content that is expected to be retrieved.
[0028] The candidate video segments may be video segments in a collection of shot slices of several videos, or may be several independent video segments.
[0029] For example, when the application scenario of video search is to retrieve a segment corresponding to a query statement in a complete video, the complete video can be sliced based on the shots, and the obtained multiple video slices are used as candidate video segments.
[0030] For example, when the application scenario of video search is to retrieve the video corresponding to the query statement in multiple different videos, the multiple different videos can be directly used as candidate video segments; in addition, the multiple different videos can be sliced separately, and a mapping relationship table between each video and the slice is constructed, and all the slices divided from the multiple different videos are used as candidate video segments.
[0031] Based on this, the similarity between the query and the candidate video segment is determined. Optionally, the similarity between the query and the candidate video segment can be directly used as the similarity score for the candidate video segment. Because the query and the candidate video segment have features from different modalities, the query and the candidate video segment can optionally be mapped to the same modality before determining their similarity.
[0032] In a feasible implementation, a model in the field of image captioning, such as BLIP (Bootstrapping Language-Image Pre-training, a pre-training model for unified visual language understanding and generation), can be used to identify several key frame images in each candidate video clip, generate descriptive text corresponding to each candidate video clip, and then calculate the text feature similarity between the query statement and the descriptive text corresponding to each candidate video clip as the similarity score between the query statement and the candidate video clip.
[0033] In other feasible implementations, a text-based graph model such as Stable Diffusion can also be used to generate an image corresponding to the query text, and then obtain the image feature similarity between the image corresponding to the query text and several key frame images corresponding to each candidate video clip as the similarity score between the query statement and the candidate video clip.
[0034] The candidate video segments are arranged in descending order of similarity scores to generate an initial similarity ranking of the candidate video segments.
[0035] Step 102: semantically annotating a preset number of candidate video clips ranked top by initial similarity using a multimodal large language model, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing result of the query statement; Optionally, the preset number is an empirical value.
[0036] It is understandable that the selected candidate video clips with the highest initial similarity ranking are the candidate video clips with the highest correlation with the query statement. In this embodiment, the selected candidate video clips with the highest initial similarity ranking are constructed as a sample set. G , which is used to perform hierarchical screening based on basic semantics based on candidate video clips in the sample set.
[0037] The basic semantics are determined based on the query's semantic analysis results and represent the query's core semantics. Alternatively, semantic analysis can be achieved by segmenting the query. For example, if the query is "children climbing outdoors," the segmentation results are "children," "outdoors," and "climbing." These three words represent the three basic semantics corresponding to the query.
[0038] On this basis, we first generate coarse-grained feedback based on a multimodal large language model: we use a multimodal large language model such as Llama-3.2-Vision to semantically annotate each candidate video clip in the sample set.
[0039] Specifically, the query and each candidate video clip in the sample set are used as the input of the multimodal language, and the semantic annotation of each candidate video clip is performed through the multimodal language to output the binary annotation result. Among them, if the basic semantics of the candidate video clip is completely correct, the candidate video clip is classified into the positive sample set. , otherwise, divided into the negative sample set .
[0040] Step 103: Verify the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; It is understandable that the candidate video clips in the negative sample set are video clips that lack basic semantic understanding of the query statement. Based on this, further fine-grained feedback generation is performed to determine the easiest semantics to understand and the most difficult semantics to understand during the retrieval process.
[0041] Specifically, for each candidate video clip in the negative sample set, a multimodal large language model is used to verify the correctness of the basic semantics it contains one by one, and annotate them, and the accuracy of each basic semantics is calculated based on the annotation results.
[0042] Preferably, for each basic semantics, the ratio between the number of candidate video segments in the negative sample set that accurately match its semantics and the number of all candidate video segments in the negative sample set is calculated as the accuracy rate of each basic semantics; The basic semantics with the lowest accuracy is regarded as the most difficult semantics, and the basic semantics with the highest accuracy is regarded as the easiest semantics.
[0043] In a specific implementation, the query statement contains three basic semantics: a 、 b and c The negative sample set contains 30 candidate video clips. The correctness of the understanding of the three basic semantics of each candidate video clip is verified one by one. Taking "children", "outdoor" and "climbing" as the three basic semantics as an example, if the content of the candidate video clip includes "children" and "climbing", but the climbing scene is indoors, then the two basic semantics of "children" and "climbing" in the candidate video clip are marked as correct, and the basic semantics of "outdoor" are marked as wrong.
[0044] The accuracy of each basic semantics is counted separately. For example, for the basic semantics of "children", 25 out of 30 candidate video clips are marked as correct, so the accuracy of this basic semantics is 25 / 30.
[0045] On this basis, it is believed that the basic semantics with the highest accuracy is the easiest semantics , the most accurate basic semantics is the most difficult semantics .
[0046] Step 104 : Using the most difficult semantics and the easiest semantics to update the initial similarity score, so as to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and generating a search result of the video clip based on the updated initial similarity score.
[0047] Optionally, in order to optimize the initial similarity scores of candidate video clips using the most difficult semantics and easiest semantics obtained through screening, so as to strengthen the retrieval process's understanding of the most difficult semantics and suppress the interference of the easiest semantics, in this embodiment, the most difficult semantics and the easiest semantics are used as pseudo-queries, and the two pseudo-queries are used to search in all candidate video clips respectively, to obtain the similarity score of the most difficult semantics corresponding to each candidate video clip and the similarity score of the easiest semantics corresponding to each candidate video clip.
[0048] Optionally, the method for obtaining the similarity scores of the candidate video segments corresponding to the two pseudo queries is the same as the method for obtaining the similarity scores used in the initial ranking of each candidate video segment.
[0049] In a feasible implementation, the updated initial similarity score is as follows: ; Where, Indicates the m The updated initial similarity scores of candidate video clips, Indicates the m The initial similarity scores of candidate video clips, Indicates the m The similarity score of the candidate video clips when the most difficult semantics is used as a pseudo query; Indicates the m The similarity score of each candidate video clip when the easiest semantics is used as a pseudo query; and is the dynamic adjustment coefficient.
[0050] Through this approach, the updated initial similarity score strengthens the scores of the most difficult semantics and suppresses the scores of the easiest semantics. Based on this, the candidate video segments are sorted in descending order based on the updated initial similarity scores to obtain the final output video segment search ranking, which serves as the optimized video segment search results.
[0051] In other feasible implementations, the similarity scores of the two pseudo queries can be used as weight coefficients to directly weight the initial similarity scores of the candidate video segments, wherein the similarity score of the most difficult semantics is used as a positive reinforcement factor, and the similarity score of the easiest semantics is used as a negative inhibition factor.
[0052] The present invention uses a multimodal large language model to semantically annotate candidate video clips, and obtains the easiest and most difficult semantics by screening the accuracy of each basic semantic of the query statement, thereby effectively mining the correct semantic information implicit in the negative samples and providing fine-grained supervision signals for video ranking optimization; on this basis, the most difficult and easiest semantics are used as pseudo-queries to enhance the discrimination ability of difficult semantics and suppress the interference of easily confused semantics, thereby improving the discrimination between positive and negative samples, significantly improving the long-tail semantic matching problem, and ultimately obtaining more accurate video clip ranking results.
[0053] In the hierarchical video search ranking optimization method of the present invention, before the step of generating an initial similarity ranking based on the descending order of similarity scores between the query statement and the candidate video segments, the method further includes: Extracting multimodal features of the candidate video clips using multiple visual language models; In this embodiment, several visual language models are used to extract multimodal features of candidate segments respectively.
[0054] Optionally, the visual language model includes one or more of CLIP, BLIP, BLIP-2, SLIP, and LaCLIP, preferably at least three.
[0055] By inputting candidate video clips into a visual language model such as CLIP, we can directly extract the cross-modal embedding vectors of text and video clips as multimodal features of the candidate video clips, which are used to calculate the similarity score between the candidate video clips and the query statement.
[0056] It is understandable that if the selected visual language model is n Then a candidate video segment is extracted n Multimodal features.
[0057] Generate a query image of the query statement using an image generation model, and extract and aggregate features of the query image and query text to obtain multimodal features of the query statement; Since in this embodiment, multimodal features that fuse text and video are extracted from the candidate video segments, it is necessary to obtain the multimodal features of the query sentence to better calculate its similarity score with each candidate video segment.
[0058] Optionally, a Stable Diffusion model is used to generate a query image corresponding to the query statement, and image features of the query image and text features of the query statement are extracted respectively. After feature aggregation, multimodal features of the query statement are obtained.
[0059] Optionally, in this embodiment, the feature aggregation method can be direct splicing or weighted fusion, as long as it can aggregate image features and text features.
[0060] Based on the similarity between the multimodal features of each candidate video segment and the multimodal features of the query sentence, a similarity score between the query sentence and the candidate video segment is determined.
[0061] Based on the above, for each candidate video clip, when only one visual language model is used, the similarity between its corresponding multimodal features and the multimodal features corresponding to the query statement can be directly calculated, and the similarity calculation value can be used as the similarity score of the candidate video clip.
[0062] When multiple visual language models are used, that is, each candidate video clip corresponds to multiple multimodal features, the similarity value between each of the multiple multimodal features of each candidate video clip and the multimodal features of the query statement can be calculated separately to obtain multiple similarity values corresponding to each candidate video clip.
[0063] On this basis, the average of multiple similarity values can be directly calculated as the similarity score between each candidate video clip and the query statement; the performance of different visual language models can be compared in advance, the corresponding weight can be determined in advance for each visual language model, and then the weighted average of multiple similarity values can be calculated as the similarity score between each candidate video clip and the query statement; the visual language model with the highest stability can also be screened out based on the calculation results of multiple candidate video clips under multiple visual language models, and its corresponding similarity value can be used as the similarity score between each candidate video clip and the query statement.
[0064] In this embodiment, there is no limitation on the method of determining the similarity score between each candidate video segment and the query statement based on multiple similarity values of each candidate video segment, and those skilled in the art can implement it.
[0065] The present invention obtains multimodal features of candidate video clips and query statements respectively, calculates the similarity values between the multimodal features to determine the similarity score between each candidate video clip and the query statement, and makes full use of the information of different modes of the candidate video clips, thereby obtaining a more accurate initial similarity ranking.
[0066] In the hierarchical video search ranking optimization method of the present invention, the step of generating an initial similarity ranking based on the descending order of similarity scores between the query statement and the candidate video clips specifically includes: The average fusion method is used to calculate the similarity scores of several visual language models for each candidate video segment as the candidate similarity score of each candidate video segment; An initial similarity ranking is generated based on the descending order of the candidate similarity scores.
[0067] In this embodiment, when using several temporal language models to extract the multimodal features of each candidate video clip, the similarity values between the several multimodal features of each candidate video clip and the multimodal features of the query statement are calculated respectively as the similarity scores of the several visual language models for each candidate video clip.
[0068] For any candidate video clip, the average fusion method is used to calculate the average similarity score of several visual language models as the candidate similarity score of the candidate video clip.
[0069] An initial similarity ranking of the candidate video segments is generated based on the descending order of the candidate similarity scores of each candidate video segment.
[0070] In the above manner, when one or more visual language models are used, an initial similarity ranking can be directly determined to perform preliminary screening of candidate video segments and obtain multiple candidate video segments that are most relevant to the query statement.
[0071] In the hierarchical video search ranking optimization method of the present invention, the step of generating search results for video clips based on the updated initial similarity scores specifically includes: Based on the semantic matching accuracy of each candidate video segment, the similarity score of each visual language model for each candidate video segment is modified, so as to determine the model weight of each visual language model based on the modified similarity score; Although multiple visual language models are used in the process of determining the initial similarity ranking, it is possible to integrate the characteristics of multiple different visual language models to obtain a more accurate initial similarity ranking, but in order to generate more accurate video clip search results, this embodiment uses weighted aggregation through the similarity value rankings of multiple visual language models to obtain the final search results.
[0072] Therefore, the model weights corresponding to each visual language model need to be determined first. At the same time, in order to fully utilize the semantic information of the query and the candidate video clips, this embodiment first corrects the similarity scores corresponding to each visual language model based on the semantic matching accuracy before calculating the model weights.
[0073] Among them, the semantic matching accuracy of the candidate video clips is the basic semantic correctness of each candidate video clip. For example, when the query sentence has three basic semantics, if the candidate video clip contains two of them based on semantics, the semantic matching accuracy of the candidate video clip is 2 / 3.
[0074] Among them, the similarity score corresponding to the visual language model That is n The visual language model is used to m The similarity score of candidate video clips, that is, m The candidate video segments use n The visual language model extracts multimodal features and calculates the similarity between the multimodal features and the multimodal features of the query sentence as the first n The visual language model is used to m The similarity scores of the candidate video clips.
[0075] Use m The semantic matching accuracy of candidate video clips is Optimize and get n The visual language model is used to mThe corrected similarity scores of candidate video clips are obtained. Based on this, the model weight of each visual language model is determined based on the corrected similarity scores. .
[0076] Optionally, based on the modified The stability of each visual language model is calculated, and more weight is given to visual language models with higher stability.
[0077] For each candidate video clip, the updated initial similarity scores determined based on each visual language model are weightedly fused using the model weight of each visual language model to obtain the re-aggregated similarity score of each candidate video clip; It is understandable that, since multiple visual language models are used in this embodiment to calculate the initial similarity score, the updated initial similarity score can be calculated by the following formula in this embodiment: ; Where, Indicates the n The visual language model is used to m The updated initial similarity scores of candidate video clips, Indicates the n The visual language model is used to m The initial similarity scores of candidate video clips, Indicates the n The visual language model is used to m The similarity scores of candidate video clips under the most difficult semantics as pseudo query; Indicates the n The visual language model is used to m The similarity scores of candidate video clips under the easiest semantics as pseudo query.
[0078] On this basis, the model weights of each visual language model are used to Perform weighted fusion to obtain the similarity score of each candidate video segment after re-aggregation: ; Where, Indicates the m The similarity score after the candidate video clips are re-aggregated, Indicates the n The weights of the visual language model, N Indicates the number of visual language models.
[0079] The search results of the video clips are generated based on the similarity scores of each candidate video clip after re-aggregation.
[0080] The candidate video clips are arranged in descending order according to the similarity scores after re-aggregation to obtain the final search results of the video clips.
[0081] In the hierarchical video search ranking optimization method of the present invention, the step of determining the model weight of each visual language model based on the semantic matching accuracy of each candidate video segment specifically includes: For each candidate video clip, calculate the product of the initial similarity score corresponding to each visual language model and the semantic matching accuracy of the candidate video clip as the modified similarity score of each candidate video clip in each visual language model; In this embodiment, for the mth candidate video segment, the product of the initial similarity score corresponding to each visual language model and the semantic matching accuracy of the candidate video segment is directly calculated as the modified similarity score of the candidate video segment under each visual language model: ; Where, Indicates the n The visual language model is used to m The corrected similarity scores of candidate video clips, Represents the semantic matching accuracy of the candidate video segment.
[0082] The inverse of the standard deviation of the corrected similarity scores of all candidate video clips in each visual language model is used as the model weight of each visual language model.
[0083] Specifically, the model weights are calculated as follows: ; Where, Indicates the n The model weight of a visual language model, G represents the sample set.
[0084] Through the above method, the higher the stability of the model and the smaller the standard deviation, the higher the corresponding model weight.
[0085] The present invention calculates the reliability weight of each visual language model by combining the modified similarity score and dynamically evaluates the overall distribution consistency of positive and negative samples to achieve intelligent weighted fusion of multi-model results. A complete ranking optimization process is as follows: Figure 2 shown.
[0086] In a specific embodiment, the ranking optimization method of the present application that integrates hierarchical screening and weighted fusion (ours in Table 1) is experimentally compared with other models in the field, IRA (Ranking Aggregation with Interactive Feedback for Collaborative Person Re-identification) and QI-IRA (Quantum-Inspired Interactive Ranking Aggregation for Person Re-identification). Specifically, based on the video sets IACC.3, V3C1, and V3C2 officially provided by TRECVID, the movies in the video set (corresponding to TV16-TV23 in Table 1 below) are sliced and segmented based on shots. A slice shot table V is generated for each movie to represent the mapping relationship between the movie and its slice. Further, combined with the scene information provided in the video set, a movie-slice and scene-slice shot mapping table is generated for TV16-TV23 respectively. The generated slice shots are used as candidate video clip sets. A query statement is used to search the candidate video sets corresponding to TV16-TV23. The results are shown in Table 1 below: Table 1
[0087] The values in Table 1 are the accuracy rates of the retrieval results obtained by using the three methods respectively to search the candidate video clip sets corresponding to each movie based on the same query statement. It can be seen that the method of the present invention has better performance in different candidate video clip sets.
[0088] The hierarchical video search ranking optimization device provided by the present invention is described below. The hierarchical video search ranking optimization device described below and the hierarchical video search ranking optimization method described above can be referenced to each other.
[0089] like Figure 3 As shown, the hierarchical video search ranking optimization device of the present invention includes an initial ranking generation module 301, a coarse-grained screening module 302, a fine-grained screening module 303 and an initial ranking update module 304; An initial ranking generating module 301 is used to generate an initial similarity ranking based on the similarity scores between the query statement and the candidate video segments in descending order; The query statement is the input search statement, which represents the video content that is expected to be retrieved.
[0090] The candidate video segments may be video segments in a collection of shot slices of several videos, or may be several independent video segments.
[0091] For example, when the application scenario of video search is to retrieve a segment corresponding to a query statement in a complete video, the complete video can be sliced based on the shots, and the obtained multiple video slices are used as candidate video segments.
[0092] For example, when the application scenario of video search is to retrieve the video corresponding to the query statement in multiple different videos, the multiple different videos can be directly used as candidate video segments; in addition, the multiple different videos can be sliced separately, and a mapping relationship table between each video and the slice is constructed, and all the slices divided from the multiple different videos are used as candidate video segments.
[0093] Based on this, the similarity between the query and the candidate video segment is determined. Optionally, the similarity between the query and the candidate video segment can be directly used as the similarity score for the candidate video segment. Because the query and the candidate video segment have features from different modalities, the query and the candidate video segment can optionally be mapped to the same modality before determining their similarity.
[0094] In a feasible implementation, a model in the field of image captioning, such as BLIP (Bootstrapping Language-Image Pre-training, a pre-training model for unified visual language understanding and generation), can be used to identify several key frame images in each candidate video clip, generate descriptive text corresponding to each candidate video clip, and then calculate the text feature similarity between the query statement and the descriptive text corresponding to each candidate video clip as the similarity score between the query statement and the candidate video clip.
[0095] In other feasible implementations, a text-based graph model such as Stable Diffusion can also be used to generate an image corresponding to the query text, and then obtain the image feature similarity between the image corresponding to the query text and several key frame images corresponding to each candidate video clip as the similarity score between the query statement and the candidate video clip.
[0096] The candidate video segments are arranged in descending order of similarity scores to generate an initial similarity ranking of the candidate video segments.
[0097] A coarse-grained screening module 302 is configured to semantically annotate a preset number of candidate video segments ranked first by initial similarity using a multimodal large language model, so as to determine candidate video segments with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing result of the query statement; Optionally, the preset number is an empirical value.
[0098] It is understandable that the selected candidate video clips with the highest initial similarity ranking are the candidate video clips with the highest correlation with the query statement. In this embodiment, the selected candidate video clips with the highest initial similarity ranking are constructed as a sample set. G , which is used to perform hierarchical screening based on basic semantics based on candidate video clips in the sample set.
[0099] The basic semantics are determined based on the query's semantic analysis results and represent the query's core semantics. Alternatively, semantic analysis can be achieved by segmenting the query. For example, if the query is "children climbing outdoors," the segmentation results are "children," "outdoors," and "climbing." These three words represent the three basic semantics corresponding to the query.
[0100] On this basis, we first generate coarse-grained feedback based on a multimodal large language model: we use a multimodal large language model such as Llama-3.2-Vision to semantically annotate each candidate video clip in the sample set.
[0101] Specifically, the query and each candidate video clip in the sample set are used as the input of the multimodal language, and the semantic annotation of each candidate video clip is performed through the multimodal language to output the binary annotation result. Among them, if the basic semantics of the candidate video clip is completely correct, the candidate video clip is classified into the positive sample set. , otherwise, divided into the negative sample set .
[0102] A fine-grained screening module 303 is configured to verify the accuracy of basic semantics for each candidate video segment in the negative sample set, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; It is understandable that the candidate video clips in the negative sample set are video clips that lack basic semantic understanding of the query statement. Based on this, further fine-grained feedback generation is performed to determine the easiest semantics to understand and the most difficult semantics to understand during the retrieval process.
[0103] Specifically, for each candidate video clip in the negative sample set, a multimodal large language model is used to verify the correctness of the basic semantics it contains one by one, and annotate them, and the accuracy of each basic semantics is calculated based on the annotation results.
[0104] Preferably, for each basic semantics, the ratio between the number of candidate video segments in the negative sample set that accurately match its semantics and the number of all candidate video segments in the negative sample set is calculated as the accuracy rate of each basic semantics; The basic semantics with the lowest accuracy is regarded as the most difficult semantics, and the basic semantics with the highest accuracy is regarded as the easiest semantics.
[0105] In a specific implementation, the query statement contains three basic semantics: a 、 b and c The negative sample set contains 30 candidate video clips. The correctness of the understanding of the three basic semantics of each candidate video clip is verified one by one. Taking "children", "outdoor" and "climbing" as the three basic semantics as an example, if the content of the candidate video clip includes "children" and "climbing", but the climbing scene is indoors, then the two basic semantics of "children" and "climbing" in the candidate video clip are marked as correct, and the basic semantics of "outdoor" are marked as wrong.
[0106] The accuracy of each basic semantics is counted separately. For example, for the basic semantics of "children", 25 out of 30 candidate video clips are marked as correct, so the accuracy of this basic semantics is 25 / 30.
[0107] On this basis, it is believed that the basic semantics with the highest accuracy is the easiest semantics , the most accurate basic semantics is the most difficult semantics .
[0108] The initial ranking updating module 304 is configured to update the initial similarity score using the most difficult semantics and the easiest semantics to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and generate a search result for the video clip based on the updated initial similarity score.
[0109] Optionally, in order to optimize the initial similarity scores of candidate video clips using the most difficult semantics and easiest semantics obtained through screening, so as to strengthen the retrieval process's understanding of the most difficult semantics and suppress the interference of the easiest semantics, in this embodiment, the most difficult semantics and the easiest semantics are used as pseudo-queries, and the two pseudo-queries are used to search in all candidate video clips respectively, to obtain the similarity score of the most difficult semantics corresponding to each candidate video clip and the similarity score of the easiest semantics corresponding to each candidate video clip.
[0110] Optionally, the method for obtaining the similarity scores of the candidate video segments corresponding to the two pseudo queries is the same as the method for obtaining the similarity scores used in the initial ranking of each candidate video segment.
[0111] In a feasible implementation, the updated initial similarity score is as follows: ; Where, Indicates the m The updated initial similarity scores of candidate video clips, Indicates the m The initial similarity scores of candidate video clips, Indicates the m The similarity score of the candidate video clips when the most difficult semantics is used as a pseudo query; Indicates the m The similarity score of each candidate video clip when the easiest semantics is used as a pseudo query; and is the dynamic adjustment coefficient.
[0112] Through this approach, the updated initial similarity score strengthens the scores of the most difficult semantics and suppresses the scores of the easiest semantics. Based on this, the candidate video segments are sorted in descending order based on the updated initial similarity scores to obtain the final output video segment search ranking, which serves as the optimized video segment search results.
[0113] In other feasible implementations, the similarity scores of the two pseudo queries can be used as weight coefficients to directly weight the initial similarity scores of the candidate video segments, wherein the similarity score of the most difficult semantics is used as a positive reinforcement factor, and the similarity score of the easiest semantics is used as a negative inhibition factor.
[0114] The present invention uses a multimodal large language model to semantically annotate candidate video clips, and obtains the easiest and most difficult semantics by screening the accuracy of each basic semantic of the query statement, thereby effectively mining the correct semantic information implicit in the negative samples and providing fine-grained supervision signals for video ranking optimization; on this basis, the most difficult and easiest semantics are used as pseudo-queries to enhance the discrimination ability of difficult semantics and suppress the interference of easily confused semantics, thereby improving the discrimination between positive and negative samples, significantly improving the long-tail semantic matching problem, and ultimately obtaining more accurate video clip ranking results.
[0115] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4As shown, the electronic device may include: a processor (processor) 410, a communication interface (Communications Interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the hierarchical video search ranking optimization method, which includes: generating an initial similarity ranking based on the similarity scores of the query statement and the candidate video clips in descending order; using a multimodal large language model to semantically annotate a preset number of candidate video clips before the initial similarity ranking, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, otherwise, they are determined as a negative sample set, wherein the basic semantics are determined based on the semantic parsing results of the query statement; verifying the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics, and the basic semantics with the highest accuracy as the easiest semantics; using the most difficult semantics and the easiest semantics to update the initial similarity score, so as to enhance the recognition ability of the most difficult semantics and suppress the interference of the easiest semantics, and generate search results for the video clip based on the updated initial similarity score.
[0116] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0117] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the hierarchical video search ranking optimization method provided by the above methods, which includes: generating an initial similarity ranking based on the similarity scores of the query statement and the candidate video clips in descending order; using a multimodal large language model to semantically annotate a preset number of candidate video clips before the initial similarity ranking, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, otherwise, they are determined as a negative sample set, wherein the basic semantics are determined based on the semantic parsing results of the query statement; verifying the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; using the most difficult semantics and the easiest semantics to update the initial similarity score to enhance the recognition ability of the most difficult semantics and suppress the interference of the easiest semantics, and generating search results for the video clips based on the updated initial similarity score.
[0118] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the hierarchical video search ranking optimization method provided by the above-mentioned methods, the method comprising: generating an initial similarity ranking based on the similarity scores of the query statement and the candidate video clips in descending order; using a multimodal large language model to semantically annotate a preset number of candidate video clips before the initial similarity ranking, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, otherwise, they are determined as a negative sample set, wherein the basic semantics are determined based on the semantic parsing results of the query statement; verifying the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics, and the basic semantics with the highest accuracy as the easiest semantics; using the most difficult semantics and the easiest semantics to update the initial similarity score, so as to enhance the recognition ability of the most difficult semantics and suppress the interference of the easiest semantics, and generate search results for the video clips based on the updated initial similarity score.
[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0120] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A hierarchical video search ranking optimization method, characterized in that: include: Generate an initial similarity ranking based on the similarity scores between the query and the candidate video clips from large to small; Using a multimodal large language model to semantically annotate a preset number of candidate video clips ranked before the initial similarity, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing result of the query statement; Verifying the accuracy of basic semantics for each candidate video segment in the negative sample set, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; The initial similarity score is updated using the most difficult semantics and the easiest semantics to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and a search result of the video clip is generated based on the updated initial similarity score.
2. The hierarchical video search ranking optimization method according to claim 1, characterized in that: Before the step of generating an initial similarity ranking based on the similarity scores between the query and the candidate video clips in descending order, the method further includes: Extracting multimodal features of the candidate video clips using multiple visual language models; Generate a query image of the query statement using an image generation model, and extract and aggregate features of the query image and query text to obtain multimodal features of the query statement; Based on the similarity between the multimodal features of each candidate video segment and the multimodal features of the query sentence, a similarity score between the query sentence and the candidate video segment is determined.
3. The hierarchical video search ranking optimization method according to claim 2, characterized in that: The step of generating an initial similarity ranking based on the descending order of similarity scores between the query statement and the candidate video clips specifically includes: The average fusion method is used to calculate the similarity scores of several visual language models for each candidate video segment as the candidate similarity score of each candidate video segment; An initial similarity ranking is generated based on the descending order of the candidate similarity scores.
4. The hierarchical video search ranking optimization method according to claim 1, characterized in that: The step of verifying the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics specifically includes: For each basic semantics, calculate the ratio between the number of candidate video segments in the negative sample set that accurately matches its semantics and the number of all candidate video segments in the negative sample set as the accuracy rate of each basic semantics; The basic semantics with the lowest accuracy is regarded as the most difficult semantics, and the basic semantics with the highest accuracy is regarded as the easiest semantics.
5. The hierarchical video search ranking optimization method according to claim 2, characterized in that: The step of generating search results for video clips based on the updated initial similarity scores specifically includes: Based on the semantic matching accuracy of each candidate video segment, the similarity score of each visual language model for each candidate video segment is modified, so as to determine the model weight of each visual language model based on the modified similarity score; For each candidate video clip, the updated initial similarity scores determined based on each visual language model are weightedly fused using the model weight of each visual language model to obtain the re-aggregated similarity score of each candidate video clip; The search results of the video clips are generated based on the similarity scores of each candidate video clip after re-aggregation.
6. The hierarchical video search ranking optimization method according to claim 5, characterized in that: The step of determining the model weight of each visual language model based on the semantic matching accuracy of each candidate video segment specifically includes: For each candidate video clip, calculate the product of the initial similarity score corresponding to each visual language model and the semantic matching accuracy of the candidate video clip as the modified similarity score of each candidate video clip in each visual language model; The inverse of the standard deviation of the corrected similarity scores of all candidate video clips in each visual language model is used as the model weight of each visual language model.
7. A hierarchical video search ranking optimization device, characterized in that: include: An initial ranking generation module is used to generate an initial similarity ranking based on the similarity scores between the query statement and the candidate video clips in descending order; A coarse-grained screening module is configured to use a multimodal large language model to semantically annotate a preset number of candidate video clips ranked before the initial similarity, so as to determine the candidate video clips with completely correct basic semantics as a positive sample set, and otherwise as a negative sample set, wherein the basic semantics are determined based on the semantic parsing results of the query statement; A fine-grained screening module is used to verify the accuracy of the basic semantics of the candidate video clips in the negative sample set one by one, so as to obtain the basic semantics with the lowest accuracy as the most difficult semantics and the basic semantics with the highest accuracy as the easiest semantics; The initial ranking update module is used to update the initial similarity score using the most difficult semantics and the easiest semantics to enhance the recognition capability of the most difficult semantics and suppress the interference of the easiest semantics, and generate search results for the video clip based on the updated initial similarity score.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the hierarchical video search ranking optimization method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the hierarchical video search ranking optimization method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the hierarchical video search ranking optimization method according to any one of claims 1 to 6 is implemented.