Multi-mode combined video retrieval method and device

By extracting common and differential features of text and visual information, the problem of inaccurate user intention understanding in multimodal combined video retrieval is solved, and more accurate video retrieval results are achieved.

CN120256679APending Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510150950.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multimodal combined video retrieval technology is difficult to accurately understand the user's true intentions, and the search results are not accurate enough because the independent comparison of text, images and video does not take into account the correlation between multimodal needs.

Method used

By extracting the text features in the text information and the visual semantic features in the visual information, obtaining common and differential features, querying matching results from the video information database with visual features, and using the differential features to filter out video search results that meet the user's intentions.

Benefits of technology

It improves the accuracy of multimodal combined video retrieval and meets the user's personalized and flexible search needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256679A_ABST
    Figure CN120256679A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-mode combined video retrieval method and device. The method comprises the following steps: acquiring text information and visual information; extracting character features from the character information; extracting visual features from the visual information; extracting visual semantic features from the visual information according to the character features; extracting common features and difference features between the character features and the visual semantic features from the character features and the visual semantic features; querying a preset video information base according to the visual features and the common features to obtain a plurality of video retrieval results matched with the visual features and the common features; and screening the plurality of video retrieval results according to the difference characteristics to obtain a screened video retrieval result. According to the method, effective information of multi-modal information can be fused, the real intention of the user can be accurately understood, and the accuracy of multi-modal combined video retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of information retrieval, and in particular, to a multi-modal combined video retrieval method and apparatus. Background Art

[0002] With the development of video retrieval and artificial intelligence technologies, multi-modal combined video retrieval technologies are also evolving continuously. Users can use text, video, and / or images, etc. as combined retrieval conditions to retrieve matching videos, providing a more flexible and personalized retrieval experience. However, when retrieving using combined retrieval conditions, the retrieval process still compares the similarities of text, images, and videos separately. The three conditions are independent of each other, without considering the relevance between the multi-modal needs of users, making it difficult to accurately understand the true intentions of users, and the retrieval results are not accurate enough. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present application is to propose a multi-modal combined video retrieval method and apparatus to solve the problem of inaccurate retrieval results.

[0004] Based on the above purpose, the embodiments of the present application provide a multi-modal combined video retrieval method, including:

[0005] Obtain text information and visual information;

[0006] Extract text features from the text information;

[0007] Extract visual features from the visual information;

[0008] According to the text features, extract visual semantic features from the visual information;

[0009] Extract the common features and different features between the text features and visual semantic features from the text features and visual semantic features;

[0010] According to the visual features and common features, query a preset video information database to obtain multiple video retrieval results that match the visual features and common features;

[0011] Screen the multiple video retrieval results according to the different features to obtain the screened video retrieval results.

[0012] Optionally, extracting text features from the text information includes:

[0013] Use a preset first large language model to extract descriptive features associated with the visual information and text semantic features other than the visual information from the text information;

[0014] Extract visual semantic features from the visual information according to the text features, including:

[0015] Use a preset visual language model to extract visual semantic features from the visual information according to the descriptive features.

[0016] Optionally, extract the common features and different features between the text features and the visual semantic features, including:

[0017] Use a preset second large language model to extract the common features and different features from the text semantic features and the visual semantic features.

[0018] Optionally, the video information library includes multiple video files and the corresponding video features and semantic features of each video file;

[0019] Query a preset video information library according to the visual features and the common features to obtain multiple video retrieval results that match the visual features and the common features, including:

[0020] Match the visual features with the video features of each video file to obtain a video matching result;

[0021] Select a predetermined first number of first video retrieval results in descending order of the video matching results;

[0022] Match the common features with the semantic features of each video file to obtain a semantic matching result;

[0023] Select a predetermined second number of second video retrieval results in descending order of the semantic matching results;

[0024] For the first video retrieval results and the second video retrieval results, calculate a matching score according to the matching degree and the corresponding weight of the video matching, and the matching degree and the corresponding weight of the semantic matching.

[0025] Re - sort the first video retrieval results and the second video retrieval results in descending order of the matching score to obtain the video retrieval results after sorting by matching degree.

[0026] Optionally, screen the multiple video retrieval results according to the different features to obtain the screened video retrieval results, including:

[0027] Sort the video retrieval results including the corresponding text information in the different features forward among the video retrieval results after sorting by matching degree, and sort the video retrieval results not including the corresponding text information in the different features backward to obtain the re - sorted video retrieval results;

[0028] Based on the re-ordered video retrieval results, select a predetermined third quantity of video retrieval results as the filtered video retrieval results.

[0029] Optionally, for the first video retrieval result and the second video retrieval result, calculate a matching score according to the matching degree and corresponding weight of video matching, and the matching degree and corresponding weight of semantic matching, including:

[0030] According to the semantic matching degree and video matching degree or image matching degree of each video retrieval result, query a preset weight table to determine the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree, or the image weight corresponding to the image matching degree;

[0031] Calculate the matching score of the video retrieval result according to the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree, or the image weight corresponding to the image matching degree.

[0032] Optionally, the method further includes:

[0033] Obtain a video file;

[0034] Extract video features and semantic features from the video file;

[0035] Based on the video file and the corresponding video features and semantic features, construct the video information library.

[0036] Optionally, the video file includes a long video file; after obtaining the video file, it further includes:

[0037] Perform scene segmentation on the long video file to obtain video segments corresponding to multiple scenes;

[0038] Extracting video features and semantic features from the video file includes:

[0039] Extract the corresponding video features and semantic features from the video segments corresponding to each scene respectively;

[0040] Based on the video file and the extracted video features and semantic features, constructing the video information library includes:

[0041] Based on the video segments corresponding to multiple scenes of the video file and the extracted video features and semantic features, construct a scene-level video information library.

[0042] Optionally, the video features include picture features and fingerprint features, the fingerprint features include key frames, texture features, color features, spatial relationship features, and voiceprint features, and the semantic features include audio features, video text features, and emotion features.

[0043] The embodiment of the present application also provides a multimodal combined video retrieval device, including:

[0044] An acquisition module, configured to acquire text information and visual information;

[0045] A first extraction module, configured to extract text features from the text information;

[0046] A second extraction module, configured to extract visual features from the visual information;

[0047] A third extraction module, configured to extract visual semantic features from the visual information according to the text features;

[0048] A fourth extraction module, configured to extract common features and differential features between the text features and the visual semantic features from the text features and the visual semantic features;

[0049] A retrieval module, configured to query a preset video information database according to the visual features and the common features, and obtain multiple video retrieval results that match the visual features and the common features;

[0050] A screening module, configured to screen the multiple video retrieval results according to the differential features to obtain the screened video retrieval results.

[0051] As can be seen from the above, the multimodal combined video retrieval method and device provided by the embodiment of the present application acquire text information and visual information, extract text features from the text information, extract visual features from the visual information, extract visual semantic features from the visual information according to the text features, extract common features and differential features between the text features and the visual semantic features from the text features and the visual semantic features, query the video information database according to the visual features and the common features to obtain multiple matching video retrieval results, and screen the multiple video retrieval results according to the differential features to obtain the screened video retrieval results. The present application can integrate the effective information of multimodal information, accurately understand the true intention of the user, and improve the accuracy of multimodal combined video retrieval. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present application;

[0054] Figure 2Schematic diagram of the method flow of another embodiment of the present application;

[0055] Figure 3 Schematic diagram of the method flow for constructing a video information library and retrieval in an embodiment of the present application;

[0056] Figure 4 Schematic diagram of the video scene generation method in an embodiment of the present application;

[0057] Figure 5 Block diagram of the device structure in an embodiment of the present application;

[0058] Figure 6 Block diagram of the electronic device structure in an embodiment of the present application. Detailed implementation manners

[0059] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0060] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0061] In the related art, when performing retrieval using combined retrieval conditions such as text, video, and / or images, there are deficiencies in semantic understanding and scene recognition, it is difficult to capture high-level semantic information in the video, various retrieval conditions are independently compared, and the effective information in multiple retrieval conditions is not integrated, it is difficult to accurately understand the true intention of the user, and the retrieval results are not accurate enough.

[0062] In view of this, an embodiment of the present application provides a multi-modal combined video retrieval method. For the combined retrieval condition including text and visual information input, visual semantic features can be extracted from the visual information through text features, the common features and differential features between the text semantic features and the visual semantic features are extracted, multiple matching video retrieval results are queried from the video information library according to the common features and the visual features extracted from the visual information, and then multiple video retrieval results that meet the user's true intention are screened out by using the differential features. By integrating the effective information between multi-modal information, the accuracy of the multi-modal combined retrieval results can be improved, and the personalized and flexible retrieval needs of users can be met.

[0063] Hereinafter, the technical solution of the present application will be further described in detail through specific embodiments.

[0064] As Figure 1 、 2 shown, an embodiment of the present application provides a multi-modal combined video retrieval method, including:

[0065] S101: Obtain text information and visual information;

[0066] In this embodiment, the user can input text information and visual information for multi-modal combined video retrieval. The visual information includes image information or video information, that is, the user can input text information and image information, or input text information and video information. Among them, the text information can be input through an input box on the interface, voice recognition of text, etc., and the image information and video information can be input through uploading, taking pictures, videos, etc. The specific input method is not limited.

[0067] S102: Extract text features from the text information;

[0068] In this embodiment, for the obtained text information, text features are extracted therefrom, and the method includes:

[0069] Use a preset first large language model to extract descriptive features associated with the visual information from the text information, as well as text semantic features other than the visual information.

[0070] In this embodiment, the text information and the visual information are input into a preset large language model, and the large language model is used to perform fine-grained parsing on the text information, and text features such as descriptive features and text semantic features are extracted therefrom. Among them, the descriptive features are the feature descriptions of the specific elements of the visual information in the text information, and the text semantic features include additional information other than the visual information, and this additional information reflects the user's additional needs other than the visual information, such as expectations of high-level abstract concepts such as ideal pictures and emotional atmospheres. Optionally, a deepseek large language model can be used to extract text features such as descriptive features and text semantic features from the text information.

[0071] For example, the text information input by the user is "I need a video of a couple walking in the forest in this style", and the uploaded video information shows a scene of a couple walking in the park (or, the uploaded image information includes a picture of a couple walking in the park); the descriptive features extracted from the text information by the deepseek large language model are "this style", and the text semantic features are "a couple walking in the forest".

[0072] S103: Extract visual features from the visual information;

[0073] In this embodiment, for the obtained visual information, visual features are extracted therefrom for subsequent retrieval and matching of video files from a preset video information library. Among them, the visual features include picture features, fingerprint features, etc. of the video or image.

[0074] S104: Extract visual semantic features from the visual information according to the text features;

[0075] In this embodiment, after extracting text features such as descriptive features from the text information, visual semantic features are extracted from the visual information according to the text features. The method includes:

[0076] Using a preset vision-language model, extract visual semantic features from the visual information according to the descriptive features.

[0077] In this embodiment, for the visual information, the descriptive features extracted from the text information and the visual information are input into a preset vision-language model, and the vision-language model is used to deeply analyze and understand the video information, and extract visual semantic features that can accurately locate the emphasis in the user's intention. Optionally, the CogVLM2 vision-language model can be used to extract visual semantic features based on the descriptive features and the visual information.

[0078] For example, the extracted descriptive feature "this style" and the uploaded video information are input into the vision-language model, and the visual semantic feature extracted by the model from the video information is "a couple walking romantically in the park".

[0079] S105: Extract the common features and different features between the text features and the visual semantic features from the text features and the visual semantic features;

[0080] In this embodiment, after extracting text semantic features from text information and visual semantic features from visual information, common features shared by both and differential features with differences between the two are further extracted from the text semantic features and visual semantic features. Among them, the common features refer to shared attributes or characteristics that can be found in multiple information sources. The common features are the intersection part between different description methods or data forms, and can reflect the core content or commonality of multiple information. For example, for two information sources of text and video, elements such as the same objects, actions, or scenes involved in both can be used as common features. The differential features refer to specific elements with differences in multiple information sources, mainly including new elements reflecting the user's true intention, modifying elements in existing information sources, or excluding unnecessary information, etc.

[0081] In some embodiments, a preset second large language model is used to extract the common features and differential features between the text semantic features and visual semantic features.

[0082] In this embodiment, the text semantic features and visual semantic features are input into a fine-tuned deepseek large language model, and the large language model outputs the extracted common features and differential features. For example, inputting the text semantic feature "a couple is walking in the forest" and the visual semantic feature "a couple is walking romantically in the park" into the large language model, the common feature output by the model is "a couple is walking romantically", and the differential features are "forest" and "park", that is, the environment shown in the video information is the park, while the environment described in the text information is the forest, and the two constitute a difference. And the text information input by the user can reflect the user's true demand for the environment, so it can be used as additional information outside the video information to more accurately understand the user's intention and retrieve results more in line with the user's intention.

[0083] In some ways, based on the already trained deepseek large language model, the model is further trained using a dataset for a specific task to obtain a fine-tuned large language model, and the fine-tuned large language model is used to better adapt to the target task.

[0084] S106: Query a preset video information library according to the visual features and common features to obtain multiple video retrieval results that match the visual features and common features;

[0085] In this embodiment, the pre-constructed video information library stores several video files and the semantic features and video features of each video file. The method of retrieving the video information library using the visual features and common features is as follows:

[0086] Match the visual features with the video features of each video file to obtain video matching results;

[0087] Select a predetermined first number of first video retrieval results in descending order of the video matching results;

[0088] Match the common features with the semantic features of each video file to obtain semantic matching results;

[0089] Select a predetermined second number of second video retrieval results in descending order of the semantic matching results;

[0090] For the first video retrieval results and the second video retrieval results, calculate the matching scores according to the matching degree and corresponding weight of video matching, and the matching degree and corresponding weight of semantic matching;

[0091] Re - sort the first video retrieval results and the second video retrieval results in descending order of the matching scores to obtain the video retrieval results after sorting by matching degree.

[0092] In this embodiment, based on the extracted visual features and common features, the video information library is retrieved. During the retrieval, the visual features of the visual information are respectively matched with the video features of each video file to obtain the video matching results with each video file, and all the video matching results are sorted in descending order. A predetermined number of video matching results are selected from the front to the back as the first video retrieval results; the common features are respectively matched with the semantic features of each video file to obtain semantic matching results, and all the semantic matching results are sorted in descending order. A predetermined number of semantic matching results are selected from the front to the back as the second video retrieval results. Combining the first video retrieval results and the second video retrieval results, a plurality of initially retrieved video retrieval results are obtained. Optionally, the first number and the second number are the same, that is, N retrieval results are retrieved according to the video features, and N retrieval results are retrieved according to the semantic features, and 2N initially retrieved retrieval results are obtained.

[0093] In some ways, the cosine similarity method can be used to calculate the similarity between the visual features and the video features of each video file in the video information library to obtain the video matching degree between the video information and the video file, or the image matching degree between the image information and the video file, and calculate the similarity between the common features and the semantic features of each video file to obtain the semantic matching degree.

[0094] In some embodiments, for the 2N retrieval results of the initial retrieval, corresponding weights can be assigned according to the video matching degree or image matching degree and semantic matching degree. If the similarity of video matching is high, a higher weight is assigned to the video matching result; if the similarity of semantic matching is high, a higher weight is assigned to the semantic matching result. Then, based on the retrieval results after weight assignment, the 2N retrieval results are recalculated and reordered. For example, if the video matching degree of N video matching results is greater than or equal to 0.85, a higher weight is assigned to the video matching results; if the semantic matching degree of N semantic matching results is greater than or equal to 0.95, a higher weight is assigned to the semantic matching results.

[0095] In some ways, stepwise weight values can be assigned according to the values of the video matching results and semantic matching results, as shown in Table 1:

[0096] Table 1 Weight Table

[0097]

[0098] For the first video retrieval result and the second video retrieval result of the initial retrieval, according to the semantic matching degree and video matching degree or image matching degree of each video retrieval result, a preset weight table (such as Table 1) is queried to determine the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree, or the image weight corresponding to the image matching degree; according to the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree, or the image weight corresponding to the image matching degree, the matching score of the video retrieval result is calculated. The calculation method is:

[0099] S match = semantic matching degree × semantic weight + image matching degree × image weight (1)

[0100] S match = semantic matching degree × semantic weight + video matching degree × video weight (2)

[0101] After calculating the matching scores of each video retrieval result according to formulas (1) and (2), the 2N retrieval results are reordered in descending order of the matching scores to obtain multiple video retrieval results after sorting by matching degree.

[0102] S107: Screen the multiple video retrieval results according to the differential features to obtain the screened video retrieval results.

[0103] In this embodiment, the multiple initially retrieved search results include the video search results of the user's true intention, as well as other video search results. The part of the corresponding text information in the differential features reflects the additional information of the user's true intention. Therefore, based on the multiple initially retrieved video search results, the differential features are further used for screening, and the screened video search results are used as the final search results that meet the user's true intention.

[0104] Among them, screening the multiple video search results according to the differential features includes:

[0105] Sort the video search results including the corresponding text information in the differential features forward among the video search results sorted by matching degree, and sort the video search results not including the corresponding text information in the differential features backward to obtain the re-sorted video search results;

[0106] Based on the re-sorted video search results, select a predetermined third number of video search results as the screened video search results.

[0107] In this embodiment, considering that the differential features include the feature part reflecting the user's true intention and the feature part actually included in the information source, when screening the results, to conform to the user's true intention, based on the multiple initially retrieved video search results, the video search results including the corresponding text information in the differential features are sorted forward, and the video search results not including the corresponding text information in the differential features are sorted backward. Or rather, the video search results conforming to the user's true intention are sorted forward, and the video search results not including the features of the user's true intention are sorted backward to obtain the re-sorted video search results, and then a predetermined number of video search results are selected from the re-sorted results as the final search results. For example, for the multiple initially retrieved video search results, the search results of the forest scene are sorted forward, and the search results of the park scene or other scenes are sorted backward to obtain the video search results that meet the user's scene requirements.

[0108] In some embodiments, the multi-modal combined video search method of the present application further includes:

[0109] Obtain a video file;

[0110] Extract video features and semantic features from the video file;

[0111] Based on the video file and the corresponding video features and semantic features, construct a video information library.

[0112] This embodiment provides a method for constructing a video information library. For the obtained video files, corresponding video features and semantic features are extracted from each video file respectively, and a video information library is constructed based on all the video files and the corresponding video features and semantic features of each video file. Among them, the video features of the video file include picture features, fingerprint features, etc.

[0113] As Figure 3 shown, the video files include short video files and long video files. For short video files, video features and semantic features can be directly extracted from them, and the short videos and the corresponding video features and semantic features are saved in the video information library.

[0114] For long video files, after obtaining the video files, it further includes:

[0115] Performing scene segmentation on the long video file to obtain video segments corresponding to multiple scenes;

[0116] Extracting video features and semantic features from the video file includes:

[0117] Respectively extracting corresponding video features and semantic features from the video segments corresponding to each scene;

[0118] Constructing a video information library based on the video file and the extracted video features and semantic features includes:

[0119] Constructing a scene-level video information library based on the video segments corresponding to multiple scenes of the video file and the extracted video features and semantic features.

[0120] This embodiment provides a method for constructing a scene-level video database of long video files. Specifically, first perform shot segmentation on the long video file to obtain video segments under the same shot as subsets of the video scene. By performing shot-level feature analysis on the video segments under the same shot, the similarity of video content within a period of time can be ensured, and related shots are found and combined into a video scene. Optionally, the TransNetv2 model can be used to perform shot segmentation on the long video file.

[0121] For the video segments of the same shot after shot segmentation, picture feature extraction is performed. The extracted picture features include position features, character features, action features, audio features, etc. Among them, the position feature is a feature related to the geographical location in the video segment, for example, geographical location identifier, geomorphic feature, etc., the character feature is all features related to people in the video segment, for example, face feature, limb feature, expression feature, etc., the action feature is a feature related to the action that the object in the video segment is performing, for example, action type, action amplitude, action speed, etc., and the audio feature is a feature related to the voice signal in the video segment, for example, frequency distribution, etc.

[0122] In some ways, a relevant feature extraction model can be used to extract the features of the video frames. For example, ResNet50, Faster R-CNN, and NaverNet networks can be used to extract features. Among them, the location feature is set as a 2048-dimensional vector, and the ResNet50 network model is pre-trained through the Places dataset to achieve feature extraction; the person feature is set as a 512-dimensional vector, and the Faster R-CNN network model is pre-trained through the CIM dataset to detect the people in the key frames, and the ResNet50 network model is pre-trained through the PIPA dataset to extract the person features, and finally the person features of the shot are represented by the mean vector; the action feature is set as a 512-dimensional vector, and the TSN network model is pre-trained through the AVA dataset to achieve feature extraction; the audio feature is set as a 512-dimensional vector, and the NaverNet network model is pre-trained through the AVA-ActiveSpeaker dataset to separate the speech signal, and then the Mel-frequency cepstral coefficients are extracted from the speech signal. The extracted location features, person features, action features, and audio features are concatenated to obtain the video frame features of a 3584-dimensional vector, which are used as the semantic representation of the shot. Through metric learning, the semantic representation of the shot is fused and dimensionally reduced to obtain a 256-dimensional feature vector, which is used to provide the vector representation of the shot in the subsequent scene segmentation.

[0123] As Figure 4 shown, a scene refers to a video segment composed of a series of shots that are continuous in time and space and semantically coherent, and these shots together describe a complete event or situation. The goal of scene segmentation is to provide a semantic-level structured information technology for video content, emphasizing maintaining the semantic integrity of each segment, so as to form scene segments containing complete semantic information. After extracting the video frame features of the video segments of the same shot after shot segmentation, a shot vector representation is constructed based on the video frame features of the video segments of the same shot, and the shots belonging to the same scene are determined based on the similarity between the shot vector representations of pairwise shots. Specifically, it includes:

[0124] Calculate the similarity matrix between the shot vector representations of pairwise shots, binarize the elements of the similarity matrix according to a preset threshold, and perform region growing processing on the high-similarity region according to the region growing algorithm. After region growing, perform an opening operation on the similarity matrix. Then, perform the core steps of scene segmentation, including: maintaining a sequence of segmentation points, gradually adding segmentation points through the beam search strategy based on the von Neumann entropy. By calculating the von Neumann entropy of the similarity matrix, the internal consistency of the scene can be judged. If the entropy is high, it means that the internal differences of the video segments are large, otherwise it means that the internal differences of the video segments are small. By calculating the entropy of a certain block diagonal matrix of the similarity matrix, the similarity judgment of the corresponding part can be obtained.

[0125] In each iteration, the sum of the von Neumann entropies of the block diagonal matrices corresponding to the video segments in this iteration is compared with the entropy sum result of the previous iteration, and several smaller segmentation points are selected and saved to enter the next iteration. When the difference between the entropy results of two consecutive iterations is less than a preset difference threshold or the iteration reaches the upper limit of the number of times, the iteration stops. After the iteration is completed, multiple segmented scene segments are obtained according to the determined segmentation points. The scene segments of the video can be used as the objects for subsequent semantic information and fingerprint feature extraction, which helps to achieve efficient storage and compact representation of features, and reduce the redundancy of feature storage in a large amount of video content.

[0126] In some embodiments, the semantic features of the video file include audio features, video text features, and emotional features. Among them, a pre-constructed audio recognition model can be used to capture the background music, dialogue content, and ambient sound of the video file, etc., to achieve the extraction of audio features, provide continuous and rich context information, help understand the overall content and context association of the video, and improve the integrity and accuracy of understanding the video content. By combining optical character recognition and natural language processing, the text in the video frame such as subtitles and bullet screens is extracted from the video file as video text features, which can deeply understand the meaning and context of the text content, and improve the accuracy and reliability of information extraction. Using a preset sentiment analysis method to infer the emotional state expressed in the video file, capture the emotional tone of the video, and extract emotional type features such as happiness, sadness, and tension, which enriches the dimension of the semantic features of the video file and is of great significance for understanding the overall emotional atmosphere of the video and the emotional response of the audience, enabling the video content analysis to penetrate to the emotional level.

[0127] In some embodiments, the fingerprint feature of a video file is the unique identifier of the video file. Even if the video is processed such as format conversion, editing, or compression, its stability and uniqueness can still be maintained, playing a role in protecting the video copyright. The fingerprint features include key frame sequences, texture features, color features, spatial relationship features, voiceprint features, etc. Among them, the key frame sequence is a representative key frame selected from the video file, which can be a frame at the scene change or a visually significant frame; the texture feature captures the structural pattern and repeatability within the capture window and is an important part of the local features of the image, which can be extracted through local binary patterns, filters, wavelet transforms; the color feature is extracted by analyzing the color information in the video, such as color histograms, dominant colors, means, variances, average brightness and its variants, etc.; the spatial relationship feature utilizes the relative position relationship between pixels, for example, obtaining spatial structure information through a block strategy or calculating the image gradient. The voiceprint feature is obtained by analyzing unique attributes such as spectral features, Mel frequency cepstral coefficients, and linear predictive coding in the background audio of the video, which is used to detect whether the video contains copyright-protected audio content, or to locate the original video even when other sounds are mixed into the background music of the video. Adding the voiceprint feature to the fingerprint feature of the video can improve the robustness and applicability of the video, effectively making up for the problem of insufficient fingerprint feature recognition due to low video quality.

[0128] In some ways, after extracting the key frames from the video file, based on image and video features such as color features, spatial relationship features, and texture features, the similarity between regions is calculated using a sliding window, and a self-similarity matrix is used to reflect the similarity relationship between different local regions. The self-similarity matrix is subjected to dimensionality reduction and quantization processing to compress the high-dimensional matrix and extract key information. Then, the dimensionality-reduced self-similarity matrix is transformed into a fixed-length compact feature vector to obtain the video fingerprint feature.

[0129] In some implementation manners, for a long video file, after shot segmentation, the frame features are extracted, and similarity comparison is performed based on the frame features of the video file of each shot to achieve scene segmentation; for each video segment of each scene after segmentation, the corresponding semantic features, fingerprint features, etc. are respectively extracted, and each long video file is stored according to the features at the scene level. As shown in Table 2, the constructed video information library at the scene level includes several video files, and the videos are classified by a preset tagging method. The short video file does not need to be scene-segmented and directly stores the short video and corresponding video features such as semantic features, frame features, and fingerprint features. The long video file stores each video segment of each scene and the corresponding semantic features and video features according to the segmented scene segments.

[0130] Table 2 Video Information Library at the Scene Level

[0131]

[0132] When a user performs combined video retrieval with cross-modal multiple conditions, text information and video information, or text information and image information can be input. When video information is input, semantic features, as well as video features such as frame features and fingerprint features, are extracted from the video information. The extracted video features are correspondingly matched with the video features of each video segment in the video information database to obtain a video matching result, which serves as the first video retrieval result. When image information is input, semantic features, as well as image features such as frame features and fingerprint features, are extracted from the image information. The extracted image features are correspondingly matched with the video features of each video segment in the video information database to obtain an image matching result, which serves as the first video retrieval result.

[0133] In some embodiments, when a user inputs a single text, image, or video retrieval condition, video retrieval can also be performed according to the single retrieval condition. That is, when text information, image information, or video information is input, the video information database can be retrieved based on the text information, image information, or video information, and the corresponding video retrieval result can be obtained.

[0134] For the multi-modal combined video retrieval method provided by the embodiments of the present application, when a multi-modal combined retrieval condition of text information and visual information is input, visual semantic features can be extracted from the visual information through the descriptive information about the visual information in the text. According to the commonalities and differences between the text semantic features and the visual semantic features, the common features and the difference features of the two are extracted. Multiple matching video retrieval results are queried from the video information database according to the common features and the visual features extracted from the visual information, and then the difference features are used to screen out multiple video retrieval results that conform to the user's true intention. By performing higher-level semantic analysis on the text and the video and comprehensively fusing and understanding the semantic information between the multi-modal information, the true intention of the user can be fully understood, the accuracy of the multi-modal combined retrieval result can be improved, and the personalized and flexible retrieval needs of the user can be met.

[0135] It should be noted that the method of the embodiments of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of the distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present application, and these multiple devices will interact with each other to complete the described method.

[0136] It should be noted that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0137] As Figure 5 shown, an embodiment of the present application further provides a multimodal combined video retrieval device, including:

[0138] An acquisition module, configured to acquire text information and visual information;

[0139] A first extraction module, configured to extract text features from the text information;

[0140] A second extraction module, configured to extract visual features from the visual information;

[0141] A third extraction module, configured to extract visual semantic features from the visual information according to the text features;

[0142] A fourth extraction module, configured to extract common features and difference features between the text features and the visual semantic features from the text features and the visual semantic features;

[0143] A retrieval module, configured to query a preset video information library according to the visual features and the common features to obtain multiple video retrieval results that match the visual features and the common features;

[0144] A screening module, configured to screen the multiple video retrieval results according to the difference features to obtain the screened video retrieval results.

[0145] For convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in one or more software and / or hardware.

[0146] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment and has the beneficial effects of the corresponding method embodiment, which will not be elaborated herein.

[0147] Figure 6Fig. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0148] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0149] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0150] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0151] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0152] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0153] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0154] The electronic device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0155] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0156] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

[0157] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present application difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0158] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0159] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the scope of protection of the present disclosure.

Claims

1. A multimodal combined video retrieval method, characterized in that Including: Obtaining text information and visual information; Extracting text features from the text information; Extracting visual features from the visual information; Extracting visual semantic features from the visual information according to the text features; Extracting common features and differential features between the text features and the visual semantic features from the text features and the visual semantic features; Querying a preset video information library according to the visual features and the common features to obtain a plurality of video retrieval results matching the visual features and the common features; Screening the plurality of video retrieval results according to the differential features to obtain the screened video retrieval results.

2. The method according to claim 1, wherein Extracting text features from the text information includes: Using a preset first large language model to extract descriptive features associated with the visual information and text semantic features other than the visual information from the text information; According to the text features, extracting visual semantic features from the visual information includes: Using a preset vision-language model to extract visual semantic features from the visual information according to the descriptive features.

3. The method according to claim 2, wherein Extracting common features and differential features between the text features and the visual semantic features from the text features and the visual semantic features includes: Using a preset second large language model to extract the common features and the differential features from the text semantic features and the visual semantic features.

4. The method according to claim 1, characterized in that The video information library includes a plurality of video files and corresponding video features and semantic features of each video file; According to the visual features and the common features, querying a preset video information library to obtain a plurality of video retrieval results matching the visual features and the common features includes: Matching the visual features with the video features of each video file to obtain a video matching result; Selecting a predetermined first number of first video retrieval results in descending order of the video matching result; Matching the common features with the semantic features of each video file to obtain a semantic matching result; Selecting a predetermined second number of second video retrieval results in descending order of the semantic matching result; For the first video retrieval results and the second video retrieval results, calculating a matching score according to the matching degree and corresponding weight of the video matching and the matching degree and corresponding weight of the semantic matching; Re-sorting the first video retrieval results and the second video retrieval results in descending order of the matching score to obtain the video retrieval results sorted by matching degree.

5. The method according to claim 4, characterized in that, Screening the plurality of video retrieval results according to the differential features to obtain the screened video retrieval results includes: Sorting the video retrieval results including the corresponding text information in the differential features forward among the video retrieval results sorted by matching degree, and sorting the video retrieval results not including the corresponding text information in the differential features backward to obtain the re-sorted video retrieval results; Based on the re-sorted video retrieval results, selecting a predetermined third number of video retrieval results as the screened video retrieval results.

6. The method according to claim 4, characterized in that, For the first video retrieval result and the second video retrieval result, calculate the matching scores according to the matching degrees and corresponding weights of video matching and the matching degrees and corresponding weights of semantic matching, including: According to the semantic matching degree and the video matching degree or image matching degree of each video retrieval result, query a preset weight table to determine the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree or the image weight corresponding to the image matching degree; Calculate the matching scores of the video retrieval results according to the semantic weight corresponding to the semantic matching degree, the video weight corresponding to the video matching degree or the image weight corresponding to the image matching degree.

7. The method according to claim 1, characterized in that, It further includes: Obtain a video file; Extract video features and semantic features from the video file; Based on the video file and the corresponding video features and semantic features, construct the video information library.

8. The method according to claim 7, wherein The video file includes a long video file; after obtaining the video file, it further includes: Perform scene segmentation on the long video file to obtain video segments corresponding to multiple scenes; Extracting video features and semantic features from the video file includes: Respectively extract corresponding video features and semantic features from the video segments corresponding to each scene; Based on the video file and the extracted video features and semantic features, constructing the video information library includes: Based on the video segments corresponding to multiple scenes of the video file and the extracted video features and semantic features, construct a scene-level video information library.

9. The method according to claim 7 or 8, characterized in that, The video features include picture features and fingerprint features, the fingerprint features include key frames, texture features, color features, spatial relationship features, and voiceprint features, and the semantic features include audio features, video text features, and emotion features.

10. A multimodal combined video retrieval device, characterized in that, It includes: An acquisition module for acquiring text information and visual information; A first extraction module for extracting text features from the text information; A second extraction module for extracting visual features from the visual information; A third extraction module for extracting visual semantic features from the visual information according to the text features; A fourth extraction module for extracting the common features and differential features between the text features and the visual semantic features from the text features and the visual semantic features; A retrieval module for querying a preset video information library according to the visual features and the common features to obtain multiple video retrieval results matching the visual features and the common features; A screening module for screening multiple video retrieval results according to the differential features to obtain the screened video retrieval results.

Citation Information

Cited By

  • Video searching method and device

    CN121071182A

  • Method and device for retrieving video and electronic equipment

    CN121166973A