A content retrieval method, device, and computer-readable storage medium
Through multimodal feature extraction and fusion, the problem of insufficient accuracy of video features in the prior art is solved, the matching rate between video content and text content is improved, and the accuracy rate of content retrieval is improved.
Patent Information
- Application Number
- CN202110733613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-06-30
AI Technical Summary
In the prior art, due to the insufficient accuracy of a single feature extraction network in video content retrieval, the semantic matching rate of video features and text is low, which in turn affects the accuracy rate of content retrieval.
Multimodal feature extraction method is used to extract multimodal feature of video content, obtain modal features of each mode, and generate more accurate video features through feature extraction and fusion to improve the matching rate with text content.
Through multimodal feature extraction and fusion, the accuracy of video features is improved, so that the video content can better express its information, thereby improving the accuracy of content retrieval.
Smart Images

Figure CN113821687B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to a content retrieval method, apparatus, and computer-readable storage medium. Background Art
[0002] In recent years, a vast amount of content has been generated on the Internet. This content can include various types, such as text and videos. To better retrieve the required content from the vast amount of content, it is usually possible to retrieve another type of content that matches a certain type of content. For example, it is possible to retrieve text content that matches the video content provided by the user. Existing content retrieval often uses a feature extraction network to directly extract video features and text features for feature matching to complete content retrieval.
[0003] In the process of researching and practicing the existing technology, the inventors of the present invention found that due to the fact that videos contain multiple modalities and complex semantics, the accuracy of the video features extracted by a single feature extraction network is insufficient, making it impossible to correspond one-to-one with the text semantics. Therefore, the accuracy of content retrieval is insufficient. Summary of the Invention
[0004] Embodiments of the present invention provide a content retrieval method, apparatus, and computer-readable storage medium, which can improve the accuracy of content retrieval.
[0005] A content retrieval method includes:
[0006] Obtaining content to be retrieved for retrieving target content;
[0007] When the content to be retrieved is video content, performing multi-modal feature extraction on the video content to obtain modal features of each modality;
[0008] Respectively performing feature extraction on the modal features of each modality to obtain modal content features of each modality;
[0009] Fusing the modal content features to obtain video features of the video content, and retrieving target text content corresponding to the video content from a preset content set according to the video features.
[0010] Correspondingly, an embodiment of the present invention provides a content retrieval apparatus, including:
[0011] An obtaining unit, configured to obtain content to be retrieved for retrieving target content;
[0012] A first extraction unit, configured to, when the content to be retrieved is video content, perform multi-modal feature extraction on the video content to obtain modal features of each modality;
[0013] A second extraction unit, configured to perform feature extraction on the modal features of each modality respectively to obtain the modal video features of each modality;
[0014] A text retrieval unit, configured to fuse the modal video features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content from the preset content set according to the video features.
[0015] Optionally, in some embodiments, the first extraction unit may specifically be configured to perform multi-modal feature extraction on the video content by using a trained content retrieval model to obtain the initial modal features of each modality in the video content; extract video frames from the video content, and perform multi-modal feature extraction on the video frames by using the trained content retrieval model to obtain the basic modal features of each video frame; screen out the target modal features corresponding to each modality from the basic modal features, and fuse the target modal features and the corresponding initial modal features to obtain the modal features of the video content of each modality.
[0016] Optionally, in some embodiments, the second extraction unit may specifically be configured to identify the target video feature extraction network corresponding to each modality in the video feature extraction network of the trained content retrieval model; use the target video feature extraction network to perform feature extraction on the modal features to obtain the modal video features of each modality.
[0017] Optionally, in some embodiments, the content retrieval device may further include a training unit, and the training unit may specifically be configured to obtain a content sample set, where the content sample set includes video samples and text samples, and the text samples include at least one text word; perform multi-modal feature extraction on the video samples by using a preset content retrieval model to obtain the sample modal features of each modality; perform feature extraction on the sample modal features of each modality respectively to obtain the sample modal content features of the video samples, and fuse the sample modal content features to obtain the sample video features of the video samples; perform feature extraction on the text samples to obtain the sample text features and the text word features corresponding to each text word, and converge the preset content retrieval model according to the sample modal video features, sample video features, sample text features, and text word features to obtain the trained content retrieval model.
[0018] Optionally, in some embodiments, the training unit may specifically be configured to determine the feature loss information of the content sample set according to the sample modal content features and text word features; determine the content loss information of the content sample set based on the sample video features and sample text features; fuse the feature loss information and the content loss information, and converge a preset content retrieval model based on the fused loss information to obtain a trained content retrieval model.
[0019] Optionally, in some embodiments, the training unit may specifically be configured to calculate the feature similarity between the sample modal content features and the text word features to obtain a first feature similarity; determine the sample similarity between the video sample and the text sample according to the first feature similarity; calculate the feature distance between the video sample and the text sample based on the sample similarity to obtain the feature loss information of the content sample set.
[0020] Optionally, in some embodiments, the training unit may specifically be configured to perform feature interaction on the sample modal content features and the text word features according to the first feature similarity to obtain the interacted video features and interacted text word features; calculate the feature similarity between the interacted video features and the interacted text word features to obtain a second feature similarity; fuse the second feature similarity to obtain the sample similarity between the video sample and the text sample.
[0021] Optionally, in some embodiments, the training unit may specifically be configured to perform normalization processing on the first feature similarity to obtain a target feature similarity; determine the association weight of the sample modal content features according to the target feature similarity, where the association weight is used to indicate the association relationship between the sample modal content features and the text word features; weight the sample modal content features based on the association weight, and update the text word features based on the weighted sample modal content features to obtain the interacted video features and interacted text word features.
[0022] Optionally, in some embodiments, the training unit may specifically be configured to use the weighted sample modal content features as the initial interacted video features, and update the text word features based on the initial interacted video features to obtain the initial interacted text word features; calculate the feature similarity between the initial interacted video features and the initial interacted text word features to obtain a third feature similarity; update the initial interacted video features and the initial interacted text word features according to the third feature similarity to obtain the interacted video features and interacted text word features.
[0023] Optionally, in some embodiments, the training unit may specifically be configured to perform feature interaction on the post-initial-interaction video features and the post-initial-interaction text word features according to the third feature similarity to obtain target post-interaction video features and target post-interaction text word features; use the target post-interaction video features as the post-initial-interaction video features and the target post-interaction text word features as the post-initial-interaction text word features; return to perform the step of calculating the feature similarity between the post-initial-interaction video features and the post-initial-interaction text word features until the number of times of feature interaction between the post-initial-interaction video features and the post-initial-interaction text word features reaches a preset number of times, so as to obtain the post-interaction video features and the post-interaction text word features.
[0024] Optionally, in some embodiments, the training unit may specifically be configured to obtain a preset feature boundary value corresponding to the content sample set; screen out a first content sample pair in which the video sample and the text sample match and a second content sample pair in which the video sample and the text sample do not match from the content sample set according to the sample similarity; calculate a feature distance between the first content sample pair and the second content sample pair based on the preset feature boundary value to obtain feature loss information of the content sample set.
[0025] Optionally, in some embodiments, the training unit may specifically be configured to screen out a content sample pair with the largest sample similarity from the second content sample pairs to obtain a target content sample pair; calculate a similarity difference between the sample similarity of the first content sample pair and the sample similarity of the target content sample pair to obtain a first similarity difference; fuse the preset feature boundary value and the first similarity difference to obtain feature loss information of the content sample set.
[0026] Optionally, in some embodiments, the training unit may specifically be configured to calculate a feature similarity between the sample video features and the text features to obtain a content similarity between the video sample and the text sample; screen out a third content sample pair in which the video sample and the text sample match and a fourth content sample pair in which the video sample and the content sample do not match from the content sample set according to the content similarity; obtain a preset content boundary value corresponding to the content sample set, and calculate a content difference between the third content sample pair and the fourth content sample pair according to the preset content boundary value to obtain content loss information of the content sample set.
[0027] Optionally, in some embodiments, the training unit may specifically be configured to calculate a similarity difference between the content similarity of the third content sample pair and the content similarity of the fourth content sample pair to obtain a second similarity difference; fuse the second similarity difference with a preset content boundary value to obtain a content difference between the third content sample pair and the fourth content sample pair; and perform a normalization process on the content difference to obtain content loss information of the content sample set.
[0028] Optionally, in some embodiments, the content retrieval device may further include a video retrieval unit. The video retrieval unit may specifically be configured to, when the content to be retrieved is text content, extract features from the text content to obtain text features of the text content; and retrieve a target video content corresponding to the text content from the preset content set according to the text features. In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores an application program, and the processor is configured to run the application program in the memory to implement the content retrieval method provided by the embodiment of the present invention.
[0029] In addition, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute steps in any content retrieval method provided by the embodiment of the present invention.
[0030] After obtaining the content to be retrieved for retrieving the target content in the embodiment of the present application, when the content to be retrieved is video content, multi-modal feature extraction is performed on the video content to obtain modal features of each modality. Feature extraction is respectively performed on the modal features of each modality to obtain modal content features of each modality. The modal content features are fused to obtain video features of the video content, and according to the video features, a target text content corresponding to the video content is retrieved from the preset content set. Since this solution first performs multi-modal feature extraction on the video content, and then extracts modal video features from the modal features corresponding to each modality, the accuracy of the modal video features in the video is improved, and the modal video features are fused to obtain video features of the video content, so that the extracted video features can better express the information in the video. Therefore, the accuracy of content retrieval can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic diagram of the scenario of the content retrieval method provided by an embodiment of the present invention;
[0033] Figure 2 It is a schematic flowchart of the content retrieval method provided by an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of extracting modal features from video content provided by an embodiment of the present invention;
[0035] Figure 4 It is a schematic diagram of training a preset content retrieval model provided by an embodiment of the present invention;
[0036] Figure 5 It is another schematic flowchart of the content retrieval method provided by an embodiment of the present invention;
[0037] Figure 6 It is a schematic diagram of the structure of the content retrieval device provided by an embodiment of the present invention;
[0038] Figure 7 It is another schematic diagram of the structure of the content retrieval device provided by an embodiment of the present invention;
[0039] Figure 8 It is another schematic diagram of the structure of the content retrieval device provided by an embodiment of the present invention;
[0040] Figure 9 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0042] An embodiment of the present invention provides a content retrieval method, device and computer-readable storage medium. Among them, the content retrieval device can be integrated in an electronic device, and the electronic device can be a server or a terminal device, etc.
[0043] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0044] For example, referring to Figure 1 , taking the case where the content retrieval device is integrated in an electronic device as an example, after the electronic device obtains the content to be retrieved for retrieving the target content, when the content to be retrieved is video content, multi-modal feature extraction is performed on the video content to obtain the modal features of each modality, feature extraction is respectively performed on the modal features of each modality to obtain the modal content features of each modality, the modal content features are fused to obtain the video features of the video content, and according to the video features, the target text content corresponding to the video content is retrieved from the preset content set, thereby improving the accuracy of content retrieval.
[0045] It should be noted that the content retrieval method provided in the embodiments of this application involves computer vision technology in the field of artificial intelligence, that is, in the embodiments of this application, the computer vision technology of artificial intelligence can be used to extract features from text content and video content, and based on the extracted features, the target content is screened out from the preset content set.
[0046] Artificial Intelligence (AI) refers to the theory, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0047] Among them, Computer Vision (CV): Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the computer process the images into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and intelligent transportation, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0048] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0049] This embodiment will be described from the perspective of a content retrieval device. The content retrieval device can be specifically integrated in an electronic device, which can be a server or a terminal device, etc.; among them, the terminal can include devices such as tablet computers, laptop computers, personal computers (PCs), wearable devices, virtual reality devices, or other intelligent devices that can perform content retrieval.
[0050] A content retrieval method includes:
[0051] Obtain the content to be retrieved for retrieving the target content. When the content to be retrieved is video content, perform multi-modal feature extraction on the video content to obtain the modal features of each modality, perform feature extraction on the modal features of each modality respectively to obtain the modal content features of each modality, fuse the modal content features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content from the preset content set according to the video features.
[0052] As Figure 2 shown, the specific process of this content retrieval method is as follows:
[0053] 101. Obtain the content to be retrieved for retrieving the target content.
[0054] Among them, the content to be retrieved can be understood as the content in the retrieval condition for retrieving the target content. There can be various types of the content to be retrieved. For example, it can be text content or video content.
[0055] Among them, there can be various ways to obtain the content to be retrieved. Specifically, it can be as follows:
[0056] For example, it is possible to directly receive the content to be retrieved sent by the user through the terminal, or it is possible to obtain the content to be retrieved from the network or a third-party database. Or, when the memory of the content to be retrieved is large or the quantity is large, receive a content retrieval request that carries the storage address of the content to be retrieved, and obtain the content to be retrieved from the memory, cache or third-party database according to the storage address.
[0057] 102. When the content to be retrieved is video content, perform multi-modal feature extraction on the video content to obtain the modal features corresponding to multiple modalities.
[0058] Among them, the modal feature can be understood as the feature information corresponding to each modality in the video content. The video content can include multiple modalities. For example, it can include actions, audio, scenes, faces, OCR speeches, and / or entities, etc.
[0059] Among them, there can be various ways to perform multi-modal feature extraction on the video content. Specifically, it can be as follows:
[0060] For example, use the trained content retrieval model to perform multi-modal feature extraction on the video content to obtain the initial modal features of each modality in the video content. Extract video frames from the video content, and use the trained content retrieval model to perform multi-modal feature extraction on the video frames to obtain the basic modal features of each video frame. Screen out the target modal features corresponding to each modality from the basic modal features, and fuse the target modal features and the corresponding initial modal features to obtain the modal features of each modality.
[0061] Among them, the video content and video frames contain multiple modalities. For different modalities, different feature extraction methods can be used to perform multi-modal feature extraction on the video content and video frames. For example, for the modality of describing actions, the S3D (an action recognition model) model pre-trained on an action recognition dataset can be used for feature extraction. For the audio modality, the pre-trained VGGish (an audio extraction model) model can be used for feature extraction. For the scene modality, the pre-trained DenseNet-161 (a deep model) model can be used for feature extraction. For the face modality, the pre-trained SSD model and ResNet50 model can be used for feature extraction. For the face modality, the Google API (a feature extraction network) can be used for feature extraction. For the entity modality, the pre-trained SENet-154 (a feature extraction network) can be used for feature extraction. The extracted initial modality features and basic modality features can both include image features, expert features, and time features, etc.
[0062] Among them, there are multiple ways to fuse the target modality features and the corresponding initial modality features. For example, the image features (F), expert features (E), and time features (T) in the target modality features and the initial modality features can be added together to obtain the modality features (Ω) of each modality. Specifically, it can be as Figure 3 shown. Or, the weighting coefficients of the target modality features and the initial modality features can also be obtained. According to the weighting coefficients, the target modality features and the initial modality features are weighted, and the weighted target modality features and the initial modality features are fused to obtain the modality features of each modality.
[0063] Among them, the trained content retrieval model can be set according to the actual application requirements. In addition, it should be noted that the trained content retrieval model can be pre-set by the maintenance personnel or can be trained by the content retrieval device itself. That is, before the step of "using the trained content retrieval model to perform multi-modal feature extraction on the video content to obtain the initial modality features of each modality in the video content", the content retrieval method can also include:
[0064] Obtain a content sample set, which includes video samples and text samples. The text samples include at least one text word. Use a preset content retrieval model to perform multi-modal feature extraction on the video samples to obtain the sample modal features of each modality. Then, perform feature extraction on the sample modal features of each modality to obtain the sample modal content features of the video samples, and fuse the sample modal content features to obtain the sample video features of the video samples. Perform feature extraction on the text samples to obtain the sample text features and the text word features corresponding to each text word. And according to the sample modal video features, sample video features, sample text features, and text word features, converge the preset content retrieval model to obtain a trained content retrieval model. Specifically, it can be as follows:
[0065] S1. Obtain a content sample set.
[0066] Among them, the content sample set includes video samples and text samples, and the text samples include at least one text word.
[0067] Among them, there are various ways to obtain the content sample set. Specifically, it can be as follows:
[0068] For example, video samples and text samples can be directly obtained to get the content sample set. Or, the original video content and original text content can be obtained, and then the original video content and original text content are sent to the annotation server. Receive the matching tags between the original video content and the original text content returned by the annotation server, and add the matching tags to the original video content and the original text content, so as to obtain video samples and text samples. Combine the video samples and text samples to obtain the content sample set. Or, when the number of content samples in the content sample set is large or the memory is large, a model training request can be received. The model training request carries the storage address of the content sample set. According to the storage address, obtain the content sample set in the memory, cache, or third-party database.
[0069] S2. Use a preset content retrieval model to perform multi-modal feature extraction on the video samples to obtain the sample modal features of each modality.
[0070] For example, use a preset content retrieval model to perform multi-modal feature extraction on the video samples to obtain the initial sample modal features of each modality in the video samples. Extract video frames from the video samples, and use a preset content retrieval model to perform multi-modal feature extraction on the video frames to obtain the basic sample modal features of each video frame. Screen out the target sample modal features corresponding to each modality from the basic sample modal features, and fuse the target sample modal features and the corresponding initial sample modal features to obtain the modal sample features of each modality. For specific details, please refer to the above, and will not be elaborated here one by one.
[0071] S3. Feature extraction is respectively performed on the sample modal features of each modality to obtain the sample modal content features of the video sample, and the sample modal content features are fused to obtain the sample video features of the video sample.
[0072] For example, according to the modality of the sample modal features, the target video feature extraction network corresponding to each modality is identified in the video feature extraction network of the preset content retrieval model. The target video feature extraction network is used to perform feature extraction on the sample modal features to obtain the sample modal content features corresponding to each modality. The sample modal content features are fused to obtain the sample video features of the video sample.
[0073] Among them, the modality of the video feature extraction network of the preset content retrieval model is fixed. Therefore, only according to the modality of the sample modal features, the video feature extraction network corresponding to this modality can be identified, and the identified video feature extraction network is used as the target video feature extraction network.
[0074] After the target video feature extraction network is identified, the target video feature extraction network can be used to perform feature extraction on the modal features. The feature extraction process can be various. For example, the target video feature extraction network can be the encoder of a modality-specific Transformer (a transformation network) to encode the sample modal features, so as to extract the sample modal content features of each modality.
[0075] After the sample modal content features are extracted, the sample modal content features can be fused. The fusion process can be various. For example, the sample modal content features of each modality can be combined to obtain the sample modal content feature set of the video sample. The sample modal content feature set is input into the Transformer for encoding to calculate the correlation weights of the sample modal content features. The sample modal content features are weighted according to the correlation weights, and the weighted sample modal content features are fused to obtain the sample video features of the video sample.
[0076] S4. Feature extraction is performed on the text sample to obtain the sample text features and the text word features corresponding to each text word. Then, according to the sample modal content features, sample video features, sample text features, and text word features, the preset content retrieval model is converged to obtain the trained content retrieval model.
[0077] For example, the text feature extraction network of the preset content retrieval model is used to perform feature extraction on the text sample to obtain the text features of the text sample and the text word features of the text words. Then, according to the sample modal content features, sample video features, sample text features, and text word features, the preset content retrieval model is converged to obtain the trained content retrieval model.
[0078] Among them, there are various ways to extract features from text samples. For example, a text encoder can be used to extract text features and text word features from text samples. There are various types of text encoders. For instance, it can include Bert (a type of text encoder) or word2vector (a word vector generation model), etc.
[0079] After extracting the text features and text word features, the preset content retrieval model can be converged based on the sample modality content features, sample video features, sample text features, and text word features. There are various ways of convergence, which can be specifically as follows:
[0080] For example, the feature loss information of the content sample set can be determined according to the sample modality content features and text word features, the content loss information of the content sample set can be determined based on the sample video features and sample text features, the feature loss information and content loss information are fused, and the preset content retrieval model is converged based on the fused loss information to obtain the trained content retrieval model. It can be specifically as follows:
[0081] (1) Determine the feature loss information of the content sample set according to the sample modality content features and text word features.
[0082] For example, the feature similarity between the sample modality content features and text word features can be calculated to obtain the first feature similarity. According to the first feature similarity, the sample similarity between the video sample and the text sample can be determined. Based on the sample similarity, the feature distance between the video sample and the text sample is calculated to obtain the feature loss information of the content sample set.
[0083] Among them, there are various ways to calculate the feature similarity between the sample modality content features and text word features. For example, the cosine similarity between the sample modality content features and text word features can be calculated, and the cosine similarity is used as the first feature similarity. Specifically, it can refer to the formula (1) shown:
[0084]
[0085] Among them, S ij is the first feature similarity, w i is the text word feature, is the sample modality video feature.
[0086] After calculating the first feature similarity, the sample similarity between the video sample and the text sample can be determined based on the first feature similarity, and there can be multiple ways to determine it. For example, based on the first feature similarity, the sample modal content features and the text word features can be feature interacted to obtain the interacted video features and the interacted text word features, calculate the feature similarity between the interacted video features and the interacted text word features to obtain the second feature similarity, and fuse the second feature similarity to obtain the sample similarity between the video sample and the text sample.
[0087] Among them, there can be multiple ways to feature interact the sample modal content features and the text word features. For example, the first feature similarity can be normalized to obtain the target feature similarity, and based on the target feature similarity, the association weight of the sample modal content features can be determined. This association weight is used to indicate the association relationship between the sample modal content features and the text word features. Based on the association weight, the sample modal content features are weighted, and the text word features are updated based on the weighted sample modal content features to obtain the interacted video features and the interacted text word features.
[0088] Among them, there can be multiple ways to normalize the first feature similarity. For example, an activation function can be used to normalize the first feature similarity. There can be multiple types of activation functions. For example, it can be ReLU (relu(x) = max(0, x)), and the normalization process can be as shown in formula (2):
[0089]
[0090] Among them, is the target feature similarity, S ij is the first feature similarity, and relu is the activation function.
[0091] Among them, there can be multiple ways to determine the association weight of the sample modal content features based on the target feature similarity. For example, a preset association parameter can be obtained, and the association parameter and the target feature similarity are fused to obtain the association weight. This association weight can also be understood as the attention weight, and specifically can be as shown in formula (3):
[0092]
[0093] Among them, a ij is the association weight, λ is the preset association parameter, and this preset association parameter can be the inverse temperature parameter of softmax, is the target feature similarity.
[0094] After determining the correlation weights of the sample modal content features, the sample modal visual content features can be weighted based on the correlation weights, and the weighted sample modal content features can be fused to obtain the weighted modal content features. The weighted modal content features are used as the initial interactive post-video features of the video sample, which can be specifically referred to as shown in formula (4):
[0095]
[0096] Among them, a i is the initial interactive post-video feature, a ij is the correlation weight, is the sample modal video feature.
[0097] After calculating the initial interactive post-video feature, the text word features can be updated based on the initial interactive post-video feature to obtain the interactive post-video feature and the interactive post-text word feature. For example, the text word features can be updated based on the initial interactive post-video feature to obtain the initial interactive post-text word feature, calculate the feature similarity between the initial interactive post-video feature and the initial interactive post-text word feature to obtain the third feature similarity, and update the initial interactive post-video feature and the initial interactive post-text word feature according to the third feature similarity to obtain the interactive post-video feature and the interactive post-text word feature.
[0098] Among them, there are various ways to update the text word features based on the initial interactive post-video feature. For example, the preset update parameters can be obtained, and the preset update parameters, the initial interactive post-video feature, and the text word features can be fused to obtain the initial interactive post-text word feature, which can be specifically as shown in formula (5):
[0099]
[0100] Among them, f9w i , a i ) is the initial interactive post-text word feature, w i is the text word feature, a i is the initial interactive post-video feature, g i is the gate operation, which is used to select the most useful information, o i is the fusion feature, which is used to enhance the interaction between the text word feature and the initial interactive post-video feature, F g , b g , F o and b o are the preset update parameters. For multi-step operations (multiple feature interactions), which involve multiple updates of the text word features, therefore, formula (5) can be integrated to obtain F a , and the formula for K times of feature interaction (cross-attention operation) can be obtained, as shown in formula (6):
[0101]
[0102] Among them, K represents the K-th time, and respectively represent the text word features of the K-th time and the (K - 1)-th time. A k is the video feature after interaction for the K-th time, and V rep is the modal video feature.
[0103] Among them, there are various ways to update the initial video feature after interaction and the initial text word feature after interaction according to the third feature similarity. For example, according to the third feature similarity, the initial video feature after interaction and the initial text word feature after interaction can be subjected to feature interaction to obtain the target video feature after interaction and the target text word feature after interaction. The target video feature after interaction is used as the initial video feature after interaction, and the target text word feature after interaction is used as the initial text word feature after interaction, and the step of calculating the feature similarity between the initial video feature after interaction and the initial text word feature after interaction is returned until the number of times of feature interaction between the initial video feature after interaction and the initial text word feature after interaction reaches the preset number of times, and the video feature after interaction and the text word feature after interaction are obtained.
[0104] Among them, the process of feature interaction can be regarded as performing multi-step cross-attention calculations to obtain the video feature after interaction and the text word feature after interaction. The number of times of feature interaction can be set according to the actual application and is usually.
[0105] After obtaining the video feature after interaction and the text word feature after interaction, the sample similarity between the video sample and the text sample can be calculated, and there are various calculation methods. For example, the feature similarity between the video feature after interaction and the text word feature after interaction can be calculated to obtain the second feature similarity, and the second feature similarity is fused to obtain the sample similarity between the video sample and the text sample, as shown in formula (7):
[0106]
[0107] Among them, S(T w , rep ) is the sample similarity, w ki is the text word feature after interaction, and a ki is the video feature after interaction.
[0108] After calculating the sample similarity, the feature distance between the video sample and the text sample can be calculated, so as to obtain the feature loss information of the content sample set. There can be multiple calculation methods. For example, the preset feature boundary value corresponding to the content sample set can be obtained, and according to the sample similarity, the first content sample pairs in which the video sample and the text sample match and the second content sample pairs in which the video sample and the text sample do not match are screened out from the content sample set. Based on the preset feature boundary value, the feature distance between the first content sample pairs and the second content sample pairs is calculated to obtain the feature loss information of the content sample set.
[0109] Among them, there can be multiple ways to screen out the first content sample pairs and the second content sample pairs from the content sample set according to the sample similarity. For example, the sample similarity can be compared with the preset similarity threshold, and the video samples and the corresponding text samples with sample similarity exceeding the preset similarity threshold are screened out from the content sample set, so that the first content sample pairs can be obtained. The video samples and the corresponding text samples with sample similarity not exceeding the preset similarity threshold are screened out from the content sample set, so that the second content sample pairs can be obtained.
[0110] After screening out the first content sample pairs and the second content sample pairs, the feature distance between the first content sample pairs and the second content sample pairs can be calculated. There can be multiple calculation methods. For example, the content sample pair with the largest sample similarity can be screened out from the second content sample pairs to obtain the target content sample pair. The similarity difference between the sample similarity of the first content sample pair and the sample similarity of the target content sample pair is calculated to obtain the first similarity difference. The preset feature boundary value is fused with the first similarity difference to obtain the feature loss information of the content sample set, as shown in formula (8):
[0111]
[0112] Among them, L Tri is the feature loss information, Δ is the preset feature boundary value, B is the number of the content sample set, b indicates that the video sample and the text sample match, and b* indicates the hard negative sample, that is, the video sample or the text sample in the content sample pair with the largest sample similarity in the second content sample. It can be found that there can be two target content sample pairs. Moreover, after fusing the preset feature boundary value and the first similarity difference, the fused similarity difference can be normalized, and the normalized similarity difference is fused again, so that the feature loss information of the content sample set can be obtained.
[0113] Among them, the feature loss information can be regarded as the loss information obtained after backpropagation and parameter update using Triplet loss (a loss function), and the feature loss information is mainly used to narrow the distance between the matching video texts in the feature space.
[0114] (2) Determine the content loss information of the content sample set based on the sample video features and the sample text features.
[0115] For example, the feature similarity between the sample video features and the text features can be calculated to obtain the content similarity between the video sample and the text sample. According to the content similarity, in the content sample set, the third content sample pairs where the video sample and the text sample match, and the fourth content sample pairs where the video sample and the content sample do not match are screened out. The preset content boundary value corresponding to the content sample set is obtained, and according to the preset content boundary value, the content difference between the third content sample pair and the fourth content sample pair is calculated to obtain the content loss information of the content sample set.
[0116] Among them, there are various ways to calculate the content difference between the third content sample pair and the fourth content sample pair to obtain the content loss information of the content sample set. For example, the similarity difference between the content similarity of the third content sample pair and the content similarity of the fourth content sample pair can be calculated to obtain the second similarity difference. The second similarity difference is fused with the preset content boundary value to obtain the content difference between the third content sample pair and the fourth content sample pair. The content difference is normalized to obtain the content loss information of the content sample set, as shown in formula (9):
[0117]
[0118] Among them, L mar is the content loss information, B is the number of the content sample set, b represents that the video sample and the text sample match, d represents the content samples in the content sample set other than b, Θ is the preset content boundary value, S h is the content similarity, is the text feature, is the video feature of the video sample matching . The content loss information can be regarded as the loss information obtained through backpropagation and parameter update using the bidirectional max-margin ranking loss (a loss function).
[0119] (3) Fuse the feature loss information and the content loss information, and based on the fused loss information, converge the preset content retrieval model to obtain the trained content retrieval model.
[0120] For example, a preset balance parameter can be obtained, and the preset balance parameter is fused with the feature loss information to obtain the balanced feature loss information. The balanced feature loss information is added to the content loss information to obtain the fused loss information, as shown in formula (10):
[0121] L = L mar + β * L Tri (10)
[0122] Where L is the fused loss information, L mar is the content loss information, L Tri is the feature loss information, and β is the preset balance parameter, which is used to balance these two loss functions on the scale.
[0123] Optionally, a weighted parameter of the feature loss information and the content loss information can also be obtained. Based on this weighted parameter, the feature loss information and the content loss information are weighted, and the weighted feature loss information and content loss information are fused to obtain the fused loss information.
[0124] After obtaining the fused loss information, the preset content retrieval model can be converged based on the fused loss information. There are various ways of convergence. For example, according to the fused loss information, the gradient descent algorithm can be used to update the network parameters in the preset content retrieval model, so as to converge the preset content retrieval model to obtain the trained content retrieval model. Or, other algorithms can also be used to update the network parameters in the preset content retrieval model with the fused loss information, so as to converge the preset content retrieval model to obtain the trained content retrieval model.
[0125] It should be noted that during the training process of the content retrieval model, the text samples and video samples go through multiple steps of cross-attention calculation and content similarity calculation, and Triplet loss and bidirectional max-margin ranking loss are respectively used for backpropagation and parameter update to obtain the trained content retrieval model, specifically as Figure 4 shown.
[0126] 103. Feature extraction is respectively performed on the modal features corresponding to each modality to obtain the modal content features corresponding to each modality.
[0127] Among them, the modal content feature can be the overall feature of each modal content, which is used to indicate the content feature in this modality.
[0128] Among them, there are various ways to perform feature extraction on the modal features, specifically as follows:
[0129] For example, according to the modality of the modality features, the target video feature extraction network corresponding to each modality can be identified in the video feature extraction network of the trained content retrieval model, and the modality features can be extracted by using the target video feature extraction network to obtain the modality content features corresponding to each modality.
[0130] Among them, the modality of the video feature extraction network of the trained content retrieval model is fixed. Therefore, only according to the modality of the modality features, the video feature extraction network corresponding to this modality can be identified, and the identified video feature extraction network is used as the target video feature extraction network.
[0131] After the target video feature extraction network is identified, the modality features can be extracted by using the target video feature extraction network. There can be various processes for feature extraction. For example, the target video feature extraction network can be the encoder of a modality-specific Transformer to encode the modality features, so as to extract the modality content features corresponding to the video content of each modality.
[0132] 104. Fuse the modality content features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content from the preset content set according to the video features.
[0133] Among them, there can be various ways to fuse the modality video features. Specifically, it can be as follows:
[0134] For example, the modality content features of each modality can be combined to obtain a modality content feature set of the video content, and the modality content feature set is input into the Transformer model for encoding to calculate the correlation weights of the modality content features, and the modality content features are weighted according to the correlation weights, and the weighted modality content features are fused to obtain the video features of the video content. Or, obtain the weighting parameters corresponding to each modality, and based on the weighting parameters, weight the modality content features, and fuse the weighted modality content features to obtain the video features of the video content. Or, directly splice the modality content features to obtain the video features of the video content.
[0135] After obtaining the video features of the video content, the target text content corresponding to the video content can be retrieved from the preset content set according to the video features. There can be various retrieval methods. For example, the feature similarity between the video features and the text features of the candidate text content in the preset content set can be calculated respectively, and according to the feature similarity, the target text content corresponding to the video content is screened out from the candidate text content.
[0136] Among them, there are various ways to extract text features from candidate text content. For example, a text encoder can be used to extract features from candidate text content to obtain the text features of the candidate text content. There are various types of text encoders. For instance, it can include Bert and word2vector. Or, it is also possible to extract the features of each text word in the candidate text content, and then calculate the correlation weights between each text word. Based on the correlation weights, the text word features are weighted to obtain the text features of the candidate text content. There are various times to extract text features from candidate text content in the preset content set. For example, it can be real-time extraction. For instance, when the content to be retrieved is video content, the text features of the candidate text content can be extracted to obtain the text features of the candidate text content. Or, it is also possible to extract the text features of the candidate text content in the preset content set before obtaining the content to be retrieved, to obtain the text features of the candidate text content, so as to realize offline calculation of the feature similarity between the text features and video features, and thus more quickly screen out the target text content corresponding to the video content from the candidate text content.
[0137] Among them, there are also various ways to calculate the feature similarity between the video features and the text features of the candidate text content. For example, the cosine similarity between the video features and the text features of the candidate text content can be calculated to obtain the feature similarity. Or, it is also possible to calculate the feature distance between the video features and the text features of the candidate text content, and determine the feature similarity between the video features and the text features according to the feature distance.
[0138] After calculating the feature similarity, the target text content corresponding to the video content can be screened out from the candidate text content according to the feature similarity. There are various screening methods. For example, among the candidate text content, the candidate text content with a feature similarity exceeding the preset similarity threshold is screened out, and the screened candidate text content is sorted, and the sorted candidate text content is used as the target text content corresponding to the video content. Or, it is also possible to sort the candidate text content according to the feature similarity, and screen out the target text content corresponding to the video content from the sorted candidate text content. The screened target text content can be one or multiple. When the number of target text contents is one, the candidate text content with the largest feature similarity to the video features can be used as the target text content. When the number of target text contents is multiple, the top N candidate text contents with the highest ranking feature similarity to the video features can be screened out from the sorted candidate text content as the target text content.
[0139] Optionally, when the content to be retrieved is text content, feature extraction can also be performed on the text content, and based on the extracted text features, the target video content corresponding to the text content can be retrieved from the preset content set. Specifically, it can be as follows:
[0140] For example, when the content to be retrieved is text content, the text feature extraction network of the trained content retrieval model is used to extract features from the text content to obtain the text features of the text content. The feature similarity between the text features and the video features of the candidate video content in the preset content set is calculated respectively, and based on the feature similarity, the target video content corresponding to the text content is screened out from the candidate video content.
[0141] Among them, there are various ways to extract features from the text content. For example, a text encoder can be used to extract the overall features in the text content to obtain text features. There are various types of text encoders. For instance, it can include Bert and word2vector. Or, the features of each text word in the text content can also be extracted, and then the correlation weights between each text word are calculated, and based on the correlation weights, the text word features are weighted to obtain the text features.
[0142] After the text features of the text content are extracted, the feature similarity between the text features and the video features can be calculated. There are various ways to calculate the feature similarity. For example, feature extraction can be performed on the candidate video content in the preset content set to obtain the video features of each candidate video content, and then the cosine similarity between the text features and the video features is calculated, so that the feature similarity can be obtained.
[0143] Among them, there are various ways to extract video features from the candidate video content. For example, the trained content retrieval model can be used to perform multi-modal feature extraction on the candidate video content to obtain the modal features corresponding to multiple modalities. Feature extraction is performed on the modal features corresponding to each modality respectively to obtain the modal video features corresponding to each modality, and the modal video features are fused to obtain the video features of each candidate video content. The time to extract the video features of the candidate video content can be various. For example, the video features of the candidate video content can be extracted in real time. For instance, every time the content to be retrieved is obtained, the video features of the candidate video content can be extracted. Or, before the content to be retrieved is obtained, feature extraction can be performed on each candidate video content in the preset content set to extract the video features, so that the feature similarity between the text features and the video features can be calculated offline, and thus the target video content corresponding to the text content can be screened out from the candidate video content faster.
[0144] Among them, there are various ways to screen out the target video content corresponding to the text content from the candidate video content according to the feature similarity. For example, screen out the candidate video content with a feature similarity exceeding a preset similarity threshold from the candidate video content, sort the screened candidate video content, and use the sorted candidate video content as the target video content corresponding to the text content. Or, according to the feature similarity, sort the candidate video content, and screen out the target video content corresponding to the text content from the sorted candidate video content. The screened target video content can be one or multiple. When the number of target video content is one, the candidate video content with the largest feature similarity to the text feature can be used as the target video content. When the number of target video content is multiple, the top N candidate video content with the highest ranking in terms of the feature similarity to the text feature can be screened out from the sorted candidate video content as the target video content. Among them, in this solution, not only better feature extraction of the multi-modal information in the video is performed, but also more important words in the retrieval text are better focused on, thus achieving better retrieval results. On the datasets MSR-VTT, LSMDC, and ActivityNet, the content retrieval performance has been greatly improved compared with the current mainstream methods, and the results are shown in Tables 1, 2, and 3. In the tables, R1, R5, R10, and R50 respectively represent the recognition rates of ranking 1, ranking 5, ranking 10, and ranking 50, and MdR and MnR are the mean and median of the recognition rates.
[0145] Table 1 Results on the MSR-VTT Dataset
[0146]
[0147] Table 2 Results on the LSMDC Dataset
[0148]
[0149] Table 3 Results on the ActivityNet Dataset
[0150]
[0151] As can be seen from the above, after obtaining the content to be retrieved for retrieving the target content in the embodiment of the present application, when the content to be retrieved is video content, multi-modal feature extraction is performed on the video content to obtain the modal features of each modality, feature extraction is respectively performed on the modal features of each modality to obtain the modal content features of each modality, the modal content features are fused to obtain the video features of the video content, and according to the video features, the target text content corresponding to the video content is retrieved from the preset content set; since this solution first performs multi-modal feature extraction on the video content, and then extracts the modal video features from the modal features corresponding to each modality, thereby improving the accuracy of the modal video features in the video, and fusing the modal video features to obtain the video features of the video content, so that the extracted video features can better express the information in the video, therefore, the accuracy of content retrieval can be improved.
[0152] According to the method described in the above embodiment, the following will be further described in detail by way of examples.
[0153] In this embodiment, it will be described by taking the content retrieval device being specifically integrated in an electronic device, and the electronic device being a server as an example.
[0154] (1) The server trains the content retrieval model
[0155] C1. The server obtains the content sample set.
[0156] For example, the server can directly obtain video samples and text samples to obtain the content sample set, or it can obtain the original video content and the original text content, and then send the original video content and the original text content to the annotation server, receive the matching tags between the original video content and the original text content returned by the annotation server, add the matching tags to the original video content and the original text content, so as to obtain video samples and text samples, combine the video samples and text samples to obtain the content sample set, or when the number of content samples in the content sample set is large or the memory is large, it can receive a model training request, and the storage address of the content sample set is carried in the model training request, and according to the storage address, obtain the content sample set in the memory, cache or third-party database.
[0157] C2. The server performs multi-modal feature extraction on the video samples by using the preset content retrieval model to obtain the sample modal features of each modality.
[0158] For example, the server uses a preset content retrieval model to extract multi-modal features from a video sample, obtaining the initial sample modal features of each modality in the video sample. The server extracts video frames from the video sample and uses the preset content retrieval model to extract multi-modal features from the video frames, obtaining the basic sample modal features of each video frame. The server filters out the target sample modal features corresponding to each modality from the basic sample modal features and fuses the target sample modal features with the corresponding initial sample modal features to obtain the modal sample features of each modality.
[0159] C3. The server respectively extracts features from the sample modal features of each modality to obtain the sample modal content features of the video sample, and fuses the sample modal content features to obtain the sample video features of the video sample.
[0160] For example, the server identifies the Transformer network corresponding to each modality in the video feature extraction network of the preset content retrieval model as the target video feature extraction network according to the modality of the sample modal features, and uses the encoder of the Transformer network to encode the sample modal features, thereby extracting the sample modal content features of each modality. The server combines the sample modal content features of each modality to obtain the sample modal content feature set of the video sample, inputs the sample modal content feature set into the overall Transformer network for encoding to calculate the correlation weights of the sample modal content features, weights the sample modal content features according to the correlation weights, and fuses the weighted sample modal content features to obtain the sample video features of the video sample.
[0161] C4. The server extracts features from the text sample to obtain the sample text features and the text word features corresponding to each text word, and converges the preset content retrieval model according to the sample modal content features, the sample video features, the sample text features, and the text word features to obtain the trained content retrieval model.
[0162] For example, the server can use text encoders such as Bert or word2vector to extract features from the text features of the text sample to obtain the text features and the text word features. The server determines the feature loss information of the content sample set according to the sample modal content features and the text word features, determines the content loss information of the content sample set based on the sample video features and the sample text features, fuses the feature loss information and the content loss information, and converges the preset content retrieval model based on the fused loss information to obtain the trained content retrieval model. Specifically, it can be as follows:
[0163] (1) The server determines the feature loss information of the content sample set according to the sample modal content features and the text word features.
[0164] For example, the server can calculate the cosine similarity between the sample modal content features and the text word features, and use the cosine similarity as the first feature similarity. Specifically, reference can be made to formula (1). The activation function is used to normalize the first feature similarity. There can be various types of activation functions. For example, it can be ReLU (relu(x) = max(0, x)). The normalization process can be as shown in formula (2), and then the normalized target feature similarity is obtained. The preset association parameter is obtained, and the association parameter is fused with the target feature similarity to obtain the association weight, which can also be understood as the attention weight. Specifically, it can be as shown in formula (3). The sample modal content features are weighted based on the association weight, and the weighted sample modal content features are fused to obtain the weighted modal video features. The weighted modal content features are used as the initial interaction posterior video features of the video sample. Specifically, reference can be made to formula (4).
[0165] After the server calculates the initial interaction posterior video features, it can obtain the preset update parameter, and fuse the preset update parameter, the initial interaction posterior video features, and the text word features to obtain the initial interaction posterior text word features. Specifically, it can be as shown in formula (5). Calculate the feature similarity between the initial interaction posterior video features and the initial interaction posterior text word features to obtain the third feature similarity. According to the third feature similarity, the initial interaction posterior video features and the initial interaction posterior text word features can be feature interacted to obtain the target interaction posterior video features and the target interaction posterior text word features. The target interaction posterior video features are used as the initial interaction posterior video features, and the target interaction posterior text word features are used as the initial interaction posterior text word features. Return to execute the step of calculating the feature similarity between the initial interaction posterior video features and the initial interaction posterior text word features until the number of times of feature interaction between the initial interaction posterior video features and the initial interaction posterior text word features reaches the preset number of times, and the interaction posterior video features and the interaction posterior text word features are obtained.
[0166] After obtaining the post-interaction video features and post-interaction text word features, the server can calculate the feature similarity between the post-interaction video features and the post-interaction text word features to obtain the second feature similarity, and fuse the second feature similarity to obtain the sample similarity between the video sample and the text sample, as shown in formula (7). Compare the sample similarity with a preset similarity threshold, and screen out the video samples and corresponding text samples whose sample similarity exceeds the preset similarity threshold in the content sample set, so as to obtain the first content sample pair. Screen out the video samples and corresponding text samples whose sample similarity does not exceed the preset similarity threshold in the content sample set, so as to obtain the second content sample pair. Obtain the preset feature boundary value corresponding to the content sample set, screen out the content sample pair with the largest sample similarity in the second content sample pair to obtain the target content sample pair, calculate the similarity difference between the sample similarity of the first content sample pair and the sample similarity of the target content sample pair to obtain the first similarity difference, and fuse the preset feature boundary value with the first similarity difference to obtain the feature loss information of the content sample set, as shown in formula (8).
[0167] (2) The server determines the content loss information of the content sample set based on the sample video features and the sample text features.
[0168] For example, the server can calculate the feature similarity between the sample video features and the text features to obtain the content similarity between the video sample and the text sample. According to the content similarity, screen out the third content sample pair where the video sample and the text sample match and the fourth content sample pair where the video sample and the content sample do not match in the content sample set, and obtain the preset content boundary value corresponding to the content sample set. Calculate the similarity difference between the content similarity of the third content sample pair and the content similarity of the fourth content sample pair to obtain the second similarity difference, fuse the second similarity difference with the preset content boundary value to obtain the content difference between the third content sample pair and the fourth content sample pair, and perform normalization processing on the content difference to obtain the content loss information of the content sample set, as shown in formula (9).
[0169] (3) The server fuses the feature loss information and the content loss information, and converges the preset content retrieval model based on the fused loss information to obtain the trained content retrieval model.
[0170] For example, the server can obtain a preset balance parameter, fuse the preset balance parameter with the feature loss information to obtain the balanced feature loss information, and add the balanced feature loss information to the content loss information to obtain the fused loss information, as shown in formula (10). Then, according to the fused loss information, the gradient descent algorithm is used to update the network parameters in the preset content retrieval model, so as to converge the preset content retrieval model to obtain the trained content retrieval model. Alternatively, other algorithms can also be used to update the network parameters in the preset content retrieval model with the fused loss information, so as to converge the preset content retrieval model to obtain the trained content retrieval model.
[0171] As Figure 5 shown, a content retrieval method has the following specific process:
[0172] 201. The server obtains the content to be retrieved for retrieving the target content.
[0173] For example, the server can directly receive the content to be retrieved sent by the user through the terminal, or can obtain the content to be retrieved from the network or a third-party database. Or, when the memory of the content to be retrieved is large or the quantity is large, a content retrieval request is received, and the storage address of the content to be retrieved is carried in the content retrieval request. According to the storage address, the content to be retrieved is obtained from the memory, cache or third-party database.
[0174] 202. When the content to be retrieved is video content, the server performs multi-modal feature extraction on the video content to obtain modal features corresponding to multiple modalities.
[0175] For example, when the content to be retrieved is video content, the server uses the trained content retrieval model to perform multi-modal feature extraction on the video content to obtain the initial modal features of each modality in the video content. Video frames are extracted from the video content, and the trained content retrieval model is used to perform multi-modal feature extraction on the video frames to obtain the basic modal features of each video frame. The target modal features corresponding to each modality are screened out from the basic modal features, and the target modal features and the corresponding initial modal features are fused to obtain the modal features of each modality.
[0176] Among them, the video content and the video frames in the video content can include multiple modalities. For the description of the action modality, the S3D model pre-trained on the action recognition dataset can be used for feature extraction. For the audio modality, the pre-trained VGGish model can be used for feature extraction. For the scene modality, the pre-trained DenseNet-161 model can be used for feature extraction. For the face modality, the pre-trained SSD model and ResNet50 model can be used for feature extraction. For the face modality, Google API can be used for feature extraction. For the entity modality, the pre-trained SENet-154 can be used for feature extraction. The extracted initial modality features and basic modality features can both include image features, expert features, and time features, etc.
[0177] 203. The server respectively extracts the modality features corresponding to each modality to obtain the modality content features corresponding to each modality.
[0178] For example, according to the modality of the modality features, the Transformer network corresponding to each modality can be identified in the video feature extraction network of the trained content retrieval model as the target video feature extraction network, and the encoder of the modality-specific Transformer is used to encode the modality features, so as to extract the modality content features corresponding to each modality.
[0179] 204. The server fuses the modality content features to obtain the video features of the video content.
[0180] For example, the server can combine the modality content features of each modality to obtain the sample modality content feature set of the video content, input the modality visual content feature set into the Transformer model for encoding to calculate the correlation weights of the modality content features, weight the modality content features according to the correlation weights, and fuse the weighted modality content features to obtain the video features of the video content. Or, obtain the weighting parameters corresponding to each modality, based on the weighting parameters, weight the modality content features, and fuse the weighted modality content features to obtain the video features of the video content. Or, directly splice the modality video features to obtain the video features of the video content.
[0181] 205. The server retrieves the target text content corresponding to the video content from the preset content set according to the video features.
[0182] For example, the server can use a text encoder such as Bert or word2vector to extract features from the candidate text content to obtain the text features of the candidate text content. Alternatively, it can also extract the features of each text word in the candidate text content, and then calculate the correlation weights between each text word. Based on the correlation weights, the text word features are weighted to obtain the text features of the candidate text content.
[0183] The server calculates the cosine similarity between the video features and the text features of the candidate text content, thereby obtaining the feature similarity. Alternatively, it can also calculate the feature distance between the video features and the text features of the candidate text content, and determine the feature similarity between the video features and the text features based on the feature distance.
[0184] The server filters out the candidate video text content in the candidate text content whose feature similarity exceeds the preset similarity threshold, sorts the filtered candidate text content, and uses the sorted candidate text content as the target text content corresponding to the video content. Alternatively, it can also sort the candidate text content according to the feature similarity, and filter out the target text content corresponding to the video content from the sorted candidate text content. The filtered target text content can be one or multiple. When the number of target text contents is one, the candidate text content with the largest feature similarity to the video features can be used as the target text content. When the number of target text contents is multiple, the top N candidate text contents with the highest feature similarity rankings to the video features can be filtered out from the sorted candidate text content as the target text content.
[0185] Among them, there are various times for extracting text features from the candidate text content in the preset content set. For example, it can be real-time extraction. For instance, when the content to be retrieved is video content, the text features of the candidate text content can be extracted to obtain the text features of the candidate text content. Alternatively, it can also extract the text features of the candidate text content in the preset content set before obtaining the content to be retrieved, so as to realize offline calculation of the feature similarity between the text features and the video features, and thus more quickly filter out the target text content corresponding to the video content from the candidate text content.
[0186] 206. When the content to be retrieved is text content, the server extracts features from the text content and retrieves the target video content corresponding to the text content from the preset content set according to the extracted text features.
[0187] For example, when the content to be retrieved is text content, the server can use text encoders such as Bert or word2vector to extract the overall features in the text content to obtain the text features of the text content. The trained content retrieval model is used to perform multi-modal feature extraction on the candidate video content to obtain the modal features corresponding to multiple modalities. Feature extraction is respectively performed on the modal features corresponding to each modality to obtain the modal video features corresponding to each modality. By fusing the modal video features, the video features of each candidate video content can be obtained. Then, the cosine similarity between the text features and the video features is calculated, and thus the feature similarity can be obtained. The candidate video content with a feature similarity exceeding the preset similarity threshold is screened out from the candidate video content, and the screened candidate video content is sorted. The sorted candidate video content is used as the target video content corresponding to the text content. Alternatively, the candidate video content can also be sorted according to the feature similarity, and the target video content corresponding to the text content is screened out from the sorted candidate video content. The screened target video content can be one or multiple. When the number of target video content is one, the candidate video content with the largest feature similarity to the text features can be used as the target video content. When the number of target video content is multiple, the top N candidate video content with the top-ranked feature similarity to the text features can be screened out from the sorted candidate video content as the target video content.
[0188] Among them, there are various times for extracting the video features of the candidate video content. For example, the video features of the candidate video content can be extracted in real time. For instance, every time the content to be retrieved is obtained, the video features of the candidate video content can be extracted. Or, before obtaining the content to be retrieved, feature extraction can be performed on each candidate video content in the preset content set to extract the video features, so as to realize the offline calculation of the feature similarity between the text features and the video features, and thus screen out the target video content corresponding to the text content from the candidate video content faster.
[0189] As can be seen from the above, after the server in the embodiment of the present application obtains the content to be retrieved for retrieving the target content, when the content to be retrieved is video content, multi-modal feature extraction is performed on the video content to obtain the modal features of each modality, feature extraction is respectively performed on the modal features of each modality to obtain the modal content features of each modality, the modal content features are fused to obtain the video features of the video content, and according to the video features, the target text content corresponding to the video content is retrieved from the preset content set. When the content to be retrieved is text content, feature extraction is performed on the text content, and according to the extracted text features, the target video content corresponding to the text content is retrieved from the preset content set. Since this solution first performs multi-modal feature extraction on the video content, and then extracts the modal video features from the modal features corresponding to each modality, the accuracy of the modal video features in the video is improved, and the modal video features are fused to obtain the video features of the video content, so that the extracted video features can better express the information in the video and realize the two-way retrieval of text and video. Therefore, the accuracy of content retrieval can be improved.
[0190] To better implement the above method, an embodiment of the present invention further provides a content retrieval device, which can be integrated in an electronic device, such as a server or a terminal, etc. The terminal may include a tablet computer, a notebook computer, and / or a personal computer, etc.
[0191] For example, as Figure 6 shown, the content retrieval device may include an acquisition unit 301, a first extraction unit 302, a second extraction unit 303, and a text retrieval unit 304, as follows:
[0192] (1) Acquisition unit 301;
[0193] The acquisition unit 301 is configured to acquire the content to be retrieved for retrieving the target content.
[0194] For example, the acquisition unit 301 may specifically be configured to receive the content to be retrieved sent by the user through the terminal, or may acquire the content to be retrieved from the network or a third-party database. Or, when the memory of the content to be retrieved is large or the quantity is large, receive a content retrieval request carrying the storage address of the content to be retrieved, and acquire the content to be retrieved from the memory, cache, or third-party database according to the storage address.
[0195] (2) First extraction unit 302;
[0196] The first extraction unit 302 is configured to perform multi-modal feature extraction on the video content when the content to be retrieved is video content to obtain the modal features of each modality.
[0197] For example, the first extraction unit 302 can be specifically used to, when the content to be retrieved is video content, extract multi-modal features from the video content using the trained content retrieval model to obtain the initial modal features of each modality in the video content, extract video frames from the video content, and extract multi-modal features from the video frames using the trained content retrieval model to obtain the basic modal features of each video frame. Then, the target modal features corresponding to each modality are screened out from the basic modal features, and the target modal features and the corresponding initial modal features are fused to obtain the modal features of the video content of each modality.
[0198] (3) The second extraction unit 303;
[0199] The second extraction unit 303 is used to extract feature extraction for the modal features of each modality to obtain the modal content features of each modality.
[0200] For example, the second extraction unit 303 can be specifically used to identify the target video feature extraction network corresponding to each modality in the video feature extraction network of the trained content retrieval model according to the modality of the modal features, and use the target video feature extraction network to extract feature extraction for the modal features to obtain the modal content features of each modality.
[0201] (4) The text retrieval unit 304;
[0202] The text retrieval unit 304 is used to fuse the modal content features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content from the preset content set according to the video features.
[0203] For example, the text retrieval unit 304 can be specifically used to combine the modal content features of each modality to obtain the sample modal content feature set of the video content, input the modal visual content feature set into the Transformer model for encoding to calculate the correlation weights of the modal content features, weight the modal content features according to the correlation weights, and fuse the weighted modal content features to obtain the video features of the video content. Then, calculate the feature similarity between the video features and the text features of the candidate text content in the preset content set respectively, and screen out the target text content corresponding to the video content from the candidate text content according to the feature similarity.
[0204] Optionally, the content retrieval device may further include a training unit 305, as Figure 7 shown, specifically as follows:
[0205] The training unit 305 is used to train the preset content retrieval model to obtain the trained content retrieval model.
[0206] For example, the training unit 305 can be specifically used to obtain a content sample set, which includes video samples and text samples. The text samples include at least one text word. The preset content retrieval model is used to perform multi-modal feature extraction on the video samples to obtain the sample modal features of each modality. Feature extraction is respectively performed on the sample modal features of each modality to obtain the sample modal content features of the video samples, and the sample modal content features are fused to obtain the sample video features of the video samples. Feature extraction is performed on the text samples to obtain the sample text features and the text word features corresponding to each text word. The preset content retrieval model is converged according to the sample modal content features, sample video features, sample text features, and text word features to obtain the trained content retrieval model.
[0207] Optionally, the content retrieval device may further include a video retrieval unit 306, as Figure 8 shown, specifically as follows:
[0208] The video retrieval unit 306 is used to perform feature extraction on the text content when the content to be retrieved is text content, and retrieve the target video content corresponding to the text content from the preset content set according to the extracted text features.
[0209] For example, the video retrieval unit 306 can be specifically used to perform feature extraction on the text content using the text feature extraction network of the trained content retrieval model when the content to be retrieved is text content, to obtain the text features of the text content. The feature similarity between the text features and the video features of the candidate video content in the preset content set is calculated respectively, and the target video content corresponding to the text content is screened out from the candidate video content according to the feature similarity.
[0210] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0211] As can be seen from the above, after the acquisition unit 301 in this embodiment acquires the content to be retrieved for retrieving the target content, when the content to be retrieved is video content, the first extraction unit 302 performs multi-modal feature extraction on the video content to obtain the modal features of each modality. The second extraction unit 303 performs feature extraction on the modal features of each modality respectively to obtain the modal content features of each modality. The text retrieval unit 304 fuses the modal content features to obtain the video features of the video content, and retrieves the target text content corresponding to the video content from the preset content set according to the video features. Since this solution first performs multi-modal feature extraction on the video content, and then extracts the modal video features from the modal features corresponding to each modality, the accuracy of the modal video features in the video is improved, and the modal video features are fused to obtain the video features of the video content, so that the extracted video features can better express the information in the video. Therefore, the accuracy of content retrieval can be improved.
[0212] An embodiment of the present invention further provides an electronic device, as Figure 9 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:
[0213] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 9 the structure of the electronic device shown in
[0214] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Among them:
[0215] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0216] The electronic device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0217] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0218] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:
[0219] Obtain the content to be retrieved for retrieving the target content. When the content to be retrieved is video content, perform multi-modal feature extraction on the video content to obtain the modal features of each modality, respectively perform feature extraction on the modal features of each modality to obtain the modal content features of each modality, fuse the modal content features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content from the preset content set according to the video features.
[0220] For example, an electronic device receives the content to be retrieved sent by a user through a terminal, or alternatively, can obtain the content to be retrieved from a network or a third-party database. Or, when the memory of the content to be retrieved is large or the quantity is large, a content retrieval request is received, and the storage address of the content to be retrieved is carried in the content retrieval request. According to the storage address, the content to be retrieved is obtained from the memory, cache or third-party database. When the content to be retrieved is video content, a trained content retrieval model is used to perform multi-modal feature extraction on the video content to obtain the initial modal features of each modality in the video content. Video frames are extracted from the video content, and a trained content retrieval model is used to perform multi-modal feature extraction on the video frames to obtain the basic modal features of each video frame. The target modal features corresponding to each modality are screened out from the basic modal features, and the target modal features and the corresponding initial modal features are fused to obtain the modal features of each modality. According to the modality of the modal features, the target video feature extraction network corresponding to each modality is identified in the video feature extraction network of the trained content retrieval model, and the target video feature extraction network is used to perform feature extraction on the modal features to obtain the modal content features corresponding to each modality. The modal content features of each modality are combined to obtain the sample modal content feature set of the video content, and the modal content feature set is input into the Transformer model for encoding to calculate the correlation weights of the modal content features. The modal content features are weighted according to the correlation weights, and the weighted modal content features are fused to obtain the video features of the video content. The feature similarities between the video features and the text features of the candidate text contents in the preset content set are calculated respectively. According to the feature similarities, the target text content corresponding to the video content is screened out from the candidate text contents.
[0221] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0222] As can be seen from the above, in the embodiment of the present invention, after obtaining the content to be retrieved for retrieving the target content, when the content to be retrieved is video content, multi-modal feature extraction is performed on the video content to obtain the modal features of each modality, feature extraction is respectively performed on the modal features of each modality to obtain the modal content features of each modality, the modal content features are fused to obtain the video features of the video content, and according to the video features, the target text content corresponding to the video content is retrieved from the preset content set; since this solution first performs multi-modal feature extraction on the video content, and then extracts the modal video features from the modal features corresponding to each modality, thereby improving the accuracy of the modal video features in the video, and fusing the modal video features to obtain the video features of the video content, so that the extracted video features can better express the information in the video, therefore, the accuracy of content retrieval can be improved.
[0223] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0224] For this reason, an embodiment of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any content retrieval method provided by the embodiment of the present invention. For example, the instructions can execute the following steps:
[0225] Obtain the content to be retrieved for retrieving the target content. When the content to be retrieved is video content, perform multi-modal feature extraction on the video content to obtain the modal features of each modality, respectively perform feature extraction on the modal features of each modality to obtain the modal content features of each modality, fuse the modal content features to obtain the video features of the video content, and retrieve the target text content corresponding to the video content in the preset content set according to the video features.
[0226] For example, receive the content to be retrieved sent by the user through the terminal, or, the content to be retrieved can be obtained from the network or a third-party database, or, when the memory of the content to be retrieved is large or the quantity is large, receive a content retrieval request, and the storage address of the content to be retrieved is carried in the content retrieval request. According to the storage address, obtain the content to be retrieved in the memory, cache or third-party database. When the content to be retrieved is video content, use the trained content retrieval model to perform multi-modal feature extraction on the video content to obtain the initial modal features of each modality in the video content, extract video frames in the video content, and use the trained content retrieval model to perform multi-modal feature extraction on the video frames to obtain the basic modal features of each video frame. Screen out the target modal features corresponding to each modality from the basic modal features, and fuse the target modal features and the corresponding initial modal features to obtain the modal features of each modality. According to the modality of the modal features, identify the target video feature extraction network corresponding to each modality in the video feature extraction network of the trained content retrieval model, and use the target video feature extraction network to perform feature extraction on the modal features to obtain the modal content features corresponding to each modality. Combine the modal content features of each modality to obtain a sample modal content feature set of the video content, input the modal content feature set into the Transformer model for encoding to calculate the correlation weights of the modal content features, weight the modal content features according to the correlation weights, and fuse the weighted modal content features to obtain the video features of the video content. Calculate the feature similarity between the video features and the text features of the candidate text content in the preset content set respectively, and screen out the target text content corresponding to the video content from the candidate text content according to the feature similarity.
[0227] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0228] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0229] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the content retrieval methods provided in the embodiments of the present invention, the beneficial effects achievable by any of the content retrieval methods provided in the embodiments of the present invention can be realized. For details, refer to the foregoing embodiments, which will not be elaborated herein.
[0230] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners in the above content retrieval aspect or the video / text bidirectional retrieval aspect.
[0231] The above has introduced in detail a content retrieval method, device and computer-readable storage medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, there will be changes in the specific implementation manners and application scopes according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A content retrieval method, characterized in that, Including: Obtaining the content to be retrieved for retrieving the target content; When the content to be retrieved is video content, performing multi-modal feature extraction on the video content to obtain the modal features of each modality, where the modal feature is the feature information corresponding to each modality in the video content, and the video content includes multiple modalities, and the multiple modalities include action description, audio, scene, face, and / or entity; Performing feature extraction on the modal features of each modality respectively to obtain the modal content features of each modality, where the modal content feature is the overall feature of each modal content and is used to indicate the content feature under each modality; Fusing the modal content features to obtain the video features of the video content, and retrieving the target text content corresponding to the video content from a preset content set according to the video features; Among them, the performing multi-modal feature extraction on the video content to obtain the modal features of each modality includes: Performing multi-modal feature extraction on the video content by using a trained content retrieval model to obtain the initial modal features of each modality in the video content; Extracting video frames from the video content, and performing multi-modal feature extraction on the video frames by using the trained content retrieval model to obtain the basic modal features of each video frame; Selecting the target modal features corresponding to each modality from the basic modal features, and fusing the target modal features and the corresponding initial modal features to obtain the modal features corresponding to the video content of each modality.
2. The content retrieval method according to claim 1, characterized in that, The performing feature extraction on the modal features of each modality respectively to obtain the modal video features of each modality includes: Identifying the target video feature extraction network corresponding to each modality in the video feature extraction network of the trained content retrieval model; Performing feature extraction on the modal features by using the target video feature extraction network to obtain the modal video features of each modality.
3. The content retrieval method according to claim 1, wherein Before the performing multi-modal feature extraction on the video content by using a trained content retrieval model to obtain the initial model features of each modality in the video content, it further includes: Obtaining a content sample set, where the content sample set includes video samples and text samples, and the text samples include at least one text word; Performing multi-modal feature extraction on the video samples by using a preset content retrieval model to obtain the sample modal features of each modality; Performing feature extraction on the sample modal features of each modality respectively to obtain the sample modal content features of the video samples, and fusing the sample modal content features to obtain the sample video features of the video samples; Performing feature extraction on the text samples to obtain the sample text features and the text word features corresponding to each text word, and converging the preset content retrieval model according to the sample modal content features, sample video features, sample text features, and text word features to obtain the trained content retrieval model.
4. The content retrieval method according to claim 3, wherein The converging the preset content retrieval model according to the sample modal content features, sample video features, sample text features, and text word features to obtain the trained content retrieval model includes: Determine the feature loss information of the content sample set according to the sample modal content features and text word features; Determine the content loss information of the content sample set based on the sample video features and sample text features; Fuse the feature loss information and content loss information, and converge the preset content retrieval model based on the fused loss information to obtain the trained content retrieval model.
5. The content retrieval method according to claim 4, characterized in that The determining the feature loss information of the content sample set according to the sample modal content features and text word features includes: Calculate the feature similarity between the sample modal content features and text word features to obtain the first feature similarity; Determine the sample similarity between the video sample and the text sample according to the first feature similarity; Based on the sample similarity, calculate the feature distance between the video sample and the text sample to obtain the feature loss information of the content sample set.
6. The content retrieval method according to claim 5, characterized in that, The determining the sample similarity between the video sample and the text sample according to the first feature similarity includes: According to the first feature similarity, perform feature interaction on the sample modal content features and text word features to obtain the interacted video features and interacted text word features; Calculate the feature similarity between the interacted video features and the interacted text word features to obtain the second feature similarity; Fuse the second feature similarity to obtain the sample similarity between the video sample and the text sample.
7. The content retrieval method according to claim 6, characterized in that, The performing feature interaction on the sample modal content features and text word features according to the first feature similarity to obtain the interacted video features and interacted text word features includes: Perform normalization processing on the first feature similarity to obtain the target feature similarity; Determine the association weight of the sample modal content features according to the target feature similarity, and the association weight is used to indicate the association relationship between the sample modal content features and text word features; Based on the association weight, weight the sample modal content features, and update the text word features based on the weighted sample modal content features to obtain the interacted video features and interacted text word features.
8. The content retrieval method according to claim 7, characterized in that The updating the text word features based on the weighted sample modal content features to obtain the interacted video features and interacted text word features includes: Take the weighted sample modal content features as the initial interacted video features, and update the text word features based on the initial interacted video features to obtain the initial interacted text word features; Calculate the feature similarity between the initial interacted video features and the initial interacted text word features to obtain the third feature similarity; Update the initial interacted video features and the initial interacted text word features according to the third feature similarity to obtain the interacted video features and interacted text word features.
9. The content retrieval method according to claim 8, wherein The updating the initial interacted video features and the initial interacted text word features according to the third feature similarity to obtain the interacted video features and interacted text word features includes: According to the third feature similarity, perform feature interaction on the video feature after the initial interaction and the text word feature after the initial interaction to obtain a target video feature after interaction and a target text word feature after interaction; Use the target video feature after interaction as the video feature after the initial interaction, and use the target text word feature after interaction as the text word feature after the initial interaction; Return to the step of calculating the feature similarity between the video feature after the initial interaction and the text word feature after the initial interaction until the number of times of feature interaction between the video feature after the initial interaction and the text word feature after the initial interaction reaches a preset number of times, and obtain the video feature after interaction and the text word feature after interaction.
10. The content retrieval method according to claim 5, wherein Calculating the feature distance between the video sample and the text sample based on the sample similarity to obtain the feature loss information of the content sample set includes: Obtain the preset feature boundary value corresponding to the content sample set; According to the sample similarity, screen out the first content sample pairs where the video sample and the text sample match and the second content sample pairs where the video sample and the text sample do not match in the content sample set; Based on the preset feature boundary value, calculate the feature distance between the first content sample pair and the second content sample pair to obtain the feature loss information of the content sample set.
11. The content retrieval method according to claim 10, characterized in that, Calculating the feature distance between the first content sample pair and the second content sample pair based on the preset feature boundary value to obtain the feature loss information of the content sample set includes: Screen out the content sample pair with the largest sample similarity in the second content sample pair to obtain the target content sample pair; Calculate the similarity difference between the sample similarity of the first content sample pair and the sample similarity of the target content sample pair to obtain the first similarity difference; Fuse the preset feature boundary value and the first similarity difference to obtain the feature loss information of the content sample set.
12. The content retrieval method according to claim 4, wherein Determining the content loss information of the content sample set based on the sample video feature and the sample text feature includes: Calculate the feature similarity between the sample video feature and the text feature to obtain the content similarity between the video sample and the text sample; According to the content similarity, screen out the third content sample pairs where the video sample and the text sample match and the fourth content sample pairs where the video sample and the content sample do not match in the content sample set; Obtain the preset content boundary value corresponding to the content sample set, and calculate the content difference between the third content sample pair and the fourth content sample pair according to the preset content boundary value to obtain the content loss information of the content sample set.
13. The content retrieval method according to claim 12, wherein Calculating the content difference between the third content sample pair and the fourth content sample pair according to the preset content boundary value to obtain the content loss information of the content sample set includes: Calculate the similarity difference between the content similarity of the third content sample pair and the content similarity of the fourth content sample pair to obtain the second similarity difference; Fuse the second similarity difference with a preset content boundary value to obtain the content difference between the third content sample pair and the fourth content sample pair; Perform a normalization process on the content difference to obtain the content loss information of the content sample set.
14. The content retrieval method according to claim 1, wherein It further includes: When the content to be retrieved is text content, perform feature extraction on the text content to obtain the text features of the text content; According to the text features, retrieve the target video content corresponding to the text content in the preset content set.
15. A content retrieval device, characterized in that, It includes: An acquisition unit for acquiring the content to be retrieved for retrieving the target content; A first extraction unit for, when the content to be retrieved is video content, performing multi-modal feature extraction on the video content to obtain the modal features of each modality, where the modal feature is the feature information corresponding to each modality in the video content, and the video content includes multiple modalities, and the multiple modalities include description of actions, audio, scenes, faces, and / or entities; A second extraction unit for respectively performing feature extraction on the modal features of each modality to obtain the modal content features of each modality, where the modal content feature is the overall feature of each modal content and is used to indicate the content feature under each modality; A text retrieval unit for fusing the modal content features to obtain the video features of the video content, and according to the video features, retrieving the target text content corresponding to the video content in the preset content set; Among them, the performing multi-modal feature extraction on the video content to obtain the modal features of each modality includes: Performing multi-modal feature extraction on the video content by using a trained content retrieval model to obtain the initial modal features of each modality in the video content; Extracting video frames from the video content, and performing multi-modal feature extraction on the video frames by using the trained content retrieval model to obtain the basic modal features of each video frame; Screening out the target modal features corresponding to each modality from the basic modal features, and fusing the target modal features and the corresponding initial modal features to obtain the modal features corresponding to the video content of each modality.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the content retrieval method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Frame-by-frame cross-modal similarity association implementation text query video clip positioning method
CN111930999A