Content label generation method and device, computer equipment and readable storage medium
By extracting the single-modal and multimodal features of multimedia content and performing multiplexed searches to recall similar content labels, the problem of low label generation accuracy in traditional methods is solved, and higher content label generation accuracy is achieved.
Patent Information
- Application Number
- CN202510237767.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional content label generation methods have limited label information and are difficult to accurately match given content, resulting in low generation accuracy.
By obtaining the single-modal and multimodal fusion features of multimedia content, single-modal and multimodal searches are performed, similar multimedia content is recalled, and prediction generation is performed based on the labels of these contents.
Improve the accuracy of content label generation, making full use of different modal features of multimedia content and semantic information of similar content.
Smart Images

Figure CN120216709A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for generating content tags. Background Art
[0002] With the development of computer technologies, content tag generation technologies have emerged. Content tag generation can also be referred to as content tag retrieval, which refers to a technology that, in fields such as information retrieval and recommendation systems, identifies and "retrieves" (searches) the content tags most relevant to a given piece of content from a large number of tags or features. The goal is to retrieve as many relevant content tags as possible while maintaining a low false retrieval rate. These content tags may contain important information about the given content, such as names of people, place names, product names, brand logos, and so on.
[0003] In traditional technologies, a commonly used content tag generation method is the retrieval method, which finds the set of tags most relevant to the query content from a large number of candidate tags through retrieval technologies. It generally includes the following steps: First, extract the content features and tag features of the given content respectively, and then calculate the feature similarity to select tags.
[0004] However, in traditional methods, since tags are usually simple words or phrases and their information content is limited, it is difficult to match the given content when directly retrieving using the extracted tag features, resulting in the problem of low accuracy in generating content tags. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for generating content tags that can improve the accuracy of content tag generation.
[0006] In a first aspect, the present application provides a method for generating content tags, including:
[0007] Obtain a first piece of multimedia content, and extract unimodal features and multimodal fusion features from the first piece of multimedia content respectively;
[0008] According to the unimodal features, retrieve a second piece of multimedia content that is similar to the first piece of multimedia content in a single modality;
[0009] According to the multimodal fusion features, retrieve a third piece of multimedia content that is similar to the first piece of multimedia content in at least two modalities;
[0010] Based on the second content tags of the retrieved multiple second pieces of multimedia content and the third content tags of the retrieved multiple third pieces of multimedia content, predict and generate a first content tag for the first piece of multimedia content.
[0011] In a second aspect, the present application also provides a content label generation device, including:
[0012] A modality feature extraction module, configured to obtain a first multimedia content, and extract unimodal features and multimodal fusion features from the first multimedia content respectively;
[0013] A unimodal feature retrieval module, configured to retrieve a second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal features;
[0014] A multimodal feature retrieval module, configured to retrieve a third multimedia content that is similar to the first multimedia content in at least two modalities according to the multimodal fusion features;
[0015] A content label prediction module, configured to predict and generate a first content label of the first multimedia content based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents.
[0016] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0017] Obtain a first multimedia content, and extract unimodal features and multimodal fusion features from the first multimedia content respectively;
[0018] Retrieve a second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal features;
[0019] Retrieve a third multimedia content that is similar to the first multimedia content in at least two modalities according to the multimodal fusion features;
[0020] Predict and generate a first content label of the first multimedia content based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents.
[0021] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0022] Obtain a first multimedia content, and extract unimodal features and multimodal fusion features from the first multimedia content respectively;
[0023] Retrieve a second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal features;
[0024] Retrieve third multimedia content that is similar to the first multimedia content in at least two modalities according to the multimodal fusion feature;
[0025] Predict and generate a first content label for the first multimedia content based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents.
[0026] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0027] Obtain a first multimedia content, and extract unimodal features and multimodal fusion features from the first multimedia content respectively;
[0028] Retrieve second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal feature;
[0029] Retrieve third multimedia content that is similar to the first multimedia content in at least two modalities according to the multimodal fusion feature;
[0030] Predict and generate a first content label for the first multimedia content based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents.
[0031] The above content label generation method, device, computer device, computer-readable storage medium and computer program product can extract multiple different modality features by obtaining a first multimedia content and extracting unimodal features and multimodal fusion features from the first multimedia content respectively. By retrieving second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal feature, and retrieving third multimedia content that is similar to the first multimedia content in at least two modalities according to the multimodal fusion feature, the second multimedia content and the third multimedia content similar to the first multimedia content can be obtained through the multi-way retrieval and recall method of unimodal retrieval and multimodal retrieval. Furthermore, the first content label of the first multimedia content can be predicted and generated based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents. In the whole process, by using multiple different modality features of the first multimedia content to indirectly recall content labels through multi-way recall retrieval of similar multimedia content, the multiple different modality features of the first multimedia content, as well as the semantic similarity between the first multimedia content and the similar multimedia content, can be fully utilized for content label generation, and the accuracy of content label generation can be improved. Description of the Drawings
[0032] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.
[0033] Figure 1 It is an application environment diagram of the content label generation method in an embodiment;
[0034] Figure 2 It is a schematic flowchart of the content label generation method in an embodiment;
[0035] Figure 3 It is a schematic diagram of unimodal feature matching in an embodiment;
[0036] Figure 4 It is a schematic diagram of modal feature extraction in an embodiment;
[0037] Figure 5 It is a schematic diagram of modal feature extraction in another embodiment;
[0038] Figure 6 It is a schematic diagram of multimodal fusion feature extraction in an embodiment;
[0039] Figure 7 It is a schematic diagram of text modal feature extraction in an embodiment;
[0040] Figure 8 It is a schematic diagram of visual modal feature extraction in an embodiment;
[0041] Figure 9 It is a schematic diagram of label classification in an embodiment;
[0042] Figure 10 It is a schematic diagram of calculating label scores in an embodiment;
[0043] Figure 11 It is an example of video labels and news labels in an embodiment;
[0044] Figure 12 It is an application scenario diagram of the content label generation method in an embodiment;
[0045] Figure 13 It is a schematic flowchart of the content label generation method in another embodiment;
[0046] Figure 14 It is a schematic diagram of multi-channel recall in an embodiment;
[0047] Figure 15 Schematic diagram for determining label scores based on recall results in one embodiment;
[0048] Figure 16 Application scenario diagram of the content label generation method in another embodiment;
[0049] Figure 17 Structural block diagram of the content label generation device in one embodiment;
[0050] Figure 18 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0052] In order to clearly describe the technical solutions of the present application and facilitate the understanding of the technical solutions of the present application, the key concepts involved in the present application will be first explained below.
[0053] 1. First multimedia content.
[0054] The first multimedia content refers to the multimedia content for which content labels need to be determined. For example, the first multimedia content can specifically refer to a video for which content labels need to be determined, or it can also refer to graphic and text content for which content labels need to be determined. By way of example, such as news, etc.
[0055] 2. Modality.
[0056] Modality refers to different types of information, and specifically can include but is not limited to visual modality, sound modality, text modality, etc. Among them, the visual modality can include the content that can be seen in the multimedia content, such as images, colors, actions, etc. The sound modality includes the sounds that can be heard in the multimedia content, such as voices, music, and background sounds, etc. The text modality refers to the text resources that can describe the video content, such as subtitles, titles, content description information for the multimedia content, etc.
[0057] 3. Single-modal feature.
[0058] A single-modal feature refers to the feature corresponding to a single modality, and specifically can include various types of features such as visual modality features corresponding to the visual modality, sound modality features corresponding to the sound modality, and text modality features corresponding to the text modality.
[0059] 4. Multi-modal fusion feature.
[0060] The multi-modal fusion feature refers to the fusion feature obtained by combining at least two modalities. Specifically, it can be various types of fusion features such as the visual-text fusion feature obtained by combining the visual modality and the text modality, the audio-text fusion feature obtained by combining the audio modality and the text modality, and the content modality fusion feature obtained by combining the visual modality, the audio modality, and the text modality.
[0061] The content label generation method provided in the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the data storage system can store the data that the server 102 needs to process. The data storage system can be set separately, integrated on the server 102, or placed on the cloud or other network servers. The server 102 obtains the first multimedia content, extracts the single-modal feature and the multi-modal fusion feature from the first multimedia content respectively, retrieves the second multimedia content similar to the first multimedia content in a single modality according to the single-modal feature, and retrieves the third multimedia content similar to the first multimedia content in at least two modalities according to the multi-modal fusion feature. Based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents, the first content label of the first multimedia content is predicted and generated. Among them, the server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0062] In an exemplary embodiment, as Figure 2 shown, a content label generation method is provided. In this embodiment, an example is given where the method is applied to a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps 202 to 208. Among them:
[0063] Step 202, obtain the first multimedia content, and extract the single-modal feature and the multi-modal fusion feature from the first multimedia content respectively.
[0064] Among them, the first multimedia content refers to the multimedia content for which a content label needs to be determined. For example, the first multimedia content can specifically refer to a video for which a content label needs to be determined, or can also refer to graphic and text content for which a content label needs to be determined. For example, news, etc. The content label of the multimedia content refers to the keyword about the multimedia content. By associating the corresponding content label with the first multimedia content, accurate classification, retrieval, and push and other processes can be performed on the first multimedia content.
[0065] Among them, modality refers to different types of information, which can specifically include but are not limited to visual modality, sound modality, and text modality, etc. Among them, visual modality can include the content that can be seen in multimedia content, such as images, colors, actions, etc., and these visual elements provide a direct visual experience. Sound modality includes the sounds that can be heard in multimedia content, such as voices, music, and background sounds, etc. Sound modality adds auditory elements to multimedia content, enhancing the immersion and emotional resonance. Text modality refers to the text resources that can describe video content, such as subtitles, titles, content description information for multimedia content, etc. Text modality provides additional information for multimedia content, helping to better understand multimedia content.
[0066] Among them, unimodal features refer to the features corresponding to a single modality, which can specifically include various types of features such as visual modality features corresponding to visual modality, sound modality features corresponding to sound modality, and text modality features corresponding to text modality. Multimodal fusion features refer to the fusion features obtained by combining at least two modalities, which can specifically be various types of fusion features such as visual-text fusion features obtained by combining visual modality and text modality, sound-text fusion features obtained by combining sound modality and text modality, and content modality fusion features obtained by combining visual modality, sound modality, and text modality.
[0067] Exemplarily, the server can obtain the first multimedia content that needs to be tag-processed, and extract unimodal features and multimodal fusion features from the first multimedia content respectively. In a specific application, the server can extract the modality information of each of at least two modalities from the first multimedia content, and then perform modality feature extraction based on the modality information of each of at least two modalities to obtain the unimodal features and multimodal fusion features of the first multimedia content.
[0068] In a specific application, taking at least two modalities including visual modality and text modality as an example, the server can extract the modality information under visual modality from the first multimedia content, that is, visual modality information, and extract the modality information under text modality from the first multimedia content, that is, text modality information, and then perform modality feature extraction based on the visual modality information and text modality information to obtain the unimodal features and multimodal fusion features of the first multimedia content.
[0069] Step 204, retrieve the second multimedia content that is similar to the first multimedia content in a single modality according to the unimodal features.
[0070] Exemplarily, the server performs unimodal feature matching based on unimodal features to retrieve second multimedia content that is similar to the first multimedia content in a single modality. In a specific application, the server separately performs unimodal feature matching on each modality feature in the unimodal features to retrieve second multimedia content that is similar to the first multimedia content in a single modality.
[0071] In a specific application, taking the unimodal features including visual modality features and text modality features as an example, the server separately performs unimodal feature matching on the visual modality features and the text modality features to retrieve second multimedia content that is similar to the first multimedia content in a single modality.
[0072] Step 206: Retrieve third multimedia content that is similar to the first multimedia content in at least two modalities based on the multimodal fusion features.
[0073] Exemplarily, the server performs multimodal feature matching based on the multimodal fusion features to retrieve third multimedia content that is similar to the first multimedia content in at least two modalities. In a specific application, taking the visual-text fusion feature as the multimodal fusion feature as an example, the server performs multimodal feature matching based on the visual-text fusion feature to retrieve third multimedia content that is similar to the first multimedia content in the visual modality and the text modality.
[0074] Step 208: Predict and generate a first content label for the first multimedia content based on the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents.
[0075] Exemplarily, the first content label is the label finally determined for the first multimedia content, and the first content label can be obtained by predicting labels based on the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents.
[0076] In a specific application, the server can combine the modality information of the first multimedia content in at least two modalities, the respective second content labels of the retrieved multiple second multimedia contents, and the respective third content labels of the retrieved multiple third multimedia contents to perform label prediction and generate the first content label for the first multimedia content.
[0077] In a specific application, the server can also screen out the first content label for the first multimedia content from the multiple second content labels and the multiple third content labels by scoring the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents to determine the label scores.
[0078] In a specific application, for the second content label, the server can obtain the unimodal feature matching results corresponding to multiple retrieved second multimedia contents to score the second content label and obtain the label score of the second content label. For the third content label, the server can obtain the multimodal feature matching results corresponding to multiple retrieved third multimedia content labels to score the third content label and obtain the label score of the third content label.
[0079] In a specific application, the server can also traverse the second content labels of each of the multiple retrieved second multimedia contents and the third content labels of each of the multiple retrieved third multimedia contents, and extract at least one identical content label as the first content label of the first multimedia content.
[0080] In a specific application, for the second content label and the third content label, when determining whether the two content labels are the same, the server can compare the semantic similarity of the two content labels. If the semantic similarity of the two content labels is greater than the label semantic similarity threshold, it is considered that the two content labels are identical content labels, and the server can select any one of the two content labels as the first content label of the first multimedia content.
[0081] Examples of identical content labels and different content labels are given below. For example, if the second content label is "digital products" and the third content label is "electronic products", by comparing the semantic similarity of the two content labels, it can be considered that the two content labels are identical content labels, and either "digital products" or "electronic products" can be selected as the first content label of the first multimedia content. If the second content label is "digital products" and the third content label is "mobile phone review", by comparing the semantic similarity of the two content labels, it can be determined that the two content labels are different content labels.
[0082] The above content label generation method can extract multiple different modal features by obtaining the first multimedia content and extracting unimodal features and multimodal fusion features from the first multimedia content respectively. By retrieving the second multimedia content similar to the first multimedia content in a single modality according to the unimodal features, and retrieving the third multimedia content similar to the first multimedia content in at least two modalities according to the multimodal fusion features, the second multimedia content and the third multimedia content similar to the first multimedia content can be obtained through the multi-way retrieval and recall method of unimodal retrieval and multimodal retrieval. Furthermore, the first content label of the first multimedia content can be predicted and generated based on the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents. In the whole process, by using multiple different modal features of the first multimedia content to indirectly recall the content labels through multi-way recall retrieval of similar multimedia contents, the multiple different modal features of the first multimedia content, as well as the semantic similarity between the first multimedia content and the similar multimedia contents, can be fully utilized for content label generation, which can improve the accuracy of content label generation.
[0083] In an exemplary embodiment, retrieving the second multimedia content similar to the first multimedia content in a single modality according to the unimodal features includes:
[0084] Determine a multimedia content library; the multimedia content library includes the unimodal sample features of multiple sample multimedia contents;
[0085] Perform unimodal feature matching between the unimodal features and the unimodal sample features of multiple sample multimedia contents respectively to obtain unimodal feature matching results;
[0086] Retrieve from the multimedia content library the second multimedia content similar to the first multimedia content in a single modality according to the unimodal feature matching results.
[0087] Among them, the multimedia content library is used to store various unimodal features and multimodal features as references, including the unimodal sample features of multiple sample multimedia contents. The sample multimedia content refers to the multimedia content for which the content label has been determined in advance. The sample multimedia content can respectively extract the corresponding unimodal sample features in various modalities. For example, the unimodal sample features can include visual modality sample features corresponding to the visual modality, text modality sample features corresponding to the text modality, or sound modality sample features corresponding to the sound modality.
[0088] Exemplarily, the server determines a multimedia content library, which includes the unimodal sample features of multiple sample multimedia contents. By separately performing unimodal feature matching between the unimodal feature and the unimodal sample features of multiple sample multimedia contents, the server can obtain the unimodal feature matching results. Based on the unimodal feature matching results, the server can determine the correlation between the first multimedia content and multiple sample multimedia contents, and thus can retrieve from the multimedia content library a second multimedia content that is similar to the first multimedia content in a single modality.
[0089] In a specific application, the server can determine a pre-constructed multimedia content library. When constructing the multimedia content library, the server can use historical multimedia contents as sample multimedia contents, extract unimodal sample features for the sample multimedia contents, and establish an association relationship between the sample multimedia contents and the unimodal sample features, so as to support retrieval of sample multimedia contents based on the unimodal sample features.
[0090] In this embodiment, based on the determination of the multimedia content library, by separately performing unimodal feature matching between the unimodal feature and the unimodal sample features of multiple sample multimedia contents, it is possible to use unimodal feature retrieval to obtain a second multimedia content that is similar to the first multimedia content in a single modality.
[0091] In an exemplary embodiment, separately performing unimodal feature matching between the unimodal feature and the unimodal sample features of multiple sample multimedia contents to obtain the unimodal feature matching results includes:
[0092] Performing same-modality feature matching between each modality feature in the unimodal feature and the same-kind modality features in the unimodal sample features of multiple sample multimedia contents respectively to obtain the same-modality feature matching results;
[0093] Based on the same-modality feature matching results, obtain the unimodal feature matching results.
[0094] Exemplarily, in the case where the unimodal feature includes multiple modality features, the server will perform same-modality feature matching between each modality feature in the unimodal feature and the same-kind modality features in the unimodal sample features of multiple sample multimedia contents respectively to obtain the same-modality feature matching results. Based on the same-modality feature matching results, the server can determine the correlation between the first multimedia content and multiple sample multimedia contents under the same-modality feature matching, and thus can obtain the unimodal feature matching results based on the same-modality feature matching results. According to the unimodal feature matching results, retrieve from the multimedia content library a second multimedia content that is similar to the first multimedia content in a single modality.
[0095] In a specific application, the same-modal feature matching result can specifically be the content similarity score between the first multimedia content and multiple sample multimedia contents under the same-modal feature matching, that is, the same-modal feature similarity. Then, based on the same-modal feature matching result, the second multimedia content that is the same-modal as the first multimedia content can be screened out.
[0096] In a specific application, taking the unimodal features including visual modal features and text modal features as an example, the server will respectively perform the same-modal feature matching between the visual modal features and the visual modal sample features in the respective unimodal sample features of multiple sample multimedia contents, and obtain the first same-modal matching result for the visual modality. At the same time, the server will respectively perform the same-modal feature matching between the text modal features and the text modal sample features in the respective unimodal sample features of multiple sample multimedia contents, and obtain the second same-modal matching result for the text modality. Based on the first same-modal matching result and the second same-modal matching result, the same-modal feature matching result is obtained.
[0097] In a specific application, for each modality feature in the unimodal features, the number of the targeted modality features can be one or multiple. When the number of the targeted modality features is multiple, the server will respectively perform the same-modal feature matching between each targeted modality feature and the same-kind modality features in the respective unimodal sample features of multiple sample multimedia contents, and obtain the same-modal feature matching result for each targeted modality feature.
[0098] In a specific application, when the number of the targeted modality features is multiple, the multiple targeted modality features are obtained based on different modality information under the same modality. For example, taking the targeted modality feature as the text modal feature as an example, the targeted modality feature can be obtained based on at least one of the multimedia content title, multimedia content voice text, and content description information under the text modality. Another example is taking the targeted modality feature as the visual modal feature as an example, the targeted modality feature can be obtained based on at least one content image set under the visual modality.
[0099] In a specific application, based on the obtained same-modal feature matching result, the server can directly use the same-modal feature matching result as the unimodal feature matching result, or can further perform cross-modal feature matching. After obtaining the cross-modal feature matching result, based on the same-modal feature matching result and the cross-modal feature matching result, the unimodal feature matching result is obtained.
[0100] In this embodiment, by respectively performing the same-modal feature matching between each modality feature in the unimodal features and the same-kind modality features in the respective unimodal sample features of multiple sample multimedia contents, the unimodal feature retrieval can be realized by using the same-modal feature retrieval, and the unimodal feature matching result is obtained.
[0101] In an exemplary embodiment, based on the same-modal feature matching result, a single-modal feature matching result is obtained, including:
[0102] Each type of modal feature in the single-modal feature is respectively subjected to cross-modal feature matching with different types of modal features in the single-modal sample features of multiple sample multimedia contents, to obtain a cross-modal feature matching result;
[0103] Based on the same-modal feature matching result and the cross-modal feature matching result, a single-modal feature matching result is obtained.
[0104] Exemplarily, when the single-modal feature includes multiple types of modal features, on the basis of performing same-modal feature matching to obtain the same-modal feature matching result, the server will respectively perform cross-modal feature matching on each type of modal in the single-modal feature with different types of modal in the single-modal sample features of multiple sample multimedia contents, to obtain a cross-modal feature matching result. Based on the cross-modal feature matching result, the relevance between the first multimedia content and multiple sample multimedia contents under cross-modal feature matching can be determined, so as to aggregate the same-modal feature matching result and the cross-modal feature matching result to obtain a single-modal feature matching result. According to the single-modal feature matching result, the second multimedia content similar to the first multimedia content in a single modality is retrieved from the multimedia content library.
[0105] In a specific application, the cross-modal feature matching result may specifically be the content similarity score between the first multimedia content and multiple sample multimedia contents under cross-modal feature matching, that is, the similarity of different types of modal features. Then, according to the cross-modal feature matching result, the second multimedia content similar to the first multimedia content in cross-modal can be screened out.
[0106] In a specific application, for each type of modal feature in the single-modal feature, the number of the targeted modal features may be single or multiple. When the number of the targeted modal features is multiple, the server will respectively perform cross-modal feature matching on each targeted modal feature with different types of modal features in the single-modal sample features of multiple sample multimedia contents, to obtain the cross-modal feature matching result of each targeted modal feature.
[0107] In this embodiment, by the method of respectively performing cross-modal feature matching on each type of modal feature in the single-modal feature with different types of modal features in the single-modal sample features of multiple sample multimedia contents, single-modal feature retrieval can be realized by combining same-modal feature retrieval and cross-modal feature retrieval, to obtain a single-modal feature matching result.
[0108] In an exemplary embodiment, the unimodal features include visual modality features and text modality features; each type of modality feature in the unimodal features is respectively subjected to cross-modal feature matching with different types of modality features in the unimodal sample features of multiple sample multimedia contents, and the cross-modal feature matching result is obtained, including:
[0109] The visual modality features are respectively subjected to cross-modal feature matching with different types of modality features in the unimodal sample features of multiple sample multimedia contents to obtain a first cross-modal matching result for the visual modality;
[0110] The text modality features are respectively subjected to cross-modal feature matching with different types of modality features in the unimodal sample features of multiple sample multimedia contents to obtain a second cross-modal matching result for the text modality;
[0111] According to the first cross-modal matching result and the second cross-modal matching result, the cross-modal feature matching result is obtained.
[0112] Exemplarily, the unimodal features include visual modality features and text modality features. When performing cross-modal feature matching, the server will respectively perform cross-modal feature matching on the visual modality features with different types of modality features in the unimodal sample features of multiple sample multimedia contents to obtain a first cross-modal matching result for the visual modality, and will respectively perform cross-modal feature matching on the text modality features with different types of modality features in the unimodal sample features of multiple sample multimedia contents to obtain a second cross-modal matching result for the text modality, and collect the first cross-modal matching result and the second cross-modal matching result to obtain the cross-modal feature matching result.
[0113] In a specific application, the unimodal sample features include visual modality sample features and text modality sample features, then the server will respectively perform cross-modal feature matching on the visual modality features with the text modality sample features of multiple sample multimedia contents to obtain a first cross-modal matching result for the visual modality, and will respectively perform cross-modal feature matching on the text modality features with the visual modality sample features of multiple sample multimedia contents to obtain a second cross-modal matching result for the text modality.
[0114] In a specific application, the process of performing unimodal feature matching between the unimodal features and the unimodal sample features of the sample multimedia content can be as Figure 3As shown, it includes two parts: intra-modal feature matching and cross-modal feature matching. The intra-modal feature matching refers to the feature matching between the visual modal features of the first multimedia content and the visual modal sample features of the sample multimedia content to obtain the intra-modal matching result for the visual modality, which can specifically be the modality feature similarity between the visual modal features and the visual modal sample features, and the feature matching between the text modal features of the first multimedia content and the text modal sample features of the sample multimedia content to obtain the intra-modal matching result for the text modality, which can specifically be the modality feature similarity between the text modal features and the text modal sample features. The cross-modal feature matching refers to the feature matching between the visual modal features of the first multimedia content and the text modal sample features of the sample multimedia content to obtain the cross-modal matching result for the visual modality, which can specifically be the modality feature similarity between the visual modal features and the text modal sample features, and the feature matching between the text modal features of the first multimedia content and the visual modal sample features of the sample multimedia content to obtain the cross-modal matching result for the text modality, which can specifically be the modality feature similarity between the text modal features and the visual modal sample features.
[0115] It should be noted that Figure 3 in the example, the number of each modality feature in the unimodal feature is taken as a single one. For each modality feature in the unimodal feature, when the number of the targeted modality feature is multiple, the server will respectively perform intra-modal feature matching between each targeted modality feature and the same-kind modality features in the unimodal sample features of multiple sample multimedia contents to obtain the intra-modal matching results for each targeted modality feature, and respectively perform cross-modal feature matching between each targeted modality feature and the different-kind modality features in the unimodal sample features of multiple sample multimedia contents to obtain the cross-modal matching results for each targeted modality feature.
[0116] In a specific application, taking the number of text modality features as two for example, the process of single-modal feature matching between the single-modal features and the single-modal sample features of the sample multimedia content also includes two parts: intra-modal feature matching and cross-modal feature matching. The intra-modal feature matching refers to the feature matching between the visual modality features of the first multimedia content and the visual modality sample features of the sample multimedia content to obtain the intra-modal matching result for the video modality, and the feature matching between the two text modality features of the first multimedia content and the text modality sample features of the sample multimedia content respectively to obtain the intra-modal matching results for the two text modality features respectively. The cross-modal feature matching refers to the feature matching between the visual modality features of the first multimedia content and the text modality sample features of the sample multimedia content to obtain the cross-modal matching result for the video modality, and the feature matching between the two text modality features of the first multimedia content and the visual modality sample features of the sample multimedia content to obtain the cross-modal matching results for the two text modality features respectively.
[0117] In this embodiment, by performing cross-modal feature matching between the visual modality features in the single-modal features and the text modality sample features of each of the multiple sample multimedia contents respectively, and performing cross-modal feature matching between the text modality features in the single-modal features and the visual modality sample features of each of the multiple sample multimedia contents respectively, cross-modal feature retrieval can be realized through the cross-modal feature matching between the visual modality and the text modality, and cross-modal feature matching results can be obtained.
[0118] In an exemplary embodiment, retrieving the third multimedia content that is similar to the first multimedia content in at least two modalities according to the multi-modal fusion feature includes:
[0119] Determine the multimedia content library; the multimedia content library includes the multi-modal sample features of each of the multiple sample multimedia contents;
[0120] Perform multi-modal feature matching between the multi-modal fusion feature and the multi-modal sample features of each of the multiple sample multimedia contents respectively to obtain multi-modal feature matching results;
[0121] Retrieve from the multimedia content library the third multimedia content that is similar to the first multimedia content in at least two modalities according to the multi-modal feature matching results.
[0122] Exemplarily, the server determines a multimedia content library, which includes multi-modal sample features of multiple sample multimedia contents. By respectively performing multi-modal feature matching between the multi-modal fusion feature and the multi-modal sample features of multiple sample multimedia contents, the server can obtain a multi-modal feature matching result. Based on the multi-modal feature matching result, the relevance between the first multimedia content and multiple sample multimedia contents can be determined, so that the third multimedia content similar to the first multimedia content in at least two modalities can be retrieved from the multimedia content library.
[0123] In a specific application, when constructing the multimedia content library and taking historical multimedia content as sample multimedia content, the server extracts multi-modal sample features for the sample multimedia content and establishes an association relationship between the sample multimedia content and the multi-modal sample features, so as to support retrieval of the sample multimedia content through the multi-modal sample features.
[0124] In a specific application, the multi-modal feature matching result can specifically be the content similarity score between the first multimedia content and multiple sample multimedia contents under multi-modal feature matching, that is, the multi-modal feature similarity. Then, based on the multi-modal feature matching result, the third multimedia content similar to the first multimedia content in at least two modalities can be filtered out.
[0125] In a specific application, taking the multi-modal fusion feature as the visual-text fusion feature as an example, by respectively performing multi-modal feature matching between the visual-text fusion feature and the visual-text fusion sample features of multiple sample multimedia contents, the server can obtain a multi-modal feature matching result. Based on the multi-modal feature matching result, the visual and text relevance between the first multimedia content and multiple sample multimedia contents can be determined, so that the third multimedia content similar to the first multimedia content in the visual modality and the text modality can be retrieved from the multimedia content library.
[0126] In this embodiment, based on the determination of the multimedia content library, by respectively performing multi-modal feature matching between the multi-modal fusion feature and the multi-modal sample features of multiple sample multimedia contents, multi-modal feature retrieval can be used to obtain the third multimedia content similar to the first multimedia content in at least two modalities.
[0127] In an exemplary embodiment, extracting the single-modal feature and the multi-modal fusion feature from the first multimedia content respectively includes:
[0128] Extracting the modal information of the first multimedia content in at least two modalities respectively;
[0129] Extract single-modal features from the modal information of each of at least two modalities respectively to obtain the single-modal features of the first multimedia content;
[0130] Extract multi-modal features by combining the modal information of each of at least two modalities to obtain the multi-modal fusion features of the first multimedia content.
[0131] Exemplarily, the server will extract the modal information of each of at least two modalities from the first multimedia content, then extract single-modal features from the modal information of each of at least two modalities respectively to obtain the single-modal features of the first multimedia content, and extract multi-modal features by combining the modal information of each of at least two modalities to obtain the multi-modal fusion features of the first multimedia content.
[0132] In a specific application, such as Figure 4 shown, taking at least two modalities including the visual modality and the text modality as an example, the server will extract the visual modality information in the visual modality from the first multimedia content, and extract the text modality information in the text modality from the first multimedia content. Extract single-modal features from the visual modality information and the text modality information respectively to obtain the single-modal features of the first multimedia content, namely the visual modality features and the text modality features. At the same time, the server will extract multi-modal features by combining the visual modality information and the text modality information to obtain the multi-modal fusion features of the first multimedia content, namely the visual-text fusion features.
[0133] It should be noted that Figure 4 the above is described by taking the number of visual modality information and text modality information as single for example. It can be understood that the number of visual modality information and text modality information can also be multiple. Then, for each visual modality information and each text modality information, single-modal features will be extracted to obtain the single-modal features of the first multimedia content, including the visual modality features corresponding to each visual modality information and the text modality features corresponding to each text modality information. And any visual modality information and text modality information can be combined to extract multi-modal features to obtain the multi-modal fusion features of the first multimedia content.
[0134] In a specific application, taking the first multimedia content as a video for example, as Figure 5 shown, the visual modality information of the video can be video frames, and the text modality information can include the video title and video description information. Then the server can extract visual modality features from the video frames to obtain visual modality features, and extract text modality features from the video title and video description information respectively to obtain two text modality features (such as Figure 5The text modality features 1 obtained based on the video title and the text modality features 2 obtained based on the video description information are shown respectively), and the video frames, video title, and video description information are combined to perform multimodal feature extraction to obtain the multimodal fusion features of the video. It should be noted that in addition to the video title and video description information, the text modality information of the video can also be video subtitles, etc., which are not shown here in this embodiment.
[0135] In a specific application, when performing multimodal feature extraction by combining the modality information of at least two modalities, the server can input the modality information of at least two modalities into a pre-trained multimodal classification model to extract the multimodal fusion features of the first multimedia content, or can first perform single-modal feature extraction on the modality information of at least two modalities respectively to obtain two single-modal features, and then fuse the two single-modal features to obtain the multimodal fusion features of the first multimedia content. Among them, the pre-trained multimodal classification model can be trained according to the actual application scenario. For example, the pre-trained multimodal classification model can specifically be a multimodal model based on Transformer. Transformer is a deep learning model for processing sequence data, mainly relying on the attention mechanism.
[0136] In a specific application, taking at least two modalities including the visual modality and the text modality as an example, the process of multimodal feature extraction can be as Figure 6 shown. The server will input the visual modality information and the text modality information into the pre-trained multimodal classification model at the same time, extract and integrate the information from different modalities, and extract the output features of the last layer (the layer before the classification head) of the pre-trained multimodal classification model as the multimodal fusion features.
[0137] Furthermore, as Figure 6 shown, the extracted features can specifically be CLS Token Embedding (CLS embedding vector). The pre-trained multimodal classification model will insert a special token [CLS] (Classification Token) at the beginning of the input sequence to represent the classification information of the entire sequence. During the training process, the embedding vector of [CLS] will learn the global semantic information of the entire sequence, that is, the multimodal fusion features. Among them, Embedding refers to the feature vector. In machine learning and natural language processing, Embedding is a method of converting high-dimensional data (such as words, sentences, or images) into low-dimensional vector representations. This representation preserves the semantic information of the data and makes it suitable for the input of machine learning models.
[0138] It can be understood that this method can effectively combine visual and text features, improve the understanding and analysis ability of the first multimedia content. Through this multi-modal feature extraction strategy, we can more comprehensively capture the semantic information of the first multimedia content and enhance the overall performance of the system.
[0139] In this embodiment, based on extracting the modal information of each of at least two modalities from the first multimedia content, by performing single-modal feature extraction on the modal information of each of at least two modalities respectively, single-modal features that can accurately represent the first multimedia content can be obtained. By jointly performing multi-modal feature extraction on the modal information of each of at least two modalities, multi-modal fusion features that can accurately represent the first multimedia content can be obtained. The entire process can achieve the full extraction of the modal features of the first multimedia content, enrich the representation form of the modal features of the first multimedia content, thereby facilitating the retrieval of similar sample multimedia content to determine content tags and improving the accuracy of content tag generation.
[0140] In an exemplary embodiment, at least two modalities include the text modality; extracting the modal information of each of at least two modalities from the first multimedia content includes:
[0141] Input the first multimedia content into a pre-trained content description model, and predict the content description information of the first multimedia content through the pre-trained content description model;
[0142] Based on the content description information, obtain the modal information of the first multimedia content in the text modality.
[0143] Among them, the pre-trained content description model refers to a model that is pre-trained to output content descriptions for given content and can be configured according to the actual application scenario. For example, the pre-trained content description model can specifically be a model based on a large language model. By pre-training on a large-scale text data, it can learn the grammar rules, semantic structures, and context relationships of the language, thereby generating natural language content.
[0144] Exemplarily, when at least two modalities include the text modality, the server will input the first multimedia content into a pre-trained content description model. Through the pre-trained content description model, using its pre-trained general language ability and context understanding ability, predict the content description information of the first multimedia content. Based on the content description information, obtain the modal information of the first multimedia content in the text modality.
[0145] In a specific application, after obtaining the content description information, if the number of modal information of the required first multimedia content in the text modality is single, the server can directly use the content description information as the modal information of the first multimedia content in the text modality, or can obtain the modal information of the first multimedia content in the text modality by splicing the content description information and other text information. Here, the other text information can specifically be the content title, content subtitle, etc. of the first multimedia content.
[0146] If the number of modal information of the required first multimedia content in the text modality is multiple, the server will use the content description information as a part of the modal information of the first multimedia content in the text modality, and then obtain the remaining part of the modal information of the first multimedia content in the text modality by additionally acquiring other text information.
[0147] In this embodiment, by inputting the first multimedia content into a pre-trained content description model, the content description information of the first multimedia content can be predicted, and then the modal information of the first multimedia content in the text modality can be determined by using the content description information. It can be understood that the content description information can use its powerful understanding ability to provide richer and more detailed text modality information, enriching the representation form of the modal features of the first multimedia content, so as to facilitate retrieving similar sample multimedia content to determine the content label and improve the accuracy of content label generation.
[0148] In an exemplary embodiment, at least two modalities include the visual modality; extracting the modal information of the first multimedia content in each of at least two modalities includes:
[0149] Extracting at least one content image set from the first multimedia content, and using the at least one content image set as the modal information of the first multimedia content in the visual modality.
[0150] Exemplarily, in the case where at least two modalities include the visual modality, the server will extract at least one content image set from the first multimedia content, and use the at least one content image set as the modal information of the first multimedia content in the visual modality.
[0151] In a specific application, if the first multimedia content is graphic-text content, the server can obtain at least one content image set by extracting pictures from the graphic-text content. The pictures in different content image sets are not completely the same. If the first multimedia content is a video, the server can obtain at least one content image set by extracting video frames from the video. The video frames in different content image sets are not completely the same. Specifically, the server can extract video frames from the video at multiple frame extraction frequencies, and then can obtain the content image sets corresponding to the respective multiple frame extraction frequencies.
[0152] In this embodiment, by extracting at least one content image set from the first multimedia content, it is possible to determine the modal information of the first multimedia content in the visual modality.
[0153] In an exemplary embodiment, the at least two modalities include a text modality and a visual modality; single-modal feature extraction is respectively performed on the modal information in the at least two modalities to obtain the single-modal features of the first multimedia content, including:
[0154] Performing text feature extraction on the modal information in the text modality to obtain text modality features, and performing visual feature extraction on at least one content image set in the modal information in the visual modality to obtain visual modality features;
[0155] According to the text modality features and the visual modality features, the single-modal features of the first multimedia content are obtained.
[0156] Exemplarily, when the at least two modalities include a text modality and a visual modality, the server performs text feature extraction on the modal information in the text modality to obtain text modality features, and performs visual feature extraction on at least one content image set in the modal information in the visual modality to obtain visual modality features, and according to the text modality features and the visual modality features, the single-modal features of the first multimedia content are obtained.
[0157] In a specific application, the server can perform text feature extraction on the modal information in the text modality through a text feature extractor (such as BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representations from Transformer)). Taking the BERT model as an example of the text feature extractor, the modal information in the text modality can be mapped to a sequence of Tokens (the smallest unit after tokenization of the modal information) and input into the BERT model, and finally the CLS Token Embedding (CLS embedding vector) is extracted as the text modality feature. It should be noted that the BERT model inserts a special marker [CLS] (Classification Token) at the beginning of the input sequence to represent the classification information of the entire sequence. During the training process, the embedding vector of [CLS] will learn the global semantic information of the entire sequence.
[0158] In a specific application, taking the modal information in the text modality including a content title and content description information as an example, such as Figure 7As shown, the server inputs the content title into the BERT model to extract the CLS Token Embedding as the text modality feature, and inputs the content description information into the BERT model to extract the CLS Token Embedding as the text modality feature. Further, as Figure 7 shown, the content description information can be obtained by inputting the first multimedia content into a pre-trained content description model.
[0159] In a specific application, the server can use a visual feature extractor (such as ViT (Vision Transformer), Swin Transformer, etc.) to extract visual features from at least one content image set to obtain visual modality features. Taking the ViT model as an example of the visual feature extractor, specifically, for each content image set, the server inputs each content image in the content image set into the ViT model for processing. For each content image, the ViT model extracts its CLS Token Embedding, and averages the CLS Token Embeddings of each content image to obtain the average value as the visual modality feature of the content image set.
[0160] In a specific application, taking the first multimedia content as a video as an example, as Figure 8 shown, the server can extract n video frames from the video to obtain a content image set, input each video frame in the content image set into the ViT model for processing. For each video frame, the ViT model extracts its CLS Token Embedding, and averages the CLS Token Embeddings of each video frame to obtain the average value as the visual modality feature of the video.
[0161] In a specific application, for the convenience of feature extraction and accurate comparison of the modality feature similarity between different modality features, the server can use the text feature extractor in the CLIP (Contrastive Language-Image Pre-Training) model for text feature extraction, and use the visual feature extractor in the CLIP model for visual feature extraction. It is a multi-modal machine learning model developed by OpenAI that can understand the relationship between pictures and text. The CLIP model is pre-trained on a large number of pictures and related description texts, aiming to let the model learn how to associate visual information and language information.
[0162] In this embodiment, by extracting text features from the modal information in the text modality, text modality features can be obtained. By extracting visual features from at least one content image set in the modal information in the visual modality, visual modality features can be obtained. Thus, based on the text modality features and the visual modality features, the determination of the unimodal features of the first multimedia content can be realized, enriching the form of the unimodal features.
[0163] In an exemplary embodiment, predicting and generating a first content label of the first multimedia content based on the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents includes:
[0164] Collect the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents to obtain a candidate content label set, and obtain the modal information of the first multimedia content in the text modality;
[0165] Based on the modal information in the text modality, classify each candidate content label in the candidate content label set to obtain the label category of each candidate content label;
[0166] Based on the label category of each candidate content label, determine the first content label of the first multimedia content.
[0167] Among them, the label category of each candidate content label is one of a set of label types predefined according to the classification method. In this embodiment, the classification method may specifically be binary classification, that is, classify whether the candidate content label is the first content label. Then, the label category of each candidate content label may specifically be the first category indicating that it is the first content label or the second category indicating that it is not the first content label.
[0168] Exemplarily, based on the retrieved multiple second multimedia contents and multiple third multimedia contents, the server will collect the respective second content labels of the retrieved multiple second multimedia contents and the respective third content labels of the retrieved multiple third multimedia contents to obtain a candidate content label set, and obtain the modal information of the first multimedia content in the text modality. Based on the modal information in the text modality, classify each candidate content label in the candidate content label set to obtain the label category of each candidate content label. Thus, based on the label category of each candidate content label, it can be determined whether each candidate content label is the first content label of the first multimedia content, and the first content label of the first multimedia content can be determined.
[0169] In a specific application, the server will extract content label features for each candidate content label in the candidate content label set based on the modality information in the text modality, obtain the label features of each candidate content label, and then classify according to the label features of each candidate content label to obtain the label category of each candidate content label.
[0170] In this embodiment, by constructing a candidate content label set and obtaining the modality information of the first multimedia content in the text modality, it is possible to use the modality information in the text modality to accurately classify each candidate content label in the candidate content label set, obtain the label category of each candidate content label, and then determine the first content label of the first multimedia content according to the label category of each candidate content label.
[0171] In an exemplary embodiment, classifying each candidate content label in the candidate content label set based on the modality information in the text modality to obtain the label category of each candidate content label includes:
[0172] Concatenate the modality information in the text modality and each candidate content label in the candidate content label set to construct a content label classification text;
[0173] Extract content label features based on the content label classification text to obtain the label features of each candidate content label;
[0174] Classify according to the label features of each candidate content label to obtain the label category of each candidate content label.
[0175] Exemplarily, the server will concatenate the modality information in the text modality and each candidate content label in the candidate content label set to construct a content label classification text, extract content label features based on the content label classification text to obtain the label category of each candidate content label, then classify according to the label features of each candidate content label to obtain the label score of each candidate content label, and finally compare the label score of each candidate content label with the first score threshold for classification to obtain the label category of each candidate content label.
[0176] Specifically, when the label score of the candidate content label is greater than the first score threshold, it is determined that the label category of the candidate content label is the first category indicating that it is the first content label, and when the label score of the candidate content label is less than or equal to the first score threshold, it is determined that the label category of the candidate content label is the second category indicating that it is not the first content label. Among them, the first score threshold can be configured according to the actual application scenario.
[0177] It can be understood that when classifying each candidate content label, specifically a binary classification, the label score is actually the probability that the candidate content label is the first content label, which can be any value between 0 and 1. Therefore, the first score threshold can be a relatively large value between 0 and 1, such as 0.8, 0.7, etc. The higher the label score, the greater the possibility that the candidate content label is the first content label.
[0178] In a specific application, the server can classify each candidate content label in the candidate content label set through a pre-trained ranking model. Among them, the pre-trained ranking model can be configured according to the actual application scenario. For example, the pre-trained ranking model can specifically be a model based on the BERT model and a binary classification layer. Then the process of classifying each candidate content label in the candidate content label set can be as Figure 9 shown Figure 9 illustrated by taking the candidate content labels as label a, label b, and label c as examples. The server will splice the modal information in the text modality and each candidate content label in the candidate content label set, and set placeholders for each candidate content label to construct the content label classification text. Then the content label classification text is input into the BERT model in the pre-trained ranking model. Through the BERT model, the modal information in the text modality and the semantics of each candidate content label in the content label classification text are deeply understood, and the Token Embedding (embedding vector) at the placeholder corresponding to each candidate content label in the output of the BERT model is extracted to classify the candidate content labels. It can be understood that the BERT model represents the content label features of the candidate content labels through Token Embedding (embedding vector).
[0179] After extracting the Token Embedding (embedding vector) at the placeholder corresponding to each candidate content label, the Token Embedding (embedding vector) at the placeholder corresponding to each candidate content label will be used as the input of the binary classification layer. Through the binary classification layer, the label score of each candidate content label can be output (as Figure 9 shown including label a score, label b score, and label c score). Furthermore, by comparing the label score of each candidate content label with the first score threshold for classification, the label category of each candidate content label can be obtained.
[0180] It should be noted that the method of classifying candidate content tags to determine the first content tag is efficient because it can process all candidate content tags simultaneously. In addition, since we perform binary classification on each candidate content tag instead of multi-classification on the entire input (the number of classes is the number of classes in the training set labels), this method supports adding new tags. This means that even for candidate content tags not seen in the training set, the server can still determine their correctness through binary classification. And this method does not require repeated training. Only one pre-training is needed to make the pre-trained ranking model adapt to the output paradigm, and it can then adapt to subsequent new candidate content tags and different business scenarios. This flexibility enables the server to maintain efficient and accurate performance in the face of constantly changing candidate content tag sets and application requirements.
[0181] In this embodiment, by concatenating the modal information in the text modality and each candidate content tag in the candidate content tag set, the construction of the content tag classification text can be achieved. Furthermore, the content tag feature extraction can be performed using the content tag classification text to obtain the tag features of each candidate content tag, and then classification can be carried out based on the tag features of each candidate content tag to obtain the tag categories of each candidate content tag, thus achieving accurate classification of each candidate content tag.
[0182] In an exemplary embodiment, predicting and generating the first content tag of the first multimedia content based on the second content tags of each of the retrieved multiple second multimedia contents and the third content tags of each of the retrieved multiple third multimedia contents includes:
[0183] Obtain the single-modal feature matching results corresponding to the retrieved multiple second multimedia contents;
[0184] Determine the tag scores of the second content tags of each of the multiple second multimedia contents according to the single-modal feature matching results;
[0185] Based on the tag scores of the second content tags of each of the multiple second multimedia contents, perform content tag screening to obtain multiple screened second content tags;
[0186] Generate the first content tag of the first multimedia content based on the multiple screened second content tags and the third content tags of each of the retrieved multiple third multimedia contents.
[0187] Exemplarily, when predicting and generating the first content label of the first multimedia content, the server obtains the unimodal matching results corresponding to multiple retrieved second multimedia contents, determines the label scores of the second content labels of the multiple second multimedia contents respectively according to the feature matching results under various modal feature matching methods in the unimodal matching results, compares the label scores of the second content labels of the multiple second multimedia contents with the second score threshold, performs content label screening, and obtains multiple screened second content labels.
[0188] Among them, the label scores of the screened second content labels are greater than the second score threshold, and the second score threshold can be configured according to the actual application scenario. Specifically, the label score represents the probability that the candidate content label is the first content label, and can specifically be any value between 0 and 1. Therefore, the second score threshold can be a relatively large value between 0 and 1, such as 0.8, 0.7, etc. The higher the label score, the greater the possibility that the candidate content label is the first content label.
[0189] In a specific application, after obtaining multiple screened second content labels, the server generates the first content label of the first multimedia content based on the multiple screened second content labels and the third content labels of the multiple retrieved third multimedia contents respectively. Specifically, the server can further score the third content labels of the multiple third multimedia contents respectively, determine the label scores of the third content labels of the multiple third multimedia contents respectively, perform content label screening based on the label scores of the third content labels of the multiple third multimedia contents respectively, obtain multiple screened third content labels, and then use the screened second content labels and the screened third content labels as the first content label of the first multimedia content.
[0190] In a specific application, the server can obtain the multimodal feature matching results corresponding to multiple retrieved third multimedia contents, determine the feature matching scores of the third content labels of the multiple third multimedia contents respectively according to the multimodal feature matching results, and use the feature matching scores as the label scores. Specifically, for each third content label, the server can first determine the associated multimedia content associated with the third content label among the multiple third multimedia contents, then determine the content similarity score of the associated multimedia content from the multimodal feature matching results, and then determine the feature matching score of the third content label based on the content similarity score of the associated multimedia content and the number of the multiple third multimedia contents.
[0191] In this embodiment, by determining the label scores of the second content labels of multiple second multimedia contents according to the unimodal feature matching results, the label scores of the second content labels of multiple second multimedia contents can be used to perform content label screening to obtain multiple screened second content labels. Furthermore, on the basis of screening out low-quality second content labels, multiple screened second content labels and the third content labels of multiple retrieved third multimedia contents can be used to accurately determine the first content label of the first multimedia content. Determining the first content label through label scoring does not require model training, is universal, supports different business scenarios, has good flexibility and scalability, can arbitrarily add more recall sources without retraining the model, and consumes less computing resources.
[0192] In an exemplary embodiment, the unimodal feature matching results include feature matching results under multiple modal feature matching methods; determining the label scores of the second content labels of multiple second multimedia contents according to the unimodal feature matching results includes:
[0193] For each modal feature matching method, according to the feature matching results under the modal feature matching method, respectively determine the feature matching scores of the second content labels of multiple second multimedia contents under the modal feature matching method;
[0194] Obtain the matching result weights corresponding to each modal feature matching method;
[0195] For each second content label, based on the feature matching scores of the second content label under each modal feature matching method and the matching result weights corresponding to each modal feature matching method, obtain the label score of the second content label.
[0196] Exemplarily, the unimodal feature matching results include feature matching results under multiple modal feature matching methods. For each modal feature matching method, the server will respectively determine the feature matching scores of the second content labels of multiple second multimedia contents under the modal feature matching method according to the feature matching results under the modal feature matching method, then obtain the matching result weights corresponding to each modal feature matching method, and for each second content label, based on the feature matching scores of the second content label under each modal feature matching method and the matching result weights corresponding to each modal feature matching method, obtain the label score of the second content label.
[0197] In a specific application, taking the unimodal features including visual modal features and text modal features as an example, such as Figure 10As shown, the multiple modality feature matching methods include four types: text-text modality feature matching (i.e., text modality feature matching with text modality sample feature), vision-vision modality feature matching (i.e., vision modality feature matching with vision modality sample feature), text-vision modality feature matching (i.e., text modality feature matching with vision modality sample feature), and vision-text modality feature matching (i.e., vision modality feature matching with text modality sample feature). The single modality feature matching result includes the feature matching results under the four modality feature matching methods (including the topK similar feature results, such as Figure 10 shown, which can also be called the topK retrieval result, where K is a positive integer greater than or equal to 1 and can be configured according to the actual application scenario), and each modality feature matching method has a corresponding matching result weight, such as Figure 10 shown as weight 1, weight 2, weight 3, and weight 4 respectively. For each second content label, the server will perform a weighted sum based on the feature matching scores of the second content label under each modality feature matching method and the matching result weights corresponding to each modality feature matching method to obtain the label score of the second content label, that is, the label result.
[0198] In a specific application, the label score of the second content label can be calculated by the following formula:
[0199] ;
[0200] where, represents the label score of label b, represents the result matching weight corresponding to the i-th modality feature matching method, represents the feature matching score under the i-th modality feature matching method, and n is the number of modality feature matching methods. For example, in Figure 10 n is 4.
[0201] In this embodiment, for each second content label, by first determining the feature matching scores of the second content label under each modality feature matching method, the accurate determination of the label score of the second content label can be achieved by combining the feature matching scores of the second content label under each modality feature matching method and the matching result weights corresponding to each modality feature matching method.
[0202] In an exemplary embodiment, according to the feature matching results under the modality feature matching method, respectively determining the feature matching scores of the second content labels of multiple second multimedia contents under the modality feature matching method includes:
[0203] For each second content label among the second content labels of multiple second multimedia contents, determine the associated multimedia content associated with the second content label in the second multimedia content set retrieved by the modal feature matching method;
[0204] Determine the content similarity score of the associated multimedia content from the feature matching results under the modal feature matching method;
[0205] Based on the content similarity score of the associated multimedia content and the number of multimedia contents in the second multimedia content set, determine the feature matching score of the second content label under the modal feature matching method.
[0206] Among them, the associated multimedia content refers to the second multimedia content in the second multimedia content set that has the second content label, and can also be understood as the second multimedia content bound with the second content label. That is, through the second content label, the associated multimedia content can be summarized and characterized. The content similarity score refers to the quantitative similarity value between the associated multimedia content and the first multimedia content, specifically, it can be the modal feature similarity between the modal features of the associated multimedia content and the modal features of the first multimedia content under the modal feature matching method.
[0207] Exemplarily, for each second content label among the second content labels of multiple second multimedia contents, the server will determine the associated multimedia content associated with the second content label in the second multimedia content set retrieved by the modal feature matching method, and then determine the content similarity score of the associated multimedia content from the feature matching results under the modal feature matching method. Based on the content similarity score of the associated multimedia content, determine the total similarity score of the second content label under the modal feature matching method, and determine the number of multimedia contents in the second multimedia content set. Calculate the ratio of the total similarity score and the number of multimedia contents, and use the ratio as the feature matching score of the second content label under the modal feature matching method.
[0208] In a specific application, when the second content label has a label coefficient, when determining the total similarity score, the server needs to perform coefficient weighting on the content similarity score of the associated multimedia content according to the label coefficient to obtain the weighted similarity score of the associated multimedia content, and then add up the weighted similarity scores of the associated multimedia content as the total similarity score. When the second content label does not have a label coefficient, the server can directly add up the content similarity scores of the associated multimedia content as the total similarity score.
[0209] In a specific application, the label coefficient may appear when the importance levels among different second content labels of the second multimedia content are different. For example, if the second content label is a news label, the importance levels among different second content labels will be different. The higher the label coefficient, the more relevant the second content label is to the second multimedia content. Another example, if the second content label is a video label, the importance levels among different second content labels may be the same, and the second content label does not have a label coefficient, or it can be understood that the label coefficients are all the same, which is 1.
[0210] In a specific application, taking the second content label having a label coefficient as an example, the feature matching score of the second content label under the i-th modal feature matching method can be calculated by the following formula:
[0211] ;
[0212] where , indicates the feature matching score of label b under the i-th modal feature matching method, k refers to the number of multimedia contents, is the label coefficient of label b, is the content similarity score of the j-th associated multimedia content under the i-th modal feature matching method, is the weighted similarity score of the j-th associated multimedia content under the i-th modal feature matching method.
[0213] In a specific application, taking the second content label not having a label coefficient as an example, the feature matching score of the second content label under the i-th modal feature matching method can be calculated by the following formula:
[0214] ;
[0215] where , indicates the feature matching score of label b under the i-th modal feature matching method, k refers to the number of multimedia contents, is the content similarity score of the j-th associated multimedia content under the i-th modal feature matching method.
[0216] In a specific application, taking the second content labels being video labels and news labels respectively as an example, such as Figure 11As shown, the feature matching results under visual-text modality feature matching and examples of the second content labels are given, including the top 1 to top K similar videos, and the labels of each of the K similar videos. In the example of video labels, the second content labels of the top 1 similar video include label a and label b, the second content labels of the top K similar videos include label b and label c, and there is no label coefficient for the video labels. In the example of news labels, the second content labels of the top 1 similar video include label a and label b with a strong correlation coefficient (represented by "strong" in Figure 11 ), and label c with a weak correlation coefficient (represented by "weak" in Figure 11 ), the second content labels of the top K similar videos include label c and label d with a strong correlation coefficient (represented by "strong" in Figure 11 ), and label b with a weak correlation coefficient (represented by "weak" in Figure 11 ). Among them, the strong correlation coefficient and the weak correlation coefficient are different and can be configured according to the actual application scenario, and the strong correlation coefficient is greater than the weak correlation coefficient.
[0217] In this embodiment, by first determining the associated multimedia content associated with the second content label, the content similarity score of the associated multimedia content and the number of multimedia content in the second multimedia content set can be used to accurately determine the feature matching score of the second content label in the modality feature matching manner.
[0218] In an exemplary embodiment, taking the content label generation method applied to the content label recall of videos as an example, the present application also provides an application scenario. This application scenario applies the above content label generation method. Specifically, the application of this content label generation method in this application scenario is as follows:
[0219] Video label recognition is an important part of video content features. By automatically generating labels for a large number of UGC (User-generated Content) videos through machines, video content features at different granularities can be provided for downstream content distribution links, such as recommendation systems and content operations, to improve the efficiency of content distribution and at the same time greatly reduce the cost of manual content review. Due to the diversity of UGC video content, the number of labels in the label library commonly used in business scenarios can reach hundreds of thousands or even millions or more, making it difficult to assign corresponding content labels to each video.
[0220] In traditional technologies, common content tag recall methods include classification methods and retrieval methods. Among them, the classification method can specifically be: inputting information such as videos, audios, texts, etc., extracting features through a neural network classification model, and outputting relevant content tags. This method requires a fixed content tag set, treating each content tag as a category, training the classification model, and outputting a multi-classification result. The retrieval method can specifically be: finding the set of content tags most relevant to the query content from a large number of candidate content tags through retrieval techniques. It usually includes the following steps: feature extraction, similarity calculation, sorting and selection, post-processing, etc. When using deep learning for video content tag retrieval, first extract video features and content tag features separately, and then calculate the feature similarity to select content tags.
[0221] However, in traditional methods, the classification method can only support the fixed categories during training and cannot recognize newly added content tags. It is necessary to re-collect data and re-train. This process is not only time-consuming but also requires a large amount of computing resources, resulting in an inability to quickly respond to the needs of new content tags. Especially for label business scenarios such as news, there are often newly added content tags such as hot events, which have high requirements for real-time performance. The retrieval method has the following problems: First, the information volume of content tag features is limited: directly extracting features from content tags, since content tags are usually simple words or phrases, the information volume is limited, resulting in insufficiently rich semantic features being extracted. Second, the retrieval effect is not good: due to the lack of semantic information in content tag features, it is difficult to accurately match relevant content directly during retrieval, affecting the accuracy and effectiveness of recall. Third, the recall source is single: traditional technologies usually rely on a single recall source, lacking diversity and being unable to make full use of information from different modalities to improve the comprehensiveness and accuracy of recall. Fourth, the sorting model is complex and lacks universality: traditional sorting models use discriminant models, which may have the problem of only supporting the content tags included in the training set and cannot adapt to new content tags.
[0222] Based on this, the present application proposes a content tag generation method to solve the above problems. By retrieving similar video samples to recall the content tags of the video, it includes how to indirectly recall the content tags of the video by retrieving similar video samples, how to make full use of multiple feature extractors such as CLIP, multi-modal classification models, etc., to extract feature vectors of different modalities, and use multiple ways to retrieve similar video samples to obtain multiple recall sources, and how to perform fusion sorting on the multiple recall sources to obtain the final content tag recall result, etc.
[0223] The content label generation method of this application has the following innovative aspects: First, the method of indirectly retrieving content labels of videos: Different from traditional video content label retrieval methods that usually rely on direct content label prediction, this application indirectly retrieves content labels by retrieving similar video samples. This method utilizes the semantic similarity between similar samples, thereby improving the accuracy and robustness of content label retrieval. Second, multi-model feature extraction: Multiple feature extractors, such as the CLIP model, multi-modal classification model, etc., are used to extract feature vectors of different modalities, such as visual, text, and fusion modalities of the video, to capture more rich and comprehensive information. Third, multi-channel retrieval and recall: Implement same-modal retrieval, cross-modal retrieval, and multi-modal retrieval. Through these retrieval methods, similar video samples are found among different modalities, increasing the diversity and accuracy of recall and obtaining multi-channel recall. Fourth, multi-channel recall fusion and sorting strategy: The recall results obtained from different recall paths are fused and sorted to ensure that the results comprehensively consider multiple information sources and are compatible with the support of new content labels. Specifically, a rule-based sorting method is provided: suitable for scenarios that do not require a large amount of computing resources, and a sorting model method is provided: using the model to sort the recall results, combining context and semantic information to provide a more intelligent and accurate sorting result. It should be noted that this method only needs to be trained once to adapt to the sorting paradigm, and subsequent support for new content labels can be directly applied without repeated training.
[0224] In an exemplary embodiment, the content label generation method of this application can be applied in a content distribution scenario. As Figure 12 shown, the video enters the content processing link from the content production link, obtains corresponding content features through a human-machine collaboration method, and enters the downstream content distribution link, such as recommendation, etc. The content label generation method provided in this embodiment belongs to the machine tagging link in the content processing of this video. Specifically, in the process of content processing for the video, the video can be tagged based on the collaboration method of machine tagging and manual tagging, that is, corresponding content labels are added to the video. It can be understood that the content label generation method of this application can be applied to multiple business scenarios, such as video platforms, news platforms, and short video platforms, etc.
[0225] In an exemplary embodiment, the flowchart of the content label generation method of this application can be as Figure 13 shown. The overall technical concept is to predict the content labels of the video based on retrieving similar video samples, as Figure 13As shown, it includes parts such as model feature extraction, multi-channel recall, and multi-channel fusion. In this application, a multimedia content library containing a large number of video samples with labeled content tags is constructed. The feature extractor is used to extract the feature vectors of the video samples, and then video samples with similar features to the target video (i.e., the first multimedia content) are retrieved from the multimedia content library, and the content tags of the target video are inferred from the content tags of these video samples. Taking at least two modalities including the visual modality and the text modality as examples, each part will be introduced specifically below. It can be understood that the sample video and the target video can be videos in any format and have video titles.
[0226] In a specific application, feature extraction can be divided into two parts. One part is the feature extraction for the target video, and the other part is the feature extraction for the video samples. Among them, for the feature extraction of the target video, it is for the visual modality information and text modality information extracted from the target video, and the same modality can also have different modality information. For example, the text modality information can be the video title, video subtitle text, video description information obtained using a large language model, etc. After these different text modality information are subjected to feature extraction, they are used as independent recall paths respectively and do not conflict with each other. For video samples, different feature extractors can also be used to obtain different recall paths. In this embodiment, for feature extraction, the use methods are introduced taking the CLIP model and the multi-modal classification model as examples. It can be understood that the feature extraction model used in this application can be from any feature extraction model, and the training of the feature extraction model is not introduced here.
[0227] For video samples, in this application, the CLIP model is used to extract visual modality features and text modality features respectively. The ViT model in the CLIP model is used to extract visual modality features, and the BERT model is used to extract text modality features. At the same time, in this application, the multi-modal classification model is used as a multi-modal feature extractor to extract the multi-modal sample features of the video samples, and the output feature of the last layer (the layer before the classification head) of the multi-modal classification model is taken.
[0228] For the target video, regarding the visual modality feature extraction part, in this application, the method of frame extraction is used to obtain video frames as visual modality information, and the multiple extracted video frames are input into the visual feature extractor of the CLIP model, that is, the ViT model, to obtain the visual modality features of the target video. The specific steps are as follows: First, n frames are extracted from the target video, and each frame is regarded as an independent image. These image frames are input into the ViT model for processing at the same time. For each frame, the ViT model extracts the Embedding of its CLS Token. Next, in order to obtain the final feature representation of the target video, we average the CLS Token Embeddings of the n frames of the same video, and this average value is used as the visual modality feature of the target video.
[0229] It can be understood that in this application, the target video can be extracted multiple times for video frames by presetting multiple frame extraction frequencies to obtain multiple visual modality information, that is, each video frame extraction can obtain a visual modality information, and different visual modality information is not exactly the same because the corresponding frame extraction frequencies are different. In this way, rich visual modality information can be extracted, and then multi-channel recall can be performed to improve the recall accuracy.
[0230] For the text modality feature extraction part, in this application, the text modality information is processed by the BERT model. The text modality information is mapped into a Token sequence and input into BERT, and finally the CLS Token Embedding is extracted as the NLP Embedding of the sample. In a specific application, in this application, the video title and the video description information predicted by the large language model can be used as two text modality information. It can be understood that the video title is usually relatively short, and the video description information predicted by the large language model can utilize its powerful understanding ability to provide richer and more detailed video information.
[0231] For the multi-modal fusion feature extraction part, in this application, the text modality information and the visual modality information are sent into the multi-modal classification model at the same time to extract and integrate information from different modalities. The visual modality information is video frames, and the text modality information is information such as video titles and video description information. This method can effectively combine visual features and text features to improve the ability to understand and analyze video content. Through this multi-modal fusion feature extraction strategy, we can capture the semantic information of the video more comprehensively and enhance the overall performance of the system.
[0232] In a specific application, the multi-channel recall part mainly involves retrieving similar video samples. Specifically, in this application, a multimedia content library is pre-constructed, and the multimedia content library includes the single-modal sample features and multi-modal sample features of multiple video samples (i.e., sample multimedia content) with manually labeled content tags. For a newly generated target video that needs to be labeled, based on extracting the single-modal features (including visual-modal features and text-modal features) and multi-modal fusion features of the target video through the above feature extraction method, we can use the single-modal features and multi-modal fusion features to retrieve similar features in the multimedia content library, and then determine similar video samples.
[0233] In a specific application, visual-modal features, text-modal features, and multi-modal fusion features are all used for video retrieval. To simplify the process, as Figure 14 shown, taking the case where only one path of input is used for visual-modal features and text-modal features respectively, the multi-channel recall in this application is described. It can be understood that this application can also be extended to multi-path input. It should be noted that since the CLIP model is trained through cross-modal contrast learning, it can not only perform modal feature matching and calculate feature similarity between features of the same modality, but also perform cross-modal feature matching and calculate feature similarity between visual-modal features and text-modal features of different modalities. The video retrieval in this application can be divided into two parts: single-modal feature retrieval and multi-modal feature retrieval. The following combines Figure 14 to describe the video retrieval in this application.
[0234] In single-modal feature retrieval, through same-modal feature retrieval and cross-modal feature retrieval, four recall results (i.e., feature matching results) can be obtained, namely Figure 14 the visual-visual recall result, text-visual recall result, text-text recall result, and visual-text recall result in Figure 14As shown, the multi-modal fusion features of the target video are used to retrieve the multi-modal sample features, and the recall result of multi-modal - multi-modal is obtained.
[0235] It should be noted that the retrieval method involved in this application has strong scalability, and the system performance can be enhanced by adding more recall sources. For example, in the case where the video title is used as the text modal information, the video description information can be generated by a large language model as the text modal information to add one more recall source. This can further improve the accuracy and comprehensiveness of the retrieval system.
[0236] In a specific application, in the multi-way recall part, similar video samples can be retrieved, and then the content labels manually annotated with the similar video samples can be determined, and the corresponding retrieval scores can be determined. The retrieval score can specifically be the content similarity score, that is, the modal feature similarity. In the multi-way recall part, each path will output the desired recall result. Taking the recall result of text - text as an example, the form of the recall result can be: NLP - NLP (text - text): [top1: {"score": 0.9, "tag": ["content label a", "content label b"]}, top2: {"score": 0.7, "tag": ["content label a", "content label c"]}, top3...]. Taking the recall result of vision - text as an example, the form of the recall result can be: CV - NLP (vision - text): [top1: {"score": 0.8, "tag": ["content label c", "content label b"]}, top2: {"score": 0.6, "tag": ["content label d"]}, top3...]. In the case of obtaining the multi-way recall results, it is necessary to further fuse and sort the multi-way recall results to obtain the final prediction result (i.e., the first content label).
[0237] In a specific application, in the multi-way fusion and sorting part, two fusion strategies are provided in this application. First, a sorting model is used to fuse and sort the multi-way recall results. Specifically, the multiple content labels of all recall paths (including multiple second content labels and multiple third content labels) can be pooled into a candidate content label set, and in combination with the text modal information, each candidate content label in the candidate content label set is input into the sorting model for processing at the same time. The processing process of the sorting model for each candidate content label can be as Figure 9 shown ( Figure 9Among them, the candidate content labels are labeled as label a, label b, and label c), and each candidate content label is analyzed by extracting Token Embedding. For each candidate content label, we extract its corresponding Token Embedding and perform binary classification to determine whether the content label is correct. Specifically, the output of the binary classification represents "yes" or "no", indicating the correctness of the candidate content label, and its corresponding score is the label score of the content label.
[0238] It should be noted that the method of classifying candidate content labels to determine the first content label is efficient because it can process all candidate content labels simultaneously. In addition, since we perform binary classification on each candidate content label instead of multi-classification on the entire input (the number of classes is the number of classes in the training set labels), this method supports adding new labels. This means that even for candidate content labels not seen in the training set, the server can still determine their correctness through binary classification. And this method does not require repeated training. Only one-time pre-training is needed to make the pre-trained ranking model adapt to the output paradigm, and then it can adapt to subsequent new candidate content labels and different business scenarios. This flexibility enables the server to maintain high efficiency and accurate performance when facing continuously changing candidate content label sets and application requirements.
[0239] In addition, a rule-based fusion ranking method is provided in this application. The specific fusion strategy can be: for N-way recall, take the top K similar video samples for each way. Each video sample is accompanied by a human-reviewed content label, and each video sample has a content similarity score Score. Different weights are set according to the revenue difference of each way (specifically, it can be recall accuracy). For example, the effect of the same-modal recall source is often better than that of the cross-modal, so the weight should be higher. The scores of each content label are weighted and summed to obtain the final score of each content label, and then threshold filtering is performed to obtain the final machine-generated content labels. Where K is a positive integer greater than or equal to 1 and can be configured according to the actual application scenario.
[0240] In a specific application, taking N-way recall as 4-way recall, and the 4-way recall being visual-visual recall, text-visual recall, text-text recall, and visual-text recall as an example, as Figure 15 shown, take the top K similar video samples for each way, that is, Figure 15 the top K retrieval results in, each video sample is accompanied by a human-reviewed second content label, and each video sample has a content similarity score Score. Different weights (i.e., matching result weights, such as Figure 15Shown are Weight 1, Weight 2, Weight 3, and Weight 4 respectively. Since the recall source effect of the same modality is often better than that of cross-modal, the weight should be higher. The scores of each content label are weighted and summed to obtain the final score of each content label, and then threshold filtering is performed to obtain the final machine-typed content label.
[0241] For each second content label, the server will perform weighted summation based on the feature matching scores recalled by each path of the second content label and the weights recalled by each path to obtain the label score of the second content label, that is, the label result. In a specific application, the label score of the second content label can be calculated by the following formula:
[0242] ;
[0243] Among them, represents the label score of label b, represents the weight of the i-th path, represents the feature matching score of the i-th path, n is the number of recall paths, and in Figure 15 n is 4.
[0244] In a specific application, when calculating the feature matching scores recalled by each path, according to whether the second content label has a label coefficient, two methods can be used for calculation respectively. In a specific application, taking the second content labels as video labels and news labels respectively as an example, as Figure 15 shown, the recall results (i.e., retrieval results) of vision-text and examples of the second content labels are given, including top1 similar videos to topK similar videos, and the labels of each of the K similar videos. In the example of video labels, the second content labels of the top1 similar video include label a and label b, and the second content labels of the topK similar videos include label b and label c, and the video labels have no label coefficients. In the example of news labels, the second content labels of the top1 similar video include label a and label b with strong correlation coefficients (denoted as strong in Figure 15 ), and label c with weak correlation coefficients (denoted as weak in Figure 15 ), and the second content labels of the topK similar videos include label c and label d with strong correlation coefficients (denoted as strong in Figure 15 ), and label b with weak correlation coefficients (denoted as weak in Figure 15 ). Among them, the strong correlation coefficient and the weak correlation coefficient are different and can be configured according to the actual application scenario, and the strong correlation coefficient is greater than the weak correlation coefficient.
[0245] In a specific application, taking the second content label having a label coefficient as an example, the feature matching score of the second content label in the i-th path can be calculated by the following formula:
[0246] ;
[0247] Among them, , which is the feature matching score of label b in the i-th path, and k is the number of multimedia contents in the i-th path, is the label coefficient of label b, is the content similarity score of the j-th associated multimedia content in the i-th path, is the weighted similarity score of the j-th associated multimedia content in the i-th path.
[0248] In a specific application, taking the second content label without a label coefficient as an example, the feature matching score of the second content label in the i-th path can be calculated by the following formula:
[0249] ;
[0250] Among them, , which is the feature matching score of label b in the i-th path, and k is the number of multimedia contents in the i-th path, is the content similarity score of the j-th associated multimedia content in the i-th path.
[0251] It should be noted that in the rule-based fusion sorting method, there are two adjustable hyperparameters, the number of samples retrieved per path topK and the weight Wi of each path. Grid Search can be used to obtain the optimal hyperparameters.
[0252] This strategy does not require training a sorting model, is general, supports different business scenarios. For example, in the news scenario, content labels need to be marked as strongly relevant and weakly relevant, and only the strong and weak correlation coefficients need to be increased. It has good flexibility and scalability, can add more recall sources arbitrarily without retraining the sorting model, and consumes less computing resources.
[0253] It should be noted that in the actual business scenario, different businesses have their own content label systems. For example, the video business has its own content label system, while the news business has another set of content label systems. The same video needs to be marked with different content labels under different business systems. For example, a video is marked with content labels under the video system in the video business scenario and with content labels under the news system in the news business scenario. Thanks to our retrieval method, there is no need to train separately under different business systems. Our link only needs to extract features once, and then collide with the multimedia content libraries of different business systems respectively, so as to realize inferring and generating multiple sets of business content labels at one time.
[0254] In a specific application, taking the business scenarios including video business scenario, news business scenario, first short video business scenario, second short video business scenario, and medium video business scenario as examples, for the target video, after obtaining the unimodal features and multimodal fusion features of the target video through the feature extraction model, by colliding with the multimedia content libraries of different business systems respectively, multiple sets of business content tags can be generated in one inference, that is Figure 16 the video business tags, news business tags, first short video business tags, second short video business tags, and medium video business tags in Figure 16 . Among them, a medium video refers to a video whose video duration is less than that of the video in the video business scenario and greater than that of the short video. The short video types of the first short video and the second short video are different.
[0255] It should be noted that this application proposes a content tag generation method based on similar sample retrieval. Based on the retrieval method, new content tags can be supported without retraining. Through indirect retrieval, the rich semantic information of the samples can be fully utilized. Through multi-way recall, the accuracy and robustness of the system can be improved. It has strong scalability and multiple recall paths can be added. The multi-way recall fusion strategy is simple, compatible with the support of new content tags, and can be seamlessly migrated to other business data without repeated consumption of training resources. Multiple sets of business content tags can be generated in one inference. Finally, the efficiency of machine review can be significantly improved and the cost of manual review can be saved. Specifically, the content tag generation method of this application mainly has the following advantages:
[0256] First, a novel method for indirectly recalling content tags is provided: in this application, content tags are indirectly recalled by retrieving similar samples, rather than directly retrieving content tags. Specifically, in this application, by constructing a multimedia content library containing a large number of labeled samples, using a feature extractor to extract the feature vectors of the first multimedia content, and then retrieving the sample multimedia content with similar features to the target video in the multimedia content library, the content tags of the first multimedia content are inferred from the content tags of these sample multimedia contents. Through this indirect retrieval method, the timeliness problem of the support of the classification method for new content tags can be avoided, and the resource consumption caused by repeated training can also be avoided. The indirect retrieval scheme can fully utilize the semantic information of the sample multimedia content and the semantic similarity between similar samples, avoid directly retrieving content tags lacking rich semantic information, and improve the accuracy and robustness of content tag recall.
[0257] Second, multi-modal feature extraction: In this application, a variety of feature extractors are used to extract different modal features such as the visual modality, text modality, and fusion modality (i.e., multi-modal) of the multimedia content. Feature extractors can use models such as CLIP and multi-modal classification models to capture features at different levels. Our method is not limited to a specific feature extractor or the number of extractors used. Different feature extractors can provide different recall paths and seamlessly integrate into our retrieval system.
[0258] Third, multi-path retrieval and recall: After using a variety of feature extractors to obtain features in this application, through intra-modal feature retrieval, cross-modal feature retrieval, and multi-modal feature retrieval, similar samples can be found between different modalities, providing more recall paths. It should be noted that we can obtain more recall paths through various means, which can be directly integrated into our retrieval system without additional adjustment. For example, using a large language model to generate content description information and extracting corresponding features to obtain new recall paths.
[0259] Fourth, a novel and simple content label fusion and ranking strategy: In this application, fusion ranking and scoring for multiple recall sources are provided. Specifically, there are two methods. One is a rule-based method that does not require training a model and uses the results of multiple recalls for fusion ranking, which is suitable for scenarios that do not require a large amount of computing resources. The other is a model-based method that ranks the recall results by combining context and semantic information. Both methods have good scalability, can add other recall sources at any time, and can directly support new content labels without retraining the model.
[0260] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps in other steps.
[0261] Based on the same inventive concept, the embodiments of this application also provide a content label generation device for implementing the content label generation method described above. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the content label generation device provided below can refer to the limitations on the content label generation method in the above text and will not be repeated here.
[0262] In an exemplary embodiment, as Figure 17 shown, a content label generation device is provided, including: a modality feature extraction module 1702, a single-modality feature retrieval module 1704, a multi-modality feature retrieval module 1706, and a content label prediction module 1708, where:
[0263] The modality feature extraction module 1702 is configured to obtain a first multimedia content, and extract single-modality features and multi-modality fusion features from the first multimedia content respectively;
[0264] The single-modality feature retrieval module 1704 is configured to retrieve a second multimedia content similar to the first multimedia content in a single modality according to the single-modality features;
[0265] The multi-modality feature retrieval module 1706 is configured to retrieve a third multimedia content similar to the first multimedia content in at least two modalities according to the multi-modality fusion features;
[0266] The content label prediction module 1708 is configured to predict and generate a first content label of the first multimedia content based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents.
[0267] The above content label generation device can extract multiple different modality features by obtaining the first multimedia content and extracting single-modality features and multi-modality fusion features from the first multimedia content respectively. By retrieving a second multimedia content similar to the first multimedia content in a single modality according to the single-modality features, and retrieving a third multimedia content similar to the first multimedia content in at least two modalities according to the multi-modality fusion features, the second multimedia content and the third multimedia content similar to the first multimedia content can be obtained through the multi-way retrieval and recall method of single-modality retrieval and multi-modality retrieval. Furthermore, the first content label of the first multimedia content can be predicted and generated based on the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents. In the whole process, by using multiple different modality features of the first multimedia content to perform multi-way recall retrieval of similar multimedia contents to indirectly recall content labels, the multiple different modality features of the first multimedia content, as well as the semantic similarity between the first multimedia content and the similar multimedia contents, can be fully utilized for content label generation, and the accuracy of content label generation can be improved.
[0268] In an exemplary embodiment, the unimodal feature retrieval module is further configured to determine a multimedia content library; the multimedia content library includes unimodal sample features of multiple sample multimedia contents, perform unimodal feature matching on the unimodal features and the unimodal sample features of the multiple sample multimedia contents respectively to obtain unimodal feature matching results, and retrieve, from the multimedia content library according to the unimodal feature matching results, a second multimedia content that is similar to the first multimedia content in a single modality.
[0269] In an exemplary embodiment, the unimodal feature retrieval module is further configured to perform same-modal feature matching on each modality feature in the unimodal features and the same-kind modality features in the unimodal sample features of the multiple sample multimedia contents respectively to obtain same-modal feature matching results, and obtain unimodal feature matching results based on the same-modal feature matching results.
[0270] In an exemplary embodiment, the unimodal feature retrieval module is further configured to perform cross-modal feature matching on each modality feature in the unimodal features and different-kind modality features in the unimodal sample features of the multiple sample multimedia contents respectively to obtain cross-modal feature matching results, and obtain unimodal feature matching results based on the same-modal feature matching results and the cross-modal feature matching results.
[0271] In an exemplary embodiment, the unimodal features include visual modality features and text modality features. The unimodal feature retrieval module is further configured to perform cross-modal feature matching on the visual modality features and different-kind modality features in the unimodal sample features of the multiple sample multimedia contents respectively to obtain a first cross-modal matching result for the visual modality, perform cross-modal feature matching on the text modality features and different-kind modality features in the unimodal sample features of the multiple sample multimedia contents respectively to obtain a second cross-modal matching result for the text modality, and obtain cross-modal feature matching results according to the first cross-modal matching result and the second cross-modal matching result.
[0272] In an exemplary embodiment, the multimodal feature retrieval module is further configured to determine a multimedia content library; the multimedia content library includes multimodal sample features of multiple sample multimedia contents, perform multimodal feature matching on the multimodal fusion features and the multimodal sample features of the multiple sample multimedia contents respectively to obtain multimodal feature matching results, and retrieve, from the multimedia content library according to the multimodal feature matching results, a third multimedia content that is similar to the first multimedia content in at least two modalities.
[0273] In an exemplary embodiment, the modality feature extraction module is further configured to extract modality information of the first multimedia content in at least two modalities respectively, perform single-modal feature extraction on the modality information in at least two modalities respectively to obtain single-modal features of the first multimedia content, and jointly perform multi-modal feature extraction on the modality information in at least two modalities to obtain multi-modal fusion features of the first multimedia content.
[0274] In an exemplary embodiment, at least two modalities include the text modality, and the modality feature extraction module is further configured to input the first multimedia content into a pre-trained content description model, predict content description information of the first multimedia content through the pre-trained content description model, and obtain modality information of the first multimedia content in the text modality based on the content description information.
[0275] In an exemplary embodiment, at least two modalities include the visual modality, and the modality feature extraction module is further configured to extract at least one set of content images from the first multimedia content, and use the at least one set of content images as the modality information of the first multimedia content in the visual modality.
[0276] In an exemplary embodiment, at least two modalities include the text modality and the visual modality, and the modality feature extraction module is further configured to perform text feature extraction on the modality information in the text modality to obtain text modality features, perform visual feature extraction on at least one set of content images in the modality information in the visual modality to obtain visual modality features, and obtain single-modal features of the first multimedia content according to the text modality features and the visual modality features.
[0277] In an exemplary embodiment, the content label prediction module is further configured to collect the second content labels of the retrieved multiple second multimedia contents and the third content labels of the retrieved multiple third multimedia contents respectively to obtain a candidate content label set, obtain the modality information of the first multimedia content in the text modality, classify each candidate content label in the candidate content label set based on the modality information in the text modality to obtain the label category of each candidate content label, and determine the first content label of the first multimedia content based on the label category of each candidate content label.
[0278] In an exemplary embodiment, the content label prediction module is further configured to splice the modality information in the text modality and each candidate content label in the candidate content label set to construct a content label classification text, perform content label feature extraction based on the content label classification text to obtain the label features of each candidate content label, and classify according to the label features of each candidate content label to obtain the label category of each candidate content label.
[0279] In an exemplary embodiment, the content label prediction module is further configured to obtain the unimodal feature matching results corresponding to a plurality of retrieved second multimedia contents, determine the label scores of the second content labels of the plurality of second multimedia contents respectively according to the unimodal feature matching results, perform content label screening based on the label scores of the second content labels of the plurality of second multimedia contents, obtain a plurality of screened second content labels, and generate the first content label of the first multimedia content based on the plurality of screened second content labels and the third content labels of the plurality of retrieved third multimedia contents.
[0280] In an exemplary embodiment, the unimodal feature matching results include the feature matching results under various modality feature matching methods. The content label prediction module is further configured to, for each modality feature matching method, respectively determine the feature matching scores of the second content labels of the plurality of second multimedia contents under the modality feature matching method according to the feature matching results under the modality feature matching method, obtain the matching result weight corresponding to each modality feature matching method, and for each second content label, obtain the label score of the second content label based on the feature matching score of the second content label under each modality feature matching method and the matching result weight corresponding to each modality feature matching method.
[0281] In an exemplary embodiment, the content label prediction module is further configured to, for each second content label among the second content labels of the plurality of second multimedia contents, determine the associated multimedia content associated with the second content label in the second multimedia content set retrieved by the modality feature matching method, determine the content similarity score of the associated multimedia content from the feature matching results under the modality feature matching method, and determine the feature matching score of the second content label under the modality feature matching method based on the content similarity score of the associated multimedia content and the number of multimedia contents in the second multimedia content set.
[0282] Each module in the above content label generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above respective modules.
[0283] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal. Taking the computer device being a server as an example, its internal structure diagram can be as Figure 18As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device is used to store data such as a multimedia content library. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements a content tagging generation method.
[0284] Those skilled in the art can understand that Figure 18 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0285] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0286] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0287] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0288] It should be noted that the data involved in the present application (including but not limited to data for analysis, stored data, displayed data, etc.) are all data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0289] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0290] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.
[0291] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating content tags, characterized in that: The method comprises: Acquire first multimedia content, and extract single-modal features and multi-modal fusion features from the first multimedia content respectively; Retrieving, according to the single modality feature, second multimedia content that is similar to the first multimedia content in a single modality; Retrieving third multimedia content similar to the first multimedia content in at least two modes according to the multimodal fusion feature; Based on the respective second content tags of the retrieved plurality of second multimedia contents and the respective third content tags of the retrieved plurality of third multimedia contents, a first content tag of the first multimedia content is predicted and generated.
2. The method according to claim 1, characterized in that The retrieving, according to the single modality feature, second multimedia content that is similar to the first multimedia content in a single modality comprises: Determine a multimedia content library; the multimedia content library includes unimodal sample features of each of a plurality of sample multimedia contents; Performing unimodal feature matching on the unimodal feature and the unimodal sample features of each of the plurality of sample multimedia contents to obtain unimodal feature matching results; According to the single-modal feature matching result, second multimedia content similar to the first multimedia content in a single modality is retrieved from the multimedia content library.
3. The method according to claim 2, characterized in that The performing unimodal feature matching on the unimodal feature and the unimodal sample features of each of the plurality of sample multimedia contents to obtain unimodal feature matching results includes: Performing homomodal feature matching on each modal feature in the unimodal features with the modal features of the same type in the unimodal sample features of the plurality of sample multimedia contents, respectively, to obtain homomodal feature matching results; Based on the homomodal feature matching result, a unimodal feature matching result is obtained.
4. The method according to claim 3, characterized in that The obtaining of a single-modal feature matching result based on the same-modal feature matching result includes: Performing cross-modal feature matching on each modality feature in the unimodal features with different types of modality features in the unimodal sample features of the plurality of sample multimedia contents, respectively, to obtain a cross-modality feature matching result; Based on the same-modal feature matching result and the cross-modal feature matching result, a unimodal feature matching result is obtained.
5. The method according to claim 4, characterized in that The unimodal features include visual modal features and textual modal features; and the cross-modal feature matching is performed on each modal feature in the unimodal features with different types of modal features in the unimodal sample features of the plurality of sample multimedia contents to obtain a cross-modal feature matching result, including: Performing cross-modal feature matching on the visual modality feature and different types of modality features in the unimodal sample features of the plurality of sample multimedia contents, respectively, to obtain a first cross-modal matching result for the visual modality; Performing cross-modal feature matching on the text modality feature and different types of modality features in the unimodal sample features of the plurality of sample multimedia contents, respectively, to obtain a second cross-modal matching result for the text modality; A cross-modal feature matching result is obtained according to the first cross-modal matching result and the second cross-modal matching result.
6. The method according to claim 1, characterized in that The retrieving third multimedia content similar to the first multimedia content in at least two modes according to the multimodal fusion feature includes: Determine a multimedia content library; the multimedia content library includes multimodal sample features of each of a plurality of sample multimedia contents; Performing multimodal feature matching on the multimodal fusion feature and the multimodal sample features of each of the plurality of sample multimedia contents to obtain a multimodal feature matching result; According to the multimodal feature matching result, third multimedia content similar to the first multimedia content in at least two modalities is retrieved from the multimedia content library.
7. The method according to claim 1, characterized in that The extracting of single-modal features and multi-modal fusion features from the first multimedia content respectively includes: Extracting modality information of at least two modalities from the first multimedia content; Extracting single-modal features from the modal information of each of the at least two modalities to obtain single-modal features of the first multimedia content; The modal information of each of the at least two modalities is combined to perform multimodal feature extraction to obtain a multimodal fusion feature of the first multimedia content.
8. The method according to claim 7, characterized in that The at least two modalities include a text modality; and extracting modality information of at least two modalities from the first multimedia content includes: Inputting the first multimedia content into a pre-trained content description model, and predicting content description information of the first multimedia content by using the pre-trained content description model; Based on the content description information, modal information of the first multimedia content in the text mode is obtained.
9. The method according to claim 7, characterized in that: The at least two modalities include a visual modality; and extracting modality information of at least two modalities from the first multimedia content includes: At least one content image set is extracted from the first multimedia content, and the at least one content image set is used as modality information of the first multimedia content in the visual modality.
10. The method according to claim 7, characterized in that The at least two modalities include a textual modality and a visual modality; and the extracting of single-modal features from the modal information of each of the at least two modalities to obtain single-modal features of the first multimedia content includes: Performing text feature extraction on the modal information in the text modality to obtain text modal features, and performing visual feature extraction on at least one content image set in the modal information in the visual modality to obtain visual modal features; A unimodal feature of the first multimedia content is obtained according to the textual modal feature and the visual modal feature.
11. The method according to any one of claims 1 to 10, characterized in that: The predicting and generating the first content tag of the first multimedia content based on the respective second content tags of the retrieved plurality of second multimedia contents and the respective third content tags of the retrieved plurality of third multimedia contents comprises: Collecting the second content tags of the plurality of retrieved second multimedia contents and the third content tags of the plurality of retrieved third multimedia contents to obtain a candidate content tag set, and acquiring modality information of the first multimedia content in text mode; Based on the modality information in the text modality, classify each candidate content tag in the candidate content tag set to obtain a tag category of each candidate content tag; Based on the tag category of each of the candidate content tags, a first content tag of the first multimedia content is determined.
12. The method according to claim 11, characterized in that The step of classifying each candidate content tag in the candidate content tag set based on the modality information in the text modality to obtain a tag category of each candidate content tag includes: Concatenating the modal information under the text modality and each candidate content tag in the candidate content tag set to construct a content tag classification text; Performing content tag feature extraction based on the content tag classification text to obtain a tag feature of each candidate content tag; Classification is performed according to the tag feature of each candidate content tag to obtain a tag category of each candidate content tag.
13. The method according to any one of claims 1 to 10, characterized in that: The predicting and generating the first content tag of the first multimedia content based on the respective second content tags of the retrieved plurality of second multimedia contents and the respective third content tags of the retrieved plurality of third multimedia contents comprises: Obtaining single-modal feature matching results corresponding to the retrieved plurality of second multimedia contents; Determining, according to the unimodal feature matching result, a label score of a second content label of each of the plurality of second multimedia contents; screening the content tags based on the tag scores of the second content tags of the plurality of second multimedia contents to obtain a plurality of screened second content tags; A first content tag for the first multimedia content is generated based on the plurality of filtered second content tags and respective third content tags of the plurality of retrieved third multimedia content.
14. The method according to claim 13, characterized in that The unimodal feature matching result includes feature matching results in multiple modal feature matching modes; and determining the label scores of the second content labels of the plurality of second multimedia contents according to the unimodal feature matching result includes: For each of the modal feature matching modes, determining a feature matching score of each of the second content tags of the plurality of second multimedia contents in the modal feature matching mode according to a feature matching result in the modal feature matching mode; Obtaining the matching result weight corresponding to each of the modal feature matching methods; For each of the second content tags, a tag score of the second content tag is obtained based on a feature matching score of the second content tag in each of the modal feature matching modes and a matching result weight corresponding to each of the modal feature matching modes.
15. The method according to claim 14, characterized in that The determining, according to the feature matching results in the modal feature matching mode, the feature matching scores of the second content tags of the plurality of second multimedia contents in the modal feature matching mode respectively, includes: For each second content tag of each of the plurality of second multimedia contents, determining, in the second multimedia content set retrieved by the modal feature matching method, associated multimedia content associated with the second content tag; Determining a content similarity score of the associated multimedia content from the feature matching results in the modal feature matching manner; Based on the content similarity scores of the associated multimedia contents and the number of multimedia contents in the second multimedia content set, a feature matching score of the second content tag in the modality feature matching manner is determined.
16. A content tag generating device, characterized in that: The device comprises: A modal feature extraction module, used to obtain first multimedia content, and extract single modal features and multimodal fusion features from the first multimedia content respectively; a single-modal feature retrieval module, configured to retrieve second multimedia content similar to the first multimedia content in a single modality according to the single-modal feature; a multimodal feature retrieval module, configured to retrieve third multimedia content similar to the first multimedia content in at least two modes according to the multimodal fusion feature; The content tag prediction module is used to predict and generate a first content tag of the first multimedia content based on the second content tags of each of the retrieved plurality of second multimedia contents and the third content tags of each of the retrieved plurality of third multimedia contents.
17. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.