Label data generation method and device, electronic equipment and storage medium
Through the automated label data generation method, candidate tags are extracted using the description text set and combined with the modal feature extraction of multimedia data, the problem of unstable artificially generated label data is solved, efficient and accurate generation and storage of label data is achieved, and the retrieval effect of multimedia data is improved.
Patent Information
- Application Number
- CN202510387404.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the generation of tag data relies on manually constructing content tags, resulting in unstable generation effects and affecting the retrieval effect and user experience of multimedia data.
By obtaining the description text set associated with hot topic content, keyword extraction is performed to obtain the candidate tag set, and when the total number of historical data of the candidate tag does not exceed the preset threshold, the candidate tag is used as the content tag, and feature extraction and storage is performed in combination with different modalities of the multimedia data, and search features are generated and stored in the tag data set of the corresponding modal.
It improves the efficiency of candidate tag generation, realizes comprehensive analysis of multimedia data, generates more appropriate search features, provides systematic tag data storage and refined classification, ensures the timeliness and accuracy of tag data, and improves the effectiveness of search basis.
Smart Images

Figure CN120256653A_ABST
Abstract
Description
Background Art
[0002] With the development of information technology, relevant objects can understand the latest hot topics by browsing the retrieved multimedia data. Based on this, in order to effectively manage multimedia data, content tags are usually added to multimedia data based on the latest tag data before the multimedia data is published. Therefore, it is crucial to update the tag data in a timely manner according to the latest hot topics during the content tag configuration of multimedia data.
[0003] Currently, when generating tag data, usually technicians frequently manually sort out content tags based on the latest hot topics, construct tag features for the sorted content tags, and store the content tags and the corresponding tag features as new tag data, so that the matching content tags can be determined by calculating the feature similarity between the multimedia data and the content tags.
[0004] However, in the above method of generating tag data, since content tags need to be manually constructed, the generation of content tags is highly dependent on the personal experience of technicians, resulting in very unstable generation effects of content tags. As a result, the generated tag data cannot adapt to the content tag annotation requirements of multimedia data, which greatly affects the retrieval effect of multimedia data and reduces the browsing experience of relevant objects for multimedia data. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, electronic device, and storage medium for generating tag data, which can generate tag data for retrieval basis in a timely and efficient manner for hot topics.
[0006] In a first aspect, a method for generating tag data is proposed, including: Obtain a set of description texts associated with a hot topic, and extract keywords from the set of description texts to obtain a corresponding set of candidate tags; When it is determined that there is at least one candidate tag in the set of candidate tags whose total number of associated historical data does not exceed a preset quantity threshold, respectively use the at least one candidate tag as a content tag, and obtain each multimedia data associated with the hot topic; each multimedia data contains data content of at least one data modality; the total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which the tag data containing the candidate tag is generated when generating content tags; For each data modality in each multimedia data, respectively perform the following storage operations: Extract the feature of the data content in a data modality from a piece of multimedia data to obtain a retrieval feature, and store the label data generated based on the retrieval feature and at least one content label into the label dataset set corresponding to the one data modality.
[0007] In a second aspect, a label data generation device is proposed, including: An acquisition unit, configured to acquire a set of description texts associated with a hot content, and perform keyword extraction on the set of description texts to obtain a corresponding set of candidate labels; A determination unit, configured to, when it is determined that there is at least one candidate label in the set of candidate labels whose total number of associated historical data does not exceed a preset quantity threshold, respectively use the at least one candidate label as a content label, and acquire each piece of multimedia data associated with the hot content; each piece of multimedia data includes data content in at least one data modality; the total number of historical data associated with a candidate label refers to the total number of historical multimedia data based on which label data including the candidate label are generated when generating content labels; A storage unit, configured to perform the following storage operations respectively for each data modality in each piece of multimedia data: Extract the feature of the data content in a data modality from a piece of multimedia data to obtain a retrieval feature, and store the label data generated based on the retrieval feature and at least one content label into the label dataset set corresponding to the one data modality.
[0008] Optionally, after performing the storage operations respectively for each data modality in each piece of multimedia data, the device further includes an execution unit, and the execution unit is configured to: Acquire data to be processed, and respectively extract corresponding features to be matched for the data to be processed according to at least one data modality covered by the data to be processed; For the features to be matched in each data modality of the data to be processed, perform the following operations respectively: in at least one label dataset having a feature alignment relationship with a feature to be matched, determine a retrieved feature whose feature similarity with the feature to be matched satisfies a similarity condition as a hit feature; Based on the feature similarity corresponding to each screened hit feature, and in combination with each content label determined corresponding to each hit feature, determine the content label matched by the data to be processed.
[0009] Optionally, when determining the content label matched by the data to be processed based on the feature similarity corresponding to each screened hit feature, and in combination with each content label determined corresponding to each hit feature, the execution unit is configured to: Based on the label data to which each of the selected hit features belongs, determine the respective content labels determined for each of the hit features; For each of the content labels, perform the following operations respectively: Among the selected retrieval features, at least one retrieval feature associated with a content label is respectively used as a target feature, and based on the feature similarity corresponding to each of the at least one target feature, in combination with the total number of features of each of the hit features, obtain the adaptation value corresponding to the content label; Take the content label whose adaptation value reaches the set value as the content label matched by the data to be processed.
[0010] Optionally, when determining the hit features by taking the retrieval features whose feature similarity with a feature to be matched meets the similarity condition among at least one label data set having a feature alignment relationship with the feature to be matched, the execution unit is configured to: Obtain the feature alignment relationship determined based on the feature extraction methods in each data modality; Based on the feature alignment relationship, determine at least one label data set that matches a feature to be matched, and for each of the matched label data sets, perform the following operations respectively: Calculate the feature similarity between a feature to be matched and each retrieval feature in a label data set respectively, and take the N retrieval features with the largest feature similarity as the hit features that meet the preset similarity condition.
[0011] Optionally, when respectively taking the at least one candidate label as the content label, the determination unit is configured to: Respectively determine the at least one candidate label as the content label; Take the candidate label with the total number of associated historical data being 0 among the at least one candidate label as a new content label and add it to the preset content label set.
[0012] Optionally, one data modality is the image modality; when extracting the retrieval features from the data content in one data modality of a multimedia data, the storage unit is configured to: Extract the image content in a multimedia data; wherein, the image content includes at least one image; Use the image feature extraction network in the pre-trained multi-modal model to respectively extract features for the at least one image to obtain the corresponding image features; Fuse at least one image feature to obtain the corresponding retrieval feature.
[0013] Optionally, when extracting the image content in a multimedia data, the storage unit is configured to perform any one of the following operations: When at least one video content is included in a piece of multimedia data, M frame images respectively extracted from each of the video contents are used as the extracted image content; When a piece of multimedia data includes video content and an original image, the M frame images extracted from the video content and the original image are used as the extracted image content.
[0014] Optionally, one data modality is a text modality; when extracting retrieval features from the data content in one data modality of a piece of multimedia data, the storage unit is configured to: Extract the text content in a piece of multimedia data; wherein, the text content covers at least one of the following text types: the audio text obtained after performing speech-to-text conversion processing; the original text directly included in the multimedia data; Use the text feature extraction network in the multimodal model to perform feature extraction based on the text content to obtain the corresponding retrieval features; wherein, there is a feature alignment relationship between the retrieval features extracted by the feature extraction networks of different modalities in the multimodal model.
[0015] Optionally, the one data modality is an audio modality; when extracting retrieval features from the data content in one data modality of a piece of multimedia data, the storage unit is configured to: Extract the audio content in a piece of multimedia data; wherein, the audio content covers at least one of the following audio types: the content audio separated from the video content; the original audio directly included in the multimedia data; Perform audio feature extraction on the audio content to obtain the corresponding retrieval features.
[0016] Optionally, before obtaining the description text set associated with a hot content, the device further includes a crawling unit, and the crawling unit is configured to: Crawl the description texts of each hot instance from each preset data web page respectively; Based on the text similarity between every two of the description texts, divide the description texts that meet the preset similarity condition into the same description text set to obtain each description text set; wherein, at least one description text is included in one description text set; For each of the description text sets, perform the following operations respectively: divide the hot instances corresponding to the description texts in one description text set into one hot content, and establish an association relationship between the one hot content and the one description text set.
[0017] Optionally, when obtaining each piece of multimedia data associated with the one hot content, the crawling unit is configured to: For each of at least one hotspot instance covered by a hotspot content, perform the following operations respectively: Based on the data webpages relied on when crawling the description text associated with a hotspot instance, obtain each data link information associated with the one hotspot instance, and respectively obtain each multimedia data belonging to the one hotspot instance according to each data link information; Use the obtained each multimedia data as each multimedia data associated with the one hotspot content.
[0018] Optionally, when storing the tag data generated based on the retrieval feature and at least one content tag into the tag dataset set for the one data modality, the storage unit is configured to: In at least one content tag, remove the content tags whose total number of associated historical data reaches the preset quantity threshold, and obtain the remaining content tags; Use the retrieval feature and the associated content tags as the generated tag data, store them into the tag dataset set for the one data modality, use the one multimedia data as the historical multimedia data associated with the content tag, and update the total number of historical data associated with the content tag.
[0019] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0020] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0021] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method is implemented.
[0022] The beneficial effects of this application are as follows: The present application proposes a method, apparatus, electronic device, and storage medium for generating label data. First, a description text set associated with a hot content is obtained, and keywords are extracted from the description text set to obtain a corresponding candidate label set. Then, when it is determined that there is at least one candidate label in the candidate label set whose total number of associated historical data does not exceed a preset quantity threshold, each of the at least one candidate label is used as a content label, and each multimedia data associated with the hot content is obtained. Each multimedia data contains data content of at least one data modality. The total number of historical data associated with a candidate label refers to the total number of historical multimedia data based on which the label data containing the candidate label are generated when generating content labels. In this way, a general content label extraction logic is proposed for hot content, enabling the generation efficiency of candidate labels to be improved by directly extracting keywords from the description text set. Moreover, when it is determined that there is at least one candidate label in the candidate label set whose total number of associated historical data does not exceed a preset quantity threshold, it can be determined that label data need to be generated based on the current hot content, and the content label for generating the label data can be determined. Consequently, it is necessary to obtain each multimedia data associated with the hot content. Furthermore, for each data modality in each multimedia data, the following storage operations are respectively performed: feature extraction is performed on the data content in a data modality of a multimedia data to obtain a retrieval feature, and the label data generated based on the retrieval feature and at least one content label are stored in the label data set set corresponding to the data modality. In this way, when constructing the label data for retrieval basis, content splitting and feature extraction are respectively performed on each associated multimedia data from the perspective of different data modalities, enabling a comprehensive analysis of the multimedia data from the perspective of different data modalities to obtain retrieval features in each covered data modality. Therefore, more appropriate retrieval features can be generated by leveraging the rich semantic information in the multimedia data. Moreover, by means of the proposed storage logic for storing label data in the database, the label data can be systematically stored according to the data modality, thereby realizing a refined classification of the label data. Based on this, label data associated with the hot content can be generated in a timely and efficient manner, thus providing a more effective and accurate retrieval basis for the subsequent label annotation process based on the label data. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of a possible application scenario in an embodiment of the present application; Figure 2 It is a schematic diagram of the generation process of label data in an embodiment of the present application; Figure 3 It is a schematic diagram of the process of obtaining each description text in an embodiment of the present application; Figure 4 Schematic diagram for displaying various multimedia data in a hot instance in an embodiment of the present application; Figure 5 Schematic diagram of the process for storing label data in an embodiment of the present application; Figure 6 Schematic diagram of the annotation process for content labels in an embodiment of the present application; Figure 7 Schematic diagram of the process for extracting features to be matched in an embodiment of the present application; Figure 8 Schematic diagram of the process for determining hit features in an embodiment of the present application; Figure 9 Schematic diagram of the process for generating label data of new hot labels in an embodiment of the present application; Figure 10 Schematic diagram of the process for annotating content labels in an embodiment of the present application; Figure 11 Schematic diagram of the content processing process in an embodiment of the present application; Figure 12 Schematic diagram of the logical structure of the label data generation device in an embodiment of the present application; Figure 13 Schematic diagram of the hardware composition structure of an electronic device applying an embodiment of the present application; Figure 14 Schematic diagram of the hardware composition structure of another electronic device applying an embodiment of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the technical solutions of the present application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the technical solutions of the present application.
[0025] Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.
[0026] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0027] The following explains some terms in the embodiments of the present application to facilitate understanding by those skilled in the art.
[0028] Large Language Model (LLM): It refers to a natural language processing model based on deep learning technology, with billions or even more parameters. The LLM is pre-trained on a large-scale corpus and can perform various language tasks, such as text generation, translation, question answering, and dialogue systems, enabling the LLM to have strong context understanding and multitasking capabilities.
[0029] In-context learning (ICL): It refers to the ability to learn and perform specific tasks during the inference process by providing a prompt template containing examples without explicitly updating the parameters of the model.
[0030] Contrastive Language-Image Pre-Training (CLIP): It is a multimodal machine learning model that can understand the relationship between pictures and text. This model is pre-trained on a large number of pictures and related descriptive texts, aiming to enable the model to learn how to connect visual information and language information.
[0031] Bidirectional Encoder Representations from Transformer (BERT): It is a pre-trained language model based on the Transformer model. Through bidirectional context understanding and the Transformer architecture, it can capture complex language structures and semantics, thus performing well in natural language processing (NLP) tasks; among them, NLP tasks include text classification, text generation, question answering, and translation, etc.
[0032] Vision Transformer (ViT): is a computer vision model based on the Transformer architecture that can achieve good processing results in visual tasks such as image classification.
[0033] Transformer model: A deep learning model for processing sequence data that mainly relies on the attention mechanism. The model processes the input data sequence through the self-attention mechanism and the multi-head attention mechanism, and can effectively capture the relationship between elements in the sequence. The Transformer model usually includes an encoder and a decoder structure and is widely used in natural language processing, image processing and other fields.
[0034] Hot content: refers to content that has attracted widespread attention, discussion, and dissemination, which may be reflected in hot topics, hot events, and hot information; hot content covers a wide range of fields, such as news, entertainment, technology, and social sciences.
[0035] Hotspot instance: refers to the specific manifestation of hot content on the data source; wherein, one hotspot instance corresponds to one description text. For example, for a hot content, there may be different content descriptions, which are reflected as different hotspot instances. For another example, suppose the description texts of each hotspot instance obtained include: Sun won the A City Championship; Sun won the A City Championship for the second time; Sun successfully reversed and won the A City Championship, these three hotspot instances describe the same hot content; wherein, the data source refers to the data source that publishes the description texts of each hotspot instance.
[0036] Description text set: a text set consisting of various description texts; one description text is used to describe one hot content; one hot content may have multiple description texts.
[0037] Data modality: refers to the type or form of data. In the embodiments of the present application, it is used to indicate data in different forms of expression. In some examples of the present application, the preset data modalities may include: image modality, text modality, and audio modality. In other embodiments of the present application, the preset data modalities may include: image modality, text modality, audio modality, and video modality.
[0038] The following is a brief introduction to the design concept of the embodiment of the present application: At present, the tag system is a technology for organizing and classifying information; through the tag system, useful information can be extracted from a large amount of data more efficiently, and thus applied to multiple fields such as content management, search, and recommendation. Moreover, many business scenarios have strong requirements for timeliness. When hot content occurs, the tag system is required to quickly generate relevant tag data, so that the multimedia data to be published can match the latest content tags, thereby efficiently realizing the tag annotation of multimedia data.
[0039] Under the existing technology, when generating tag data, usually technicians frequently manually sort out content tags based on the latest hot content, and store the sorted content tags, as well as the content tags and their corresponding tag features as new tag data, so that by calculating the feature similarity between the multimedia data and the content tags, the matching content tags can be determined.
[0040] For example, technicians collect hot topics or hot events, screen and rewrite appropriate content tags, and supplement them to the content tag set; at the same time, establish the corresponding relationship between the newly determined content tags and their corresponding tag features.
[0041] However, in the above way of generating tag data, since content tags need to be manually constructed, the generation of content tags is very dependent on the personal experience of technicians, resulting in very unstable generation effects of content tags. Furthermore, the generated tag data cannot adapt to the content tag annotation needs of multimedia data, thus greatly affecting the retrieval effect of multimedia data and reducing the browsing experience of relevant objects for multimedia data.
[0042] In view of this, the present application proposes a method, an apparatus, an electronic device, and a storage medium for generating tag data. First, a description text set associated with a hot content is obtained, and keyword extraction is performed on the description text set to obtain a corresponding candidate tag set. Then, when it is determined that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed a preset number threshold, the at least one candidate tag is respectively used as a content tag, and each multimedia data associated with a hot content is obtained; each multimedia data contains data content of at least one data modality; the total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which the tag data containing a candidate tag is generated when generating content tags. In this way, a general content tag extraction logic is proposed for hot content, so that by directly performing keyword extraction on the description text set, the generation efficiency of candidate tags can be improved; moreover, when it is determined that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed a preset number threshold, it can be determined that tag data needs to be generated based on the current hot content, and the content tag for generating tag data can be determined, and then each multimedia data associated with a hot content needs to be obtained. Furthermore, for each data modality in each multimedia data, the following storage operations are respectively performed: feature extraction is performed on the data content in a data modality of a multimedia data to obtain a retrieval feature, and the tag data generated based on the retrieval feature and at least one content tag is stored in the tag data set set for a corresponding data modality. In this way, when constructing the tag data for retrieval, content splitting and feature extraction are respectively performed on each associated multimedia data from the perspective of different data modalities, so that a comprehensive analysis of the multimedia data can be performed from the perspective of different data modalities respectively, and retrieval features in the covered data modalities can be obtained. Therefore, more appropriate retrieval features can be generated by means of the rich semantic information in the multimedia data; moreover, by means of the proposed storage logic for storing tag data in the library, the tag data can be systematically stored according to the data modality, so as to realize the refined classification of the tag data; based on this, the tag data associated with the hot content can be generated in a timely and efficient manner, so that a more effective and accurate retrieval basis can be provided for the subsequent tag annotation process based on the tag data.
[0043] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0044] Refer to Figure 1As shown, it is a schematic diagram of a possible application scenario in an embodiment of the present application. In this application scenario schematic diagram, it includes a client device 110 and a processing device 120.
[0045] In some feasible embodiments of the present application, in the generation stage of tag data, the processing device 120 can crawl the description texts associated with each hot content, and optionally, by comparing the similarity of each description text, integrate the similar description texts to obtain the description text set associated with each hot content; then, the processing device 120 can perform the generation processing of tag data for each hot content respectively. In the specific generation process, after extracting the candidate tag set based on the description text set associated with a hot content, when it is determined that the total number of associated historical data in the candidate tag set does not exceed the preset quantity threshold for at least one candidate tag, it can be determined that tag data needs to be constructed for the current hot content. Specifically, at least one candidate tag can be respectively determined as the content tag, and each multimedia data associated with a hot content can be obtained; each multimedia data contains data content of at least one data modality; then, for each data modality in each multimedia data, the following storage operations are respectively performed: extract the feature of the data content in one data modality of a multimedia data to obtain the retrieval feature, and store the tag data generated based on the retrieval feature and at least one content tag into the tag data set set for the corresponding data modality.
[0046] In some other feasible embodiments of the present application, in the tag configuration stage, the processing device 120 can obtain the data to be processed sent by the client device 110, and respectively extract the to-be-matched features of the data to be processed in each covered data modality; further, the processing device 120 can respectively perform feature retrieval in the tag data sets having a feature alignment relationship with the to-be-matched feature for each to-be-matched feature, and screen out each hit feature whose feature similarity meets the similarity condition; then, based on the respective feature similarities of each hit feature and the total number of each hit feature, screen out the content tags matched by the data to be processed. Usually, the number of content tags matched by the data to be processed may be one or more.
[0047] In addition, the data to be processed sent by the client device 110 can be sent based on any one of a mini-program application, a client application, and a web application, and the present application does not make specific limitations on this.
[0048] The client device 110 includes but is not limited to a mobile phone, a tablet computer, a notebook, an e-book reader, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc.
[0049] The processing device 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0050] In the embodiments of the present application, the client device 110 and the processing device 120 can communicate through a wired network or a wireless network. In the following description, the generation process of the tag data is described only from the perspective of the processing device 120.
[0051] The following is a schematic description of the generation process of the tag data in combination with possible application scenarios: Application scenario 1: Generate tag data for retrieval basis in the content review scenario.
[0052] Specifically, in order to effectively review hot content with very strong timeliness, corresponding tag data can be generated for sensitive hot content (hereinafter referred to as sensitive content); furthermore, after obtaining the data to be processed, it is possible to retrieve whether the content tag corresponding to the sensitive content is hit based on the constructed tag data, so as to effectively manage the content to be processed and prevent the data to be processed related to the sensitive content from being published.
[0053] Application scenario 2: Generate tag data for new tag addition in the content management scenario.
[0054] Specifically, in order to adapt to hot content with very strong timeliness, in scenarios such as e-commerce shopping recommendations or multimedia data recommendations, corresponding tag data can be generated for the hot content, and the newly added content tags can be determined; furthermore, the newly added content tags can be added to the matching e-commerce items or multimedia data.
[0055] For example, if it is determined that the new content tag includes tag B, and tag A and tag B are in a candidate tag set, then the processed data marked with tag A can be added with tag B.
[0056] For example, assuming that the emerging tag is "super Zhang", and "Zhang Mou" and "super Zhang" are included in a candidate tag set, then the released data marked with the tag "Zhang Mou" can be added with the tag "super Zhang".
[0057] Application scenario 3: Generate tag data for retrieval basis in the content publishing scenario.
[0058] Specifically, in order to adapt to hot content with very strong timeliness, in scenarios such as multimedia data publishing, in order to label appropriate content tags for the multimedia data to be published, corresponding tag data can be generated for the hot content; furthermore, appropriate content tags can be labeled for the multimedia data to be published (i.e., the data to be processed).
[0059] In addition, it should be understood that in the specific implementation manners of the present application, regarding the generation process of tag data, when the embodiments described in the present application are applied to specific products or technologies, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0060] The following will explain the generation process of tag data from the perspective of a processing device in conjunction with the accompanying drawings: Refer to Figure 2 As shown, it is a schematic diagram of the generation process of tag data in an embodiment of the present application. The following will explain the generation process of tag data in conjunction with the attached Figure 2 drawings: Step 201: The processing device obtains a set of description texts associated with a hot content, and extracts keywords from the set of description texts to obtain a corresponding candidate tag set.
[0061] In some feasible embodiments of the present application, the processing device can obtain a description text associated with a hot content provided by other devices and perform processing.
[0062] In other feasible embodiments of the present application, before generating tag data based on a hot content, the processing device can, for the description texts of each hot instance crawled, determine at least one set of description texts divided according to each description text and at least one hot content divided corresponding to each hot instance by calculating the text similarity between every two description texts, so as to obtain a set of description texts associated with each hot content.
[0063] When specifically obtaining a set of description texts associated with each hot content, the processing device can respectively crawl the description texts of each hot instance from each preset data web page; then, based on the text similarity between every two description texts, divide the description texts that meet the preset similarity conditions into the same set of description texts to obtain each set of description texts; where a set of description texts contains at least one description text; after that, for each set of description texts, the following operations are respectively performed: divide the hot instances corresponding to each description text in a set of description texts into one hot content, and establish an association relationship between this hot content and this set of description texts.
[0064] It should be noted that in the embodiments of the present application, the preset data web pages may come from different data sources; each data web page displays the description text of each hot instance and the associated data link information. The preset similarity condition may be that the text similarity reaches a preset similarity threshold, where the value of the similarity threshold is set according to actual processing needs, and the present application does not make specific restrictions on this.
[0065] It should be understood that in some feasible embodiments of the present application, the processing device can periodically crawl the description text, and based on the crawled description texts, organize the description text sets associated with each hot content respectively, and for each description text set associated with a hot content, execute the process of generating tag data, where the crawling period can be determined according to the update period of the content in the preset data web pages; or, it can be configured according to actual processing needs; in other feasible embodiments of the present application, the processing device can crawl the description text from the preset data web pages at a specified data crawling time, and based on the crawled description texts, organize the description text sets associated with each hot content respectively, and for each description text set associated with a hot content, execute the process of generating tag data.
[0066] For example, assume that there are three preset data web pages, denoted as Data Web Page 1 - 3. The update period of the content in Data Web Page 1 is 1 minute, the update period of the content in Data Web Page 2 is 30 seconds, and the update period of the content in Data Web Page 3 is 2 minutes; then, any one of the three update periods can be selected as the crawling period. For example, the crawling period is selected as 2 minutes to achieve one round of crawling of the description text and generation of tag data every two minutes.
[0067] For another example, according to actual processing experience, the crawling period can be set to 10 minutes to achieve one round of crawling of the description text and generation of tag data every 10 minutes.
[0068] For another example, the processing device can be set to perform one round of crawling of the description text and generation of tag data every 10 minutes between 5:00 and 24:00 every day.
[0069] In addition, in the embodiments of the present application, since the generation of tag data has nothing to do with specific hot content, specific representation of hot content may not be carried out. Specifically, for the purpose of differentially representing different hot contents, different hot contents can be identified.
[0070] For example, the description text of each hot instance can specifically be the entry text of each hot search entry published by each data source; where one hot search entry is one hot instance.
[0071] For example, refer to Figure 3As shown, it is a schematic diagram of the process of obtaining each description text in the embodiment of the present application. According to the attached Figure 3 As can be seen from the content shown, assuming that the preset data web pages are data web pages 1-3, the processing device can obtain the description texts of each hot instance from data web pages 1-3 respectively, and then integrate each description text by calculating the text similarity between every two description texts, so as to obtain the description text sets respectively associated with each hot content.
[0072] In this way, it is possible to obtain the description text sets respectively associated with each hot content currently integrated, thereby providing a processing basis for the generation of label data.
[0073] Furthermore, the processing device can respectively obtain the description text sets associated with each hot content, and based on each obtained description text set, execute the process of generating label data. In the related description of the present application, only taking the example of obtaining the description text set associated with one hot content, the process of generating label data is schematically described.
[0074] After obtaining the description text set associated with one hot content, the processing device can obtain the corresponding candidate label set by extracting keywords from the description text set.
[0075] When specifically extracting keywords from the description text set, in some feasible implementation manners, the processing device can rely on the context learning ability of the large language model, and in combination with a pre-designed prompt template containing examples, extract keywords from each description text respectively, and use the deduplicated keywords as each candidate label and store them in the candidate label set; among them, when deduplicating the extracted keywords, various feasible deduplication methods can be used for processing, and the present application does not make specific limitations on this. For example, a deduplication method based on a hash table or a deduplication method based on a set can be used, and the corresponding deduplication function is used for processing.
[0076] It should be noted that the prompt template is used to guide the large language model to rewrite complex sentences or short sentences into concise candidate labels. The prompt template contains a task description and examples to help the model understand the rewriting target; the examples are pre-annotated and included in the prompt template to assist the large language model in learning the rewriting pattern, thereby improving the accuracy of candidate label generation.
[0077] For example, a pre-designed prompt template could be: # Prompt Template You are a chatbot. Please extract key trending tags based on the topic information. The topic content is generally in the form of a sentence or a short phrase. The trending tags should be more concise word forms and represent the main content of the topic. If there are no keywords, do not extract them, but do not fabricate them out of thin air. Example 1: Trending topic: The "Village BA" is in full swing, and the candidate tags are: Village BA; Example 2: Trending topic: Talking about "working together with the same ball" again, the candidate tags are: working together with the same ball.
[0078] Another example, in order to use a large language model to extract candidate tags from a descriptive text: "In the A City Championship, 'Super Sun' came from behind to win 4-2 after being 0-2 down", the content input into the language model specifically includes: "# Prompt Template You are a chatbot. Please extract key trending tags based on the topic information. The topic content is generally in the form of a sentence or a short phrase. The trending tags should be more concise word forms and represent the main content of the topic. If there are no keywords, do not extract them, but do not fabricate them out of thin air. Example 1: Trending topic: The 'Village BA' is in full swing, trending tags are: Village BA Example 2: Trending topic: Talking about 'working together with the same ball' again, trending tags are: working together with the same ball. Trending topic: In the A City Championship,'Super Sun' came from behind to win 4-2 after being 0-2 down, candidate tags are:". Then, the candidate tags output by the large language model based on a descriptive text can be obtained, where the number of candidate tags is at least one.
[0079] In this way, by jointly inputting the prompt template containing various examples and the descriptive text to be processed into the large language model, the processing effect of the large language model can be improved by means of the ICL technology, so that the large language model can be guided to learn the problem-solving methods in the examples without updating the parameters of the large language model, and then make accurate predictions for the descriptive text of the question.
[0080] When specifically extracting keywords from a descriptive text set, in some other feasible implementation methods, the processing device can use a keyword extraction model for each descriptive text in the descriptive text set to extract keywords respectively, and use the deduplicated keywords as candidate tags respectively and store them in the candidate tag set.
[0081] Among them, when constructing the keyword extraction model, any model structure that can implement keyword extraction can be used. For example, a bidirectional encoder representation model based on Transformer (Bidirectional Encoder Representations from Transformers, BERT) can be used. During the training process, with the help of training samples with marked keywords, the keyword extraction model is trained through multiple rounds of iteration. And during each round of iterative training, the model parameters are adjusted with the help of the cross-entropy loss function. Furthermore, during the specific processing based on the trained keyword extraction model, the keyword extraction model is used to encode the description text, and keywords are selected through the attention mechanism.
[0082] In this way, it is possible to extract various possible candidate tags based on a set of description texts associated with a hot content.
[0083] Step 202: When the processing device determines that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed the preset quantity threshold, each of the at least one candidate tag is used as a content tag, and each multimedia data associated with a hot content is obtained; where each multimedia data contains data content of at least one data modality; the total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which the tag data containing the candidate tag is generated when generating content tags.
[0084] After the processing device obtains the corresponding candidate tag set based on a set of description texts, it can determine whether it is necessary to construct tag data for the candidate tags by judging the relationship between the total number of historical data associated with each candidate tag in the candidate tag set and the preset quantity threshold. The value of the preset quantity threshold is set according to actual processing needs; moreover, the preset quantity threshold is used to limit the maximum value of the historical multimedia data included in the generated tag data that contains a candidate tag.
[0085] It should be noted that in the embodiments of the present application, the total number of historical data associated with a candidate tag is updated globally. Specifically, when generating tag data for different hot contents, for a candidate tag (assumed to be candidate tag X) extracted during the processing of different hot contents, the total number of historical data associated with candidate tag X is continuously updated during the generation of tag data for different hot contents; the total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which the tag data containing the candidate tag is generated when generating content tags.
[0086] For example, the total number of historical data associated with a currently determined candidate tag refers to the total number of historical multimedia data based on which, for each of the currently generated tag data, when generating tag data whose generated content tag contains the candidate tag.
[0087] In an embodiment of the present application, when determining whether there is at least one candidate tag in the candidate tag set whose associated total number of historical data does not exceed a preset quantity threshold, the processing device may perform the following operations for each candidate tag in the candidate tag set: obtain the total number of historical data associated with a candidate tag, and compare the value relationship between the historical data threshold and the preset quantity threshold. Subsequently, when it is determined that there is at least one candidate tag whose total number of historical data has not reached the preset quantity threshold, the at least one candidate tag may be respectively determined as content tags, and it may be determined that the generation process of tag data can continue.
[0088] In addition, in an embodiment of the present application, after the processing device determines that there is at least one candidate tag in the candidate tag set whose associated total number of historical data does not exceed the preset quantity threshold, when respectively using the at least one candidate tag as content tags, the processing device may respectively determine the at least one candidate tag as content tags; at the same time, the candidate tag among the at least one candidate tag whose associated total number of historical data is 0 is used as a new content tag and added to the preset content tag set.
[0089] It should be noted that in the case where the total number of historical data associated with a candidate tag is 0, it can be determined that tag data has not been generated based on this candidate tag, and thus it can be determined that this candidate tag is a new content tag that has not been extracted previously; furthermore, the content tag set can be updated according to the new content tag, where the preset content tag set stores non-repeated content tags; a content tag is a tag that can be used to label multimedia data; a candidate tag refers to a tag directly extracted based on a description text set.
[0090] In this way, the newly generated content tags can be added to the content tag set in a timely manner, realizing the timely update of the content tag set.
[0091] Particularly, when it is determined that the total number of historical data associated with each candidate tag in the candidate tag set has reached the preset quantity threshold, the generation process of tag data for the current hot content is stopped, and a description text set associated with other hot content is retrieved again, and the generation process of tag data is executed.
[0092] Further, after the processing device determines that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed a preset quantity threshold, the processing device respectively determines the at least one candidate tag as a content tag, and obtains each multimedia data associated with a current hot content; wherein, each multimedia data contains data content of at least one data modality.
[0093] It should be noted that the data type corresponding to any one multimedia data can be any one or combination of the following: text; image; video; audio. For a pure text multimedia data, the data modality it contains is the text modality. For a multimedia data containing video, according to different division requirements, the data modalities it contains may be the image modality, the audio modality, and the text modality; or, the data modalities it contains may be the image modality, the audio modality, the text modality, and the video modality. For a pure audio multimedia data, the data modalities it contains are the audio modality and the text modality.
[0094] When obtaining each multimedia data associated, the processing device can respectively link and determine each multimedia data for each description text in the description text set, and select each multimedia data associated with a hot content from the determined multimedia data by linking.
[0095] Specifically, in the case where a description text set associated with a hot content is crawled from each preset data web page and the description text set includes description texts of at least one hot instance, the processing device can respectively perform the following operations for at least one hot instance covered by a hot content: based on the data web page relied on when crawling the description text associated with a hot instance, obtain each data link information associated with a hot instance, and respectively obtain each multimedia data belonging to a hot instance according to each data link information; then use the obtained each multimedia data as each multimedia data associated with a hot content.
[0096] It should be noted that each hot instance in each data web page is respectively associated with web page link information. By means of any one web page link information, it is possible to link to a content web page containing each multimedia data, and the content web page contains data link information of each multimedia data, so that the acquisition or download of multimedia data can be realized by means of the data link information; based on this, it can be understood that a hot instance is associated with data link information of each multimedia data.
[0097] In addition, for a hot instance, when obtaining multimedia data belonging to a hot instance according to the respective data link information associated therewith, the processing device may obtain a specified number of multimedia data according to the respective data link information as the multimedia data belonging to a hot instance, or may obtain all the multimedia data according to the respective data link information as the multimedia data belonging to a hot instance; it should be understood that the total number of multimedia data obtained for a hot content does not exceed a preset quantity threshold configured for the total number of historical data.
[0098] Optionally, when obtaining the respective multimedia data associated with a hot content, for at least one content label whose total number of historical data does not exceed a preset quantity threshold, the total number of multimedia data to be obtained may be determined according to the relationship between the minimum value of the total number of historical data in the at least one content label and the value of the preset quantity threshold.
[0099] For example, the total number of historical data associated with content label 1 is 10, the total number of historical data associated with content label 2 is 20, and the preset quantity threshold is 50. Then, the difference between the minimum value of the total number of historical data and the preset quantity threshold is 40. Then, for a hot content, at most 40 multimedia data may be obtained; in addition, in the case where a hot content includes multiple hot instances, a specified number of multimedia data may be evenly obtained for each hot instance; or, the total number of multimedia data obtained for each hot instance may not be specifically configured, and it is only necessary to ensure that the total number of multimedia data obtained as a whole does not exceed 40.
[0100] For another example, refer to Figure 4 shown, which is a schematic diagram of the display of the respective multimedia data under a hot instance in an embodiment of the present application. According to the content shown in the appendix Figure 4 it can be seen that when obtaining multimedia data for a hot instance, a web page link information associated with a hot instance may be triggered to jump to the web page shown in the appendix Figure 4 shown, and the multimedia data may be obtained according to the respective data link information corresponding to the respective multimedia data in the web page.
[0101] In this way, for each hot instance belonging to a hot content, the corresponding multimedia data can be obtained respectively, so that each multimedia data obtained for a hot content can be mapped to the content label extracted for the current hot content.
[0102] Step 203: For each data modality in each piece of multimedia data, the processing device respectively performs the following storage operations: extracting features from the data content under a data modality in a piece of multimedia data to obtain retrieval features, and storing the tag data generated based on the retrieval features and at least one content tag into the tag dataset set for the corresponding data modality.
[0103] After the processing device obtains the associated multimedia data corresponding to a current hot content, for each piece of multimedia data, it can generate tag data respectively under each data modality in the multimedia data, or generate tag data respectively under each data modality covered by the multimedia data.
[0104] Among them, in some feasible implementation manners, a piece of tag data includes: retrieval features under a data modality, and content tags screened from at least one content tag according to the total number of the latest historical data; in some other feasible implementation manners, a piece of tag data includes: retrieval features under a data modality, and at least one content tag. The at least one content tag is screened from the candidate tag set based on the total number of associated historical data before obtaining the multimedia data associated with a hot content; the data modalities in a piece of multimedia data are at least one.
[0105] When extracting retrieval features under different data modalities, in some feasible embodiments, a feature extraction network matching each data modality can be used to extract features from the data content under different data modalities respectively. At this time, there is no feature alignment relationship between the retrieval features under different data modalities, and the network structure of the feature extraction network is not specifically limited in this application; in some other feasible embodiments, a pre-trained multi-modal model can be used to extract features from the data content under different data modalities. At this time, there is a feature alignment relationship between the retrieval features under different data modalities; in some other feasible embodiments, for some data modalities, a pre-trained multi-modal model can be used to extract features, and for other data modalities, a corresponding feature extraction network can be used to extract features. At this time, there is a feature alignment relationship between the retrieval features under each data modality extracted by using the multi-modal model.
[0106] In this way, according to the actual processing requirements and combined with the actual deployment cost, the extraction method of retrieval features under different data modalities can be flexibly selected.
[0107] In some feasible embodiments of this application, the preset data modality range for processing can be: text modality, image modality, and audio modality; in some other feasible embodiments of this application, the preset data modality range for processing can be: text modality, image modality, video modality, and audio modality.
[0108] In this way, according to the actual processing needs, for video content, it can be processed from the perspective of the video modality, or it can be processed from the perspective of the image modality.
[0109] In some feasible embodiments of the present application, the selected multimodal model can be a model capable of jointly learning image features and text features, so as to promote the integration of computer vision and natural language processing; in this case, the multimodal model can be any one of the following models: CLIP, A Large-scale Image and Noisy-text embedding (ALIGN) model, Full-Length Multimodal Video-Audio-Vision-Language Alignment (FLAVA) model, UNiversal Image-TExt Representation (UNITER) model, BERT-based Image-Text Pretraining (BEiT-3) model, Florence (a multimodal foundation model), and Object-Semantics Aligned Pre-training (Oscar) model, etc.
[0110] Taking the CLIP model as an example of the multimodal model, with the help of the CLIP model, computer vision (CV) and natural language processing (NLP) features can be extracted. Moreover, by aligning visual features and text features, the CLIP model can understand the relationship between visual information and text information; specifically, the CLIP model uses the Transformer structure, can adopt ViT as the image feature extraction network, and adopt Bert as the text feature extraction network, which can map image features and text features to the same feature space, so as to realize cross-modal similarity comparison; the multimodal model adopted by the processing device can be a general CLIP model self-developed and pre-trained with a large amount of Chinese data, so as to extract general features for image content and text content, and support different services without retraining; the training process of the CLIP model is not specifically described in the present application.
[0111] For example, in the process of pre-training the feasible CLIP model, training samples can be generated by constructing image-text pairs, and then through contrastive learning, the feature distance between the matching images and texts can be shortened, and the feature distance between the non-matching images and texts can be enlarged to adjust the parameters of the model.
[0112] In some other feasible embodiments of the present application, the selected multi-modal model can be a model capable of jointly learning image features, text features, and audio features; in this case, the multi-modal features can be any one of the following models: FLAVA, Multimodal Event Representation Learning with Transformers (MERLOT) model, Audio-Visual-Language Network (AVLNet) model, Cross-Modal Multitask Multitower Model (CM3) model, Poly-modal Vision Transformer (PolyViT) model, General Multimodal Perception (Perceiver IO) model, Audio-Image-Text CLIP model (AudioCLIP), Multimodal Language-Audio Network (MuLaN), and Vision-Audio-Language Omni-Perception Pretraining (VALOR) model, etc.
[0113] Optionally, in the case where the data modality further includes the video modality, the selected multi-modal model can be a model capable of jointly learning image features, video features, text features, and audio features; in this case, the multi-modal features can be any one of the following models: FLAVA, PolyViT, and Perceiver IO model, etc.
[0114] Taking the feature extraction by means of the feature extraction network in the multi-modal model as an example, the extraction methods of the retrieval features in each data modality are described below.
[0115] Extraction method 1: Extract the retrieval features in the image modality.
[0116] In the embodiment of the present application, when the data modality covered by a piece of multimedia data includes the image modality, the processing device can extract the image content in the piece of multimedia data when extracting the retrieval features in the image modality; where the image content includes at least one image; then, use the image feature extraction network in the pre-trained multi-modal model to perform feature extraction on at least one image respectively to obtain the corresponding image features; afterwards, fuse at least one image feature to obtain the corresponding retrieval feature.
[0117] Among them, the multimodal model based on which image features are extracted can specifically be a model capable of jointly learning image features and text features; or, it can be a model capable of jointly learning image features, text features, and audio features; or, it can be a model capable of jointly learning image features, video features, text features, and audio features.
[0118] It should be noted that in the embodiments of the present application, in the case where there is only a video modality alone, when at least one original image is included in a piece of multimedia data, the processing device can determine that a piece of multimedia data covers the image modality; in the case where there is no video modality, when a piece of multimedia data contains any one of video or original image, the processing device can determine that a piece of multimedia data covers the image modality.
[0119] In the case where the preset data modalities do not include the video modality, in some feasible embodiments of extracting image content, when only a few original images are included in the multimedia data, the processing device can use each of the original images included in the multimedia data as the image content; in some other feasible embodiments of extracting image content, when at least one video content is included in a piece of multimedia data, M frames of images respectively extracted from each video content are used as the extracted image content; in some other feasible embodiments of extracting image content, when a piece of multimedia data contains video content and original images, the M frames of images extracted from the video content and the original images are used as the extracted image content; among them, the present application does not specifically limit the number of video contents included. In feasible embodiments, a piece of multimedia data may include multiple video contents, and only M frames of images need to be extracted from each video content respectively; moreover, the present application does not specifically limit the value of M.
[0120] It should be noted that taking the extraction of M frames of images from a video content as an example, the processing device can evenly extract M frames of images from the multimedia data. Specifically, it can determine the sampling positions of the M frames of images according to the content duration of the video content, and sample the images from the corresponding positions.
[0121] For example, the sampling positions of each image can be determined according to the ratio relationship between the content duration and the value of M, so as to sample each image. Suppose the content duration of a video content is 25S and the value of M is 10, then image sampling can be performed every 2.5S to obtain M images.
[0122] In this way, it is possible to extract image content from the perspective of the image modality for the multimedia data.
[0123] Furthermore, after separately extracting image features from each image in the image content, various feasible feature fusion methods can be adopted to obtain fused features based on the image features, and the obtained fused features are used as retrieval features in the image modality. Among them, the present application does not specifically limit the method of fusing each image feature. For example, the retrieval features can be obtained by calculating the average value of the feature content at each position in each image feature.
[0124] In this way, when a piece of multimedia data currently processed covers the image modality, relevant image content can be comprehensively extracted for the multimedia data, and then the retrieval features in the image modality can be obtained.
[0125] Extraction method two: Extract the retrieval features in the text modality.
[0126] In the embodiments of the present application, when the text content is included in a piece of multimedia data, it can be determined that the data modality covered by the piece of multimedia data includes the text modality. Furthermore, when the processing device extracts the retrieval features in the text modality, it can first extract the text content in the piece of multimedia data. The text content covers at least one of the following text types: the audio text obtained after performing speech-to-text conversion processing; the original text directly included in the multimedia data. Then, the text feature extraction network in the multi-modal model is used to perform feature extraction based on the text content to obtain the corresponding retrieval features. Among them, there is a feature alignment relationship between the retrieval features extracted by the feature extraction networks of different modalities in the multi-modal model.
[0127] Among them, the multi-modal model adopted can specifically be a model capable of jointly learning image features and text features; or it can be a model capable of jointly learning image features, text features, and audio features; or it can be a model capable of jointly learning image features, video features, text features, and audio features.
[0128] For example, for a piece of pure text multimedia data, the multimedia data may specifically be an article. At this time, the text content specifically refers to the title and content of the article.
[0129] In the embodiments of the present application, when the multimedia data contains audio content, the audio content can be converted into audio text by means of speech-to-text conversion (ASR). Among them, the audio content can be any one or a combination of the following contents: the original audio directly included in the multimedia data; the content audio separated from the video content. The present application does not specifically limit the technical means adopted during speech-to-text conversion, and can be processed according to various tools capable of implementing speech-to-text conversion. In addition, the present application does not specifically limit the method adopted for separating audio content from video content, and various feasible audio separation tools can be used for processing.
[0130] For example, assume that a piece of multimedia data only contains one video content. Then, the title of the video content and the audio text obtained by converting the video content can be used as the text content obtained from the video content.
[0131] Furthermore, the processing device can use the text feature extraction network in the multi-modal model to extract text features based on the text content, and use the text features as the retrieval features in the text modality.
[0132] In this way, when a piece of multimedia data being processed currently covers the text modality, all relevant text content can be comprehensively extracted for the multimedia data, and then the retrieval features in the text modality can be effectively extracted.
[0133] Extraction method three: Extract the retrieval features in the audio modality.
[0134] In the embodiments of the present application, when a piece of multimedia data includes audio content, it can be determined that the data modalities covered by the multimedia data include the audio modality. Furthermore, when the processing device extracts the retrieval features in the audio modality, it can extract the audio content in the multimedia data; where the audio content covers at least one of the following audio types: the content audio separated from the video content; the original audio directly included in the multimedia data; then perform audio feature extraction on the audio content to obtain the corresponding retrieval features.
[0135] Among them, the processing device can perform audio feature extraction by means of the audio feature extraction network in the multi-modal model; the multi-modal model adopted can be a model capable of jointly learning image features, text features, and audio features; or, it can be a model capable of jointly learning image features, video features, text features, and audio features.
[0136] It should be noted that in the embodiments of the present application, the method used for separating audio content from the video content is not specifically limited, and various feasible audio separation tools can be used for processing. Optionally, when multiple segments of audio are extracted from a piece of multimedia data, audio feature extraction can be performed on each segment of audio, and the retrieval features can be obtained by fusing the audio features of each segment of audio; where one segment of audio comes from a video content, or from an uninterrupted original audio.
[0137] In this way, when a piece of multimedia data being processed currently covers the audio modality, all relevant audio content can be comprehensively extracted for the multimedia data, and then the retrieval features in the audio modality can be extracted.
[0138] Extraction method four: Extract the retrieval features in the video modality.
[0139] In an embodiment of the present application, when video content is included in a piece of multimedia data, it can be determined that the data modalities covered by a piece of multimedia data include the video modality. Furthermore, the processing device can use the video feature extraction network in a preset multi-modal model to extract video features based on the image frame sequence corresponding to the video content. Specifically, in the case of multiple video contents, the video features obtained for multiple image frame sequences can be feature-fused to obtain retrieval features in the video modality.
[0140] Among them, the multi-modal model adopted can be a model capable of jointly learning image features, video features, text features, and audio features.
[0141] In this way, when the video modality is covered by a piece of multimedia data being processed currently, retrieval features in the video modality can be effectively extracted for the multimedia data.
[0142] Furthermore, taking the processing of a piece of multimedia data as an example, for a piece of multimedia data, after the retrieval features are extracted in any one of the covered data modalities, the tag data generated based on the retrieval features and at least one content tag can be stored in the tag dataset set for this data modality.
[0143] In some feasible embodiments, the processing device can use the retrieval features in a data modality and at least one associated content tag as a piece of generated tag data and store it in the tag dataset corresponding to the data modality.
[0144] In some other feasible embodiments, the processing device can remove the content tags whose total number of associated historical data reaches a preset quantity threshold from at least one content tag to obtain the remaining content tags, generate a piece of tag data based on the remaining content tags and the retrieval features in a data modality, and store the tag data in the tag dataset corresponding to the data modality.
[0145] Specifically, the processing device can remove the content tags whose total number of associated historical data reaches a preset quantity threshold from at least one content tag to obtain the remaining content tags; then use the retrieval features and the associated remaining content tags as the generated tag data, store it in the tag dataset set for a corresponding data modality, use a piece of multimedia data as the historical multimedia data associated with the remaining content tags, and update the total number of historical data associated with the remaining content tags.
[0146] It should be noted that the total number of historical data associated with a content tag is updated globally; moreover, for a content tag, when counting the total number of each historical multimedia data including the content tag in the generated tag data, the counted historical multimedia data are not repeated.
[0147] It should be understood that when generating tag data for each multimedia data associated with a hot content, as the tag data is generated based on each multimedia data, the total number of historical data associated with the content tag is also updated accordingly. Then, during the process of generating tag data for each multimedia data, the total number of historical data associated with the content tag can be updated; among them, this application does not limit the total number of content tags in a piece of tag data.
[0148] In particular, when generating tag data in different data modalities for each multimedia data associated with a hot content respectively, when processing a piece of multimedia data, if it is determined that among at least one content tag, there is no content tag whose associated total number of historical data reaches the preset quantity threshold, the process of generating tag data for the current hot content can be ended.
[0149] For example, refer to Figure 5 as shown, which is a schematic diagram of the process of storing tag data in an embodiment of this application. According to the content Figure 5 shown, assume that for a hot content, the determined content tag extracted is {super Zhang}; it is preset that the multimedia data is divided into three data modalities, namely text modality, audio modality, and image modality; for a piece of multimedia data 1 being processed currently, the data modalities covered by the multimedia data 1 include text modality, audio modality, and image modality; then, when generating tag data, when the processing device extracts text content from the perspective of the text modality, the original text directly included in the multimedia data 1 and the audio text converted from the video content 1 and the audio content 1 are used as the text content; when extracting image content from the perspective of the image modality, M images are extracted from the video content 1; when extracting audio content from the perspective of the audio modality, the audio content 1 can be used as the audio content. Furthermore, the feature extraction network in different data modalities can be used to extract the features of the data content in the corresponding modality respectively to obtain the retrieval features in different data modalities. After that, the new tag data constructed based on the retrieval features in different data modalities are stored in the tag data sets in the corresponding data modalities respectively.
[0150] In this way, it is possible to generate corresponding tag data specifically for content tags whose associated total number of historical data does not reach the preset quantity threshold, so that each content tag can be distributed in multiple pieces of tag data.
[0151] Similarly, the processing device can generate labeled data for other multimedia data that needs to be processed and is associated with a hot content, and similarly complete the generation of labeled data in each covered data modality.
[0152] For example, assume that for a hot content (denoted as hot content 1), the selected content labels include: content label 1, content label 2, and content label 3. Moreover, the total number of historical data associated with content label 1 is 35, the total number of historical data associated with content label 2 is 40, and the total number of historical data associated with content label 3 is 45. Assume that the preset data threshold is 50; then, in some feasible implementation methods, 15 pieces of multimedia data can be obtained for hot content 1; in other feasible implementation methods, the total number of multimedia data obtained for hot content 1 can be unrestricted. After that, in a feasible solution, for the first 5 pieces of multimedia data processed for hot content 1, the content labels associated with the generated labeled data are: content label 1-3; when processing the 6th to 10th pieces of multimedia data, the content labels associated with the generated labeled data are: content label 1 and 2; and, when processing the 11th to 15th pieces of multimedia data, the content labels associated with the generated labeled data are: content label 1.
[0153] In this way, with the help of the total number of associated historical data, the participation of content label 1 in the process of generating labeled data can be restricted, avoiding a content label being distributed in too many labeled data, thereby reducing the excessive occupation of storage space.
[0154] Generally speaking, for a hot content, new content labels can be determined, and the labeled data sets of each data modality can be updated; moreover, the labeled data sets in each data modality can be understood as different retrieval feature libraries. Compared with directly extracting label features, by extracting retrieval features in this application, more abundant content information can be integrated, so as to obtain more abundant retrieval features.
[0155] Further, the processing device can label appropriate content labels for the multimedia data to be labeled (abbreviated as data to be processed) according to the labeled data sets of each data modality.
[0156] Refer to Figure 6 shown in the figure, which is a schematic diagram of the annotation process of content labels in an embodiment of this application. The following combines the appendix Figure 6 , and specifically describes the process executed when annotating the content labels of the data to be processed: Step 601: The processing device obtains the data to be processed, and extracts corresponding to-be-matched features for the data to be processed respectively according to at least one data modality covered by the data to be processed.
[0157] Specifically, after obtaining the data to be processed, the processing device can perform feature extraction on the data to be processed respectively in each data modality covered by the data to be processed, and obtain the features to be matched of the data to be processed in each covered data modality, where the data modalities covered by the data to be processed are at least one.
[0158] Among them, when specifically extracting the features to be matched, the same processing method as when constructing the retrieval features above can be used for processing, that is, first, according to the covered data modalities, obtain the data content in the corresponding data modality, and then, based on the extracted content, extract the features to be matched; moreover, for any one data modality, the feature extraction method used when extracting the features to be matched is the same as the feature extraction method used when extracting the retrieval features, and this application will not expand on this here.
[0159] For example, refer to Figure 7 As shown, it is a schematic diagram of the process of extracting the features to be matched in the embodiment of this application. According to the content shown in the appendix Figure 7 As can be seen, assume that the data to be processed is a video content with a title, and the preset data modalities (or preset data modality ranges) include text modality, image modality, and audio modality. Then, it can be determined that the data to be processed covers all data modalities. In the text modality, the extracted text content includes: the title and the audio text converted from the video content. After that, text feature extraction can be performed on the text content to obtain the feature to be matched 1 in the text modality; in the image modality, the extracted image content is M images extracted from the video content. After that, image feature extraction can be performed on each of the M images respectively to obtain the feature to be matched 2 in the image modality obtained by fusing M image features; in the audio modality, the extracted audio content is the audio content separated from the video content. After that, audio feature extraction can be performed on the audio content to obtain the feature to be matched 3 in the audio modality.
[0160] Step 602: The processing device respectively performs the following operations on the features to be matched of the data to be processed in each data modality: respectively in at least one labeled dataset having a feature alignment relationship with a feature to be matched, determine the retrieval features whose feature similarity with a feature to be matched meets the similarity condition as the hit features.
[0161] In the embodiments of the present application, after the processing device obtains the to-be-matched features of the to-be-processed data in each covered data modality, taking the retrieval based on the to-be-matched features in one data modality as an example, the processing device may first obtain the feature alignment relationship determined based on the feature extraction methods in each data modality; then, based on the feature alignment relationship, determine at least one labeled dataset that matches a to-be-matched feature, and for each matched labeled dataset, respectively perform the following operations: calculate the feature similarity between a to-be-matched feature and each retrieval feature in a labeled dataset, and respectively use the N retrieval features with the largest feature similarity as the hit features that meet the preset similarity condition.
[0162] Among them, when the feature extraction networks in different data modalities exist in a multi-modal model and are synchronously trained with the pre-training of the multi-modal model, it can be determined that there is a feature alignment relationship between the retrieval features in different data modalities. Based on this, when the feature extraction methods in different data modalities are known, whether there is a feature alignment relationship between the retrieval features in different data modalities is known.
[0163] It should be noted that in the embodiments of the present application, for any to-be-matched feature, there is at least one labeled dataset with feature alignment; moreover, the value of N is set according to actual processing needs, and the present application does not make specific limitations on this. For example, the value of N is 2; taking the determination of hit features in a labeled dataset as an example, after calculating the feature similarity between a to-be-matched feature and each retrieval feature in the labeled data, they can be arranged from largest to smallest in terms of feature similarity, and the top N retrieval features with the largest feature similarity are respectively used as the hit features, where each hit feature is associated with at least one content label.
[0164] For example, refer to Figure 8 shown, which is a schematic diagram of the process for determining hit features in the embodiments of the present application. According to the content Figure 8 shown in the figure, it is assumed that the preset data modalities include a text modality, an image modality, and an audio modality; for a to-be-processed data, the to-be-matched feature 1 in the text modality, the to-be-matched feature 2 in the image modality, and the to-be-matched feature 3 in the audio modality are extracted; moreover, the feature extraction networks in the text modality, the image modality, and the audio modality are pre-trained in a multi-modal model; the value of M is 2. Based on this, for the to-be-matched feature 1, the labeled dataset 1 in the text modality, the labeled dataset 2 in the image modality, and the labeled dataset 3 in the audio modality can all be regarded as the labeled datasets with feature alignment with the to-be-matched feature 1. Then, when processing the to-be-matched feature 1, the retrieval features with the top 2 feature similarities can be respectively screened out from the labeled datasets in the three data modalities as the hit features; similarly, the to-be-matched features 2 and 3 can be processed.
[0165] In this way, by calculating the feature similarity, for each feature to be matched, the hit features that meet the requirements of similarity can be found respectively in the labeled dataset with aligned features.
[0166] Step 603: The processing device determines the content labels matched by the data to be processed based on the feature similarities corresponding to the filtered hit features and the content labels determined for the corresponding hit features.
[0167] In some feasible processes of determining the content labels matched by the data to be processed in this application, the processing device can first determine at least one content label associated with each hit feature with the highest feature similarity, and then determine each content label as the content label matched by the data to be processed respectively.
[0168] In other feasible processes of determining the content labels matched by the data to be processed in this application, the processing device can first determine the content labels associated with the corresponding hit features, and then screen out the content labels matched by the data to be processed from the content labels according to the feature similarity between the content labels and the corresponding hit features and the total number of hit features.
[0169] In the embodiment of this application, when executing step 603, the processing device can determine the content labels determined for the corresponding hit features based on the labeled data to which the filtered hit features belong respectively; then, for each content label, the following operations are performed respectively: at least one retrieval feature associated with a content label among the filtered retrieval features is used as the target feature respectively, and based on the feature similarities corresponding to the at least one target feature and the total number of features of the hit features, an adaptation value corresponding to a content label is obtained; then, the content label with the adaptation value reaching the set value is used as the content label matched by the data to be processed.
[0170] It should be noted that in the embodiment of this application, after the processing device determines the hit features corresponding to a data to be processed and the feature similarities corresponding to the hit features respectively through similarity calculation, at least one content label associated with each hit feature can be used as the content labels to be matched; among them, the value of the set value is set according to actual processing requirements, and this application does not make specific limitations on this.
[0171] When specifically determining the adaptation value corresponding to a content label, for each content label, the weighted average result can be calculated based on the feature similarities of the corresponding hit features to obtain the adaptation value.
[0172] For example, assume that there are two preset data modalities, namely text modality and image modality; assume that a data to be processed only covers the text modality, and the features to be matched are extracted in the text modality. Additionally, assume that there is a feature alignment relationship between the retrieval features in the text modality and the image modality. For the features to be matched, in the retrieval dataset in the text modality, the relevant information of the two identified hit features is as follows: the first-ranked hit feature: {"score": 0.9, "tag": ["Open Tournament in Country R", "Super Zhang"]}, the second-ranked hit feature: {"score": 0.7, "tag": ["Zhang", "Super Zhang"]}; in the retrieval dataset in the image modality, the relevant information of the two identified hit features is as follows: the first-ranked hit feature {"score": 0.8, "tag": ["Zhang", "Super Zhang"]}, the second-ranked hit feature: {"score": 0.6, "tag": ["Zhang"]}; where, score represents the corresponding feature similarity; tag represents the corresponding content labels.
[0173] Continuing with the above example, the identified content labels include "Open Tournament in Country R", "Super Zhang", and "Zhang"; then, for the content label "Super Zhang", the calculated adaptation value is: (0.9 + 0.7 + 0.8 + 0) / 4 = 0.6; for the content label "Zhang", the calculated adaptation value is: (0 + 0.7 + 0.8 + 0.6) / 4 = 0.525; for the content label "Open Tournament in Country R", the calculated adaptation value is: (0.9 + 0 + 0 + 0) / 4 = 0.225. Furthermore, assume that the set value is 0.55, then the content label "Super Zhang" can be determined as the content label that the data to be processed matches.
[0174] In this way, by calculating the corresponding adaptation values for each content label covered by the hit features respectively, it is possible to screen the content labels within the selected range of content labels, thereby effectively determining the most suitable content label. Moreover, combined with the generation process of the label data, since this application can achieve the timely update of the label data, the data to be processed can be labeled with the latest content labels, improving the robustness of the label determination process.
[0175] Moreover, overall, when determining the content tags matching the data to be processed, by extracting the features to be matched in different data modalities and retrieving in the label datasets of different data modalities, cross-modal pairwise retrieval can be achieved between multiple data modalities, enabling multi-way recall, thereby improving the determination effect of content tags. Moreover, by screening in the candidate tag sets selected in multiple ways and combining the screening strategy of multiple-way tags based on the adaptation value, appropriate content tags can be determined for the data to be processed, making the tag annotation process highly robust.
[0176] The following will take generating new hot tags for hot content as an example and describe the process of generating the tag data for retrieval basis in conjunction with the accompanying drawings: Refer to Figure 9 As shown, it is a schematic diagram of the process of generating the tag data for new hot tags in an embodiment of the present application. According to the content Figure 9 shown in the schematic, the processing device can regard each hot instance crawled as a hot content, and regard the description text of a hot instance as the description text associated with a hot content. Furthermore, a large language model can be used to extract a candidate tag set based on the description text of the hot content and a preset prompt template. Among them, there may only be new hot tags in the candidate tag set that do not hit the preset content tag set. After that, after determining that new hot tags (or new content tags) are extracted, the new hot tags are stored in the content tag set to update the content tag set.
[0177] Continuing to describe in conjunction with the attached Figure 9 drawings, assuming that the preset data modalities are the image modality and the text modality, after determining that new hot tags (assuming they are content tag q) are extracted, the multimedia data associated with the hot content can be obtained. Among them, the multimedia data may include articles or videos. Furthermore, for each multimedia data, the corresponding retrieval features can be extracted respectively according to the covered data modality. The content Figure 9 shown in the schematic is that at most, based on one multimedia data, the retrieval features in the image modality and the text modality can be extracted respectively. After that, the retrieval features associated with the content tag q are stored in the corresponding label datasets according to the corresponding data modalities, thereby realizing the storage of the new hot tags and related tag data in the library.
[0178] Refer to Figure 10 As shown, it is a schematic diagram of the process of annotating content tags in an embodiment of the present application. Against the background of the attached Figure 9 drawings, continue to perform the attached Figure 10For the description, assume that the two preset data modalities are the text modality under natural language processing and the image modality under computer vision, and there is a feature alignment relationship between the retrieval features in the text modality and the image modality. After extracting the features to be matched in the text modality and the features to be matched in the image modality for the data to be processed, similar features are retrieved from the retrieval database, which includes the label dataset in the text modality and the label dataset in the image modality. Furthermore, four results can be obtained: text modality - text modality, text modality - image modality, image modality - text modality, and image modality - image modality.
[0179] For example, text modality - text modality means using the feature w to be matched in the text modality of the data to be processed to retrieve in the label dataset of the text modality, obtaining the feature similarity between the feature w to be matched of the data to be processed and each retrieval feature in the label dataset, then arranging them from largest to smallest, and taking the top k retrieval features. Each retrieval feature has at least one corresponding content label, and the feature similarity is used as the score. Based on this, four retrieval results can be obtained. The number of retrieval features in each path is k, and each retrieval feature has a corresponding score and content label. Furthermore, by taking the weighted average of each content label, the final adaptation value of each content label is obtained, and finally, screening is performed according to the set value to obtain the content labels matched by the data to be processed.
[0180] Refer to Figure 11 As shown, it is a schematic diagram of the content processing process in an embodiment of the present application. According to the content Figure 11 shown, the processing process of the data to be processed can be divided into three stages. The first stage is content production, in which the data to be processed generated is obtained. The second stage is content processing, specifically, according to the Figure 9 in which the latest label datasets of each data modality are generated, and the Figure 10 content label annotation method shown in, the appropriate content labels are determined. The third stage is content distribution, that is, the data to be processed is distributed to the target objects according to the content labels. Specifically, the data to be processed can be classified according to the content labels determined for the data to be processed, so that the data to be processed can be specifically distributed to the target objects in need.
[0181] Generally speaking, the present application provides a simple and general implementation method, which can automatically discover and store new and popular labels, and perform automatic label annotation. Therefore, the technical solution proposed in the present application has extremely high timeliness and can quickly respond to hot spot changes; by reducing or even eliminating the dependence on manual intervention, the labor cost is significantly saved, and at the same time, the efficiency of machine annotation of new and popular labels is greatly improved.
[0182] When specifically generating tag data, based on the descriptive text of the crawled hot content, new hot tags suitable for inclusion in the tag system are automatically generated. At the same time, various associated multimedia data are automatically collected, retrieval features are extracted, and stored in the retrieval database, providing a processing basis for subsequent automatic content tag annotation. This effectively avoids consuming a large amount of manual labor to collect new hot tags and samples, and improves the real-time performance of data storage.
[0183] In addition, during the process of automatically tagging new hot tags, for the data to be processed with tags to be added online, by extracting the features to be matched (Embedding) in each covered data modality, and then retrieving similar Embeddings in each tag dataset with feature alignment, the content tags corresponding to the hit features can be obtained, and then the matching content tags can be screened out. Moreover, through the retrieval method, there is no need to retrain the model like in the classification task to support the tagging of new hot tags, effectively improving the real-time performance of new hot tag tagging. In addition, by retrieving in each tag dataset to map out the adapted content tags, the retrieval object is the retrieval feature extracted from the multimedia data rather than the tag feature, which has the following advantages: compared with the limited semantic information contained in the tag as a word, the video or article can extract richer semantic features, so it is more suitable as the retrieval object; the retrieval feature extracted from the multimedia data covers more abundant and comprehensive information content. In this way, the present application uses advanced natural language processing and machine learning technologies to automatically analyze and identify potential new hot tags, which are then systematically stored in the content tag set. Based on the constructed tag data, automatic tagging can be performed on the data to be processed. This process not only ensures the accuracy and relevance of the annotated content tags, but also makes information retrieval more efficient.
[0184] Furthermore, the technical solution proposed in the present application has very high versatility and can adapt to different application scenarios and industry requirements. Whether in social media or news platforms, it can effectively improve content distribution and user experience. Through this innovative automated process, market trends can be better captured and resource allocation can be optimized.
[0185] Based on the same inventive concept, refer to Figure 12 As shown, it is a schematic logical structure diagram of the tag data generation device in the embodiment of the present application. The tag data generation device 1200 includes an acquisition unit 1201, a determination unit 1202, and a storage unit 1203, where The acquisition unit 1201 is used to acquire a set of descriptive texts associated with a hot content, and extract keywords from the set of descriptive texts to obtain a corresponding candidate tag set; A determination unit 1202, configured to, when it is determined that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed a preset quantity threshold, respectively use the at least one candidate tag as a content tag, and obtain various multimedia data associated with a hot content; each multimedia data includes data content of at least one data modality; the total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which each tag data including a candidate tag is generated when generating content tags; A storage unit 1203, configured to perform the following storage operations respectively for each data modality in each multimedia data: Extract features from the data content of a data modality in a multimedia data to obtain retrieval features, and store the tag data generated based on the retrieval features and at least one content tag into the tag data set set for a corresponding data modality.
[0186] Optionally, after performing the storage operations respectively for each data modality in each multimedia data, the apparatus further includes an execution unit 1204, and the execution unit 1204 is configured to: Obtain data to be processed, and respectively extract corresponding to-be-matched features for the data to be processed according to at least one data modality covered by the data to be processed; For the to-be-matched features of the data to be processed in each data modality, respectively perform the following operations: respectively in at least one tag data set having a feature alignment relationship with a to-be-matched feature, determine the retrieval features whose feature similarity with a to-be-matched feature meets the similarity condition as hit features; Based on the feature similarities corresponding to the respective hit features selected, and in combination with the respective content tags determined for the respective hit features, determine the content tags matched by the data to be processed.
[0187] Optionally, when determining the content tags matched by the data to be processed based on the feature similarities corresponding to the respective hit features selected, and in combination with the respective content tags determined for the respective hit features, the execution unit 1204 is configured to: Based on the tag data to which the respective hit features selected belong, determine the respective content tags determined for the respective hit features; For each content tag, respectively perform the following operations: respectively use at least one retrieval feature associated with a content tag among the retrieved features selected as target features, and based on the feature similarities corresponding to the at least one target feature, in combination with the total number of features of the respective hit features, obtain an adaptation value corresponding to a content tag; Use the content tags whose adaptation values reach a set value as the content tags matched by the data to be processed.
[0188] Optionally, when determining a retrieved feature whose feature similarity with a feature to be matched meets the similarity condition as a hit feature in at least one tag dataset that has a feature alignment relationship with a feature to be matched, the execution unit 1204 is configured to: Obtain a feature alignment relationship determined based on a feature extraction method in each data modality; Based on the feature alignment relationship, determine at least one tag dataset that matches a feature to be matched, and for each matched tag dataset, respectively perform the following operations: Calculate the feature similarity between a feature to be matched and each retrieved feature in a tag dataset respectively, and use the N retrieved features with the largest feature similarity as the hit features that meet the preset similarity condition respectively.
[0189] Optionally, when respectively using at least one candidate tag as a content tag, the determination unit 1202 is configured to: Determine at least one candidate tag as a content tag respectively; Use a candidate tag with a total number of associated historical data of 0 among at least one candidate tag as a new content tag and add it to a preset content tag set.
[0190] Optionally, when a data modality is an image modality; when extracting features from the data content in a data modality of a multimedia data to obtain a retrieved feature, the storage unit 1203 is configured to: Extract the image content in a multimedia data; wherein, the image content includes at least one image; Use an image feature extraction network in a pre-trained multi-modal model to extract features from at least one image respectively to obtain corresponding image features; Fuse at least one image feature to obtain a corresponding retrieved feature.
[0191] Optionally, when extracting the image content in a multimedia data, the storage unit 1203 is configured to perform any one of the following operations: When a multimedia data includes at least one video content, use M frames of images respectively extracted from each video content as the extracted image content; When a multimedia data includes video content and an original image, use M frames of images extracted from the video content and the original image as the extracted image content.
[0192] Optionally, when a data modality is a text modality; when extracting features from the data content in a data modality of a multimedia data to obtain a retrieved feature, the storage unit 1203 is configured to: Extract the text content from a piece of multimedia data; wherein, the text content covers at least one of the following text types: the audio text obtained after performing speech-to-text conversion processing; the original text directly included in the multimedia data; Use the text feature extraction network in the multimodal model to perform feature extraction based on the text content to obtain the corresponding retrieval features; wherein, there is a feature alignment relationship between the retrieval features extracted by the feature extraction networks of different modalities in the multimodal model.
[0193] Optionally, when a data modality is the audio modality and the data content in a data modality of a piece of multimedia data is used for feature extraction to obtain retrieval features, the storage unit 1203 is used for: Extract the audio content from a piece of multimedia data; wherein, the audio content covers at least one of the following audio types: the content audio separated from the video content; the original audio directly included in the multimedia data; Perform audio feature extraction on the audio content to obtain the corresponding retrieval features.
[0194] Optionally, before obtaining the description text set associated with a hot content, the device further includes a crawling unit 1205, and the crawling unit 1205 is used for: Respectively crawl the description texts of each hot instance from each preset data web page; Based on the text similarity between every two description texts, divide the description texts that meet the preset similarity conditions into the same description text set to obtain each description text set; wherein, at least one description text is included in a description text set; For each description text set, perform the following operations respectively: divide the hot instances corresponding to the description texts in a description text set into a hot content, and establish an association relationship between a hot content and a description text set.
[0195] Optionally, when obtaining each multimedia data associated with a hot content, the crawling unit 1205 is used for: For at least one hot instance covered by a hot content, perform the following operations respectively: based on the data web page relied on when crawling the description text associated with a hot instance, obtain the data link information associated with a hot instance, and respectively obtain each multimedia data belonging to a hot instance according to each data link information; Use the obtained multimedia data as each multimedia data associated with a hot content.
[0196] Optionally, when storing the tag data generated based on the retrieval features and at least one content tag into the tag data set set for a corresponding data modality, the storage unit 1203 is used for: In at least one content tag, remove the total number of associated historical data for the content tags that reach a preset quantity threshold, to obtain the remaining content tags; Use the retrieval features and the associated content tags as the generated tag data, store it in the tag dataset set for a corresponding data modality, and use a piece of multimedia data as the historical multimedia data associated with the content tag, and update the total number of historical data associated with the content tag.
[0197] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.
[0198] After introducing the method and apparatus for generating tag data according to the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.
[0199] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0200] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. Refer to Figure 13 As shown, it is a schematic diagram of the hardware composition structure of an electronic device applying the embodiment of the present application. In one embodiment, the electronic device can be Figure 1 The processing device 120 shown. In this embodiment, the structure of the electronic device can be as Figure 13 Shown, including a memory 1301, a communication module 1303, and one or more processors 1302.
[0201] The memory 1301 is used to store the computer program executed by the processor 1302. The memory 1301 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.
[0202] The memory 1301 can be a volatile memory, such as a random-access memory (RAM); the memory 1301 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1301 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1301 can be a combination of the above memories.
[0203] The processor 1302 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1302 is used to implement the above method for generating prompt label data when calling the computer program stored in the memory 1301.
[0204] The communication module 1303 is used to communicate with the client device and the server.
[0205] In the embodiments of the present application, the specific connection medium between the above memory 1301, communication module 1303, and processor 1302 is not limited. In the embodiments of the present application Figure 13 it is described that the memory 1301 and the processor 1302 are connected through a bus 1304, and the bus 1304 is described in thick lines in Figure 13 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1304 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 13 only a thick line is used to describe it in
[0206] The memory 1301 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the method for generating prompt label data in the embodiments of the present application. The processor 1302 is used to execute the above method for generating label data, as Figure 2 shown.
[0207] In another embodiment, the electronic device can also be other electronic devices. Refer to Figure 14 shown, which is a schematic diagram of the hardware composition structure of another electronic device applying the embodiments of the present application. The electronic device can specifically be Figure 1 the client device 110 shown. In this embodiment, the structure of the electronic device can be asFigure 14 As shown, it includes components such as a communication component 1410, a memory 1420, a display unit 1430, a camera 1440, a sensor 1450, an audio circuit 1460, a Bluetooth module 1470, and a processor 1480.
[0208] The communication component 1410 is used to communicate with a server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short - range wireless transmission technology, and through the WiFi module, the electronic device can help users send and receive information.
[0209] The memory 1420 can be used to store software programs and data. The processor 1480 executes various functions and data processing of the client device 110 by running the software programs or data stored in the memory 1420. In this application, the memory 1420 can store an operating system and various application programs, and can also store computer programs related to the generation method of the prompt label data in the embodiments of this application.
[0210] The display unit 1430 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the client device 110. Specifically, the display unit 1430 may include a display screen 1432 disposed on the front of the client device 110. The display unit 1430 can be used to display pages and the like.
[0211] The display unit 1430 can also be used to receive input numerical or character information, generating signal inputs related to the user settings and function control of the client device 110. Specifically, the display unit 1430 may include a touch screen 1431 disposed on the front of the client device 110, which can collect touch operations of the user on or near it.
[0212] Among them, the touch screen 1431 can cover the display screen 1432, or the touch screen 1431 and the display screen 1432 can be integrated to implement the input and output functions of the client device 110. After integration, it can be simply called a touch display screen. In this application, the display unit 1430 can display application programs and corresponding operation steps.
[0213] The camera 1440 can be used to capture static images, and users can post comments on the images captured by the camera 1440 through an application. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1480 to be converted into a digital image signal.
[0214] The client device 110 may further include at least one sensor 1450, such as an acceleration sensor 1451, a distance sensor 1452, a fingerprint sensor 1453, and a temperature sensor 1454. The client device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0215] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface between the user and the client device 110. The audio circuit 1460 can transmit the electrical signal converted from the received audio data to the speaker 1461, and the speaker 1461 converts it into a sound signal for output. On the other hand, the microphone 1462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1460 and converted into audio data, and then the audio data is output to the communication component 1410 to be sent to, for example, another client device, or the audio data is output to the memory 1420 for further processing.
[0216] The Bluetooth module 1470 is used to interact with other Bluetooth devices with Bluetooth modules through the Bluetooth protocol.
[0217] The processor 1480 is the control center of the client device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 1420 and calling data stored in the memory 1420, it executes various functions of the client device and processes data. In some embodiments, the processor 1480 may include at least one processing unit; the processor 1480 may also integrate an application processor and a baseband processor. In this application, the processor 1480 can run an operating system, application programs, user interface display and touch response, and methods related to the generation of tag data in the embodiments of this application. In addition, the processor 1480 is coupled to the display unit 1430.
[0218] In some possible implementations, aspects of the method for generating tag data provided in this application can also be implemented in the form of a computer program product, which includes a computer program. When the computer program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the method for generating tag data according to various exemplary implementations of this application described above in this specification. For example, the electronic device can execute the steps as shown in Figure 2 shown.
[0219] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0220] The program product of the implementation of this application can adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can run on an electronic device. However, the program product of this application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.
[0221] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0222] The computer program included on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0223] The computer program for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's electronic device, partially on the user's electronic device, executed as an independent software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external electronic device (for example, by connecting through the Internet using an Internet service provider).
[0224] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more of the above-described units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0225] In addition, although the operations of the method of this application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0226] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable computer programs.
[0227] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.
[0228] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0229] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for generating label data, characterized in that, Including: Obtain a set of descriptive texts associated with a hot content, and perform keyword extraction on the set of descriptive texts to obtain a corresponding set of candidate tags; When it is determined that there is at least one candidate tag in the set of candidate tags whose total number of associated historical data does not exceed a preset quantity threshold, respectively use the at least one candidate tag as a content tag, and obtain each multimedia data associated with the hot content; Each multimedia data contains data content of at least one data modality; The total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which each tag data containing the candidate tag is generated when generating content tags; For each data modality in each multimedia data, respectively perform the following storage operations: Extract features from the data content of a data modality in a multimedia data to obtain retrieval features, and store the tag data generated based on the retrieval features and at least one content tag into the tag data set set for the data modality.
2. The method according to claim 1, wherein After respectively performing the storage operations for each data modality in each multimedia data, the method further includes: Obtain data to be processed, and respectively extract corresponding features to be matched for the data to be processed according to at least one data modality covered by the data to be processed; For the features to be matched of the data to be processed in each data modality, respectively perform the following operations: In at least one tag data set having a feature alignment relationship with a feature to be matched, determine the retrieval features whose feature similarity with the feature to be matched meets the similarity condition as the hit features; Based on the feature similarities corresponding to the respective hit features selected, and in combination with the respective content tags determined for the respective hit features, determine the content tags matched by the data to be processed.
3. The method according to claim 2, wherein The determining the content tags matched by the data to be processed based on the feature similarities corresponding to the respective hit features selected, and in combination with the respective content tags determined for the respective hit features, includes: Based on the tag data to which the respective hit features selected belong, determine the respective content tags determined for the respective hit features; For each of the content tags, respectively perform the following operations: Take at least one of the retrieval features selected that is associated with a content tag as a target feature, and based on the feature similarities corresponding to the at least one target feature, and in combination with the total number of features of the respective hit features, obtain an adaptation value corresponding to the content tag; Take the content tags whose adaptation values reach the set value as the content tags matched by the data to be processed.
4. The method according to claim 2, wherein The determining the retrieval features whose feature similarity with the feature to be matched meets the similarity condition as the hit features in at least one tag data set having a feature alignment relationship with a feature to be matched includes: Obtain the feature alignment relationship determined based on the feature extraction methods in each data modality; Based on the feature alignment relationship, determine at least one tag data set that matches a feature to be matched, and for each of the matched tag data sets, respectively perform the following operations: Calculate the feature similarity between a feature to be matched and each retrieval feature in a label dataset respectively, and use the N retrieval features with the largest feature similarity as the hit features that meet the preset similarity conditions respectively.
5. The method according to claim 1, characterized in that The step of respectively using the at least one candidate label as a content label includes: Determine the at least one candidate label as a content label respectively; Use the candidate label with the total number of associated historical data being 0 in the at least one candidate label as a new content label and add it to a preset content label set.
6. The method according to any one of claims 1-5, characterized in that, One data modality is the image modality; the step of extracting retrieval features from the data content in one data modality of a multimedia data includes: Extract the image content in a multimedia data; where the image content includes at least one image; Use the image feature extraction network in a pre-trained multi-modal model to extract features from each of the at least one image respectively to obtain corresponding image features; Fuse at least one image feature to obtain a corresponding retrieval feature.
7. The method according to claim 6, wherein When extracting the image content in a multimedia data, perform any one of the following operations: When a multimedia data includes at least one video content, use the M frames of images extracted from each of the video contents respectively as the extracted image content; When a multimedia data includes video content and original images, use the M frames of images extracted from the video content and the original images as the extracted image content.
8. The method according to any one of claims 1-5, characterized in that One data modality is the text modality; the step of extracting retrieval features from the data content in one data modality of a multimedia data includes: Extract the text content in a multimedia data; where the text content covers at least one of the following text types: the audio text obtained by performing speech-to-text conversion processing; the original text directly included in the multimedia data; Use the text feature extraction network in the multi-modal model to extract features based on the text content to obtain corresponding retrieval features; where there is a feature alignment relationship between the retrieval features extracted by the feature extraction networks of different modalities in the multi-modal model.
9. The method according to any one of claims 1-5, characterized in that, One data modality is the audio modality; the step of extracting retrieval features from the data content in one data modality of a multimedia data includes: Extract the audio content in a multimedia data; where the audio content covers at least one of the following audio types: the content audio separated from the video content; the original audio directly included in the multimedia data; Perform audio feature extraction on the audio content to obtain corresponding retrieval features.
10. The method according to any one of claims 1-5, characterized in that, Before obtaining the description text set associated with a hot content, the method further includes: Crawl the description texts of each hot instance from each preset data web page respectively; Based on the text similarity between every two of the description texts, divide the description texts that meet the preset similarity conditions into the same description text set to obtain each description text set; where one description text set includes at least one description text; For each of the described text sets, the following operations are performed separately: classify the hot instances corresponding to each described text in a described text set into a hot content, and establish an association relationship between the one hot content and the one described text set.
11. The method according to claim 10, wherein The obtaining of the multimedia data associated with the one hot content includes: For at least one hot instance covered by a hot content, the following operations are performed separately: based on the data web page according to which the described text associated with a hot instance is crawled, obtain the data link information associated with the one hot instance, and respectively obtain the multimedia data belonging to the one hot instance according to the data link information; Use the obtained multimedia data as the multimedia data associated with the one hot content.
12. The method according to any one of claims 1-5, characterized in that, The storing of the tag data generated based on the retrieval feature and at least one content tag into the tag data set set for the one data modality includes: Among the at least one content tag, remove the content tags whose total number of associated historical data reaches the preset quantity threshold, to obtain the remaining content tags; Use the retrieval feature and the associated content tags as the generated tag data, store them into the tag data set set for the one data modality, use the one multimedia data as the historical multimedia data associated with the content tag, and update the total number of historical data associated with the content tag.
13. A device for generating label data, characterized in that, Includes: An obtaining unit, configured to obtain a described text set associated with a hot content, and perform keyword extraction on the described text set to obtain a corresponding candidate tag set; A determining unit, configured to when it is determined that there is at least one candidate tag in the candidate tag set whose total number of associated historical data does not exceed the preset quantity threshold, respectively use the at least one candidate tag as a content tag, and obtain the multimedia data associated with the one hot content; each multimedia data includes data content of at least one data modality; The total number of historical data associated with a candidate tag refers to the total number of historical multimedia data based on which the tag data including the candidate tag is generated when generating the content tags; A storing unit, configured to perform the following storing operations separately for each data modality in each multimedia data: Extract features from the data content under a data modality in a multimedia data to obtain a retrieval feature, and store the tag data generated based on the retrieval feature and at least one content tag into the tag data set set for the one data modality.
14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1-12 is implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the method described in any one of claims 1-12 is implemented.
16. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1-12 is implemented.
Citation Information
Cited By
Label acquisition method and device based on LLM model, and medium
CN121303112A