A label generation method and device, computer equipment and a storage medium
By using a feature library to store the output data of each network node in the multimedia content feature extraction network, the problem of low label generation efficiency in the existing technology is solved, and more efficient feature extraction and label generation are achieved.
Patent Information
- Application Number
- CN202211620304.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-12-15
AI Technical Summary
Existing technologies are inefficient in the process of extracting features from multimedia content, resulting in time-consuming and cumbersome tag generation.
A target feature extraction network based on multi-layer network nodes is adopted. The output data of each layer of network nodes is stored in a feature library, and the required feature information is directly obtained from the feature library to avoid duplicate output.
It improved the efficiency of label generation, reduced feature extraction time, and optimized the feature extraction process.
Smart Images

Figure CN115827904B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and in particular, to a label generation method and device, a computer device and a storage medium. BACKGROUND
[0002] For multimedia content such as videos and web pages, there are more specific information contained therein. In order to more easily label the multimedia content for classification and arrangement or targeted push, a label can be determined for the multimedia content to clearly and easily express the information reflected by the multimedia content by using a short label.
[0003] When generating a label for multimedia content, feature extraction can be performed on the multimedia content, and the result of the feature extraction can be used to generate a label. However, the current feature extraction is complicated and time-consuming, which results in low efficiency of generating a label. SUMMARY
[0004] The present disclosure provides at least a label generation method and device, a computer device and a storage medium.
[0005] In a first aspect, the present disclosure provides a label generation method, including: obtaining multimedia content to be processed, and extracting information to be processed under a plurality of content dimensions from the multimedia content; inputting the information to be processed into a first layer network node of a target feature extraction network for feature extraction to obtain first feature information, and storing the first feature information in a feature library; the target feature extraction network is a multi-layer network constructed based on a dependency relationship between a plurality of network nodes; using at least an Nth layer network node of the target feature extraction network, feature extraction is performed on feature information output by an (N-1)th layer network node stored in the feature library, respectively, and the extracted feature information is stored in the feature library; wherein N is a positive integer greater than 1; based on an obtained feature configuration list, target feature information is extracted from the feature library, and a label production result corresponding to the multimedia content is determined based on the target feature information.
[0006] In an optional implementation, before the information to be processed is input into the first layer network node of the target feature extraction network for feature extraction to obtain the first feature information, the method further includes: based on a resource type of the multimedia content, determining the target feature extraction network from a plurality of preset feature extraction networks.
[0007] In an optional implementation, the target feature extraction network is determined from a plurality of preset feature extraction networks based on a resource type of the multimedia content, including: determining a plurality of labeling dimensions corresponding to the multimedia content based on the resource type of the multimedia content; and determining the target feature extraction network corresponding to each of the labeling dimensions from the plurality of preset feature extraction networks.
[0008] In an optional implementation, the extracted feature information is stored in the feature library, including: storing the feature information and a feature identifier corresponding to the feature information in the feature library, where the feature identifier corresponding to the feature information is associated with a network node extracting the feature information and / or a feature type of the feature information.
[0009] In an optional implementation, the target feature extraction network is trained in the following manner: sample multimedia content corresponding to the target feature extraction network is obtained, the sample multimedia content includes sample to-be-processed information and a sample label corresponding to the sample multimedia content; and the target feature extraction network is trained by using the sample multimedia content to obtain the target feature extraction network.
[0010] In an optional implementation, the label production result includes a plurality of generated labels and a confidence corresponding to each label, and the method further includes: determining a target label matched with the multimedia content based on the plurality of labels, the confidence corresponding to each label, and a label production strategy matched with the target business party.
[0011] In an optional implementation, the to-be-processed multimedia content is obtained, including: receiving a plurality of multimedia contents, and selecting the to-be-processed multimedia content from the plurality of multimedia contents based on a preset screening condition.
[0012] In a second aspect, the embodiments of the present disclosure further provide a label generation apparatus, comprising: an acquisition module, configured to acquire multimedia content to be processed, and extract to-be-processed information under a plurality of content dimensions from the multimedia content; a first processing module, configured to input the to-be-processed information into a first layer network node of a target feature extraction network for feature extraction, to obtain first feature information, and store the first feature information in a feature library; the target feature extraction network is a multi-layer network constructed based on a dependency relationship between a plurality of network nodes; a second processing module, configured to use at least an Nth layer network node of the target feature extraction network to perform feature extraction on feature information output by an (N-1) th layer network node of the feature library respectively, and store the extracted feature information in the feature library; wherein N is a positive integer greater than 1; and a generation module, configured to extract target feature information from the feature library based on an acquired feature configuration list, and determine a label production result corresponding to the multimedia content based on the target feature information.
[0013] In an optional implementation, before the first processing module inputs the to-be-processed information into the first layer network node of the target feature extraction network for feature extraction to obtain the first feature information, the first processing module is further configured to determine the target feature extraction network from a plurality of preset feature extraction networks based on a resource type of the multimedia content.
[0014] In an optional implementation, when the first processing module determines the target feature extraction network from the plurality of preset feature extraction networks based on the resource type of the multimedia content, the first processing module is configured to determine a plurality of labeling dimensions corresponding to the multimedia content based on the resource type of the multimedia content, and determine the target feature extraction network corresponding to each labeling dimension from the plurality of preset feature extraction networks based on the determined plurality of labeling dimensions.
[0015] In an optional implementation, when the first processing module stores the extracted feature information in the feature library, the first processing module is configured to store the feature information and a feature identifier corresponding to the feature information in the feature library; wherein the feature identifier corresponding to the feature information is associated with a network node for extracting the feature information and / or a feature type of the feature information.
[0016] In an optional implementation, the target feature extraction network is trained in the following manner: a sample multimedia content corresponding to the target feature extraction network is acquired; the sample multimedia content comprises sample to-be-processed information and a sample label corresponding to the sample multimedia content; the target feature extraction network to be trained is trained using the sample multimedia content, to obtain the target feature extraction network.
[0017] In an optional implementation, the label production result includes a plurality of generated labels and a corresponding confidence of each label; and the device further includes a third processing module configured to determine a target label matched with the multimedia content based on the plurality of labels and the corresponding confidence of each label, and the acquired label production strategy matched with the target business party.
[0018] In an optional implementation, the acquisition module, when acquiring the multimedia content to be processed, is configured to receive a plurality of multimedia contents, and select the multimedia content to be processed from the plurality of multimedia contents based on a preset screening condition.
[0019] In a third aspect, the optional implementation of the present disclosure further provides a computer device, a processor and a memory, the memory stores machine readable instructions executable by the processor, and the processor is configured to execute the machine readable instructions stored in the memory, and the machine readable instructions are executed by the processor to perform the steps of the first aspect or any possible implementation of the first aspect.
[0020] In a fourth aspect, the optional implementation of the present disclosure further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed to perform the steps of the first aspect or any possible implementation of the first aspect.
[0021] The label generation method provided by the embodiments of the present disclosure can store the data output by the network nodes in each layer of the multi-layer network in the feature library when extracting features from the information to be processed in the multimedia content, so that the feature information required by multiple network nodes can be directly obtained from the feature library when the feature information required by multiple network nodes is the same, and the output of the same feature information by the network nodes outputting the feature information does not need to be repeated multiple times, so that the time required for feature extraction can be reduced, and the efficiency of generating labels is improved.
[0022] In order to make the above objectives, characteristics and advantages of the present disclosure more apparent and understandable, the following preferred embodiments are specifically described below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. The drawings incorporated into the specification and form a part of the specification, which show the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the specification. It should be understood that the following drawings only show some of the embodiments of the present disclosure, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.
[0024] Figure 1 A flow chart of a label generation method provided by an embodiment of the present disclosure is shown;
[0025] Figure 2 A schematic diagram of a target feature extraction network provided by an embodiment of the present disclosure is shown;
[0026] Figure 3 A schematic diagram of a modular system for generating multimedia content provided by an embodiment of the present disclosure is shown;
[0027] Figure 4 A schematic diagram of a label generation device provided by an embodiment of the present disclosure is shown;
[0028] Figure 5 A schematic diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0029] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will combine the drawings in the embodiments of the present disclosure to make a clear and complete description of the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present disclosure.
[0030] It is found through research that after generating a label for multimedia content, the label can be used to express information contained in the multimedia content and further used for classification and arrangement or pushing. When generating a label for multimedia content, specific feature extraction is performed by a fusion model composed of multiple feature extraction models, and there is a need to share features as input data between some of the feature extraction models. In order to enable these feature extraction models to obtain the features as input data for feature extraction, the features need to be generated multiple times. This way of generating the same features multiple times as input data for different feature extraction models is cumbersome and time-consuming, resulting in low efficiency of generating labels.
[0031] Based on the above research, the present disclosure provides a label generation method. When performing feature extraction on to-be-processed information in multimedia content, for a multi-layer network in a target feature extraction network used for feature extraction, data output by network nodes in each layer network is stored in a feature library. If there are multiple network nodes that need to input the same feature information, the feature information can be directly obtained from the feature library, and the network nodes that output the feature information do not need to repeatedly output the same feature information multiple times. Therefore, the time required for feature extraction can be reduced, thereby improving the efficiency of generating labels.
[0032] The above-mentioned defects are the result of the inventors' practice and careful research, and therefore the discovery process of the above-mentioned problems and the solutions proposed by the present disclosure to solve the above-mentioned problems should be the contribution of the inventors to the present disclosure.
[0033] It should be noted that similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0034] To facilitate understanding of the present embodiment, first, a label generation method disclosed by the present embodiment is introduced in detail. The execution subject of the label generation method provided by the present embodiment is generally a computer device with certain computing power, which includes, for example, a terminal device or a server or other processing device. The terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the label generation method can be realized by a processor calling computer readable instructions stored in a memory.
[0035] The label generation method provided by the embodiments of the present disclosure is described below. The label generation method provided by the embodiments of the present disclosure can be used to generate corresponding labels for videos, web pages, audios, live broadcasts and other multimedia contents. The labels of the multimedia contents described herein can represent the classification or content of the multimedia contents. When the classification is labeled by using the labels, different labels can be obtained under different classification standards, such as classification into high-heat content, high-quality content and the like according to the consumption ability of the multimedia content, or classification into entertainment content, important news content and the like according to the information category. When the classification is labeled by using the content, it can also be classified into home category, makeup category, clothing category and the like to specifically reflect the specific content involved in the multimedia content.
[0036] After generating the corresponding labels for the multimedia content, the labels can be used to summarize and identify the multimedia video, and further used to classify and arrange the multimedia content and push and the like. For example, for a video playing platform, if a user selects to play a certain type of video, or it can be determined that the user is interested in a certain type of video, other videos corresponding to the labels of the videos of the type can be selected from the currently existing multiple videos for playing, so that the labels can be used for screening and targeted pushing.
[0037] Referring to Figure 1 FIG. 1 shows a flowchart of a label generation method provided by the embodiments of the present disclosure, and the method includes steps S101-S104, wherein:
[0038] S101: acquiring multimedia content to be processed, and extracting information to be processed in multiple content dimensions from the multimedia content;
[0039] S102: inputting the information to be processed into a first layer network node of a target feature extraction network for feature extraction to obtain first feature information, and storing the first feature information in a feature library; the target feature extraction network is a multi-layer network constructed based on the dependency relationship between multiple network nodes;
[0040] S103: using at least one Nth layer network node of the target feature extraction network to respectively extract feature information output by an (N-1) th layer network node stored in the feature library, and storing the extracted feature information in the feature library; wherein N is a positive integer greater than 1;
[0041] S104: extracting target feature information from the feature library based on the obtained feature configuration list, and determining a label production result corresponding to the multimedia content based on the target feature information.
[0042] The above S101-S104 are described in detail below.
[0043] For the above S101, the multimedia content to be processed can specifically include the video, webpage, audio, live broadcast and other different multimedia content described above. The following takes the multimedia content including video as an example for description. In a possible case, for the application environment such as short video platform, the multimedia content is constantly and massively produced, such as continuously producing short videos. These constantly produced videos can form a video stream, for example. Due to the large number of videos under the video stream, the way of adding tags to each video one by one will consume a long time, and not all videos need to be determined by the model to add tags, such as some videos that have been pre-labeled when uploading, or some special category videos such as advertising videos. Since these videos will be widely distributed, they can also not be added with tags. Therefore, for the obtained video stream, the filtering processing can be performed first to select the video to be processed from it as the multimedia content to be processed described in the embodiments of the present disclosure.
[0044] In a specific implementation, a plurality of multimedia contents can be received, and the multimedia content to be processed can be selected from the plurality of multimedia contents based on a preset filtering condition. Here, the plurality of multimedia contents are, for example, the video stream described above. The preset filtering condition can be determined in advance, such as determining the preset filtering condition according to the actual needs of the target business party. The target business party refers to a business party that uses the tags of the multimedia content and can further utilize the tags of the multimedia content to group, push and process the multimedia content. The actual needs can be, for example, not to determine the tags for the multimedia content with existing tags, or not to determine the tags for the advertising videos or the promotional videos under the platform activities. According to the actual needs listed above, the preset filtering condition can be determined, such as the preset filtering condition including filtering out the videos with tags and the advertising videos from the video stream.
[0045] In this way, for the video stream with a large cardinality, the part of the videos that specifically need to generate tags can be obtained from it by filtering, as the multimedia content to be processed, to reduce the number of videos that need to generate tags, and also to reduce the time waste and resource waste that may be caused by repeated labeling of tags and unnecessary labeling of tags.
[0046] In the case of obtaining the multimedia content to be processed, in order to generate the corresponding label for the multimedia content, the to-be-processed information that can express the characteristics of the multimedia content also needs to be obtained. Taking the case that the multimedia content includes a video as an example, the information that can be used for feature extraction can specifically include a video title, a video caption, author information, and the like of the video, among which the video title and the author information can be directly obtained, and the video caption can be obtained online or offline through optical character recognition. For the case that the multimedia content includes a webpage, the information that can be used for feature extraction can specifically include text information, image information, source information, and the like in the webpage, among which the text information is specifically recorded in the webpage, the image information can include picture annotations of pictures, and the source information can include author information and related information of a website to which the webpage belongs.
[0047] That is, for different types of multimedia content, the specific content that can reflect the characteristics of the multimedia content is different, but can be respectively extracted from a plurality of corresponding content dimensions under different types of multimedia content. For example, for the video explained in the above example, the plurality of content dimensions can be determined as the video title, the video caption, and the author information.
[0048] For example, if the multimedia content to be processed is a video v1, for example, “two cars released this year” can be extracted in the content dimension of the video title, “car A is launched in January and is a small car, and car B is launched in February and is a medium-sized car” can be extracted in the content dimension of the video caption, and “car blogger”, “one million fans”, and the like can be extracted in the content dimension of the author information. Here, the information extracted in a plurality of different content dimensions, that is, the to-be-processed information used for feature extraction.
[0049] In this way, according to the above steps, the multimedia information to be processed to generate the label can be determined, and the to-be-processed information in the screened multimedia information that is used for feature extraction to determine the label can also be screened.
[0050] For S102 and S103 described above, the to-be-processed information of the multimedia content determined in the above steps can be subjected to feature extraction through the target feature extraction network to obtain target feature information used for generating the label corresponding to the multimedia content.
[0051] Here, the target feature extraction network is specifically a multi-layer network constructed based on the dependency relationship between multiple network nodes. The network nodes can be specifically network models for feature extraction, such as deep learning models, etc. The network nodes have a dependency relationship, which can be specifically that the input data of one network node is derived from the output data of multiple network nodes, or the input data of multiple network nodes is derived from the output data of one network node. The dependency relationship between the network nodes can form a directed acyclic graph (DAG). According to the dependency relationship between the network nodes, the target feature extraction network can be divided into a multi-layer network, and the network nodes between adjacent two layers have a dependency relationship.
[0052] That is, in the embodiments of the present disclosure, the to-be-processed information extracted from the multimedia content flows in one direction through multiple network nodes, and has corresponding output data in the multiple network nodes. The downstream network nodes specifically depend on the output data of the upstream nodes as the input data of the downstream network nodes, and continue the feature extraction process.
[0053] In specific implementation, since there can be annotation requirements of multimedia content under different platforms and label production requirements of different categories or contents under the same platform in actual application, there can be multiple feature extraction networks for label generation. For a specific multimedia content, the corresponding feature extraction network can be selected for feature extraction of the to-be-processed information contained therein.
[0054] Specifically, the target feature extraction network can be determined from the preset multiple feature extraction networks based on the resource type of the multimedia content. Here, the resource type of the multimedia content can be specifically a resource type related to a business, such as a resource type corresponding to multiple different play platforms respectively; or it can also be a resource type corresponding to different genres such as video, webpage, audio, and live broadcast. When the resource types of the multimedia content are different, the focus of the multimedia content when adding labels is different, or the content dimension of the to-be-processed information used is different, so the feature extraction network used is also different.
[0055] In addition, according to the above description, under a specific resource type, the specific classification under the household, makeup, clothing, and the like can also be determined. In the embodiments of the present disclosure, these classifications or similar multi-classification methods are referred to as annotation dimensions. There are corresponding feature extraction networks under different annotation dimensions to obtain feature information for generating labels under the corresponding annotation dimensions. Therefore, in the case of determining a specific annotation dimension of the multimedia content, the feature extraction network corresponding to the determined specific annotation dimension can be determined from the respective feature extraction networks under each annotation dimension as the target feature extraction network.
[0056] After determining the target feature extraction network, the target feature extraction network can be used to extract features from the to-be-processed information of the multimedia content to obtain target feature information for generating labels. For ease of illustration, a simplified example of a target feature extraction network is provided, which includes three network layers, denoted as L1, L2, and L3, respectively. The first network layer has one network node, i.e., the first network node has 1. The second network layer and the third network layer each have two network nodes, i.e., the second network node and the third network node each have 2.
[0057] Referring to Figure 2 The target feature extraction network provided by the embodiments of the present disclosure is shown in the schematic diagram, in which the network nodes of each network layer are represented by circles and the corresponding network nodes are represented by the English letters marked therein. The dependency relationship between the network nodes is represented by the connection arrows between the circles corresponding to the network nodes. For example, the first network layer L1 includes the network node a, the first network layer L2 includes the network nodes b and c, and the third network layer L3 includes the network nodes d and e. The second network nodes b and c depend on the first network node a, and the third network nodes d and e depend on the second network node b.
[0058] For the to-be-processed information extracted from the above-mentioned to-be-processed multimedia content, the first network node of the target feature extraction network is first input, for example, the first network node a in the above-mentioned example. The feature information output by the first network node a is referred to as first feature information in the embodiments of the present disclosure. The first feature information is stored in a feature library so that the second network node depending on the first network node can call the first feature information as input data from the feature library.
[0059] Here, the feature library is, for example, a preset storage space for storing the feature information generated by the plurality of network nodes of the target feature extraction network. In order to facilitate storage, searching and reading of the feature information generated by each network node, when the feature information is stored in the feature information library, the feature information and the feature identifier corresponding to the feature information can be stored in the feature library, so as to quickly determine the feature information by using the feature identifier. When the corresponding feature identifier is determined for the feature information, the feature identifier can be associated with the network node extracting the feature information and / or the feature type of the feature information.
[0060] The above two different ways of determining the feature identifier are exemplified as follows. In one possible case, the network node can be used as the feature identifier, such as determining the feature identifier as "a" for the feature information output by the network node a. In this way, for the downstream network node, the feature information can be obtained from the feature library by using the feature identifier with the same name as the network node in the case of determining the upstream dependent network node.
[0061] In another possible case, the feature type of the feature information can also be used as the feature identifier. Specifically, in each stage of feature extraction by the network node, the feature type of the feature information generated by each network node is different, and therefore the feature identifier associated with the feature type can be selected as the feature identifier used when the feature information is stored in the feature library. In this way, for the downstream network node, the feature information of the corresponding feature type under the actual processing requirement can be obtained from the feature library according to the actual processing requirement when the feature information is processed by the network node.
[0062] The above two ways are only different storage manners of the feature information from the two angles of the dependency relationship between the network nodes and the feature type of the feature information when the network node processes data. Other labeling manners with similar storage logic are also within the protection scope of the embodiments of the present disclosure, and are not enumerated one by one here.
[0063] For other networks except the first layer network node, when the input data is obtained, the feature information output by the dependent last network node is extracted from the feature library in the embodiments of the present disclosure, and the obtained feature information is further stored in the feature library.
[0064] Continuing the above example, for the second layer network nodes b and c, which depend on the first layer network node a, the second layer network nodes b and c perform feature extraction on the feature information output by the first layer network node a from the feature library as input information, and store the respective output feature information in the feature library. For the third layer network nodes d and e, which depend on the second layer network nodes b, the third layer network nodes d and e perform feature extraction on the feature information output by the second layer network nodes from the feature library as input information, and store the respective output feature information in the feature library.
[0065] That is, in a specific implementation, at least one Nth layer network node of the target feature extraction network performs feature extraction on the feature information output by the (N-1) th layer network node stored in the feature library, and stores the extracted feature information in the feature library; where N is a positive integer greater than 1.
[0066] In this way, since for network nodes other than the first layer network node, when obtaining input data, the feature information can be directly obtained from the feature library, compared with the way of directly passing the feature information output by the previous layer network node to the adjacent layer network node having a dependency relationship in the traditional way, when there are multiple downstream network nodes depending on the same upstream network node, the need to pass the feature information output by the upstream network node to multiple downstream network nodes is avoided, and the need to generate multiple transmission channels of feature information is avoided, and the need for the upstream network node to repeatedly perform the feature extraction process multiple times to obtain multiple feature information transmitted to different downstream network nodes is avoided. That is, the way of obtaining input feature information for network nodes provided by the embodiments of the present disclosure can save the process of repeatedly performing feature extraction multiple times by some nodes in the target feature extraction network, thereby effectively saving time and improving efficiency. Since the feature library only needs to store the feature information output by each layer network node, the actual storage amount is small, and the storage space is not occupied.
[0067] In this way, through the multiple network nodes in the target feature extraction network, feature extraction of the to-be-processed information in the to-be-processed multimedia content can be completed, and the feature information output by each network node can be stored in the feature library.
[0068] In another embodiment of the present disclosure, the target feature extraction network can also be trained in the following way: obtaining sample multimedia content corresponding to the target feature extraction network; the sample multimedia content includes sample to-be-processed information and a sample label corresponding to the sample multimedia content; training the target feature extraction network to be trained using the sample multimedia content to obtain the target feature extraction network.
[0069] In order to train the target feature extraction network, sample data used for training, referred to as sample multimedia content in the present disclosure, can be obtained. In one possible case, the sample multimedia content includes sample to-be-processed information, which can have multiple content dimensions similar to the to-be-processed information in the above examples. In addition, the target feature extraction network can be supervised training, and therefore the sample multimedia content can also have corresponding labels. Therefore, the sample multimedia content used for training the target feature extraction network is obtained by splicing information in multiple different dimensions. The target feature extraction network to be trained is trained using the sample multimedia content, and the target feature extraction network actually used for extracting the to-be-processed information can be obtained.
[0070] For S104 described above, the feature information extracted from the to-be-processed information of the multimedia content can be further used to generate a label corresponding to the multimedia content. Specifically, the target feature information can be extracted from the feature library based on the obtained feature configuration list, and the label production result corresponding to the multimedia content can be determined based on the target feature information.
[0071] In a specific implementation, when the target feature information is extracted, a pre-determined feature configuration list can be determined first, and the feature configuration list includes feature identifiers corresponding to multiple feature information used to determine the label production result. In one possible case, the multiple feature information used to determine the label production result can not only come from the last network node in the target feature extraction network, but also some feature information generated in the intermediate network layer in the feature extraction process can be used to determine the label production result. Since in the present disclosure, the feature data generated by each network node is stored in the feature library, after determining which feature information needs to be used, the feature information can be obtained from the feature library according to the feature identifiers corresponding to the feature information.
[0072] In this way, the feature information can be easily obtained from the feature library through the feature configuration list. If the output feature data from the network node is directly selected, the output feature data can need to be processed by the next network node, and therefore the feature information needs to be obtained again to determine the label production result, which will waste processing time and reduce the efficiency of producing labels.
[0073] After the target feature information is extracted from the feature library, a label production result corresponding to the multimedia content can be determined based on the target feature information. Here, the label production result includes a plurality of generated labels and a confidence degree corresponding to each label. For example, the target feature extraction network can predict three labels, i.e., food, automobile, and travel. After feature extraction is performed on the to-be-processed information in the to-be-processed multimedia content, the confidence degrees corresponding to the three labels are obtained as food-40%, automobile-90%, and travel-60%, respectively. The confidence degree of any label indicates the probability that the to-be-processed multimedia content actually has the label, i.e., the probability that the to-be-processed information in the multimedia content actually contains the corresponding content under the label.
[0074] Here, the obtained label production result, i.e., the feature processing result that can be obtained by the model, cannot directly express the actual label possessed by the multimedia content, but only numerically reference the label that can be determined by the confidence degree. Therefore, in the embodiments of the present disclosure, the following manner can also be used to determine the target label matched to the multimedia content: based on the plurality of labels and the confidence degree corresponding to each label, and the label production strategy matched to the target business party, the target label matched to the multimedia content is determined.
[0075] Here, the target business party is the business party that actually uses the label, such as a software platform that plays and pushes multimedia content. The label labeling manner of the target business party to the multimedia content generally has actual requirements, such as setting a plurality of labels to make the multimedia content more easily searched by search information, or selecting a more accurate and most representative label for the multimedia content to achieve accurate classification and targeted pushing. In the above two examples, the former pays more attention to the number of labels added to the multimedia content, and the latter pays more attention to the accuracy of the labels added. Therefore, when the actual labeling requirements of the target business party to the label are different, the matched label production strategy is different.
[0076] Continuing the above example, in a possible case, if the labeling requirements of the target business party include determining two labels for the multimedia content, the matched label production strategy can include selecting the two labels with higher confidence degrees as the target labels corresponding to the multimedia content. Then, the labels can be sorted according to the confidence degrees of the labels, and the order of the labels is determined as automobile, travel, and food. Then, the two labels with higher confidence degrees, i.e., automobile and travel, are selected as the target labels.
[0077] In another possible case, if the labeling requirement of the target business party includes determining accurate labels for the multimedia content, the accuracy can be measured by whether the confidence corresponding to the set label exceeds a certain threshold, for example, if the confidence corresponding to the label exceeds 85%, it can be considered that the label can accurately label the multimedia content. Among the above three labels, the label "car" corresponding to the confidence exceeding 85% is selected as the target label.
[0078] In the embodiments of the present disclosure, the process of determining the label production result is separated from the process of determining the target label, that is, the model production process is separated from the business production process. In this way, the adjustment of the different labeling requirements of the target business party in actual application can be made without affecting the model production process. In this way, it is more flexible when different label production strategies are replaced.
[0079] In another embodiment of the present disclosure, a specific embodiment of a modular system for generating a target label corresponding to multimedia content is also provided. Referring to Figure 3 The specific embodiment of the modular system for generating a target label corresponding to multimedia content is shown in the schematic diagram. In this example, the multimedia content includes a video. The modular system actually reflects the overall process composed of multiple steps when generating a label for the multimedia content.
[0080] In the schematic diagram, the data stream is the video stream that can be obtained in the embodiments of the present disclosure. Each video in the video stream has a corresponding identification code (ID), and the specific video can be obtained according to the identification code. The video stream flows to the label flow entry (Tag Entry). The trigger is configured in the label flow entry. The trigger can obtain the video according to the identification code corresponding to each video in the video stream, and filter the video through the pre-configured filtering function, for example, filter out the advertisement video contained in the video stream, and take the filtered video as the multimedia content to be processed. For the obtained multimedia content to be processed, the trigger can also obtain the to-be-processed information in multiple content dimensions, which can be field information such as video title, author information, and pictures therein.
[0081] In addition, the trigger can also determine a target feature extraction network for extracting features of the to-be-processed information in the video from a plurality of preset feature extraction networks. In the modular system, different types of multimedia content need to be extracted, so the feature extraction network has multiple, which can be manifested as that the feature provider (Feature Provider) in the modular system includes multiple networks.
[0082] The networks included in the feature generator are generated through dependencies between multiple feature extraction models. Each feature extraction model can be used to acquire multimodal features, retrieve nearest neighbor information, and store, retrieve, and update general features and / or results. When configuring the feature generator, due to the dependencies between the multiple feature extraction models, these dependencies are configured, which is represented as network nodes in a directed acyclic graph. Furthermore, since different networks are actually selected in the feature generator when extracting features from different multimedia content, the networks are also named. For the triggers described above, the specific network used in the feature generator can be identified through its naming.
[0083] When the feature generator processes the information to be processed, the feature information generated by each node is stored in the feature bank. For downstream nodes, when processing the output data of the upstream nodes they depend on as input data, they also retrieve the corresponding feature information from the feature bank. To effectively store the feature information generated by each node, it can be named according to the corresponding network node or the feature type. This allows for easy retrieval of the relevant feature information from the feature bank based on its name. Only one feature bank can be set up, and different feature extraction networks in the feature generator can store the feature information output from their respective network nodes in the same feature bank. Furthermore, the feature bank can perform encoding processing on the feature information for standardized storage, and also support other basic functions such as version management.
[0084] Features stored in the feature library can be extracted by the Feature Assembler module. Specifically, it extracts the corresponding stored feature information from the feature library based on a pre-configured Feature List. This process can also be called feature consumption. The configuration of the Feature Assembler module can be used as a global data configuration, meaning that the data for the Fusion Pack, Fusion Train, and Fusion Predict modules described below are all under the same data configuration. Furthermore, the feature information obtained by the Feature Assembler module from the Feature List can be further passed to the Fusion Predictor module to obtain the label generation results.
[0085] The feature data packaging module mentioned here can package the information to be processed under each content dimension in the sample multimedia content to obtain multi-dimensional data for training; it concurrently calls the feature assembly module to generate the data format for training the fusion model. The fusion training module supports processing of multiple feature types and multiple feature fusion methods. These feature types can include numerical features such as label model scores, scalar features such as video duration and author / follower counts, identifier features such as author identifiers, and neighbor information features such as lists obtained based on nearest neighbor retrieval. During feature fusion, methods such as Neighbor Pooling, Deep Neural Network (DNN), and Transformer self-attention fusion can be used, and the specific method can be selected according to the actual situation, which will not be elaborated here. Furthermore, since the feature data packaging module and the fusion training module involve the training process rather than the actual application process, they are represented by dashed borders.
[0086] For the fusion prediction module, the tag generation module (Tag Map) can be triggered. The tag generation module determines the target tags (Tag Result) for the multimedia content based on the tag production strategy matched by the target business party. Based on the connection between the fusion prediction module and the tag generation module as two independent modules, it's easy to determine that the model production process and the business production process are separate; that is, the tag generation module specifically expresses the business logic rather than the feature processing logic. Here, the target tags can be obtained through the tag generation module, thus completing the tag production process for the multimedia content.
[0087] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0088] Based on the same inventive concept, this disclosure also provides a label generation device corresponding to the label generation method. Since the principle of the device in this disclosure for solving the problem is similar to that of the label generation method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0089] Reference Figure 4 The diagram shown is a schematic representation of a label generation device provided in an embodiment of this disclosure. The device includes: an acquisition module 41, a first processing module 42, a second processing module 43, and a generation module 44; wherein,
[0090] The acquisition module 41 is used to acquire the multimedia content to be processed and extract the information to be processed from the multimedia content under multiple content dimensions;
[0091] The first processing module 42 is used to input the information to be processed into the first layer network node of the target feature extraction network for feature extraction, obtain the first feature information, and store the first feature information in the feature library; the target feature extraction network is a multi-layer network constructed based on the dependency relationship between multiple network nodes;
[0092] The second processing module 43 is used to extract features from the feature information output by the (N-1)th layer network node stored in the feature library using at least one Nth layer network node of the target feature extraction network, and to store the extracted feature information in the feature library; where N is a positive integer greater than 1.
[0093] The generation module 44 is used to extract target feature information from the feature library based on the acquired feature configuration list, and determine the tag production result corresponding to the multimedia content based on the target feature information.
[0094] In an optional implementation, before inputting the information to be processed into the first layer network node of the target feature extraction network for feature extraction to obtain the first feature information, the first processing module 42 is further configured to: determine the target feature extraction network from a preset plurality of feature extraction networks based on the resource type of the multimedia content.
[0095] In one optional implementation, when the first processing module 42 determines the target feature extraction network from a set of preset feature extraction networks based on the resource type of the multimedia content, it is configured to: determine a set of annotation dimensions corresponding to the multimedia content based on the resource type of the multimedia content; and determine the target feature extraction network corresponding to each annotation dimension from the set of preset feature extraction networks based on the determined annotation dimensions.
[0096] In one optional implementation, when the first processing module 42 stores the extracted feature information in the feature library, it is used to: store the feature information and the feature identifier corresponding to the feature information in the feature library; wherein, the feature identifier corresponding to the feature information is associated with the network node from which the feature information was extracted and / or the feature type of the feature information.
[0097] In one optional implementation, the target feature extraction network is trained as follows: sample multimedia content corresponding to the target feature extraction network is obtained; the sample multimedia content includes sample information to be processed and sample labels corresponding to the sample multimedia content; the target feature extraction network to be trained is trained using the sample multimedia content to obtain the target feature extraction network.
[0098] In one optional implementation, the tag production result includes multiple generated tags and the confidence level corresponding to each tag; the device further includes a third processing module 45, used to: determine the target tag matching the multimedia content based on the multiple tags and the confidence level corresponding to each tag, and the obtained tag production strategy matching the target business party.
[0099] In one optional implementation, when acquiring multimedia content to be processed, the acquisition module 41 is configured to: receive multiple multimedia contents and select the multimedia content to be processed from the multiple multimedia contents based on preset filtering conditions.
[0100] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0101] This disclosure also provides a computer device, such as... Figure 5 The diagram shown is a schematic representation of a computer device structure provided in an embodiment of this disclosure, including:
[0102] Processor 10 and memory 20; the memory 20 stores machine-readable instructions executable by processor 10, and processor 10 executes the machine-readable instructions stored in memory 20. When the machine-readable instructions are executed by processor 10, processor 10 performs the following steps:
[0103] The process involves: acquiring multimedia content to be processed and extracting information from the multimedia content across multiple content dimensions; inputting the information to be processed into the first layer network node of the target feature extraction network for feature extraction to obtain first feature information, and storing the first feature information in a feature library; the target feature extraction network is a multi-layer network constructed based on the dependency relationships between multiple network nodes; using at least one Nth layer network node of the target feature extraction network, performing feature extraction on the feature information output by the (N-1)th layer network node stored in the feature library, and storing the extracted feature information in the feature library; where N is a positive integer greater than 1; extracting target feature information from the feature library based on the acquired feature configuration list, and determining the tag production result corresponding to the multimedia content based on the target feature information.
[0104] The aforementioned memory 20 includes a main memory 210 and an external memory 220; the main memory 210, also known as internal memory, is used to temporarily store the computational data in the processor 10, as well as the data exchanged with external memory 220 such as a hard disk. The processor 10 exchanges data with the external memory 220 through the main memory 210.
[0105] The specific execution process of the above instructions can be referred to the steps of the tag generation method described in the embodiments of this disclosure, and will not be repeated here.
[0106] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the tag generation method described in the above method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0107] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the tag generation method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0108] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0112] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A label generation method, characterized in that, include: Acquire multimedia content to be processed, and extract information to be processed from the multimedia content across multiple content dimensions; The information to be processed is input into the first layer network node of the target feature extraction network for feature extraction to obtain the first feature information, and the first feature information is stored in the feature library; the target feature extraction network is a multi-layer network constructed based on the dependency relationship between multiple network nodes; Using at least one Nth layer network node of the target feature extraction network, feature extraction is performed on the feature information output by the (N-1)th layer network node stored in the feature library, and the extracted feature information is stored in the feature library; where N is a positive integer greater than 1; Based on the acquired feature configuration list, target feature information is extracted from the feature library, and the tag production result corresponding to the multimedia content is determined based on the target feature information.
2. The method according to claim 1, characterized in that, Before inputting the information to be processed into the first layer network node of the target feature extraction network for feature extraction to obtain the first feature information, the method further includes: Based on the resource type of the multimedia content, the target feature extraction network is determined from a set of multiple preset feature extraction networks.
3. The method according to claim 2, characterized in that, Based on the resource type of the multimedia content, the target feature extraction network is determined from a set of preset feature extraction networks, including: Based on the resource type of the multimedia content, determine multiple annotation dimensions corresponding to the multimedia content; Based on the determined multiple annotation dimensions, the target feature extraction network corresponding to each of the preset multiple feature extraction networks is determined respectively.
4. The method according to claim 1, characterized in that, The extracted feature information is stored in the feature library, including: The feature information and the feature identifier corresponding to the feature information are stored in the feature library; wherein, the feature identifier corresponding to the feature information is associated with the network node from which the feature information is extracted and / or the feature type of the feature information.
5. The method according to claim 1, characterized in that, The target feature extraction network is trained in the following manner: Obtain the sample multimedia content corresponding to the target feature extraction network; the sample multimedia content includes sample information to be processed and sample labels corresponding to the sample multimedia content. The target feature extraction network is trained using the multimedia content of the sample to obtain the target feature extraction network.
6. The method according to claim 1, characterized in that, The label production result includes multiple generated labels and the confidence score corresponding to each label; the method further includes: Based on multiple tags and the confidence level of each tag, and the tag production strategy obtained that matches the target business party, target tags that match the multimedia content are determined.
7. The method according to claim 1, characterized in that, Acquire the multimedia content to be processed, including: Receive multiple multimedia content items, and select the multimedia content to be processed from the multiple multimedia content items based on preset filtering conditions.
8. A label generating apparatus, characterized in that, include: The acquisition module is used to acquire multimedia content to be processed and extract information to be processed from the multimedia content under multiple content dimensions. The first processing module is used to input the information to be processed into the first layer network node of the target feature extraction network for feature extraction, obtain the first feature information, and store the first feature information in the feature library; the target feature extraction network is a multi-layer network constructed based on the dependency relationship between multiple network nodes; The second processing module is used to extract features from the feature information output by the (N-1)th layer network node stored in the feature library using at least one Nth layer network node of the target feature extraction network, and to store the extracted feature information in the feature library; where N is a positive integer greater than 1. The generation module is used to extract target feature information from the feature library based on the acquired feature configuration list, and determine the tag production result corresponding to the multimedia content based on the target feature information.
9. A computer device, characterized in that, include: The processor and the memory, the memory storing machine-readable instructions executable by the processor, the processor executing the machine-readable instructions stored in the memory, wherein when the machine-readable instructions are executed by the processor, the processor performs the steps of the tag generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the tag generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Feature tag generation method, device, electronic equipment and storage medium
CN113836146A
Label determination method and device for multimedia content
CN115187837A