A method for generating a plot tag of media content and an electronic device
By conducting multi-dimensional unsupervised comparative adversarial training on the plot tag generation network, the problem of semantic coherence in media asset tag generation in existing technologies is solved, thereby improving the accuracy and efficiency of tag generation.
Patent Information
- Application Number
- CN202411188318.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-08-28
Smart Images

Figure CN119583910B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a plot label generation method for media assets and an electronic device. BACKGROUND
[0002] As an important tool for media asset management and content retrieval, the media label system needs to cope with diversified content classification requirements as media content grows rapidly and user search needs diversify. For most media content, manual tagging is still the mainstream way, which has problems such as low efficiency, high cost, and low standardization.
[0003] With the continuous development of artificial intelligence technology, intelligent label generation technology has begun to be applied in the field of media management. These technologies automatically tag media content through image recognition, speech recognition, and natural language processing, greatly improving the efficiency and accuracy of label generation.
[0004] In the prior art, the method of generating media labels through artificial intelligence mostly adopts entity recognition or TF-IDF (Term Frequency_Inverse Document Frequency) keyword extraction method. These methods extract key information from the original text as labels, lack understanding of the global semantics of the text, and the extracted labels lack semantic coherence, resulting in low accuracy of the generated labels. SUMMARY
[0005] The present application provides a plot label generation method for media assets and an electronic device, which trains a plot label generation network through multi-modal information and using unsupervised contrastive learning adversarial training, so that the plot labels obtained through the plot label generation network have better semantic coherence, improving the accuracy of the generated plot labels, and saving the cost of manual annotation and improving the efficiency.
[0006] In a first aspect, an embodiment of the present application provides a plot label generation method for media assets, comprising:
[0007] For any first media asset, a first media asset introduction text corresponding to the first media asset is input into a pre-trained plot label generation network to obtain each plot label corresponding to the first media asset, wherein the first media asset introduction text is used to introduce the media content of the first media asset.
[0008] The pre-trained plot label generation network is obtained by the following method:
[0009] input a second media asset introduction text corresponding to a second media asset into the plot tag generation network to obtain each plot tag corresponding to the second media asset and a text encoding vector corresponding to the second media asset introduction text, wherein the text encoding vector is obtained by text encoding the second media asset introduction text; and
[0010] add noise to the cover image corresponding to the second media asset by using the negative example generation network to obtain a plurality of negative sample vectors;
[0011] perform adversarial training on the plot tag generation network and the negative example generation network by using each plot tag corresponding to the second media asset, the text encoding vector, and the plurality of negative sample vectors to obtain the pre-trained plot tag generation network.
[0012] The second aspect of the present application provides an electronic device, comprising a processor and a memory, the processor and the memory are connected through a bus;
[0013] The memory stores a computer program, and the processor is configured to execute the following operations based on the computer program:
[0014] For any one first media asset, input a first media asset introduction text corresponding to the first media asset into the pre-trained plot tag generation network to obtain each plot tag corresponding to the first media asset, wherein the first media asset introduction text is used to introduce the media content of the first media asset;
[0015] The pre-trained plot tag generation network is obtained by the following method:
[0016] input a second media asset introduction text corresponding to a second media asset into the plot tag generation network to obtain each plot tag corresponding to the second media asset and a text encoding vector corresponding to the second media asset introduction text, wherein the text encoding vector is obtained by text encoding the second media asset introduction text; and
[0017] add noise to the cover image corresponding to the second media asset by using the negative example generation network to obtain a plurality of negative sample vectors;
[0018] perform adversarial training on the plot tag generation network and the negative example generation network by using each plot tag corresponding to the second media asset, the text encoding vector, and the plurality of negative sample vectors to obtain the pre-trained plot tag generation network.
[0019] According to the third aspect provided by the embodiments of the present application, a computer storage medium is provided, which stores a computer program for executing the method according to the first aspect.
[0020] In the above embodiments of the present application, the plot label of the media asset is determined by the plot label generation network. Since the plot label generation network is trained by using multi-dimensional media asset information and adopting unsupervised contrast learning adversarial training, the cost of manual labeling is saved, and the trained plot label generation network can understand the media asset from multiple dimensions, rather than determining the label only by keywords, so that the obtained plot label has more semantic coherence, and the accuracy of the generated plot label is improved. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] Figure 1 An application scenario provided by an embodiment of the present application is exemplarily shown;
[0023] Figure 2 One of the flowcharts of a plot label generation network training method provided by an embodiment of the present application is exemplarily shown;
[0024] Figure 3 A structure diagram of the plot label generation network provided by an embodiment of the present application is exemplarily shown;
[0025] Figure 4 A flowchart for determining a plurality of negative sample vectors provided by an embodiment of the present application is exemplarily shown;
[0026] Figure 5 A flowchart for training the plot label generation network provided by an embodiment of the present application is exemplarily shown;
[0027] Figure 6 A flowchart for determining a current model loss value provided by an embodiment of the present application is exemplarily shown;
[0028] Figure 7 A structure diagram of the overall model corresponding to the model training provided by an embodiment of the present application is exemplarily shown;
[0029] Figure 8 Another flowchart of the plot label generation network training method provided by an embodiment of the present application is exemplarily shown;
[0030] Figure 9 A schematic diagram of a plot label generation network training device provided by an embodiment of the present application is exemplarily shown;
[0031] Figure 10 An exemplary structure diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0032] For the purpose of making the purpose, implementation and advantages of the present application more clear, the exemplary implementation of the present application will be described clearly and completely in the following with reference to the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, but not all the embodiments.
[0033] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those skilled in the art without making creative efforts fall within the scope of the claims of the present application. In addition, although the disclosure in the present application is introduced according to one or several examples, it should be understood that each aspect of the disclosure can also constitute a complete embodiment independently.
[0034] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described embodiments, but is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0035] The terms "first", "second", and the like in the specification of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to the clearly listed components, but can include other components not clearly listed or inherent to these products or devices.
[0036] The term "module" used in the present application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware or / and software code capable of performing functions associated with the element.
[0037] The idea of the embodiments of the present application is summarized as follows.
[0038] At present, the method of generating media tags by artificial intelligence mostly adopts entity recognition or TF-IDF keyword extraction method. These methods extract the key information from the original text as tags, lack the understanding of the global semantics of the text, and the extracted tags lack semantic coherence, resulting in low accuracy of the generated tags.
[0039] Existing technologies often suffer from a lack of semantic coherence in extracted tags, leading to low accuracy in generated tags. This application provides a method for generating narrative tags for media assets. This method uses a narrative tag generation network to determine the narrative tags for media assets. Because the narrative tag generation network is trained using multi-dimensional media asset information and an unsupervised contrastive learning adversarial training method, the cost of manual annotation is eliminated. Furthermore, the trained narrative tag generation network can understand media assets from multiple dimensions, not just relying on keywords to determine tags, thus resulting in more semantically coherent narrative tags and improving the accuracy of the generated narrative tags.
[0040] like Figure 1 The diagram illustrates an application scenario for a method of generating plot tags for media assets. This scenario includes a server 110 and a terminal device 120. The application scenario is illustrated using an electronic device as the server.
[0041] In one possible application scenario, for any first media asset, server 110 inputs the first media asset description text corresponding to the first media asset into a pre-trained plot tag generation network to obtain each plot tag corresponding to the first media asset, and sends each plot tag corresponding to the first media asset to terminal device 120 for display.
[0042] The server 110 obtains the pre-trained plot tag generation network in the following ways: inputting the second media asset's synopsis text into the plot tag generation network to obtain plot tags corresponding to the second media asset and text encoding vectors corresponding to the second media asset's synopsis text, wherein the text encoding vectors are obtained by text encoding the second media asset's synopsis text; adding noise to the cover image corresponding to the second media asset using a negative example generation network to obtain multiple negative sample vectors; and using the plot tags corresponding to the second media asset, the text encoding vectors, and the multiple negative sample vectors to perform adversarial training on the plot tag generation network and the negative example generation network to obtain the pre-trained plot tag generation network.
[0043] in, Figure 1 The server 110 and the terminal device 120 can exchange information through a communication network. The communication network can be either wireless or wired.
[0044] Exemplarily, the terminal device 120 can access the network and communicate with the server 110 through a cellular mobile communication technology, such as a 5th Generation Mobile Networks (5G) technology.
[0045] Optionally, the terminal device 120 can access the network and communicate with the server 110 through a short-distance wireless communication mode, such as a Wireless Fidelity (Wi-Fi) technology.
[0046] In the description in the present application, only a single server 110 and a single terminal device 120 are described in detail, but those skilled in the art should understand that the server 110 and the terminal device 120 shown are intended to represent the operation of the server 110 and the terminal device 120 involved in the technical solution of the present application. It is not implied that there is a limitation on the number, type or location of the server 110 and the terminal device 120. It should be noted that if additional modules are added to the illustrated environment or individual modules are removed therefrom, the underlying concept of the example embodiments of the present application will not change.
[0047] It should be noted that the plot tag generation method of media content proposed in the present application is not only applicable to Figure 1 the application scenarios shown, but also applicable to any plot tag generation device of media content.
[0048] The plot tag generation method of media content of the example embodiments of the present application will be described below in combination with the above-described application scenarios and with reference to the accompanying drawings. It should be noted that the above-described application scenarios are only shown for the purpose of facilitating the understanding of the method and principle of the present application, and the embodiments of the present application are not limited in this respect.
[0049] Before introducing the plot tag generation method of media content in the present application, first introduce the training method of the plot tag generation network in the present application. As Figure 2 shown, it is a flowchart of the training method of the plot tag generation network, which can specifically include the following steps:
[0050] Step 201: input the second media introduction text corresponding to the second media content into the plot tag generation network to obtain each plot tag corresponding to the second media content and a text encoding vector corresponding to the second media introduction text, wherein the text encoding vector is obtained by text encoding the second media introduction text.
[0051] As Figure 3 shown, it is a structural diagram of the plot tag generation network in the embodiments of the present application, from Figure 3As can be seen, the plot tag generation network 300 includes a text encoder 301 and a text decoder 302. The following is a detailed description of step 201 in the embodiments of this application, in conjunction with the structure of the plot tag generation network:
[0052] The text encoder 301 is used to encode the second media asset description text to obtain the text encoding vector; and the text decoder 302 is used to decode the text encoding vector to obtain the plot tags corresponding to the second media asset.
[0053] The text decoder in this embodiment uses a beam search strategy to decode the text encoding vector. However, this embodiment does not limit the text decoder and can be configured according to specific circumstances.
[0054] In this embodiment, there are multiple second media assets. Any second media asset is input into the plot tag generation network for training in the same way. This embodiment does not limit the number of second media assets. The number of second media assets in this embodiment can be set according to the specific actual situation.
[0055] Step 202: Use a negative sample generation network to add noise to the cover image corresponding to the second media asset to obtain multiple negative sample vectors;
[0056] It should be noted that the execution order of steps 201 and 202 in this embodiment is not limited. Step 201 can be executed first, followed by step 202. Alternatively, step 202 can be executed first, followed by step 201. Steps 201 and 202 can also be executed simultaneously.
[0057] The following describes the method for determining multiple negative sample vectors in step 202, such as... Figure 4 The diagram shown illustrates the process for determining multiple negative sample vectors, which may include the following steps:
[0058] Step 401: Input the cover image into a preset image encoder to obtain an image encoding vector; and randomly sample multiple noise vectors from a preset Gaussian probability distribution function;
[0059] Step 402: Add the image encoding vector to the plurality of noise vectors respectively to obtain a plurality of initial negative sample vectors;
[0060] In this embodiment, the number of initial negative samples is the same as the number of noise vectors.
[0061] Step 403: Input the plurality of initial negative sample vectors into the negative sample generation network respectively to obtain the plurality of negative sample vectors, wherein the number of the initial negative sample vectors is the same as the number of negative samples.
[0062] The negative example generation network in this application embodiment can be a fully connected network, an RNN (Recurrent Neural Network) network, or a CNN (Convolutional Neural Network), etc. This application embodiment does not limit the structure of the negative example generation network. The specific structure of the negative example generation network in this application embodiment can be set according to the actual situation.
[0063] Step 203: Using the plot tags corresponding to the second media asset, the text encoding vector, and the multiple negative sample vectors, perform adversarial training on the plot tag generation network and the negative sample generation network to obtain the pre-trained plot tag generation network.
[0064] The network training method in step 203 of this application embodiment will be described in detail below, such as... Figure 5 The diagram shown illustrates the process of training a narrative tag generation network, which may include the following steps:
[0065] Step 501: Based on the plot tag vector corresponding to each plot tag, the text encoding vector, and the multiple negative sample vectors, obtain the current model loss value, wherein the plot tag vector corresponding to any plot tag is obtained by inputting the plot tag into a preset text encoder;
[0066] The following describes the method for determining the current model loss value in step 501, such as... Figure 6 The diagram shown illustrates the process for determining the current model loss value, which may include the following steps:
[0067] Step 601: For any plot tag among the plot tags corresponding to the second media asset, based on the plot tag vector corresponding to the plot tag, the text encoding vector, and the multiple negative sample vectors, obtain the sub-loss value corresponding to the plot tag; wherein, the sub-loss value corresponding to the plot tag is obtained through formula (1):
[0068]
[0069] Where loss_t is the sub-loss value corresponding to the plot label t. The text encoding vector, Let t be the plot tag vector corresponding to the plot tag t, and τ be a preset hyperparameter. is the rth negative sample vector, k is the total number of the plurality of negative sample vectors, is an image encoding vector, and the image encoding vector is obtained by inputting a cover image corresponding to the second media asset into an image encoder.
[0070] Step 602: obtaining a first loss value according to the respective sub-loss values of the respective plot labels.
[0071] In an embodiment, step 602 can be specifically implemented as: adding the respective sub-loss values of the respective plot labels to obtain the first loss value.
[0072] Step 603: obtaining a second loss value according to the respective plot label vectors of the respective plot labels and the total number of the respective plot labels; wherein the second loss value is obtained by formula (2):
[0073]
[0074] wherein loss_b is the second loss value, M is the total number of the respective plot labels, is a plot label vector corresponding to an i th plot label in the respective plot labels, is a plot label vector corresponding to a j th plot label in the respective plot labels.
[0075] Step 604: obtaining a current model loss value according to the first loss value and the second loss value.
[0076] In an embodiment, step 604 can be specifically implemented as: multiplying the first loss value by a length penalty factor, and then dividing by the second loss value to obtain the current model loss value, wherein the length penalty factor is obtained based on the respective lengths of the respective plot labels. Wherein the current model loss value can be obtained by formula (3):
[0077]
[0078] wherein LOSS is the current model loss value, L is the length penalty factor, and loss_a is the first loss value.
[0079] Next, the determination manner of the length penalty factor in the embodiments of the present application is described. In an embodiment, the length penalty factor is obtained by the following manner:
[0080] obtaining an average length of the respective plot labels according to the respective lengths of the respective plot labels and the total number of the respective plot labels; dividing the average length by a preset maximum length to obtain the length penalty factor. Wherein the length penalty factor can be obtained by formula (4):
[0081]
[0082] wherein, L is the length penalty factor, len i is the length of the i-th plot tag, M is the total number of the plot tags, and N is a preset maximum length.
[0083] It should be noted that the maximum length in the embodiments of the present application is 5, but the specific value of the maximum length is not limited, and the specific value of the maximum length in the embodiments of the present application can be set according to actual conditions.
[0084] Step 502: determining whether the difference between the current model loss value and the last model loss value is not greater than a first specified threshold, and the current model loss value is not greater than the last model loss value; if not, executing step 503, and if yes, executing step 504;
[0085] It should be noted that the first specified threshold in the embodiments of the present application is not limited herein, and the first specified threshold in the embodiments of the present application can be set according to specific actual conditions.
[0086] Step 503: after adjusting the first model parameter in the plot tag generation network, returning to execute the step of inputting the second media asset corresponding to the second media asset introduction text into the plot tag generation network;
[0087] The model parameter in the plot tag generation network is referred to as the first model parameter in the embodiments of the present application, and the first model parameter can be set according to specific actual conditions, and the first model parameter is not limited herein in the embodiments of the present application.
[0088] Step 504: ending the training of the plot tag generation network;
[0089] In the embodiments of the present application, the model training is first frozen for the negative example generation network, and then the model loss value is used to optimize the plot tag generation network, that is, to minimize the semantic similarity between the plot tag vector corresponding to each plot tag, the text encoding vector and the image encoding vector. The similarity between the plot tag vector corresponding to each plot tag and the noise vector is maximized. The generated plot tag is more similar in semantics to the media asset introduction text and the cover image, and is more distant in semantics from the negative sample. The accuracy of the model is improved, so that the generated plot tag is more accurate in actual use. At the same time, the generated plot tag cannot be too long, and the plot tag needs to have diversity.
[0090] Step 505: determining whether the current model loss value and the third loss value satisfy a specified condition, if not, executing step 506, and if yes, executing step 507;
[0091] The specified condition in the embodiments of the present application includes the following two conditions:
[0092] Condition 1: the current model loss value is greater than the last model loss value, and the difference between the current model loss value and the last model loss value is not greater than the first specified threshold value;
[0093] Condition 2: the third loss value is greater than the last obtained third loss value, and the difference between the third loss value and the last obtained third loss value is not greater than the second specified threshold value.
[0094] It should be noted that the first specified threshold value and the second specified threshold value in the embodiments of the present application can be the same or different, and the embodiments of the present application do not limit the specific values of the first specified threshold value and the second specified threshold value.
[0095] In one embodiment, the third loss value can be obtained by formula (5):
[0096]
[0097] wherein loss c is the third loss value, is the pth negative sample vector in the plurality of negative sample vectors, is the qth negative sample vector in the plurality of negative sample vectors, and k is the total number of the plurality of negative sample vectors.
[0098] Step 506: adjusting the second model parameter in the negative example generation network, and returning to the step of inputting the second media content corresponding to the second media content introduction text into the plot tag generation network;
[0099] The model parameter in the negative example generation network is referred to as the second model parameter, but the second model parameter is not limited, and the second model parameter in the embodiments of the present application can be set according to specific actual conditions.
[0100] Therefore, in the embodiments of the present application, after the plot tag generation network is trained each time, the plot generation network is frozen, and then the negative example generation network is optimized to maximize the current model loss value and the third loss value, which will generate more difficult negative samples to make the generated negative samples closer to the plot tags, thereby forcing the plot tag generation network to generate plot tags that are more similar to the media content introduction text of the media content when optimizing the plot tag generation network. Further improve the accuracy of the generated plot tags.
[0101] Step 507: determining the trained plot tag generation network as the pre-trained plot tag generation network.
[0102] As Figure 7As shown, a structural diagram of the overall model corresponding to the model training is shown, and the mode of the model training is introduced in the following in combination with the structural diagram. From Figure 7 It can be seen that the overall model includes a first text encoder 701, a text decoder 702, an image encoder 703, a second text encoder 704, and a negative example generation network 705. Among them:
[0103] The second media asset introduction text corresponding to the second media asset is input into the plot tag generation network, the first text encoder 701 in the plot tag generation network is used to encode the second media asset introduction text, and a text encoding vector is obtained. Then the text decoder 702 is used to decode the text encoding vector, and each plot tag corresponding to the second media asset is obtained. Then the second text encoder 704 is used to encode each plot tag corresponding to the second media asset, and a plot tag vector corresponding to each plot tag is obtained.
[0104] The cover image of the second media asset is input into the image encoder 703 to obtain an image encoding vector; and a plurality of noise vectors are randomly sampled in a preset Gaussian probability distribution function; and the plurality of noise vectors and the image encoding vector are input into the negative example generation network 705 to obtain a plurality of negative sample vectors. Finally, the plot tag generation network and the negative example generation network are adversarially trained using the plot tag vector corresponding to each plot tag of the second media asset, the text encoding vector, the image encoding vector, and the plurality of negative sample vectors, to obtain the pre-trained plot tag generation network.
[0105] After introducing the model training method in the embodiment of the application, the plot tag generation method of the media asset in the embodiment of the application is introduced. In one embodiment, for any first media asset, the first media asset introduction text corresponding to the first media asset is input into the pre-trained plot tag generation network to obtain each plot tag corresponding to the first media asset. The first media asset introduction text is used to introduce the media content of the first media asset.
[0106] In the present application, the plot tag generation network is used to determine the plot tag of the media asset. Since the plot tag generation network is trained by using multi-dimensional media asset information and adopting unsupervised contrastive learning adversarial training method, the cost of manual annotation is saved. Moreover, the trained plot tag generation network can understand the media asset from multiple dimensions, and the plot tag is not determined only by keywords, so that the obtained plot tag has more semantic coherence, and the accuracy of the generated plot tag is improved.
[0107] In order to further understand the plot tag generation method of the media asset in the embodiment of the application, as shown in Figure 8As shown, a flowchart of a method for generating a plot tag of media content is shown, which can include the following steps:
[0108] Step 801: input a second media content introduction text corresponding to a second media content into a plot tag generation network to obtain each plot tag corresponding to the second media content and a text encoding vector corresponding to the second media content introduction text, wherein the text encoding vector is obtained by text encoding the second media content introduction text;
[0109] Step 802: input the cover image into a preset image encoder to obtain an image encoding vector, and randomly sample a plurality of noise vectors in a preset Gaussian probability distribution function;
[0110] Step 803: add the image encoding vector to each of the plurality of noise vectors to obtain a plurality of initial negative sample vectors;
[0111] Step 804: input each of the plurality of initial negative sample vectors into the negative example generation network to obtain a plurality of negative sample vectors, wherein the number of initial negative sample vectors is the same as the number of negative samples;
[0112] Step 805: perform adversarial training on the plot tag generation network and the negative example generation network using each plot tag corresponding to the second media content, the text encoding vector, the image encoding vector, and the plurality of negative sample vectors to obtain the pre-trained plot tag generation network;
[0113] Step 806: for any first media content, input a first media content introduction text corresponding to the first media content into the pre-trained plot tag generation network to obtain each plot tag corresponding to the first media content, wherein the first media content introduction text is used to introduce the media content of the first media content.
[0114] Based on the same inventive concept, the plot tag generation method of media content as described above can also be implemented by a plot tag generation device of media content. The plot tag generation device of media content has similar effects to the plot tag generation method described above, and will not be described here.
[0115] Figure 9 A structural diagram of a plot tag generation device of media content according to an embodiment of the present disclosure is shown.
[0116] As Figure 9 shown, the plot tag generation device 900 of media content of the present disclosure can include a plot tag determination module 910 and a model training module 920.
[0117] The plot label determination module 910 is configured to input a first media introduction text corresponding to any first media into a pre-trained plot label generation network to obtain each plot label corresponding to the first media, the first media introduction text being used to introduce media content of the first media.
[0118] The model training module 920 is configured to obtain the pre-trained plot label generation network in the following manner:
[0119] input a second media introduction text corresponding to a second media into the plot label generation network to obtain each plot label corresponding to the second media and a text encoding vector corresponding to the second media introduction text, wherein the text encoding vector is obtained by text encoding the second media introduction text; and
[0120] add noise to a cover image corresponding to the second media by using the negative example generation network to obtain a plurality of negative sample vectors;
[0121] perform adversarial training on the plot label generation network and the negative example generation network by using each plot label corresponding to the second media, the text encoding vector, and the plurality of negative sample vectors to obtain the pre-trained plot label generation network.
[0122] In an embodiment, the plot label generation network comprises a text encoder and a text decoder.
[0123] The model training module 920 performs the inputting of the second media introduction text corresponding to the second media into the plot label generation network to obtain each plot label corresponding to the second media and the text encoding vector corresponding to the second media introduction text, and is specifically configured to:
[0124] encode the second media introduction text by using the text encoder to obtain the text encoding vector;
[0125] decode the text encoding vector by using the text decoder to obtain each plot label corresponding to the second media.
[0126] In an embodiment, the model training module 920 performs the adding of noise to the cover image corresponding to the second media by using the negative example generation network to obtain the plurality of negative sample vectors, and is specifically configured to:
[0127] input the cover image into a preset image encoder to obtain an image encoding vector, and randomly sample a plurality of noise vectors in a preset Gaussian probability distribution function;
[0128] add the image encoding vector to each of the plurality of noise vectors to obtain a plurality of initial negative sample vectors;
[0129] input the plurality of initial negative sample vectors into the negative example generation network respectively to obtain the plurality of negative sample vectors, wherein the number of the initial negative sample vectors is the same as the number of the negative samples.
[0130] In one embodiment, the model training module 920 performs the adversarial training of the plot tag generation network and the negative example generation network by using the respective plot tags corresponding to the second media asset, the text encoding vector, and the plurality of negative sample vectors, to obtain the pre-trained plot tag generation network, and specifically for:
[0131] obtaining a current model loss value based on the respective plot tag vectors corresponding to the respective plot tags, the text encoding vector, and the plurality of negative sample vectors, wherein the plot tag vector corresponding to any plot tag is obtained by inputting the any plot tag into a preset text encoder;
[0132] If the difference between the current model loss value and the last model loss value is greater than a first specified threshold, or the current model loss value is greater than the last model loss value, then adjusting the first model parameters in the plot tag generation network, and returning to the step of inputting the second media asset introduction text corresponding to the second media asset into the plot tag generation network until the current model loss value is not greater than the last model loss value, and the difference between the current model loss value and the last model loss value is not greater than the first specified threshold, then ending the training of the plot tag generation network; and,
[0133] adjusting the second model parameters in the negative example generation network, and returning to the step of inputting the second media asset introduction text corresponding to the second media asset into the plot tag generation network until the current model loss value and the third loss value satisfy a specified condition, then ending the training of the negative example generation network, wherein the third loss value is obtained based on the plurality of negative sample vectors; and,
[0134] increasing the number of model training by a specified value, and determining whether the current number of model training is greater than a specified number, if not, returning to the step of obtaining the current model loss value based on the respective plot tag vectors corresponding to the respective plot tags, the text encoding vector, and the plurality of negative sample vectors until the current number of model training is greater than the specified number, then determining the trained plot tag generation network as the pre-trained plot tag generation network;
[0135] If the difference between the current model loss value and the last model loss value is not greater than the first specified threshold, and the current model loss value is not greater than the last model loss value, the step of adjusting the second model parameter in the negative example generation network is performed.
[0136] In one embodiment, the model training module 920 performs the obtaining of the current model loss value based on the plot tag vectors corresponding to the respective plot tags, the text encoding vector, and the plurality of negative sample vectors, specifically for:
[0137] For any one of the respective plot tags corresponding to the second media asset, a sub-loss value corresponding to the plot tag is obtained based on the plot tag vector corresponding to the plot tag, the text encoding vector, and the plurality of negative sample vectors;
[0138] A first loss value is obtained according to the respective sub-loss values of the respective plot tags;
[0139] A second loss value is obtained according to the respective plot tag vectors of the respective plot tags and the total number of the respective plot tags;
[0140] The current model loss value is obtained according to the first loss value and the second loss value.
[0141] In one embodiment, the model training module 920 performs the obtaining of the sub-loss value corresponding to the plot tag based on the plot tag vector corresponding to the plot tag, the text encoding vector, and the plurality of negative sample vectors, specifically for:
[0142] The sub-loss value corresponding to the plot tag is obtained by the following formula:
[0143]
[0144] Wherein, loss t is the sub-loss value corresponding to the plot tag t, is the text encoding vector, is the plot tag vector corresponding to the plot tag t, τ is a preset hyperparameter, is the rth negative sample vector, k is the total number of the plurality of negative sample vectors, is the image encoding vector, and the image encoding vector is obtained by inputting the cover image corresponding to the second media asset into an image encoder;
[0145] The model training module 920 performs the obtaining of the first loss value according to the respective sub-loss values of the respective plot tags, specifically for:
[0146] The respective sub-loss values of the respective plot tags are added to obtain the first loss value.
[0147] The model training module 920 performs the obtaining of the second loss value according to the respective plot label vectors of the respective plot labels and the total number of the respective plot labels, and specifically for:
[0148] The second loss value is obtained by the following formula:
[0149]
[0150] Wherein, loss_b is the second loss value, M is the total number of the respective plot labels, is the plot label vector corresponding to the i-th plot label in the respective plot labels, is the plot label vector corresponding to the j-th plot label in the respective plot labels;
[0151] The model training module 920 performs the obtaining of the current model loss value according to the first loss value and the second loss value, and specifically for:
[0152] The first loss value is multiplied by a length penalty factor and then divided by the second loss value to obtain the current model loss value, wherein the length penalty factor is obtained based on the respective lengths of the respective plot labels.
[0153] In one embodiment, the device further comprises:
[0154] The length penalty factor determination module 930 is configured to obtain the length penalty factor by the following manner:
[0155] Obtaining the average length of the respective plot labels according to the respective lengths of the respective plot labels and the total number of the respective plot labels;
[0156] Dividing the average length by a preset maximum length to obtain the length penalty factor.
[0157] In one embodiment, the device further comprises:
[0158] The loss value determination module 940 is configured to obtain the third loss value by the following manner:
[0159]
[0160] Wherein, loss_c is the third loss value, is the p-th negative sample vector in the plurality of negative sample vectors, is the q-th negative sample vector in the plurality of negative sample vectors, and k is the total number of the plurality of negative sample vectors.
[0161] In one embodiment, the specified condition comprises the following two:
[0162] the current model loss value is greater than the last model loss value, and a difference between the current model loss value and the last model loss value is not greater than the first specified threshold value;
[0163] the third loss value is greater than the last obtained third loss value, and a difference between the third loss value and the last obtained third loss value is not greater than the second specified threshold value.
[0164] After introducing the media asset plot tag generation method and device of one embodiment of the present application, next, the electronic device according to another exemplary embodiment of the present application is introduced.
[0165] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method or a program product. Therefore, various aspects of the present application can be embodied as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.
[0166] In some possible embodiments, the electronic device according to the present application can at least include at least one processor and at least one computer storage medium. The computer storage medium stores program codes, which, when executed by the processor, cause the processor to perform the steps in the media asset plot tag generation method according to various exemplary embodiments of the present application described above in the specification. For example, the processor can perform steps 801-806 as shown in Figure 8
[0167] The electronic device 1000 according to this embodiment of the present application will be described below with reference to Figure 10 Figure 10 The display electronic device 1000 is only an example and should not bring any limitation to the function and use range of the embodiments of the present application.
[0168] As shown in Figure 10 The components of the electronic device 1000 can include, but are not limited to, the above-mentioned at least one processor 1001, the above-mentioned at least one computer storage medium 1002, and a bus 1003 connecting different system components including the computer storage medium 1002 and the processor 1001.
[0169] The bus 1003 represents one or more of several types of bus structures, including a computer storage medium bus or a computer storage medium controller, a peripheral bus, a processor bus, or a local bus using any of the bus structures.
[0170] Computer storage media 1002 can include readable media in the form of volatile computer storage media (e.g., random access computer storage media (RAM) 1021 and / or cache memory 1022), and / or non-volatile computer storage media (e.g., read only memory (ROM) 1023).
[0171] Computer storage media 1002 can further include a program / utility 1025 having a set of program modules 1024 such as an operating system, one or more application programs, other program modules, and program data, each of which
[0172] The electronic device 1000 can also communicate with one or more external devices 1004 such as a keyboard or pointing device, using an input / output (I / O) interface 1005. Further, the electronic device 1000 can communicate with one or more devices using an output interface 1007. In general, the input and output device interfaces 1005 and 1007 can enable the electronic device 1000 to communicate with one or more devices using electrical, electromagnetic, or optical
[0173] In some possible embodiments, various aspects of the method for generating a plot tag of a media asset provided by the present application can also be implemented as a program product, which includes program codes for causing a computer device to execute the steps of the method for generating a plot tag of a media asset according to various exemplary embodiments of the present application described above in the specification when the program product is run on the computer device.
[0174] Obviously, persons skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for generating a plot tag of media, characterized in that, The method includes: For any first media asset, the first media asset description text corresponding to the first media asset is input into a pre-trained plot tag generation network to obtain plot tags corresponding to the first media asset. The first media asset description text is used to introduce the media asset content of the first media asset. The pre-trained story tag generation network is obtained in the following way: The second media asset's description text is input into a plot tag generation network to obtain plot tags corresponding to the second media asset and text encoding vectors corresponding to the second media asset's description text. The text encoding vectors are obtained by text encoding the second media asset's description text. Noise is added to the cover image corresponding to the second media asset using a negative sample generation network to obtain multiple negative sample vectors; Using the plot tags corresponding to the second media asset, the text encoding vector, and the multiple negative sample vectors, the plot tag generation network and the negative sample generation network are subjected to adversarial training to obtain the pre-trained plot tag generation network.
2. The method of claim 1, wherein, The plot tag generation network includes a text encoder and a text decoder; The step of inputting the second media asset's description text into the plot tag generation network to obtain each plot tag corresponding to the second media asset and the text encoding vector corresponding to the second media asset's description text includes: The second media asset description text is encoded using the text encoder to obtain the text encoding vector; The text encoding vector is decoded using the text decoder to obtain the plot tags corresponding to the second media asset.
3. The method of claim 1, wherein, The negative sample generation network adds noise to the cover image corresponding to the second media asset, resulting in multiple negative sample vectors, including: The cover image is input into a preset image encoder to obtain an image encoding vector; and multiple noise vectors are randomly sampled from a preset Gaussian probability distribution function; The image encoding vector is added to the plurality of noise vectors respectively to obtain a plurality of initial negative sample vectors; The plurality of initial negative sample vectors are respectively input into the negative sample generation network to obtain the plurality of negative sample vectors, wherein the number of the initial negative sample vectors is the same as the number of the negative samples.
4. The method according to claim 1, characterized in that, The step of using the plot tags corresponding to the second media asset, the text encoding vector, and the multiple negative sample vectors to perform adversarial training on the plot tag generation network and the negative sample generation network to obtain the pre-trained plot tag generation network includes: Based on the plot tag vectors corresponding to each plot tag, the text encoding vector, and the multiple negative sample vectors, the current model loss value is obtained, wherein the plot tag vector corresponding to any plot tag is obtained by inputting the plot tag into a preset text encoder; If the difference between the current model loss value and the previous model loss value is greater than a first specified threshold, or if the current model loss value is greater than the previous model loss value, then the first model parameters in the plot tag generation network are adjusted, and the process returns to the step of inputting the second media asset description text corresponding to the second media asset into the plot tag generation network, until the current model loss value is no greater than the previous model loss value, and the difference between the current model loss value and the previous model loss value is no greater than the first specified threshold, then the training of the plot tag generation network ends; and... The second model parameters in the negative example generation network are adjusted, and the step of inputting the second media asset description text corresponding to the second media asset into the plot tag generation network is returned until the current model loss value and the third loss value meet the specified conditions. Then, the training of the negative example generation network ends, wherein the third loss value is obtained based on the multiple negative sample vectors; and, After increasing the number of training iterations of the model by a specified value, it is determined whether the current number of training iterations of the model is greater than the specified number of iterations. If not, the step of obtaining the current model loss value based on the plot tag vector corresponding to each plot tag, the text encoding vector, and the multiple negative sample vectors is returned until the current number of training iterations of the model is greater than the specified number of iterations. Then, the trained plot tag generation network is determined as the pre-trained plot tag generation network. If the difference between the current model loss value and the previous model loss value is not greater than the first specified threshold, and the current model loss value is not greater than the previous model loss value, then the step of adjusting the second model parameters in the negative example generation network is executed.
5. The method according to claim 4, characterized in that, The process of obtaining the current model loss value based on the plot tag vector corresponding to each plot tag, the text encoding vector, and the multiple negative sample vectors includes: For any plot tag among the plot tags corresponding to the second media asset, the sub-loss value corresponding to the plot tag is obtained based on the plot tag vector corresponding to the plot tag, the text encoding vector, and the multiple negative sample vectors; The first loss value is obtained based on the sub-loss values of each plot tag; The second loss value is obtained based on the plot tag vector of each plot tag and the total number of plot tags; The current model loss value is obtained based on the first loss value and the second loss value.
6. The method according to claim 4 or 5, characterized in that, The process of obtaining the sub-loss value corresponding to the plot tag based on the plot tag vector, the text encoding vector, and the multiple negative sample vectors includes: The sub-loss value corresponding to the plot tag is obtained using the following formula: Where loss_t is the sub-loss value corresponding to the plot label t. The text encoding vector, Let t be the plot tag vector corresponding to the plot tag t, and τ be a preset hyperparameter. Let r be the r-th negative sample vector, and k be the total number of the multiple negative sample vectors. The image encoding vector is obtained by inputting the cover image corresponding to the second media asset into the image encoder; The process of obtaining the first loss value based on the sub-loss values of each plot tag includes: The first loss value is obtained by summing the sub-loss values of each plot tag; The step of obtaining the second loss value based on the plot tag vector of each plot tag and the total number of plot tags includes: The second loss value is obtained using the following formula: Where loss_b is the second loss value, and M is the total number of each plot tag. Let be the plot tag vector corresponding to the i-th plot tag among all plot tags. Let be the plot tag vector corresponding to the j-th plot tag among the plot tags; The process of obtaining the current model loss value based on the first loss value and the second loss value includes: The first loss value is multiplied by the length penalty factor and then divided by the second loss value to obtain the current model loss value, wherein the length penalty factor is obtained based on the length of each plot tag.
7. The method according to claim 6, characterized in that, The length penalty factor is obtained in the following way: The average length of each plot tag is obtained based on the length of each plot tag and the total number of plot tags. The length penalty factor is obtained by dividing the average length by the preset maximum length.
8. The method according to claim 4, characterized in that, The third loss value is obtained in the following way: Where loss_c is the third loss value. Let p be the p-th negative sample vector among the plurality of negative sample vectors. Let q be the q-th negative sample vector among the plurality of negative sample vectors, and k be the total number of the plurality of negative sample vectors.
9. The method according to claim 4, characterized in that, The specified conditions include the following two: The current model loss value is greater than the previous model loss value, and the difference between the current model loss value and the previous model loss value is not greater than the first specified threshold. The third loss value is greater than the previously obtained third loss value, and the difference between the third loss value and the previously obtained third loss value is not greater than the second specified threshold.
10. An electronic device, characterized in that, It includes a processor and a memory, which are connected via a bus; The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program: For any first media asset, the first media asset description text corresponding to the first media asset is input into a pre-trained plot tag generation network to obtain plot tags corresponding to the first media asset. The first media asset description text is used to introduce the media asset content of the first media asset. The pre-trained story tag generation network is obtained in the following way: The second media asset's description text is input into a plot tag generation network to obtain plot tags corresponding to the second media asset and text encoding vectors corresponding to the second media asset's description text. The text encoding vectors are obtained by text encoding the second media asset's description text. Noise is added to the cover image corresponding to the second media asset using a negative sample generation network to obtain multiple negative sample vectors; Using the plot tags corresponding to the second media asset, the text encoding vector, and the multiple negative sample vectors, the plot tag generation network and the negative sample generation network are subjected to adversarial training to obtain the pre-trained plot tag generation network.
Citation Information
Patent Citations
Film and television label determination method, device and equipment and storage medium
CN109670080A
Cover generation method and device and computer storage medium
CN113986407A