Video description text generation method and device, equipment and storage medium
By means of modal association and filtering, highly correlated modal data are screened out to generate short video description text, which solves the problem of low generation accuracy in the existing technology and achieves higher quality description text generation.
Patent Information
- Application Number
- CN202210265068.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-03-17
AI Technical Summary
The existing technology has the problem of low generation accuracy when generating short video description text, which cannot meet the needs.
By obtaining multiple modal data of the video data to be processed, using the modal association network and the description text generation network, the modal data with a high degree of correlation with the video theme is screened out, filtered and fused, and finally the description text is generated.
The accuracy of descriptive text generation is improved, the generated text is more in line with the subject content of the video data, noise interference is eliminated, and the generation quality is improved.
Smart Images

Figure CN114817629B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for generating video description text. Background Art
[0002] In related technologies, given multimodal data (including visual modality, sound modality, text modality, etc.) are usually aligned and fused to generate a detailed description text of the video content.
[0003] However, related technologies are usually suitable for data with small differences in modal information between samples, but the diversity between short video data samples varies greatly. The use of solutions in related technologies results in low accuracy in generating description text, which cannot meet the generation requirements of short video description text. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for generating video description text, which at least addresses the problem in related art of low description text generation accuracy and inability to meet the generation requirements of short video description text. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a method for generating video description text is provided, comprising:
[0006] Obtain at least two modal data corresponding to the video data to be processed;
[0007] Inputting the at least two modal data into a modal association network to obtain modal association results corresponding to each of the at least two modal data; the modal association results represent the degree of association between each of the at least two modal data and the subject content of the video data to be processed; the modal association network is obtained by training a first neural network on the degree of association based on at least two sample modal data corresponding to sample video data, wherein each of the at least two sample modal data is annotated with a sample degree of association with a sample description text of the sample video data; the sample description text is used to describe the sample subject content of the sample video data;
[0008] Filtering the at least two modal data based on the modal association result to obtain filtered modal data;
[0009] The filtered modal data is input into a description text generation network to obtain a description text of the video data to be processed; the description text is used to describe the subject content; the description text generation network is obtained by training a second neural network for text generation based on the filtered sample modal data, and the filtered sample modal data is obtained by filtering the at least two sample modal data based on the degree of sample association.
[0010] In an exemplary embodiment, filtering the at least two modal data based on the modal association result to obtain filtered modal data includes:
[0011] Based on the modality association result, determining modality data that meets a preset condition from the at least two modality data;
[0012] The modal data that meets the preset conditions is filtered to obtain the filtered modal data.
[0013] In an exemplary embodiment, inputting the filtered modal data into a description text generation network to obtain a description text of the video data to be processed includes:
[0014] When the number of the filtered modal data is at least two, fusing the at least two filtered modal data to obtain fused modal data;
[0015] The fused modality data is input into the description text generation network to obtain the description text.
[0016] In an exemplary embodiment, inputting the filtered modal data into a description text generation network to obtain a description text of the video data to be processed includes:
[0017] When the number of the filtered modal data is one, the filtered modal data is input into the description text generation network to obtain the description text.
[0018] In an exemplary embodiment, the method further includes:
[0019] Acquiring the at least two sample modal data;
[0020] The first neural network is trained on a degree of association based on the at least two sample modal data until a preset condition is satisfied between an output result of the first neural network and the sample degree of association, thereby obtaining the modality association network.
[0021] In an exemplary embodiment, the method further includes:
[0022] filtering the at least two sample modal data based on the sample association degree to obtain the filtered sample modal data;
[0023] The second neural network is trained for text generation based on the filtered sample modal data to obtain the descriptive text generation network.
[0024] In an exemplary embodiment, the performing text generation training on the second neural network based on the filtered sample modal data to obtain the descriptive text generation network includes:
[0025] When the number of the filtered sample modal data is at least two, fusing at least two of the filtered sample modal data to obtain fused sample modal data;
[0026] The second neural network is trained for text generation based on the fused sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
[0027] In an exemplary embodiment, the performing text generation training on the second neural network based on the filtered sample modal data to obtain the descriptive text generation network includes:
[0028] When the number of the filtered sample modal data is one, the second neural network is trained for text generation based on the filtered sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
[0029] In an exemplary embodiment, obtaining the at least two types of sample modality data includes:
[0030] Get historical search log data;
[0031] extracting the sample video data and the search text matching the sample video data from the historical search log data;
[0032] Using the search text as the sample description text;
[0033] marking the sample description text on the sample video data to obtain the sample video data marked with the sample description text;
[0034] The at least two sample modality data are extracted from the sample video data annotated with the sample description text.
[0035] According to a second aspect of an embodiment of the present disclosure, a device for generating a video description text is provided, comprising:
[0036] A modality data acquisition module is configured to acquire at least two modality data corresponding to the video data to be processed;
[0037] A modality association result determination module is configured to input the at least two modality data into a modality association network to obtain modality association results corresponding to each of the at least two modality data; the modality association results represent the degree of association between each of the at least two modality data and the subject content of the video data to be processed; the modality association network is obtained by training a first neural network on the degree of association based on at least two sample modality data corresponding to sample video data, wherein each of the at least two sample modality data is annotated with a sample degree of association with a sample description text of the sample video data; the sample description text is used to describe the sample subject content of the sample video data;
[0038] a modality data filtering module configured to filter the at least two modality data based on the modality association result to obtain filtered modality data;
[0039] The description text determination module is configured to input the filtered modal data into a description text generation network to obtain a description text of the video data to be processed; the description text is used to describe the subject content; the description text generation network is obtained by training a second neural network for text generation based on the filtered sample modal data, and the filtered sample modal data is obtained by filtering the at least two sample modal data based on the sample association degree.
[0040] In an exemplary embodiment, the modal data filtering module includes:
[0041] a modality data determining unit configured to determine, based on the modality association result, modality data satisfying a preset condition from the at least two modality data;
[0042] The modal data filtering unit is configured to filter the modal data that meets the preset conditions to obtain the filtered modal data.
[0043] In an exemplary embodiment, the description text determination module includes:
[0044] a fused modal data determining unit configured to, when the number of the filtered modal data is at least two, fuse at least two of the filtered modal data to obtain fused modal data;
[0045] The description text determination unit is configured to input the fused modality data into the description text generation network to obtain the description text.
[0046] In an exemplary embodiment, the description text determination module is configured to input the filtered modal data into the description text generation network to obtain the description text when the number of the filtered modal data is one.
[0047] In an exemplary embodiment, the apparatus further comprises:
[0048] a sample modality data acquisition module, configured to acquire the at least two types of sample modality data;
[0049] The modal association network training module is configured to perform association degree training on the first neural network based on the at least two sample modal data until the output result of the first neural network and the sample association degree meet a preset condition, thereby obtaining the modal association network.
[0050] In an exemplary embodiment, the apparatus further comprises:
[0051] a sample modality data filtering module configured to filter the at least two sample modality data based on the sample association degree to obtain the filtered sample modality data;
[0052] A generation network training module is configured to perform text generation training on the second neural network based on the filtered sample modal data to obtain the description text generation network.
[0053] In an exemplary embodiment, generating a network training module includes:
[0054] a fused sample modal data determining unit configured to, when the number of the filtered sample modal data is at least two, fuse at least two of the filtered sample modal data to obtain fused sample modal data;
[0055] The generation network generation unit is configured to perform text generation training on the second neural network based on the fused sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
[0056] In an exemplary embodiment, the generation network training module is configured to perform text generation training on the second neural network based on the filtered sample modal data when the number of the filtered sample modal data is one, until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
[0057] In an exemplary embodiment, the sample modality data acquisition module includes:
[0058] A log data acquisition unit configured to acquire historical search log data;
[0059] an extraction unit configured to extract the sample video data and a search text matching the sample video data from the historical search log data;
[0060] a sample description text determining unit, configured to use the search text as the sample description text;
[0061] A labeling unit configured to label the sample description text on the sample video data to obtain the sample video data labeled with the sample description text;
[0062] The sample modality data extraction unit is configured to extract the at least two types of sample modality data from the sample video data annotated with the sample description text.
[0063] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0064] processor;
[0065] A memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method for generating video description text as described in any of the above embodiments.
[0066] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device executes the video description text generation method as described in any of the above embodiments.
[0067] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for generating a video description text as described above is implemented.
[0068] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0069] The embodiment of the present disclosure obtains modal association results corresponding to at least two modal data respectively of the at least two modal data by inputting at least two modal data corresponding to the video data to be processed into a modal association network; filters the at least two modal data based on the modal association results to obtain filtered modal data; and inputs the filtered modal data into a description text generation network to obtain a description text of the video data to be processed. This achieves that before the modal data is input into the description text generation network, the modal data with low quality, high noise, and low correlation with the subject of the video data to be processed is discarded by filtering, and the modal data with high correlation with the subject of the video data to be processed is screened out, thereby eliminating the interference of noise from certain modal data, and generating a description text that is more in line with the subject content of the video data to be processed, thereby improving the generation accuracy of the description text.
[0070] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0072] Figure 1 The figure is a schematic diagram showing an implementation environment of a method for generating video description text according to an exemplary embodiment.
[0073] Figure 2 The figure is a flowchart of a method for generating video description text according to an exemplary embodiment.
[0074] Figure 3 The figure is a flowchart of a method for training a modality association network according to an exemplary embodiment.
[0075] Figure 4 The figure is a flowchart of obtaining at least two types of sample modality data according to an exemplary embodiment.
[0076] Figure 5 is a schematic diagram showing a method of filtering modal data according to an exemplary embodiment.
[0077] Figure 6 The figure is a flowchart showing a method for training a description text generation network according to an exemplary embodiment.
[0078] Figure 7 The figure is a flowchart showing a method for obtaining a description text generation network according to an exemplary embodiment.
[0079] Figure 8The present invention is a flowchart showing a method of obtaining description text of video data to be processed according to an exemplary embodiment.
[0080] Figure 9 The figure is a block diagram of a device for generating video description text according to an exemplary embodiment.
[0081] Figure 10 The present invention is a block diagram showing an electronic device for generating video description text according to an exemplary embodiment. DETAILED DESCRIPTION
[0082] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0083] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0084] Figure 1 FIG. 1 is a schematic diagram showing an implementation environment of a method for generating video description text according to an exemplary embodiment. Figure 1 As shown, the implementation environment may include at least a terminal 01 and a server 02. The terminal 01 and the server 2 may be directly or indirectly connected via wired or wireless communication, which is not limited in the present disclosure.
[0085] Specifically, the terminal can be used to collect video data to be processed. Optionally, the terminal 01 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a smart watch, etc., but is not limited thereto.
[0086] Specifically, the server 02 can obtain at least two modal data corresponding to the video data to be processed; and input the at least two modal data into a modal association network to obtain modal association results corresponding to the at least two modal data; and filter the at least two modal data based on the modal association results to obtain filtered modal data; and input the filtered modal data into a description text generation network to obtain description text of the video data to be processed. Optionally, the server 02 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0087] It should be noted that Figure 1 This is merely an example. In other scenarios, other implementation environments may also be included. For example, the implementation environment may include a terminal, which is provided with a modality association network and a description text generation network, and collects the video data to be processed through the terminal, and generates the description text of the video data to be processed through the terminal.
[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0089] Figure 2 The flowchart of a method for generating video description text according to an exemplary embodiment includes the following steps.
[0090] In step S11, at least two modality data corresponding to the video data to be processed are obtained.
[0091] Optionally, the video data to be processed includes, but is not limited to, short videos and long videos. Short videos refer to frequently pushed video content played on various new media platforms, suitable for viewing on the go or in short leisurely moments, ranging from a few seconds to several minutes. Long videos generally refer to videos with a playback duration of more than half an hour, primarily films and TV series.
[0092] Optionally, the video data to be processed may correspond to at least two modal data, and the at least two modal data may be at least two of visual modal data, sound modal data, and text modal data.
[0093] Optionally, in the above step S11, the present disclosure may obtain at least two modal data corresponding to the video data to be processed in a variety of ways, which are not specifically limited here.
[0094] In one embodiment, a modality extraction network may be pre-trained, and the video data to be processed may be input into the modality extraction network to obtain at least two modality data corresponding to the video data to be processed.
[0095] In another embodiment, a modality extraction module may be provided to extract at least two modality data from the video data to be processed.
[0096] In step S13, the at least two modal data are input into the modal association network to obtain modal association results corresponding to the at least two modal data; the modal association results represent the degree of association between the at least two modal data and the subject content of the video data to be processed; the modal association network is obtained by training the first neural network on the degree of association based on at least two sample modal data corresponding to the sample video data, and the at least two sample modal data are each annotated with a sample degree of association with the sample description text of the sample video data; the sample description text is used to describe the sample subject content of the sample video data.
[0097] First, let’s introduce the training process of the modal association network:
[0098] Figure 3 FIG. 1 is a flow chart showing a method for training a modality association network according to an exemplary embodiment. Figure 3 As shown, training the modality association network may include:
[0099] In step S21 , the at least two types of sample modal data are obtained.
[0100] In step S23, the first neural network is trained for correlation based on the at least two sample modal data until the output result of the first neural network and the sample correlation degree meet a preset condition, thereby obtaining the modal correlation network.
[0101] In an optional embodiment, in the above step S21, sample video data may be randomly extracted from a sample video database, and corresponding sample description text may be annotated on the sample video data to obtain sample video data annotated with the sample description text.
[0102] In another optional embodiment, Figure 4 FIG. 1 is a flow chart showing a method of obtaining at least two types of sample modality data according to an exemplary embodiment. Figure 4 As shown, in the above step S21, the above obtaining of the above at least two sample modal data may include:
[0103] In step S211 , historical search log data is acquired.
[0104] In step S213 , the sample video data and the search text matching the sample video data are extracted from the historical search log data.
[0105] In step S215, the search text is used as the sample description text.
[0106] In step S217 , the sample description text is annotated on the sample video data to obtain sample video data annotated with the sample description text.
[0107] In step S219, the at least two sample modality data are extracted from the sample video data annotated with the sample description text.
[0108] Optionally, in the above steps S211-S215, massive historical search log data can be used to obtain a large amount of sample video data-search text matching data, and the search text can be cleaned (for example, removing expressions, some punctuation marks, etc.), and the cleaned search text can be used as the sample description text of the matching sample video data, and the sample description text is used to describe the subject content of the sample video data.
[0109] Optionally, in the above steps S217 to S219, corresponding sample description text may be annotated on the sample video data to obtain sample video data annotated with the sample description text, and at least two sample modal data may be extracted from the sample video data annotated with the sample description text.
[0110] In a feasible embodiment, the process of extracting at least two types of sample modality data from the sample video data is similar to the process of obtaining at least two types of modality data corresponding to the video data to be processed in step S11, and will not be repeated here.
[0111] The disclosed embodiments extract sample video data and search text matching the sample video data from massive historical search log data, and use the search text as sample description text, thereby improving the accuracy of determining the sample description text. Using the high-precision sample description text can improve the accuracy of determining the degree of sample association. In addition, directly obtaining sample video data and sample description text from massive historical search log data can also reduce the cost of obtaining the sample video data and sample description text, thereby improving the training accuracy of the modal association network and reducing the training cost of the modal association network.
[0112] In a feasible embodiment, in the above step S21, after obtaining at least two types of sample modal data, the degree of sample association between the at least two types of sample modal data and the sample description text may be annotated.
[0113] In one embodiment, the sample association degree can be "the sample modal data is associated with the sample description text" or "the sample modal data is not associated with the sample description text." For example, if the sample modal data is text modal data, the text modal data is the lyrics of background music, and the sample description text (i.e., the subject content) of the sample video data is about making a certain dish, then the sample association degree between the two is "the sample modal data is not associated with the sample description text." For another example, if the sample modal data is visual modal data, the visual modal data is a picture of the steps for making a dish, and the sample description text (i.e., the subject content) of the sample video data is about making a certain dish, then the sample association degree between the two is "the sample modal data is associated with the sample description text."
[0114] In another embodiment, the sample association degree can be a correlation score of "1" or "0", where "0" indicates that the sample modal data is not associated with the sample description text, and "1" indicates that the sample modal data is associated with the sample description text. For example, if the sample modal data is text modal data, the text modal data is the lyrics of the background music, and the sample description text (i.e., the theme content) of the sample video data is about making a certain dish, then the sample association degree between the two is "0". For another example, if the sample modal data is visual modal data, the visual modal data is a step-by-step picture of making a dish, and the sample description text (i.e., the theme content) of the sample video data is about making a certain dish, then the sample association degree between the two is "1".
[0115] In an optional embodiment, in the above step S23, at least two sample modal data can be input into the first neural network to train the first neural network, and the loss value between the output result of the first neural network and the sample association degree is calculated. When the loss value meets the preset conditions, the network obtained by the current training is used as the modal association network. When the loss value does not meet the preset conditions, the model parameters of the first neural network are adjusted, and the above training process is repeated until the loss value meets the preset conditions, and the network obtained by the current training is used as the modal association network.
[0116] In an optional embodiment, in the above step S23, in order to improve the training accuracy of the modal association network, the sample modal features corresponding to at least two types of sample modal data can be pre-extracted, and the sample modal features corresponding to at least two types of sample modal data can be input into the first neural network to train the first neural network, and the obtained network is used as the modal association network.
[0117] Optionally, to improve the training accuracy of the modal association network, the log-likelihood loss function can be used to optimize the network. The log-likelihood loss function works by finding a set of estimated values that maximizes the probability of the observed value when the unknown parameter takes that set of estimated values. Specifically, when the likelihood function reaches its maximum value, the model's predicted value is relatively close to the true value, resulting in a smaller loss function.
[0118] In the embodiment of the present disclosure, since the sample description text and at least two sample modal data can be derived from massive historical search log data, the acquisition cost of the sample description text and at least two sample modal data is low and the determination accuracy is high. Training the first neural network based on at least two sample modal data marked with sample association degrees can improve the training accuracy of the modal association network and reduce the training cost of the modal association network, so that the modal association network can accurately and cost-effectively output the degree of association between the modal data of a certain video data and the subject content of the video data.
[0119] In a feasible embodiment, in step S13, after the modal association network is trained, at least two modal data can be directly input into the modal association network, and the modal association results corresponding to the at least two modal data can be output. The modal association results represent the degree of association between the at least two modal data and the subject content of the video data to be processed.
[0120] In a feasible embodiment, in step S13, after the modal association network is trained, in order to improve the determination accuracy of the modal association results, the modal features corresponding to at least two modal data can be pre-extracted, the modal features corresponding to at least two modal data can be input into the modal association network, and the modal association results corresponding to at least two modal data can be output.
[0121] In one embodiment, in the above step S13, the modality association result may be "the modality data is associated with the subject content of the video data to be processed" or "the modality data is not associated with the subject content of the video data to be processed".
[0122] In another embodiment, in the above-mentioned step S13, the modal association result can be a score between "0-1", "0" indicates that the modal data is not associated with the subject content of the video data to be processed, and "1" indicates that the modal data is associated with the subject content of the video data to be processed. For the values between 0-1, the closer the value is to 1, the higher the degree of association between the modal data and the subject content of the video data to be processed, and the farther the value is from 1, the lower the degree of association between the modal data and the subject content of the video data to be processed.
[0123] In step S15 , the at least two modal data are filtered based on the modal association result to obtain filtered modal data.
[0124] In an optional embodiment, Figure 5 is a schematic diagram showing a method of filtering modal data according to an exemplary embodiment. Figure 5 As shown, in the above step S15, filtering the above at least two modal data based on the above modal association result to obtain filtered modal data may include:
[0125] In step S151 , based on the modality association result, modality data that meets a preset condition is determined from the at least two modality data.
[0126] In step S153 , the modal data satisfying the preset conditions are filtered to obtain the filtered modal data.
[0127] In one embodiment, when the modal association result is "the modal data is associated with the subject content of the video data to be processed" or "the modal data is not associated with the subject content of the video data to be processed", the modal data that is not associated with the subject content of the video data to be processed can be used as the modal data that meets the preset conditions, and the modal data that meets the preset conditions is filtered to obtain the filtered modal data. For example, the video data to be processed corresponds to text modal data (lyrics of background music) and visual modal data (step-by-step images of dish preparation), and the subject content of the video data to be processed is the preparation of a certain dish. In this case, the modal association result between the text modal data and the subject content is "not associated", and the modal association result between the visual modal data and the subject content is "associated". In this case, the text modal data is filtered, and the visual modal data is retained to obtain the filtered modal data.
[0128] In another way, when the modal association result is a "score between 0-1", a score threshold can be set. If the modal association result is greater than the score threshold, it is retained. If the modal association result is less than or equal to the score threshold, it is deleted to obtain filtered modal data. For example, the video data to be processed corresponds to text modal data (lyrics of background music) and visual modal data (step-by-step pictures of dish preparation). The theme content of the video data to be processed is the preparation of a certain dish. Then the modal association result between the text modal data and the theme content is "0", and the modal association result between the visual modal data and the theme content is "1". Then the text modal data is filtered and the visual modal data is retained to obtain filtered modal data.
[0129] In an embodiment of the present disclosure, based on the modal association results, modal data with low quality, high noise, and low correlation with the subject of the video data to be processed are discarded from at least two modal data, and modal data with a high correlation with the subject content of the video data to be processed are screened out. This can eliminate the interference of noise in some modal data, thereby generating a descriptive text that is more in line with the subject content of the video data to be processed, thereby improving the generation accuracy of the descriptive text.
[0130] In step S17, the filtered modal data is input into a description text generation network to obtain a description text of the video data to be processed; the description text is used to describe the subject content; the description text generation network is obtained by training a second neural network for text generation based on the filtered sample modal data, and the filtered sample modal data is obtained by filtering the at least two sample modal data based on the sample association degree.
[0131] First, let’s introduce the training process of the description text generation network:
[0132] Figure 6 FIG. 1 is a flowchart of a training description text generation network according to an exemplary embodiment. Figure 6 The training description text generation network shown may include:
[0133] In step S31 , the at least two sample modal data are filtered based on the sample association degree to obtain filtered sample modal data.
[0134] In step S33, the second neural network is trained for text generation based on the filtered sample modal data to obtain the descriptive text generation network.
[0135] In an optional embodiment, in step S31, since the sample correlation degree between the sample modal data and the sample description text is pre-labeled, sample modal data that satisfies a preset condition can be determined from at least two types of sample modal data based on the sample correlation degree, and the sample modal data that satisfies the preset condition is filtered to obtain filtered sample modal data. The filtering process for sample modal data is similar to the filtering process for modal data. For details, please refer to steps S151-S153 above, and will not be repeated here.
[0136] In an optional embodiment, in the above step S33, the above-mentioned filtered sample modal data can be input into the second neural network, and the loss value between the output of the second neural network and the sample description text is calculated. When the loss value meets the preset conditions, the currently trained neural network is used as the description text generation network; when the loss value does not meet the preset conditions, the model parameters of the neural network are adjusted until the loss value meets the preset conditions, and the currently trained neural network is used as the description text generation network.
[0137] In the embodiment of the present disclosure, since sample modal data with low quality and high noise are discarded during the sample modal data training process, the filtered sample modal data can better fit the subject content of the sample video data. Using the filtered sample modal data to train the second neural network can improve the training accuracy of the description text generation network, so that the trained description text generation network can accurately output the description text of a certain video data.
[0138] In one embodiment, Figure 7 is a flowchart showing a method for obtaining a description text generation network according to an exemplary embodiment. Figure 7 As shown, in the above step S33, the text generation training of the second neural network based on the above filtered sample modal data to obtain the above description text generation network may include:
[0139] In step S331 , when the number of the filtered sample modal data is at least two, the at least two filtered sample modal data are fused to obtain fused sample modal data.
[0140] In step S333, the second neural network is trained for text generation based on the fused sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
[0141] Optionally, when the number of the filtered sample modal data is at least two (for example, the sample visual modal data and the sample text modal data are retained at the same time), before entering the second neural network, the at least two filtered sample modal data can be fused to obtain fused sample modal data, and the fused sample modal data can be input into the second neural network to train the second neural network until the output result of the second neural network and the sample description text meet the preset conditions to obtain a description text generation network.
[0142] For example, a cross-attention module can be used for modal fusion. The cross-attention mechanism can effectively interact the information within the modal data with the information between the modalities. The following describes the process of fusing sample audio modal data and sample visual modal data using the cross-attention mechanism:
[0143] First, generate a sample sound-sample visual modality pair (Fv, Fa) and calculate a pair of matching matrices to represent the cross-modal information. The calculation formula is as follows:
[0144]
[0145] Then calculate M through the Softmax function row by row av and M va The probability distribution of each utterance is calculated, that is, the attention weight of the context is calculated, where N represents the correlation between the sample sound modality feature and the sample visual modality feature. The larger the value, the stronger the interaction between the two features, and the more important the fusion information is. The calculation formula is as follows:
[0146] N1=soft max(Mav), N2=soft max(Mva);
[0147] Then, the modal attention representation is calculated on the multimodal multi-discourse attention matrix. O1O2 represents the final attention representation matrix obtained by multiplying the attention score and the feature matrix. The formula is as follows:
[0148] O1=N1·Fa,O2=N2·Fv;
[0149] Next, a multiplication gating mechanism is used between each individual modality and the multimodal discourse-specific representations of other modalities to obtain the mutual attention information matrix A1A2 between the two modalities. This element-by-element multiplication helps to focus on the important components of multiple modalities and discourses. The formula is as follows:
[0150]
[0151] Finally, the attention matrices A1 and A2 are concatenated to obtain a new feature representation matrix as the fusion information feature representation F of the sample sound-sample visual modality pair. av , the formula is as follows:
[0152] Fav=concat[A1,A2];
[0153] In the above formula, · represents matrix multiplication, Represents element-by-element matrix multiplication.
[0154] In an embodiment of the present disclosure, when the number of filtered sample modal data is at least two, the at least two filtered sample modal data are fused to obtain fused sample modal data. Since the fused sample modal data fully contains the internal information of the at least two filtered sample modal data, the second neural network is trained with the fused sample modal data, which can improve the training accuracy of the description text generation network, so that the trained description text generation network can accurately output the description text of a certain video data.
[0155] In another embodiment, in step S33, the text generation training of the second neural network based on the filtered sample modal data to obtain the description text generation network may include:
[0156] When the number of filtered sample modal data is one, the second neural network is trained for text generation based on the filtered sample modal data until the output result of the second neural network and the sample description text meet preset conditions, thereby obtaining a description text generation network.
[0157] Optionally, when the number of filtered sample modal data is one (for example, the sample visual modal data or the sample text modal data is retained), the second neural network is trained directly based on the filtered sample modal data until the output result of the second neural network and the sample description text meet the preset conditions, thereby obtaining a description text generation network.
[0158] In the embodiment of the present disclosure, when the number of filtered sample modal data is one, the above-mentioned second neural network is trained directly based on the filtered sample modal data, which can improve the training accuracy and efficiency of the description text generation network, so that the trained description text generation network can accurately and efficiently output the description text of a certain video data.
[0159] In another embodiment, in the above-mentioned step S33, the above-mentioned training of the second neural network based on the above-mentioned filtered sample modal data to obtain the above-mentioned description text generation network may include: if the number of filtered sample modal data is equal to 0 (for example, both the sample visual modal data and the sample text modal data are discarded), then the data is discarded and the description text generation network is not trained, thereby reducing the consumption of system resources by the training process.
[0160] In an optional embodiment, Figure 8 FIG. 1 is a flowchart showing a method for obtaining a description text of video data to be processed according to an exemplary embodiment. Figure 8As shown, in the above step S17, the above filtered modal data is input into the description text generation network to obtain the description text of the above video data to be processed, which may include:
[0161] In step S171 , when the number of the filtered modal data is at least two, the at least two filtered modal data are fused to obtain fused modal data.
[0162] In step S173, the fused modal data is input into the description text generation network to obtain the description text.
[0163] Optionally, when the number of filtered modal data is at least two (for example, both visual modal data and sample text modal data are retained), before entering the trained description text generation network, the at least two filtered modal data can be fused to obtain fused modal data, and the fused modal data can be input into the trained description text generation network to obtain description text.
[0164] For example, a cross-attention module can be used for modality fusion.
[0165] In an embodiment of the present disclosure, when the number of filtered modal data is at least two, the at least two filtered modal data are fused to obtain fused modal data. Since the fused modal data fully contains the internal information of the at least two filtered modal data, the fused modal data is input into a trained description text generation network, which can accurately output a description text of the video data. The description text can be used in scenarios such as matching with search terms or recommending with topic texts. In addition, the description text generation method in the embodiment of the present disclosure is a one-stage method, that is, at least two modal data corresponding to the video to be processed are input into the description text generation network, which can summarize and extract the topic content of the video data to be processed to obtain a description text. The generated description text is relatively short (for example, about 20 words) and does not require secondary processing. It can be more appropriately applied to subsequent search or recommendation scenarios, thereby achieving the purpose of efficiently generating description text for the video data to be processed.
[0166] In another optional embodiment, in step S17, the step of inputting the filtered modal data into the description text generation network to obtain the description text of the video data to be processed may include:
[0167] When the number of the filtered modal data is one, the filtered modal data is input into the description text generation network to obtain the description text.
[0168] Optionally, when the number of modal data after filtering is one (for example, visual modal data or sample text modal data is retained), the filtered modal data is directly input into the trained description text generation network to obtain description text, which can improve the accuracy of description text generation. In addition, the description text generation method in the embodiment of the present disclosure is a one-stage method, that is, the modal data is input into the description text generation network, and the subject content of the video data to be processed can be summarized and extracted to obtain description text. The length of the description text generated is relatively short (for example, about 20 words), and it does not require secondary processing and can be more appropriately applied to subsequent search or recommendation scenarios, thereby achieving the purpose of efficiently generating description text for the video data to be processed.
[0169] In another embodiment, in the above-mentioned step S17, the above-mentioned filtered modal data is input into the description text generation network to obtain the description text of the above-mentioned video data to be processed, which may include: if the amount of filtered modal data is equal to 0 (for example, both visual modal data and textual modal data are discarded), then the data is discarded and the description text is not generated, thereby reducing system resource consumption.
[0170] Figure 9 FIG. 1 is a block diagram of a device for generating video description text according to an exemplary embodiment. Figure 9 The device includes a modal data acquisition module 41, a modal association result determination module 43, a modal data filtering module 45 and a description text determination module 47.
[0171] The modality data acquisition module 41 is configured to acquire at least two modality data corresponding to the video data to be processed.
[0172] The modal association result determination module 43 is configured to execute inputting the above-mentioned at least two modal data into the modal association network to obtain the modal association results corresponding to each of the above-mentioned at least two modal data; the above-mentioned modal association results represent the degree of association between each of the above-mentioned at least two modal data and the subject content of the above-mentioned video data to be processed; the above-mentioned modal association network is obtained by training the first neural network on the degree of association based on at least two sample modal data corresponding to the sample video data, and the above-mentioned at least two sample modal data are each annotated with the sample degree of association between the sample and the sample description text of the above-mentioned sample video data; the above-mentioned sample description text is used to describe the sample subject content of the above-mentioned sample video data.
[0173] The modal data filtering module 45 is configured to filter the at least two modal data based on the modal association result to obtain filtered modal data.
[0174] The description text determination module 47 is configured to execute inputting the above-mentioned filtered modal data into the description text generation network to obtain the description text of the above-mentioned video data to be processed; the above-mentioned description text is used to describe the above-mentioned subject content; the above-mentioned description text generation network is obtained by training the second neural network for text generation based on the filtered sample modal data, and the above-mentioned filtered sample modal data is obtained by filtering the above-mentioned at least two sample modal data based on the above-mentioned sample association degree.
[0175] In an exemplary embodiment, the modal data filtering module includes:
[0176] The modal data determining unit is configured to determine the modal data that meets the preset conditions from the at least two modal data based on the modal association result.
[0177] The modal data filtering unit is configured to filter the modal data that meets the preset conditions to obtain the filtered modal data.
[0178] In an exemplary embodiment, the description text determination module includes:
[0179] The fused modal data determining unit is configured to, when the number of the filtered modal data is at least two, fuse at least two filtered modal data to obtain fused modal data.
[0180] The description text determination unit is configured to input the above-mentioned fusion modality data into the above-mentioned description text generation network to obtain the above-mentioned description text.
[0181] In an exemplary embodiment, the description text determination module is configured to input the filtered modal data into the description text generation network to obtain the description text when the number of the filtered modal data is one.
[0182] In an exemplary embodiment, the apparatus further comprises:
[0183] The sample modality data acquisition module is configured to acquire the at least two types of sample modality data.
[0184] The modal association network training module is configured to perform association degree training on the above-mentioned first neural network based on the above-mentioned at least two sample modal data until the output result of the above-mentioned first neural network and the above-mentioned sample association degree meet the preset conditions, thereby obtaining the above-mentioned modal association network.
[0185] In an exemplary embodiment, the apparatus further comprises:
[0186] The sample modal data filtering module is configured to filter the at least two sample modal data based on the sample association degree to obtain filtered sample modal data.
[0187] The generation network training module is configured to perform text generation training on the second neural network based on the filtered sample modal data to obtain the above-mentioned description text generation network.
[0188] In an exemplary embodiment, the generation of the network training module includes:
[0189] The fused sample modal data determining unit is configured to, when the number of the filtered sample modal data is at least two, fuse at least two filtered sample modal data to obtain fused sample modal data.
[0190] The generation network generation unit is configured to perform text generation training on the above-mentioned second neural network based on the above-mentioned fused sample modal data until the output result of the above-mentioned second neural network and the above-mentioned sample description text meet the preset conditions, thereby obtaining the above-mentioned description text generation network.
[0191] In an exemplary embodiment, the above-mentioned generation network training module is configured to perform text generation training on the above-mentioned second neural network based on the above-mentioned filtered sample modal data when the number of the above-mentioned filtered sample modal data is one, until the output result of the above-mentioned second neural network and the above-mentioned sample description text meet a preset condition, thereby obtaining the above-mentioned description text generation network.
[0192] In an exemplary embodiment, the sample modality data acquisition module includes:
[0193] The log data acquisition unit is configured to acquire the historical search log data.
[0194] The extraction unit is configured to extract the sample video data and the search text matching the sample video data from the historical search log data.
[0195] The sample description text determining unit is configured to use the search text as the sample description text.
[0196] The labeling unit is configured to label the sample description text on the sample video data to obtain the sample video data labeled with the sample description text.
[0197] The sample modality data extraction unit is configured to extract the at least two types of sample modality data from the sample video data annotated with the sample description text.
[0198] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0199] In an exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of any video description text generation method in the above embodiments when executing the instructions stored in the memory.
[0200] The electronic device may be a terminal, a server or a similar computing device. For example, the electronic device is a server. Figure 10 This is a block diagram of an electronic device for generating video description text according to an exemplary embodiment. The electronic device 50 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 51 (the central processing unit 51 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 53 for storing data, and one or more storage media 52 for storing application programs 523 or data 522 (for example, one or more mass storage devices). Among them, the memory 53 and the storage medium 52 can be temporary storage or permanent storage. The program stored in the storage medium 52 may include one or more modules, each module may include a series of instruction operations in the electronic device. Furthermore, the central processing unit 51 can be configured to communicate with the storage medium 52 to execute a series of instruction operations in the storage medium 52 on the electronic device 50. The electronic device 50 may also include one or more power supplies 56, one or more wired or wireless network interfaces 55, one or more input and output interfaces 54, and / or one or more operating systems 521, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0201] The input / output interface 54 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the electronic device 50. In one embodiment, the input / output interface 54 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In an exemplary embodiment, the input / output interface 54 can be a radio frequency (RF) module for wireless communication with the Internet.
[0202] It can be understood by those skilled in the art that Figure 10 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 10 More or fewer components than shown, or with Figure 10 Different configurations shown.
[0203] In an exemplary embodiment, a computer-readable storage medium is further provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of any of the methods for generating video description text in the above embodiments.
[0204] In an exemplary embodiment, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the method for generating a video description text provided in any one of the above embodiments is implemented.
[0205] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, which can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present disclosure can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0206] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0207] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for generating video description text, characterized in that: include: Obtain at least two modal data corresponding to the video data to be processed; Inputting the at least two modal data into a modal association network to obtain modal association results corresponding to each of the at least two modal data; the modal association results represent the degree of association between each of the at least two modal data and the subject content of the video data to be processed; the modal association network is obtained by training a first neural network on the degree of association based on at least two sample modal data corresponding to sample video data, wherein each of the at least two sample modal data is annotated with a sample degree of association with a sample description text of the sample video data; the sample description text is used to describe the sample subject content of the sample video data; Based on the modality association result, determining modality data that meets a preset condition from the at least two modality data; Filtering the modal data that meets the preset conditions to obtain filtered modal data; the modal data that meets the preset conditions is modal data that is not associated with the subject content; The filtered modal data is input into a description text generation network to obtain a description text of the video data to be processed; the description text is used to describe the subject content; the description text generation network is obtained by training a second neural network for text generation based on the filtered sample modal data, and the filtered sample modal data is obtained by filtering the at least two sample modal data based on the degree of sample association.
2. The method for generating video description text according to claim 1, wherein: The step of inputting the filtered modal data into a description text generation network to obtain a description text of the video data to be processed comprises: When the number of the filtered modal data is at least two, fusing the at least two filtered modal data to obtain fused modal data; The fused modality data is input into the description text generation network to obtain the description text.
3. The method for generating video description text according to claim 1, wherein: The step of inputting the filtered modal data into a description text generation network to obtain a description text of the video data to be processed comprises: When the number of the filtered modal data is one, the filtered modal data is input into the description text generation network to obtain the description text.
4. The method for generating video description text according to any one of claims 1 to 3, characterized in that: The method further comprises: Acquiring the at least two sample modal data; The first neural network is trained on a degree of association based on the at least two sample modal data until a preset condition is satisfied between an output result of the first neural network and the sample degree of association, thereby obtaining the modality association network.
5. The method for generating video description text according to claim 4, wherein: The method further comprises: filtering the at least two sample modal data based on the sample association degree to obtain the filtered sample modal data; The second neural network is trained for text generation based on the filtered sample modal data to obtain the descriptive text generation network.
6. The method for generating video description text according to claim 5, characterized in that: The performing text generation training on the second neural network based on the filtered sample modal data to obtain the description text generation network includes: When the number of the filtered sample modal data is at least two, fusing at least two of the filtered sample modal data to obtain fused sample modal data; The second neural network is trained for text generation based on the fused sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
7. The method for generating video description text according to claim 5, wherein: The performing text generation training on the second neural network based on the filtered sample modal data to obtain the description text generation network includes: When the number of the filtered sample modal data is one, the second neural network is trained for text generation based on the filtered sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
8. The method for generating video description text according to claim 4, wherein: The acquiring of the at least two types of sample modal data includes: Get historical search log data; extracting the sample video data and the search text matching the sample video data from the historical search log data; Using the search text as the sample description text; marking the sample description text on the sample video data to obtain the sample video data marked with the sample description text; The at least two sample modality data are extracted from the sample video data annotated with the sample description text.
9. A device for generating video description text, characterized in that: include: A modality data acquisition module is configured to acquire at least two modality data corresponding to the video data to be processed; A modality association result determination module is configured to input the at least two modality data into a modality association network to obtain modality association results corresponding to each of the at least two modality data; the modality association results represent the degree of association between each of the at least two modality data and the subject content of the video data to be processed; the modality association network is obtained by training a first neural network on the degree of association based on at least two sample modality data corresponding to sample video data, wherein each of the at least two sample modality data is annotated with a sample degree of association with a sample description text of the sample video data; the sample description text is used to describe the sample subject content of the sample video data; a modality data filtering module configured to filter the at least two modality data based on the modality association result to obtain filtered modality data; The modal data filtering module includes: a modal data determining unit configured to determine, based on the modal association result, modal data that meets a preset condition from the at least two modal data; a modal data filtering unit configured to filter the modal data that meets the preset condition to obtain the filtered modal data; the modal data that meets the preset condition is modal data that is not associated with the subject content; The description text determination module is configured to input the filtered modal data into a description text generation network to obtain a description text of the video data to be processed; the description text is used to describe the subject content; the description text generation network is obtained by training a second neural network for text generation based on the filtered sample modal data, and the filtered sample modal data is obtained by filtering the at least two sample modal data based on the sample association degree.
10. The video description text generation device according to claim 9, characterized in that: The description text determination module includes: a fused modal data determining unit configured to, when the number of the filtered modal data is at least two, fuse at least two of the filtered modal data to obtain fused modal data; The description text determination unit is configured to input the fused modality data into the description text generation network to obtain the description text.
11. The video description text generation device according to claim 9, characterized in that: The description text determination module is configured to input the filtered modal data into the description text generation network to obtain the description text when the number of the filtered modal data is one.
12. The video description text generation device according to any one of claims 9 to 11, characterized in that: The device further comprises: a sample modality data acquisition module, configured to acquire the at least two types of sample modality data; The modal association network training module is configured to perform association degree training on the first neural network based on the at least two sample modal data until the output result of the first neural network and the sample association degree meet a preset condition, thereby obtaining the modal association network.
13. The video description text generation device according to claim 12, characterized in that: The device further comprises: a sample modality data filtering module configured to filter the at least two sample modality data based on the sample association degree to obtain the filtered sample modality data; A generation network training module is configured to perform text generation training on the second neural network based on the filtered sample modal data to obtain the description text generation network.
14. The video description text generation device according to claim 13, characterized in that: The generation network training module includes: a fused sample modal data determining unit configured to, when the number of the filtered sample modal data is at least two, fuse at least two of the filtered sample modal data to obtain fused sample modal data; The generation network generation unit is configured to perform text generation training on the second neural network based on the fused sample modal data until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
15. The video description text generation device according to claim 13, characterized in that: The generation network training module is configured to perform text generation training on the second neural network based on the filtered sample modal data when the number of the filtered sample modal data is one, until the output result of the second neural network and the sample description text meet a preset condition, thereby obtaining the description text generation network.
16. The video description text generation device according to claim 12, characterized in that: The sample modal data acquisition module includes: A log data acquisition unit configured to acquire historical search log data; an extraction unit configured to extract the sample video data and a search text matching the sample video data from the historical search log data; a sample description text determining unit, configured to use the search text as the sample description text; A labeling unit configured to label the sample description text on the sample video data to obtain the sample video data labeled with the sample description text; The sample modality data extraction unit is configured to extract the at least two types of sample modality data from the sample video data annotated with the sample description text.
17. An electronic device, characterized in that: include: processor; A memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the video description text generation method according to any one of claims 1 to 8. 18 . A computer-readable storage medium, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device executes the method for generating video description text according to claim 1 .
19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating video description text according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Information obtaining method and device, electronic equipment and storage medium
CN113792166A
Personalized recommendation online and offline integration method based on multi-modal feature association
CN114139051A