A method, apparatus, device and medium for generating a semantic vector of a video
By analyzing user behavior logs and training models with video frame sequences, video semantic vectors are generated, which solves the problem of insufficient semantic representation in video recommendation models, achieves more accurate and interpretable video semantic vectors, and improves the performance of video recommendation and other services.
Patent Information
- Application Number
- CN202210467951.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-04-29
AI Technical Summary
Existing video recommendation models lack effective semantic representation of video content information, resulting in low interpretability and accuracy of the extracted video semantic vectors.
By acquiring user behavior logs from sample videos, analyzing the statistical values of various user behavior indicators, and combining them with video frame sequences to input into a pre-set model for training, semantic vectors of the videos are generated. The model is then optimized using the statistical values of various user behavior indicators and video frame sequences. Self-supervised learning and data augmentation techniques are employed to improve the model's accuracy and interpretability.
The generated video semantic vectors have higher interpretability and accuracy, better matching the needs of downstream tasks and improving the accuracy of video recommendation and other related services.
Smart Images

Figure CN114842382B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a method, apparatus, device, and medium for generating semantic vectors of video. Background Technology
[0002] A video's semantic vector is a vector representing the content information of the video, essentially a quantization of the video. Video semantic vectors are of great significance in video-related businesses such as video recommendation.
[0003] In video recommendation scenarios, recommendation models often fail to fully utilize the content information of videos, lacking effective semantic representation. To obtain the semantic vector of a video in this context, the output of the intermediate layers of the recommendation model is typically used. However, this method may result in low interpretability and accuracy of the extracted semantic vector. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention provide a method, apparatus, device and medium for generating semantic vectors of videos, so as to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of this invention provides a method for generating semantic vectors of a video, the method comprising:
[0006] Obtain sample videos and user behavior logs of the sample videos;
[0007] The user behavior logs are analyzed to obtain statistical values of various user behavior indicators for the sample videos;
[0008] The statistical values of the various user behavior indicators and the video frame sequence of the sample video are input into the preset model to be trained to obtain the semantic vector of the sample video output by the preset model to be trained.
[0009] Optionally, analyzing the user behavior logs to obtain statistical values of various user behavior indicators for the sample video further includes:
[0010] Based on the user behavior information required by the application side of the sample video, various user behavior indicators of the sample video are determined.
[0011] Optionally, the statistical values of the various user behavior indicators and the video frame sequence of the sample video are input into a preset model to be trained to obtain the semantic vector of the sample video output by the preset model to be trained, including:
[0012] The statistical values of the various user behavior indicators and the video frame sequence of the sample video are input into the preset model to be trained to obtain the predicted values of the various user behavior indicators output by the preset model to be trained.
[0013] The first loss function value is obtained based on the statistical values and corresponding predicted values of the various user behavior indicators.
[0014] The model parameters of the preset model to be trained are updated based on the first loss function value.
[0015] Optionally, the method further includes:
[0016] The video frame sequence of the sample video is subjected to strong data augmentation and weak data augmentation to obtain strong video frame sequence and weak video frame sequence, respectively.
[0017] The statistical values of the various user behavior indicators and the video frame sequence of the sample video are input into a preset model to be trained to obtain the semantic vector of the sample video output by the preset model to be trained, including:
[0018] The statistical values of the various user behavior indicators and the weak video frame sequence are input into the preset model to be trained to obtain the predicted values of the various user behavior indicators output by the preset model to be trained.
[0019] The strong video frame sequence and the weak video frame sequence are input into the preset model to be trained to obtain the semantic vectors of the strong video frame sequence and the weak video frame sequence output by the preset model to be trained.
[0020] The first loss function value is obtained based on the statistical values and corresponding predicted values of the various user behavior indicators.
[0021] The semantic vector of the weak video frame sequence and the semantic vector of the weak video frame sequence are used to obtain the second loss function value;
[0022] The model parameters of the preset model to be trained are updated based on the first loss function value and the second loss function value.
[0023] Optionally, the preset model to be trained includes: a vector generation module, an indicator prediction module concatenated after the vector generation module, and a self-supervised module concatenated after the vector generation module and arranged in parallel with the indicator prediction module. The indicator prediction module is used to output the predicted values of the various user behavior indicators, and the self-supervised module is used to output the semantic vectors of the strong video frame sequence and the weak video frame sequence, respectively. After the preset model to be trained is completed, the method further includes:
[0024] The trained vector generation module is used as the semantic vector generation model;
[0025] Acquire the target video;
[0026] The target video is input into the semantic vector generation model to obtain the semantic vector of the target video output by the semantic vector generation model.
[0027] A second aspect of the present invention provides an apparatus for generating semantic vectors of a video, the apparatus comprising:
[0028] The sample video acquisition module is used to acquire sample videos and user behavior logs of the sample videos;
[0029] The statistical value acquisition module is used to analyze the user behavior logs and obtain statistical values of various user behavior indicators for the sample video.
[0030] The semantic vector acquisition module is used to input the statistical values of the various user behavior indicators and the video frame sequence of the sample video into the preset model to be trained, and obtain the semantic vector of the sample video output by the preset model to be trained.
[0031] Optionally, before analyzing the user behavior logs to obtain statistical values of the user behavior metrics for the sample video, the device further includes:
[0032] The behavior indicator determination module is used to determine various user behavior indicators of the sample video based on the user behavior information required by the application end of the sample video.
[0033] Optionally, the semantic vector acquisition module includes:
[0034] The prediction value acquisition unit is used to input the statistical values of the various user behavior indicators and the video frame sequence of the sample video into the preset model to be trained, and obtain the predicted values of the various user behavior indicators output by the preset model to be trained.
[0035] The first loss function generation unit is used to obtain the first loss function value based on the statistical values and corresponding predicted values of the various user behavior indicators;
[0036] The model parameter update unit is used to update the model parameters of the preset model to be trained based on the first loss function value.
[0037] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for generating semantic vectors of video as disclosed in the embodiments of this application.
[0038] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method for generating semantic vectors of a video as disclosed in the embodiments of this application.
[0039] The embodiments of the present invention have the following advantages:
[0040] In this embodiment of the invention, sample videos and user behavior logs of the sample videos are acquired; the user behavior logs are analyzed to obtain statistical values of various user behavior indicators for the sample videos; the statistical values of the various user behavior indicators and the video frame sequence of the sample videos are input into a preset model to be trained to obtain the semantic vector of the sample videos output by the preset model to be trained. Thus, training the preset model using video frame sequences allows the preset model to fully utilize the content information of the video, resulting in highly interpretable semantic vectors of the output videos. Simultaneously, training the model using statistical values of multiple user behavior indicators avoids the potential inaccuracies of statistical values for a single user behavior indicator or the hindering effect of overly simplistic user behavior indicators on the performance of the trained model, making the semantic vectors of the videos output by the model more accurate. Furthermore, training the model using statistical values of multiple user behavior indicators helps downstream users perform tasks related to user behavior indicators using the semantic vectors of videos, ensuring a better match between the semantic vectors of the videos used and the semantic vectors of the videos generated by the preset model. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating the steps of a method for generating semantic vectors for a video according to an embodiment of the present invention.
[0043] Figure 2 This is a schematic diagram of the structure of the preset model to be trained in an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of the semantic vector generation model in an embodiment of the present invention;
[0045] Figure 4 This is a schematic diagram of the structure of a device for generating semantic vectors of video according to an embodiment of the present invention. Detailed Implementation
[0046] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] To address the technical problem of low interpretability and accuracy of semantic vectors extracted from videos in related technologies, the applicant proposes to train the model using statistical values of various user behavior indicators and video frame sequences, so that the semantic vectors of the videos output by the trained model have high interpretability and accuracy.
[0048] It should be noted that the method for generating semantic vectors for videos proposed in this application is applicable not only to video recommendation scenarios but also to other scenarios. The generated semantic vectors for videos can be used not only for video recommendation but also for various video-related services such as video retrieval, video recognition, video classification, video recall, and video ranking.
[0049] Reference Figure 1 The diagram illustrates a flowchart of the steps involved in generating semantic vectors for a video according to an embodiment of the present invention. Figure 1 As shown, the method for generating semantic vectors for videos may specifically include the following steps:
[0050] Step S11: Obtain the sample video and the user behavior logs of the sample video.
[0051] Sample videos refer to videos used to train a pre-defined model. User behavior logs for sample videos record various user behaviors related to the sample videos, including search behavior, browsing behavior, browsing duration, click behavior, playback behavior, and playback duration. These logs can record user behavior for all time periods or only for a specific time segment.
[0052] Step S12: Analyze the user behavior logs to obtain statistical values of various user behavior indicators for the sample video.
[0053] Analyzing user behavior logs of sample videos can yield statistical values for various user behavior metrics, which may include average dwell time, retention rate within a fixed number of seconds, and the percentage of users whose play rate exceeds a fixed threshold.
[0054] The average dwell time of a video is the ratio of the length of time a user spends watching the video to the number of users who have watched the video. The fixed-second retention rate of a video is the ratio of the number of users who spend more than a fixed number of seconds watching the video to the number of users who have watched the video. The percentage of users whose play rate exceeds a fixed threshold is the ratio of the length of time they play the video to the total length of the video to a fixed threshold to the number of users who have watched the video.
[0055] Step S13: Input the statistical values of the various user behavior indicators and the video frame sequence of the sample video into the preset model to be trained to obtain the semantic vector of the sample video output by the preset model to be trained.
[0056] The process involves extracting video frame sequences from sample videos and inputting these sequences, along with statistical values of various user behavior metrics, into a pre-defined model. This model extracts features from the video frame sequences and outputs semantic vectors for the sample videos based on these features. The model is then trained using the statistical values of various user behavior metrics as optimization objectives. This trained model is used to generate semantic vectors for videos; inputting a video into the trained model yields the semantic vectors for that video.
[0057] The technical solution of this application involves acquiring sample videos and user behavior logs of those videos; analyzing the user behavior logs to obtain statistical values of various user behavior indicators for the sample videos; and inputting the statistical values of these indicators and the video frame sequence of the sample videos into a preset model to be trained, thereby obtaining the semantic vector of the sample videos output by the preset model. In this way, training the preset model using video frame sequences allows the model to fully utilize the content information of the videos, resulting in highly interpretable semantic vectors. Simultaneously, training the model using statistical values of multiple user behavior indicators avoids the potential inaccuracies of statistical values from a single user behavior indicator or the hindering effect of overly simplistic user behavior indicators on the performance of the trained model, making the semantic vectors of the videos output by the model more accurate. Furthermore, training the model using statistical values of multiple user behavior indicators helps downstream users perform tasks related to user behavior indicators using the semantic vectors of videos, ensuring a better match between the semantic vectors of the videos used and those generated by the preset model.
[0058] Optionally, based on the above technical solution, before analyzing the user behavior logs to obtain the statistical values of the user behavior indicators of the sample video, the method further includes: determining multiple user behavior indicators of the sample video based on the user behavior information required by the application end of the sample video.
[0059] To align the semantic vectors of the generated video with those of the video required by downstream tasks, the user behavior metrics of the sample video to be acquired can be determined based on the user behavior information required by the downstream tasks. The application end of the sample video is the terminal or server that executes the downstream task.
[0060] For example, the application terminal of the sample video is a terminal that sorts videos by click-through rate and search volume. Since the business performed by the application terminal of the sample video is to sort videos by click-through rate and search volume, when using the statistical values of user behavior indicators of the sample video to train the preset model, the video click-through rate and search volume can be used to train the preset model.
[0061] In this way, the semantic vectors of the videos generated by the pre-trained model are more consistent with the semantic vectors of the videos required by the business processes executed by the application. When using the semantic vectors of the videos generated by the pre-trained model for downstream tasks, the accuracy of the downstream tasks can be improved, and the semantic gap between the semantic vectors used by the downstream tasks and the semantic vectors generated by the pre-trained model can be avoided.
[0062] Optionally, based on the above technical solution, training the model using statistical values of multiple user behavior indicators and video frame sequences of sample videos may include the following steps: inputting the statistical values of the multiple user behavior indicators and the video frame sequences of the sample videos into a preset model to be trained to obtain the predicted values of the multiple user behavior indicators output by the preset model to be trained; obtaining a first loss function value based on the statistical values of the multiple user behavior indicators and the corresponding predicted values; and updating the model parameters of the preset model to be trained based on the first loss function value.
[0063] The statistical values of various user behavior indicators and the video frame sequences of sample videos are input into the preset model to be trained. The preset model extracts the features of the video frame sequences, determines the semantic vector of the sample video based on the features of the video frame sequences, and determines the predicted values corresponding to the statistical values of various user behavior indicators based on the features of the video frame sequences.
[0064] The preset model to be trained includes an MMOE (Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts) module. Using MMOE can avoid the impact of mutual interference between the statistical values of multiple user behavior indicators on model training.
[0065] The MMOE module of the pre-defined model to be trained is followed by multiple two-layer MLP (Multilayer Perceptron) modules. Each two-layer MLP module can predict the value of a user behavior metric. The number of two-layer MLP modules connected after the MMOE module is consistent with the number of statistical values of the user behavior metrics used.
[0066] Based on the differences between the statistical values and corresponding predicted values of various user behavior indicators, a first loss function is established, and its value is obtained. With the objective of minimizing the differences between the statistical values and corresponding predicted values of various user behavior indicators, a pre-defined model is trained based on the first loss function value. The parameters of the pre-defined model are then updated, resulting in a trained pre-defined model. This trained pre-defined model is used to generate semantic vectors for videos. By inputting a video into the trained pre-defined model, the semantic vectors of the video can be obtained.
[0067] Optionally, based on the above technical solutions, in order to make the preset model have a certain degree of interpretability and play a certain role in regularization, a self-supervised learning method can also be used as an auxiliary training method for the preset model.
[0068] Strong and weak data augmentation were performed on the video frame sequences of the sample video to obtain strong and weak video frame sequences, respectively. Weak data augmentation can involve random spatial cropping, random horizontal flipping, or uniform scaling to a fixed scale. Strong data augmentation can involve random spatial cropping, random temporal cropping (i.e., randomly selecting a playback starting point), random horizontal flipping, random color enhancement (including brightness, contrast, and saturation), random grayscale processing, and random Gaussian blurring.
[0069] The weak video frame sequence, the strong video frame sequence, and statistical values of various user behavior indicators are input into the preset model to be trained. The vector generation module of the preset model to be trained generates semantic vectors F for the weak video frame sequence based on the weak video frame sequence and the strong video frame sequence, respectively. w The semantic vector F of a strong video frame sequence s The preset model's index prediction module predicts the index based on the semantic vector F of the weak video frame sequence. w The system predicts the values of multiple user behavior indicators and derives the first loss function value based on the differences between the predicted values and their corresponding statistical values. Considering the significant differences between the strong video frame sequence and the sample video's video frame sequence, the predicted values of multiple user behavior indicators using the strong video frame sequence would be inaccurate in reflecting the predicted values of user behavior indicators from the sample video. Therefore, the strong video frame sequence is not used to predict the values of multiple user behavior indicators.
[0070] The self-supervised module of the pre-defined model to be trained obtains semantic vectors for self-supervised learning from the semantic vectors of weak and strong video frame sequences, respectively. These include weakly self-supervised semantic vectors for weak video frame sequences and strongly self-supervised semantic vectors for strong video frame sequences. Based on the difference between the weakly and strongly self-supervised semantic vectors obtained from self-supervised learning, a second loss function is established, and the value of the second loss function is obtained.
[0071] Based on the first loss function value and the second loss function value, the model parameters of the preset model to be trained are updated to obtain the trained preset model.
[0072] The technical solution of this application embodiment, in addition to using the first loss function value to train the model, also uses the second loss function value to assist in training the model, so that the trained preset model has better performance.
[0073] Optionally, based on the above technical solution, Figure 2 The diagram illustrates the structure of a pre-defined model to be trained. This model includes a vector generation module, an indicator prediction module concatenated after the vector generation module, and a self-supervised module concatenated after and parallel to the indicator prediction module. The indicator prediction module outputs predicted values of various user behavior indicators, and the self-supervised module outputs semantic vectors for the strong and weak video frame sequences. The indicator prediction module may include the aforementioned MMOE module and multiple two-layer MLP modules; the self-supervised module may include two weight-shared two-layer MLP modules. A first loss function is established based on the predicted values of various user behavior indicators output by the indicator prediction module and the statistical values of these indicators input to the pre-defined model. A second loss function is established based on the semantic vectors for self-supervised learning from the strong and weak video frame sequences output by the self-supervised module.
[0074] During the training process, the preset model will gradually update the model parameters in the vector generation module. After training, the indicator prediction module and the self-supervised module can be removed from the preset model, and the trained vector generation module can be used as the semantic vector generation model. Figure 3 The diagram illustrates the structure of a semantic vector generation model. By inputting a sequence of video frames into the semantic vector generation model, the semantic vectors of the video output by the model can be obtained.
[0075] Optionally, a video frame extraction module can be added to the semantic vector generation model. The video is directly input into the semantic vector generation model, the video frame extraction module extracts the video frame sequence, and the vector generation module generates the semantic vector of the video based on the video frame sequence extracted by the video frame extraction module.
[0076] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0077] Figure 4 This is a schematic diagram of the structure of a device for generating semantic vectors of video according to an embodiment of the present invention, as shown below. Figure 4 As shown, an apparatus for generating semantic vectors from videos includes a sample video acquisition module, a statistical value acquisition module, and a semantic vector acquisition module, wherein:
[0078] The sample video acquisition module is used to acquire sample videos and user behavior logs of the sample videos;
[0079] The statistical value acquisition module is used to analyze the user behavior logs and obtain statistical values of various user behavior indicators for the sample video.
[0080] The semantic vector acquisition module is used to input the statistical values of the various user behavior indicators and the video frame sequence of the sample video into the preset model to be trained, and obtain the semantic vector of the sample video output by the preset model to be trained.
[0081] Optionally, as an embodiment, before analyzing the user behavior logs to obtain statistical values of the user behavior metrics for the sample video, the device further includes:
[0082] The behavior indicator determination module is used to determine various user behavior indicators of the sample video based on the user behavior information required by the application end of the sample video.
[0083] Optionally, as an embodiment, the semantic vector acquisition module includes:
[0084] The prediction value acquisition unit is used to input the statistical values of the various user behavior indicators and the video frame sequence of the sample video into the preset model to be trained, and obtain the predicted values of the various user behavior indicators output by the preset model to be trained.
[0085] The first loss function generation unit is used to obtain the first loss function value based on the statistical values and corresponding predicted values of the various user behavior indicators;
[0086] The model parameter update unit is used to update the model parameters of the preset model to be trained based on the first loss function value.
[0087] Optionally, as an embodiment, the apparatus further includes:
[0088] The enhancement module is used to perform strong data enhancement and weak data enhancement on the video frame sequence of the sample video to obtain strong video frame sequence and weak video frame sequence, respectively.
[0089] The semantic vector acquisition module includes:
[0090] The prediction value acquisition unit is used to input the statistical values of the various user behavior indicators and the weak video frame sequence into the preset model to be trained, and obtain the predicted values of the various user behavior indicators output by the preset model to be trained.
[0091] The sequence semantic vector obtaining unit is used to input the strong video frame sequence and the weak video frame sequence into the preset model to be trained, and obtain the semantic vectors of the strong video frame sequence and the weak video frame sequence output by the preset model to be trained.
[0092] The first loss function obtaining unit is used to obtain the first loss function value based on the statistical values and corresponding predicted values of the various user behavior indicators;
[0093] The second loss function obtaining unit is used to obtain the second loss function value from the semantic vector of the weak video frame sequence and the semantic vector of the weak video frame sequence.
[0094] The model update unit is used to update the model parameters of the preset model to be trained based on the first loss function value and the second loss function value.
[0095] Optionally, as an embodiment, the preset model to be trained includes: a vector generation module, an indicator prediction module concatenated after the vector generation module, and a self-supervised module concatenated after the vector generation module and arranged in parallel with the indicator prediction module. The indicator prediction module is used to output the predicted values of the various user behavior indicators, and the self-supervised module is used to output the semantic vectors of the strong video frame sequence and the weak video frame sequence, respectively. After the preset model to be trained is completed, the device further includes:
[0096] The model generation module is used to take the trained vector generation module as a semantic vector generation model.
[0097] The video acquisition module is used to acquire the target video;
[0098] The semantic vector generation module is used to input the target video into the semantic vector generation model and obtain the semantic vector of the target video output by the semantic vector generation model.
[0099] It should be noted that the device embodiments are similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the method embodiments.
[0100] This invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method for generating semantic vectors of video disclosed in this application.
[0101] This invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the method for generating semantic vectors for videos disclosed in this application.
[0102] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0103] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0107] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0108] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0109] The present application provides a detailed description of a method, apparatus, device, and medium for generating semantic vectors of videos. Specific examples have been used to illustrate the principles and implementation methods of the present application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present application. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present application. Therefore, the content of this specification should not be construed as a limitation of the present application.
Claims
1. A method of generating semantic vectors for a video, the method comprising: The method comprises: obtaining a sample video and a user behavior log of the sample video; analyzing the user behavior log to obtain statistical values of a plurality of user behavior indicators of the sample video; inputting the statistical values of the plurality of user behavior indicators and a video frame sequence of the sample video into a preset model to be trained to obtain a semantic vector of the sample video output by the preset model to be trained; The method further comprises: performing strong data enhancement and weak data enhancement on the video frame sequence of the sample video to obtain a strong video frame sequence and a weak video frame sequence, respectively; inputting the statistical values of the plurality of user behavior indicators and the video frame sequence of the sample video into the preset model to be trained to obtain a semantic vector of the sample video output by the preset model to be trained, comprising: inputting the statistical values of the plurality of user behavior indicators and the weak video frame sequence into the preset model to be trained to obtain predicted values of the plurality of user behavior indicators output by the preset model to be trained; inputting the strong video frame sequence and the weak video frame sequence into the preset model to be trained to obtain respective semantic vectors of the strong video frame sequence and the weak video frame sequence output by the preset model to be trained; obtaining a first loss function value according to the statistical values of the plurality of user behavior indicators and corresponding predicted values; obtaining a second loss function value according to the semantic vector of the weak video frame sequence and the semantic vector of the weak video frame sequence; updating model parameters of the preset model to be trained according to the first loss function value and the second loss function value.
2. The method of claim 1, wherein, The method further comprises: determining the plurality of user behavior indicators of the sample video according to user behavior information required by an application end of the sample video.
3. The method of claim 1, wherein, inputting the statistical values of the plurality of user behavior indicators and the video frame sequence of the sample video into the preset model to be trained to obtain a semantic vector of the sample video output by the preset model to be trained, comprising: inputting the statistical values of the plurality of user behavior indicators and the weak video frame sequence into the preset model to be trained to obtain predicted values of the plurality of user behavior indicators output by the preset model to be trained; obtaining a first loss function value according to the statistical values of the plurality of user behavior indicators and corresponding predicted values; updating model parameters of the preset model to be trained according to the first loss function value.
4. The method according to any of claims 1 to 3, characterized in that, The preset model to be trained comprises a vector generation module, an indicator prediction module connected in series after the vector generation module, and a self-supervised module connected in series after the vector generation module and arranged in parallel with the indicator prediction module, the indicator prediction module being configured to output predicted values of the plurality of user behavior indicators, and the self-supervised module being configured to output respective semantic vectors of the strong video frame sequence and the weak video frame sequence; after the preset model to be trained is trained, the method further comprises: using the trained vector generation module as a semantic vector generation model; obtaining a target video; input the target video into the semantic vector generation model to obtain a semantic vector of the target video output by the semantic vector generation model.
5. An apparatus for generating semantic vectors of a video, the apparatus comprising: The device comprises: a sample video acquisition module configured to acquire a sample video and a user behavior log of the sample video; a statistical value acquisition module configured to analyze the user behavior log to obtain statistical values of a plurality of user behavior indicators of the sample video; a semantic vector acquisition module configured to input the statistical values of the plurality of user behavior indicators and a video frame sequence of the sample video into a preset model to be trained to obtain a semantic vector of the sample video output by the preset model to be trained; The steps further comprise: performing strong data enhancement and weak data enhancement on the video frame sequence of the sample video to obtain a strong video frame sequence and a weak video frame sequence, respectively; inputting the statistical values of the plurality of user behavior indicators and the video frame sequence of the sample video into the preset model to be trained to obtain a semantic vector of the sample video output by the preset model to be trained, comprising: inputting the statistical values of the plurality of user behavior indicators and the weak video frame sequence into the preset model to be trained to obtain predicted values of the plurality of user behavior indicators output by the preset model to be trained; inputting the strong video frame sequence and the weak video frame sequence into the preset model to be trained to obtain respective semantic vectors of the strong video frame sequence and the weak video frame sequence output by the preset model to be trained; obtaining a first loss function value according to the statistical values of the plurality of user behavior indicators and corresponding predicted values; obtaining a second loss function value according to the semantic vector of the weak video frame sequence and the semantic vector of the weak video frame sequence; updating model parameters of the preset model to be trained according to the first loss function value and the second loss function value.
6. The apparatus of claim 5, wherein, Before analyzing the user behavior log to obtain statistical values of user behavior indicators of the sample video, the device further comprises: a behavior indicator determination module configured to determine a plurality of user behavior indicators of the sample video according to user behavior information of an application end requirement of the sample video.
7. The apparatus of claim 5, wherein, The semantic vector acquisition module comprises: a predicted value acquisition unit configured to input the statistical values of the plurality of user behavior indicators and the video frame sequence of the sample video into the preset model to be trained to obtain predicted values of the plurality of user behavior indicators output by the preset model to be trained; a first loss function generation unit configured to obtain a first loss function value according to the statistical values of the plurality of user behavior indicators and corresponding predicted values; a model parameter updating unit configured to update model parameters of the preset model to be trained according to the first loss function value.
8. An electronic device, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method for generating a semantic vector of a video according to any one of claims 1 to 4.
9. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enable the electronic device to perform the method of Claim 1 to 4 for generating a semantic vector of a video.
Citation Information
Patent Citations
Video recommendation method and device
CN110717069A
Video classification method, device and equipment and readable storage medium
CN114299321A