Video information processing method, device, computer equipment and storage medium

By performing topic mapping, fusion and adjustment on content features of multiple description dimensions during video processing, multimodal content description information is generated, which solves the problem of inaccurate video topic description information in the existing technology and improves the accuracy of video recommendation.

CN115114477BActive Publication Date: 2025-09-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210751947.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-09-26
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

The accuracy of video topic description information generated by existing technologies is low, resulting in a decrease in the accuracy of video recommendations.

Method used

By obtaining the descriptive content features of the video to be processed in multiple different description dimensions, performing topic mapping and fusion processing, generating multimodal descriptive content features, and adjusting the initial video topic features, the multimodal content description information is finally decoded.

Benefits of technology

The accuracy of video recommendations is improved and more accurate topic description information is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114477B_ABST
    Figure CN115114477B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a video information processing method, apparatus, computer equipment and storage medium; the embodiments of the present application can obtain descriptive content features of a video to be processed in multiple different descriptive dimensions; perform topic mapping on the descriptive content features in each descriptive dimension to obtain initial video topic features of the video to be processed in multiple different descriptive dimensions; fuse the descriptive content features of multiple different descriptive dimensions to obtain multimodal descriptive content features; based on the multimodal descriptive content features, adjust the initial video topic features corresponding to each descriptive dimension to obtain target video topic features of multiple different descriptive dimensions of the video to be processed; decode the target video topic features of different descriptive dimensions and the multimodal descriptive content features to obtain multimodal content description information of the video to be processed, and generate accurate subject description information, thereby improving the accuracy of video recommendations for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for processing video information. Background Art

[0002] With the development of computer technology, multimedia applications have become increasingly widespread, and the number of videos has also increased dramatically. To facilitate users to quickly find the videos they want to watch from the vast amount of videos, many video websites and video applications typically generate topic descriptions for videos and then match this topic description with user-entered search keywords to recommend videos to users. The inventors of this application discovered through practical research on existing technologies that these technologies suffer from low accuracy in generating topic descriptions for videos, thereby reducing the accuracy of video recommendations for users. Summary of the Invention

[0003] The embodiments of the present application propose a video information processing method, apparatus, computer device, and storage medium, which can generate accurate subject description information, thereby improving the accuracy of video recommendations for users.

[0004] The present invention provides a method for processing video information, including:

[0005] Obtaining description content features of the video to be processed in multiple different description dimensions;

[0006] Performing topic mapping on the description content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions;

[0007] fusing the description content features of the multiple different description dimensions to obtain a multimodal description content feature;

[0008] Based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted to obtain target video theme features of multiple different description dimensions for the video to be processed;

[0009] The target video theme features of the different description dimensions and the multimodal description content features are decoded to obtain multimodal content description information of the video to be processed.

[0010] Accordingly, an embodiment of the present application further provides a video information processing device, including:

[0011] An acquisition unit, configured to acquire description content features of a video to be processed in multiple different description dimensions;

[0012] A topic mapping unit, configured to perform topic mapping on the description content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions;

[0013] a fusion unit, configured to fuse the description content features of the multiple different description dimensions to obtain a multimodal description content feature;

[0014] An adjustment unit, configured to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features, to obtain target video theme features for a plurality of different description dimensions of the video to be processed;

[0015] The decoding unit is used to decode the target video theme features of the different description dimensions and the multimodal description content features to obtain the multimodal content description information of the video to be processed.

[0016] In one embodiment, the topic mapping unit may include:

[0017] A first mapping subunit is configured to map the description content features on each description dimension to a preset subject feature space, and obtain feature distribution information of the description content features on each description dimension in the preset subject feature space;

[0018] A probability calculation subunit, configured to calculate probability fitting information of each description content feature based on feature distribution information of the description content feature on each description dimension in a preset subject feature space;

[0019] The feature determination subunit is used to determine the initial video theme features of the video to be processed in multiple different description dimensions based on the probability fitting information of each description content feature.

[0020] In one embodiment, the fusion unit may include:

[0021] A convolution operation subunit, configured to perform a convolution operation on the description content features of multiple different description dimensions to obtain a convolution operation result corresponding to the description content feature of each description dimension;

[0022] The cross attention fusion subunit is used to perform cross attention fusion on the convolution operation results of each description dimension to obtain the cross attention features corresponding to each description dimension;

[0023] The first splicing subunit is used to splice the cross-attention features corresponding to each description dimension to obtain the spliced ​​features;

[0024] The fully connected subunit is used to perform fully connected processing on the spliced ​​features to obtain the multimodal description content features.

[0025] In one embodiment, the adjusting unit may include:

[0026] A second splicing subunit is configured to splice the multimodal content feature with the initial video theme feature corresponding to the current description dimension to obtain a spliced ​​feature corresponding to the current description dimension;

[0027] A second mapping subunit is configured to map the concatenated features corresponding to the current description dimension using the topic mapping information corresponding to the current description dimension to obtain mapped features;

[0028] A nonlinear conversion processing subunit, configured to perform nonlinear conversion processing on the mapped features to obtain nonlinear converted features;

[0029] The feature fusion subunit is used to fuse the nonlinearly converted features with the initial video theme features to obtain the target video theme features.

[0030] In one embodiment, the decoding unit may include:

[0031] a third splicing subunit, configured to splice the target video theme features of different description dimensions and the multimodal description content features to obtain a spliced ​​content feature, wherein the spliced ​​content feature is composed of a multidimensional content vector;

[0032] An information prediction subunit, configured to perform information prediction on the content vector of the concatenated content feature to obtain an information unit corresponding to each dimension of the content vector;

[0033] The combining subunit is configured to combine the information units to obtain the multimodal content description information.

[0034] In one embodiment, the information processing device may further include:

[0035] a content acquisition unit, configured to acquire description content of at least one video segment of a video to be processed in multiple different description dimensions;

[0036] a feature extraction unit, configured to extract features of the description content of the video clip in a plurality of different description dimensions, and obtain initial description content features of the video clip in the plurality of different description dimensions;

[0037] The feature normalization mapping unit is used to perform feature normalization mapping on the initial description content features of the video clips in multiple different description dimensions to obtain the description content features of at least one video clip of the video to be processed in multiple different description dimensions.

[0038] In one embodiment, the information processing device may further include:

[0039] A model acquisition unit, used to obtain descriptions of the information processing model to be trained and the video samples in multiple different description dimensions;

[0040] a processing unit, configured to process the description content of the video sample in multiple different description dimensions using the information processing model to be trained, to obtain description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample;

[0041] a loss calculation unit, configured to calculate target loss information based on description content features of the video sample in a plurality of different description dimensions, target video subject features of the video sample in a plurality of different description dimensions, and multimodal content description information of the video sample;

[0042] A training unit is used to train the information processing model to be trained using the target loss information to obtain the information processing model.

[0043] In one embodiment, the loss calculation unit may include:

[0044] A first loss calculation subunit, configured to calculate description content feature loss information based on description content features of the video sample in a plurality of different description dimensions;

[0045] A second loss calculation subunit is configured to calculate topic mapping constraint loss information based on the description content features of the video sample in multiple description dimensions and the multimodal content description information;

[0046] A third loss calculation subunit is configured to calculate topic feature loss information based on target video topic features of the video sample in a plurality of different description dimensions;

[0047] a fourth loss calculation subunit, configured to calculate multimodal loss information based on the multimodal content description information;

[0048] The loss fusion subunit is used to fuse the description content feature loss information, the topic mapping constraint loss information, the topic feature loss information and the multimodal loss information to obtain the target loss information.

[0049] In one embodiment, the second loss calculation subunit may include:

[0050] The information acquisition module is used to obtain the latent space information corresponding to the information processing model to be trained;

[0051] a distribution calculation module, configured to perform distribution calculation on the latent space information, the descriptive content features, and the multimodal content description information to obtain a latent space distribution corresponding to the latent space information, a descriptive content feature distribution corresponding to the descriptive content features, and a multimodal content description information distribution corresponding to the multimodal content description information;

[0052] A loss calculation module, configured to calculate distribution loss information between the latent space distribution and the content description feature distribution;

[0053] A statistical operation module, configured to perform statistical operations on the distribution of the multimodal content description information to obtain statistical information;

[0054] An arithmetic operation module is used to perform an arithmetic operation on the distribution loss information and the statistical information to obtain the topic mapping constraint loss information.

[0055] In one embodiment, the third loss calculation subunit may include:

[0056] A first exponentiation module, configured to perform an exponentiation operation using the target video theme features in the multiple different description dimensions as exponents and a preset base to obtain first exponentiation information;

[0057] A second exponentiation module is configured to perform an exponentiation operation using the target video theme feature and the target video theme feature corresponding to the negative sample as an exponent and a preset base to obtain information after the second exponentiation operation;

[0058] A comparison operation module is used to perform a comparison operation on the first information after the power operation and the second information after the power operation to obtain the topic feature loss information.

[0059] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in various optional embodiments of the above-mentioned aspect.

[0060] Correspondingly, an embodiment of the present application further provides a storage medium, which stores instructions, and when the instructions are executed by a processor, implements any video information processing method provided in the embodiment of the present application.

[0061] The embodiment of the present application can obtain the descriptive content features of the video to be processed in multiple different description dimensions; perform topic mapping on the descriptive content features in each description dimension to obtain the initial video topic features of the video to be processed in multiple different description dimensions; fuse the descriptive content features of multiple different description dimensions to obtain multimodal descriptive content features; based on the multimodal descriptive content features, adjust the initial video topic features corresponding to each description dimension to obtain target video topic features for multiple different description dimensions of the video to be processed; decode the target video topic features of different description dimensions and the multimodal descriptive content features to obtain multimodal content description information of the video to be processed. The embodiment of the present application can generate accurate subject description information, thereby improving the accuracy of video recommendations for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0063] Figure 1 Schematic diagram of a video information processing method according to an embodiment of the present application;

[0064] Figure 2 1 is a flow chart of a method for processing video information provided by an embodiment of the present application;

[0065] Figure 3 This is another scenario diagram of the video information processing method provided in an embodiment of the present application;

[0066] Figure 4 This is another scenario diagram of the video information processing method provided in an embodiment of the present application;

[0067] Figure 5 This is another flowchart of the video information processing method provided in an embodiment of the present application;

[0068] Figure 6 Schematic diagram of the structure of the video information processing device provided in an embodiment of the present application;

[0069] Figure 7 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. However, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0071] The embodiments of the present application provide a method for processing video information. The method can be performed by a video information processing device, which can be integrated into a computer device. The computer device can include at least one of a terminal and a server. In other words, the method can be performed by a terminal, a server, or a terminal and a server that can communicate with each other.

[0072] Among them, terminals may include but are not limited to smartphones, tablets, laptops, personal computers (PCs), smart home appliances, wearable electronic devices, VR / AR devices, vehicle-mounted terminals, intelligent voice interaction devices, etc.

[0073] The server can be an intercommunication server or background server between multiple heterogeneous systems, an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms, etc.

[0074] It should be noted that the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0075] In one embodiment, if Figure 1As described, the video information processing device can be integrated into a computer device such as a terminal or a server to implement the video information processing method proposed in the embodiment of the present application. Specifically, the server 11 can obtain the descriptive content features of the video to be processed in multiple different descriptive dimensions; perform topic mapping on the descriptive content features on each descriptive dimension to obtain the initial video topic features of the video to be processed in multiple different descriptive dimensions; fuse the descriptive content features of multiple different descriptive dimensions to obtain multimodal descriptive content features; based on the multimodal descriptive content features, adjust the initial video topic features corresponding to each descriptive dimension to obtain target video topic features of multiple different descriptive dimensions for the video to be processed; decode the target video topic features of different descriptive dimensions and the multimodal descriptive content features to obtain the multimodal content description information of the video to be processed. Then, the terminal 10 can match the search keywords input by the user with the multimodal content description information of the video, thereby screening out videos of interest to the user and recommending them to the user.

[0076] The following are detailed descriptions of each embodiment. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0077] The embodiments of the present application will be described from the perspective of a video information processing device, which can be integrated into a computer device, which can be a server, a terminal, or other device.

[0078] like Figure 2 A video information processing method is provided, and the specific process includes:

[0079] 101. Obtain description content features of the video to be processed in multiple different description dimensions.

[0080] In one embodiment, a video is generally composed of content in multiple different description dimensions. For example, a video is composed of multiple video frames and audio information. The multiple video frames may be described in the image description dimension, while the audio information may be described in the audio description dimension. Furthermore, if the audio information includes conversational content, the conversational content in the audio information can be converted into textual information, thereby obtaining a description of the video in the text description dimension.

[0081] In one embodiment, feature extraction may be performed on the description content of the to-be-processed video in each description dimension to obtain description content features of the to-be-processed video in multiple different description dimensions.

[0082] There are multiple methods for extracting content features of the video to be processed at different description dimensions, and obtaining description content features of the video to be processed at multiple different description dimensions.

[0083] For example, artificial intelligence algorithms such as machine learning or deep learning can be used to extract content features of the video to be processed at different description dimensions to obtain descriptive content features of the video to be processed at multiple different description dimensions.

[0084] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0085] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0086] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration. Reinforcement learning is a field within machine learning that emphasizes how to act based on the environment to maximize expected benefits. Deep reinforcement learning combines deep learning and reinforcement learning, applying deep learning techniques to solve reinforcement learning problems.

[0087] For example, the source encoder can be an artificial intelligence algorithm such as Convolutional Neural Networks (CNN), De-Convolutional Neural Networks (DN), Deep Neural Networks (DNN), Deep Convolutional Inverse Graphics Networks (DCIGN), Region-based Convolutional Networks (RCNN), Self-Attentive Sequential Recommendation (SASRec), Faster Region-based Convolutional Networks (Faster RCNN) and Bidirectional Encoder Representations from Transformers (BERT) model to extract features of the description content of the processed video in each description dimension, and obtain the description content features of the processed video in multiple different description dimensions.

[0088] For another example, in an embodiment of the present application, multimodal content description information of a video to be processed is generated by combining the descriptive content features of the video to be processed in multiple different description dimensions. In order to improve the confidence of the multimodal content description information, feature extraction can be performed on the descriptive content of the video to be processed in multiple different description dimensions to obtain initial descriptive content features of the video to be processed in multiple different description dimensions. Then, feature normalization mapping is performed on the initial descriptive content features of the video to be processed in multiple different dimensions to obtain descriptive content features of at least one video segment of the video to be processed in multiple different description dimensions.

[0089] Specifically, the step of "obtaining description content features of the video to be processed in multiple different description dimensions" may include:

[0090] Obtain description content of the video to be processed in multiple different description dimensions;

[0091] Extract features of the description content of the processed video in multiple different description dimensions to obtain initial description content features of the video clip in multiple different description dimensions;

[0092] The initial description content features of the video to be processed in multiple different description dimensions are subjected to feature normalization mapping to obtain the description content features of the video to be processed in multiple different description dimensions.

[0093] Among them, feature normalization mapping of the initial description content features of the processed video at multiple different description dimensions can refer to mapping the initial description content features at different description dimensions into the same feature space, thereby solving the semantic gap problem between the initial description content features at different description dimensions.

[0094] In one embodiment, after obtaining the video to be processed, content extraction may be performed on the video to be processed to obtain description content of the video to be processed in multiple different description dimensions.

[0095] For example, audio extraction tools can be used to extract audio information from the video being processed, thereby obtaining descriptions in the audio description dimension. Automatic speech recognition (ASR) technology can then be used to convert the audio descriptions into text descriptions. Furthermore, the video being processed can be split into multiple frames to obtain descriptions in the image description dimension.

[0096] In one embodiment, feature extraction may be performed on the description content of the video to be processed in multiple different description dimensions to obtain initial description content features of the video clip in the multiple different description dimensions.

[0097] Specifically, different methods can be used to extract features for descriptions of different description dimensions. For example, for descriptions of image descriptions, a trained separable 3D convolutional neural network (S3D) can be used to extract features to obtain initial description features of the image description dimension. For another example, for descriptions of text descriptions, a trained BERT network can be used to extract features to obtain initial description features of the text dimension.

[0098] In one embodiment, feature normalization mapping may be performed on the initial description content features of the video to be processed at multiple different description dimensions to obtain the description content features of the video to be processed at multiple different description dimensions.

[0099] For example, a feature normalization mapping function can be defined to construct a semantic space. This feature normalization function can then be used to map the initial description content features at different description dimensions into the same semantic space, thereby obtaining description content features of the processed video at multiple different description dimensions to address the semantic gap. The feature normalization mapping function can be a multilayer perceptron (MLP).

[0100] In one embodiment, to improve the accuracy of generating multimodal content description information for a video to be processed, when the audio information in the video to be processed includes language, the video to be processed can be split into multiple video segments based on the audio information. Then, descriptive content features of at least one video segment of the video to be processed in multiple different descriptive dimensions can be obtained, based on the video segments. Specifically, the step of "obtaining descriptive content features of the video to be processed in multiple different descriptive dimensions" can include:

[0101] Obtain description content features of at least one video segment of a to-be-processed video in multiple different description dimensions.

[0102] For example, the video to be processed may be split into n (n is a positive integer greater than 1) video segments, and then the description content features of each video segment of the video to be processed in multiple different description dimensions are obtained.

[0103] In one embodiment, the step of “obtaining descriptive content features of at least one video segment of the video to be processed in multiple different description dimensions” may include:

[0104] Obtaining description content of at least one video segment of a video to be processed in multiple different description dimensions;

[0105] Extracting features of the description content of the video clip in multiple different description dimensions to obtain initial description content features of the video clip in multiple different description dimensions;

[0106] Perform feature normalization mapping on the initial description content features of the video clips in multiple different description dimensions to obtain description content features of at least one video clip of the video to be processed in multiple different description dimensions.

[0107] In one embodiment, after obtaining the video to be processed, an ASR tool can be used to extract the audio information from the video to be processed and convert the audio information into text information. When the ASR tool is used to convert the audio information into text information, the text information can also include the time information corresponding to the text information in the video to be processed. The video to be processed can then be cut into several video segments based on the time information corresponding to the text information in the video to be processed.

[0108] For example, the video to be processed is a one-minute variety show clip. After using an ASR tool to convert the audio information in the video to text, the resulting text includes text messages such as "Let's play a game," "The rules of the game are XXX," and "The winners of the game are Xiaoming and Xiaohong." The time information for "Let's play a game" runs from the 0th to the 3rd second, the time information for "The rules of the game are XXX" runs from the 4th to the 15th second, and the time information for "The winners of the game are Xiaoming and Xiaohong" runs from the 16th to the 30th second, and so on. The video can then be segmented into several segments based on the corresponding time information in the text messages. For example, the 0th to the 3rd second segment of the video to be processed is divided into the first segment, the 4th to the 15th second segment is divided into the second segment, and the 16th to the 30th second segment is divided into the third segment, and so on.

[0109] Next, feature extraction is performed on the descriptive content of each video clip across multiple different descriptive dimensions to obtain initial descriptive content features for the video clip across multiple different descriptive dimensions. Next, feature normalization mapping is performed on the initial descriptive content features of the video clip across multiple different descriptive dimensions to obtain descriptive content features for at least one video clip of the video to be processed across multiple different descriptive dimensions. The process of generating descriptive content features for the video clips can be referenced to the process of generating descriptive content features for the video to be processed and will not be fully described here.

[0110] In one embodiment, the information processing model can also be used to obtain the descriptive content features of the video to be processed in multiple different descriptive dimensions. In addition, the information processing model can also be used to obtain the descriptive content features of at least one video segment of the video to be processed in multiple different descriptive dimensions.

[0111] The information processing model can be a model built based on an artificial intelligence algorithm. For example, the information processing module can be a model built based on ASR, MLP, Transformer, DNN, and variational autoencoder. The information processing module can include a topic mapping module, an information encoding module, and an information decoding module.

[0112] The topic mapping module may include ASR, MLP, and variational autoencoder. The topic mapping module may be used to perform topic mapping on the description content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions.

[0113] The information encoding module may include a Transformer and a DNN. The information encoding module can be used to fuse descriptive content features from multiple different descriptive dimensions to generate multimodal descriptive content features. Furthermore, the information encoding module can be used to adjust the initial video theme features corresponding to each descriptive dimension based on the multimodal descriptive content features to generate target video theme features for the multiple different descriptive dimensions of the video being processed.

[0114] The information decoding module may include a Transformer, which can be used to decode target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video to be processed.

[0115] 102. Perform topic mapping on the description content features in each description dimension to obtain initial video topic features of the to-be-processed video in multiple different description dimensions.

[0116] In one embodiment, in order to improve the accuracy of the multimodal content description information of the video, the embodiment of the present application can perform topic mapping on the content features on each description dimension to obtain the initial video topic features of the video to be processed on multiple different description dimensions. Then, the description content features of multiple different description dimensions are fused to obtain multimodal description content features. By performing topic mapping on the content features on each description dimension, the content features on each description dimension can be mined more deeply to capture more fine-grained information. Then, the description content features of multiple different description dimensions are fused to obtain multimodal description content features, which can take into account global information and local information while mining local fine-grained features, not only to have an understanding from a global perspective, but also to facilitate the capture of some fine-grained objects, understand the occurrence of actions, and even enhance the semantic understanding of the video by capturing the relationship between objects.

[0117] In one embodiment, there are multiple methods for performing topic mapping on the description content features at each description dimension to obtain initial video topic features of the to-be-processed video at multiple different description dimensions.

[0118] For example, the information processing model can be used to perform topic mapping on the descriptive content features on each description dimension to obtain the initial video topic features of the video to be processed on multiple different description dimensions. For example, the topic mapping module in the information processing model can be used to perform topic mapping on the descriptive content features on each description dimension to obtain the initial video topic features of the video to be processed on multiple different description dimensions. For example, the topic mapping module can include a variational autoencoder, which can perform topic mapping on the descriptive content features on each description dimension to obtain the initial video topic features of the video to be processed on multiple different description dimensions.

[0119] Variational autoencoders are a variational version of autoencoders, excelling at modeling semantic spatial distributions and capturing high-level concepts. Specifically, a topic-based variational autoencoder first uses the encoder to model latent variables using input data. The latent vectors are then used to map content features to initial video topic features.

[0120] For another example, the step of "performing topic mapping on the description content features in each description dimension to obtain initial video topic features of the to-be-processed video in multiple different description dimensions" may include:

[0121] Mapping the description content features on each description dimension to a preset subject feature space to obtain feature distribution information of the description content features on each description dimension in the preset subject feature space;

[0122] Calculate the probability fitting information of each description content feature based on the feature distribution information of the description content feature on each description dimension in the preset subject feature space;

[0123] According to the probability fitting information of each descriptive content feature, the initial video theme features of the video to be processed in multiple different description dimensions are determined.

[0124] Among them, the preset theme feature space is a pre-trained space that constructs a mapping relationship between the description content features and the initial video theme features.

[0125] In one embodiment, a preset theme feature space matrix can be used to map the descriptive content features on each description dimension to the preset theme feature space, thereby obtaining feature distribution information of the descriptive content features on each description dimension in the preset theme feature space. For example, a dot product operation can be performed on the preset theme feature space matrix and the content features on each description dimension, thereby obtaining feature distribution information of the descriptive content features on each description dimension in the preset theme feature space.

[0126] Then, the probability fitting information of each descriptive content feature can be calculated based on the feature distribution information of the descriptive content feature in each descriptive dimension in the preset topic feature space. For example, the probability fitting information of each descriptive content feature can be calculated using a softmax function based on the feature distribution information of the descriptive content feature in each descriptive dimension in the preset topic feature space.

[0127] Then, the initial video theme features of the to-be-processed video in multiple different description dimensions can be determined based on the probability fitting information of each description content feature. The probability fitting information of the description content feature can indicate which feature best fits the description content feature.

[0128] In one embodiment, when the video to be processed is split into multiple video segments, the step of "performing topic mapping on the descriptive content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions" may include:

[0129] The description content features of each video clip in each description dimension are subjected to topic mapping to obtain the initial video topic features of each video clip in the video to be processed in multiple different description dimensions.

[0130] Among them, referring to the above method, multiple methods can also be used to perform topic mapping on the description content features of each description dimension of each video clip to obtain the initial video topic features of each video clip in multiple different description dimensions of the video to be processed.

[0131] For example, an information processing model can be used to perform topic mapping on the descriptive content features of each video clip at each description dimension to obtain initial video topic features for each video clip in the processed video at multiple different description dimensions. For example, the descriptive content features at different description dimensions include text content features and image content features. The variational autoencoder can include a first topic variational autoencoder and a second topic variational autoencoder.

[0132] Among them, the first topic variational autoencoder can use the encoder Model the latent variable zs through text content features. For example, the second topic variational autoencoder can use the encoder Modeling latent variables z through image content features v .

[0133] Then, the softmax function can be used to map it to the initial video topic features.

[0134] For example, the initial video topic features in the text description dimension can be expressed as:

[0135]

[0136] in, It can represent the initial video theme features corresponding to the i-th video clip in the text description dimension, It can represent the weight matrix.

[0137] It is worth noting that after being projected into a common space, the semantics between image content features and text content features can be considered measurable. Furthermore, considering that the central theme of a segment and the central theme of a sentence are necessarily consistent, a shared weight matrix is ​​used here to obtain the initial video topic features in the image description dimension. For example, the initial video topic features in the image description dimension can be expressed as:

[0138]

[0139] in, It can represent the initial video theme features corresponding to the i-th video clip in the image description dimension.

[0140] 103. The description content features of multiple different description dimensions are fused to obtain multimodal description content features.

[0141] In one embodiment, after obtaining initial video theme features of a video to be processed in multiple different description dimensions, in order to generate unified information describing the video theme for the video to be processed, the descriptive content features of the multiple different description dimensions can be fused to obtain a multimodal descriptive content feature. Then, a multimodal descriptive content feature is generated based on the multimodal descriptive content features.

[0142] In one embodiment, an information processing model may be used to fuse description content features of multiple different description dimensions to obtain multimodal description content features.

[0143] For example, the information encoding module in the information processing model can be used to fuse the description content features of multiple different description dimensions to obtain multimodal description content features. For example, the Transformer in the information processing module can be used to fuse the description content features of multiple different description dimensions to obtain multimodal description content features.

[0144] In one embodiment, the step of “fusing description content features of multiple different description dimensions to obtain multimodal description content features” may include:

[0145] Performing a convolution operation on the description content features of multiple different description dimensions to obtain a convolution operation result corresponding to the description content feature of each description dimension;

[0146] Perform cross-attention fusion on the convolution operation results of each description dimension to obtain the cross-attention features corresponding to each description dimension;

[0147] Concatenate the cross-attention features corresponding to each description dimension to obtain the concatenated features;

[0148] The concatenated features are fully connected to obtain multimodal description content features.

[0149] In one embodiment, a single convolution kernel can be used to perform convolution operations on descriptive content features of multiple different description dimensions. This convolution operation can fuse the information of the descriptive content features of each description dimension, thereby enabling the interaction of local information. For example, multiple one-dimensional convolution kernels can be used to perform convolution operations on descriptive content features of different description dimensions, with the final convolution kernel outputting the convolution operation results corresponding to the descriptive content features of each description dimension.

[0150] In one embodiment, the convolution operation results of each description dimension can be cross-attended and fused based on the attention mechanism to obtain the cross-attention features corresponding to each description dimension. By performing cross-attention fusion on the convolution operation results of each description dimension, the interaction between information of different description dimensions can be deepened. For example, the convolution operation results including three description dimensions can use the attention mechanism to operate on the convolution operation results of the first and second description dimensions. In addition, the attention mechanism can also be used to operate on the convolution operation results of the first and third description dimensions. Then, the two attention operation results corresponding to the first description dimension are spliced ​​to obtain the cross-attention features of the first description dimension. In this way, the cross-attention features of the first description dimension not only include its own features, but also integrate the features of the second description dimension and the third description dimension, thereby deepening the interaction between local information.

[0151] In one embodiment, the cross-attention features corresponding to each description dimension can be concatenated to obtain concatenated features. The concatenated features are then predicted to obtain multimodal description content features. For example, a fully connected layer can be used to predict the concatenated features to obtain multimodal description content features.

[0152] In one embodiment, when the video to be processed is divided into multiple video segments, the step of “fusing the description content features of multiple different description dimensions to obtain multimodal description content features” may include:

[0153] The description content features of the video clips of the video to be processed in multiple description dimensions are fused to obtain the multimodal description content features corresponding to each video clip.

[0154] For example, Figure 3 As shown in FIG, Transformer can be used to fuse the description content features of each video clip of the video to be processed in multiple description dimensions to obtain the multimodal description content features corresponding to each video clip.

[0155] For example, the multimodal description content features can be expressed as follows:

[0156]

[0157] in, It can represent the multimodal description content features of the i-th video clip.

[0158] Then, the multimodal description content features of each video clip can be spliced ​​together to obtain the multimodal description content features of the video to be processed.

[0159] 104. Based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted to obtain target video theme features of multiple different description dimensions for the video to be processed.

[0160] In one embodiment, the initial video theme features corresponding to each description dimension can be adjusted based on the multimodal description content features to obtain target video theme features for multiple different description dimensions of the video to be processed. Adjusting the initial video theme features corresponding to each description dimension based on the multimodal description content features can refer to optimizing the initial video theme features corresponding to each description dimension using the multimodal description content features, thereby improving the representational capability of the target video theme features, so that the target video theme features corresponding to each description dimension can more accurately express the theme of the video to be processed.

[0161] In one embodiment, the information processing model can be used to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features of multiple different description dimensions for the video to be processed. For example, the information encoding module in the information processing model can be used to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features of multiple different description dimensions for the video to be processed. For example, the DNN can be used to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features of multiple different description dimensions for the video to be processed, and so on.

[0162] In one embodiment, the step of “adjusting the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features for multiple different description dimensions of the video to be processed” may include:

[0163] The multimodal content features and the initial video theme features corresponding to the current description dimension are spliced ​​together to obtain the spliced ​​features corresponding to the current description dimension;

[0164] Using the topic mapping information corresponding to the current description dimension, the concatenated features corresponding to the current description dimension are mapped to obtain the mapped features;

[0165] Perform nonlinear transformation on the mapped features to obtain nonlinear transformed features;

[0166] The nonlinearly converted features and the initial video theme features are fused to obtain the target video theme features.

[0167] For example, suppose that the initial video topic features of the text description dimension and the initial video topic features of the image description dimension are included. The multimodal content features and the initial video topic features of the text description dimension can be spliced ​​together to obtain the spliced ​​features corresponding to the text description dimension. Then, the topic mapping information corresponding to the text description dimension can be used to map the spliced ​​features corresponding to the text description dimension to obtain the mapped features. The topic mapping information corresponding to the text description dimension can be the network parameters corresponding to the neural network corresponding to the text description dimension. For example, the DNN can be used to map the spliced ​​features corresponding to the current description dimension to obtain the mapped features. The DNN corresponding to the text description dimension can be used to map the spliced ​​features corresponding to the text description dimension to obtain the mapped features.

[0168] Then, the mapped features corresponding to the text description dimension can be subjected to nonlinear transformation processing to obtain nonlinear transformed features. For example, the mapped features can be subjected to nonlinear transformation processing using a nonlinear function to obtain nonlinear transformed features. For example, the mapped features can be subjected to nonlinear transformation processing using a softmax function, a Relu function, or a sigmoid function to obtain nonlinear transformed features.

[0169] Then, the nonlinearly transformed features and the initial video theme features can be fused to obtain the target video theme features corresponding to the current description dimension. For example, the nonlinearly transformed features of the text description dimension and the initial video theme features can be multiplied to obtain the target video theme features corresponding to the text description dimension.

[0170] In one embodiment, when the video to be processed is split into multiple video segments, the step of "adjusting the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features for multiple different description dimensions of the video to be processed" may include:

[0171] Based on the multimodal description content features corresponding to the video clips, the initial video theme features corresponding to each description dimension of the video clips are updated to obtain the target video theme features of each video clip in multiple different description dimensions;

[0172] The target video theme features corresponding to the video clips in each description dimension are fused to obtain the target video theme features of the video to be processed in multiple different description dimensions.

[0173] For example, the multimodal description content features and the initial video theme features corresponding to each video clip can be spliced ​​together, and then the similarity between the two can be calculated using a neural network, and then the target video theme features can be obtained after softmax normalization.

[0174] For example, the target video topic features in the image description dimension can be expressed as follows:

[0175]

[0176] Among them, W2 can represent the parameters of the neural network, h v It can represent the target video theme features corresponding to the image description dimension.

[0177] For another example, the target video theme features in the text description dimension can be expressed as follows:

[0178]

[0179] Among them, W3 can represent the parameters of the neural network, h s It can represent the target video theme features corresponding to the text description dimension.

[0180] 105. Decode the target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video to be processed.

[0181] In one embodiment, information decoding may be performed by combining target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video to be processed.

[0182] There are multiple ways to combine the target video theme features of different description dimensions and the multimodal description content features to perform information decoding and obtain the multimodal content description information of the video to be processed.

[0183] For example, the step of “decoding the target video theme features and the multimodal description content features of different description dimensions to obtain multimodal content description information of the video to be processed” may include:

[0184] The target video theme features and multimodal description content features of different description dimensions are spliced ​​together to obtain a spliced ​​content feature, wherein the spliced ​​content feature is composed of a multidimensional content vector;

[0185] Perform information prediction on the content vector of the concatenated content features to obtain the information unit corresponding to each dimension of the content vector;

[0186] The information units are combined to obtain multimodal content description information.

[0187] For example, the concatenated content feature is generally a multi-dimensional matrix, which is composed of multiple vectors, each of which is a content vector.

[0188] Then, information prediction can be performed on each content vector of the concatenated content features to obtain the information unit corresponding to each content vector. For example, the distribution of the content vectors can be calculated, and then, based on the distribution of the content vectors, the probability of which character the content vector represents can be calculated, thereby obtaining the character corresponding to each content vector, i.e., the information unit. Another example is that the distribution of the content vectors can be calculated, and then, based on the distribution of the content vectors, the probability of which word the content vector represents can be calculated, thereby obtaining the word corresponding to each content vector, i.e., the information unit. The information units can then be combined to obtain multimodal content description information.

[0189] For another example, the information processing model can be used to decode the target video theme features and multimodal description content features of different description dimensions to obtain the multimodal content description information of the video to be processed. For example, the information decoding module in the information processing model can be used to decode the target video theme features and multimodal description content features of different description dimensions to obtain the multimodal content description information of the video to be processed. For example, the Transformer can be used to decode the target video theme features and multimodal description content features of different description dimensions to obtain the multimodal content description information of the video to be processed.

[0190] For example, the multimodal content description information can be expressed as follows:

[0191] Y′=TransformerDecoder([x c ;h v ;h s ])

[0192] Here, Y′ may represent multimodal content description information.

[0193] In one embodiment, when implementing the method proposed in the embodiment of the present application using an information processing model, an information processing model to be trained may be obtained, and the information processing model to be trained may be trained to obtain an information processing model. Specifically, the method proposed in the present application may further include:

[0194] Obtaining descriptions of the information processing model to be trained and the video samples in multiple different description dimensions;

[0195] Using the information processing model to be trained, the description content of the video sample in multiple different description dimensions is processed to obtain description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample;

[0196] Calculating target loss information based on the description content features of the video sample in multiple different description dimensions, the target video theme features of the video sample in multiple different description dimensions, and the multimodal content description information of the video sample;

[0197] The target loss information is used to train the information processing model to be trained to obtain an information processing model.

[0198] In one embodiment, the description content of a video sample in multiple description dimensions is processed using the information processing model to be trained to obtain description content features of the video sample in multiple description dimensions, target video theme features of the video sample in multiple description dimensions, and multimodal content description information of the video sample, which may include:

[0199] Extracting features of the description content of the video sample in multiple description dimensions using the information processing model to be trained, and obtaining initial description content features of the video sample in multiple description dimensions;

[0200] Using the information processing model to be trained, the initial description content features in multiple different description dimensions are normalized and mapped to obtain the description content features of the video sample in multiple different description dimensions;

[0201] Performing topic mapping on the description content features in each description dimension using the information processing model to be trained, to obtain initial video topic features of the video sample in multiple different description dimensions;

[0202] The information processing model to be trained is used to fuse the initial video theme features of multiple different description dimensions to obtain the multimodal description content features corresponding to the video sample;

[0203] Using the information processing model to be trained based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted and processed to obtain target video theme features for multiple different description dimensions of the video sample;

[0204] The target video theme features and multimodal description content features of different description dimensions are decoded to obtain the multimodal content description information of the video sample.

[0205] Among them, the detailed process of the information processing model to be trained processing the description content of the video sample in multiple different description dimensions can be referred to the above description and will not be repeated here.

[0206] In one embodiment, to improve the model performance of the information processing model and thus improve video information processing, the present application embodiment can calculate target loss information based on multiple pieces of information generated when the information processing model to be trained processes description content on multiple different description dimensions. By calculating target loss information based on multiple pieces of information and then using the target loss information to train the information processing model to be trained, the information processing model to be trained can receive feedback from multiple aspects of information and, based on this information, improve its own model parameters, thereby improving model performance.

[0207] In one embodiment, the step of “calculating target loss information based on the description content features of the video sample in multiple different description dimensions, the target video theme features of the video sample in multiple different description dimensions, and the multimodal content description information of the video sample” may include:

[0208] Calculate description content feature loss information based on the description content features of the video sample in multiple different description dimensions;

[0209] Calculating topic mapping constraint loss information based on the description content features of the video sample in multiple different description dimensions and the multimodal content description information;

[0210] Calculate the topic feature loss information based on the target video topic features of the video samples in multiple different description dimensions;

[0211] Calculating multimodal loss information based on multimodal content description information;

[0212] The target loss information is obtained by fusing the description content feature loss information, the topic mapping constraint loss information, the topic feature loss information and the multimodal loss information.

[0213] In one embodiment, the information processing model to be trained can be used to perform feature normalization mapping on the initial description content features on multiple different description dimensions to obtain the description content features of the video sample on multiple different description dimensions. For example, the information processing model to be trained can include a feature normalization mapping function, and then the feature normalization mapping function is used to perform feature normalization mapping on the initial description content features on multiple different description dimensions to obtain the description content features of the video sample on multiple different description dimensions. The feature normalization mapping function can be an MLP. For example, the MLP can be used to map all the initial description content features to a cross-modal common space constructed by the MLP, and then the description content features can be determined in the cross-modal common space.

[0214] Among them, in order to constrain the cross-modal common space, the Euclidean distance can be used to calculate the description content feature loss information, so that the cross-modal common space can accurately express the description content features belonging to different modalities.

[0215] For example, suppose the descriptive content feature x includes the text description dimension s and the descriptive content feature x of the image description dimension v Among them, the description of content feature loss information can be expressed as follows:

[0216] L vs =||x v - s ||

[0217] Among them, L vs It can represent the loss information describing the content characteristics.

[0218] In one embodiment, the information processing model to be trained can be used to perform topic mapping on the descriptive content features on each descriptive dimension to obtain the initial video topic features of the video sample in multiple different description dimensions. For example, the information processing model can include a variational autoencoder, and the variational autoencoder can be used to perform topic mapping on the descriptive content features on each description dimension to obtain the initial video topic features of the video sample in multiple different description dimensions. For example, the variational autoencoder includes an encoder and a decoder. Among them, the encoder in the variational autoencoder can first model the latent space information based on the descriptive content features, and then the decoder attempts to reconstruct the distribution of the multimodal content description information based on the latent space information. That is, the variational autoencoder can be constrained in combination with the descriptive content features and the multimodal content description information, so that the variational autoencoder can construct accurate latent space information, thereby using the latent space information to generate accurate initial video topic features.

[0219] Specifically, the step of “calculating topic mapping constraint loss information based on the description content features of the video sample in multiple different description dimensions and the multimodal content description information” may include:

[0220] Obtain latent space information corresponding to the information processing model to be trained;

[0221] Performing distribution calculation on the latent space information, the descriptive content features, and the multimodal content description information to obtain a latent space distribution corresponding to the latent space information, a descriptive content feature distribution corresponding to the descriptive content features, and a multimodal content description information distribution corresponding to the multimodal content description information;

[0222] Calculate the distribution loss information between the latent space distribution and the distribution of content features;

[0223] Performing statistical operations on the distribution of multimodal content description information to obtain statistical information;

[0224] Perform arithmetic operations on the distribution loss information and statistical information to obtain the topic mapping constraint loss information.

[0225] For example, assuming that the latent space information of the variational autoencoder is z, the latent space information is distributed and the latent space distribution p(z) corresponding to the latent space information is obtained. The latent space information and the descriptive content features are jointly distributed and the descriptive content feature distribution q(z|x s ). The latent space information, the description content features and the multimodal content description information are jointly distributed to obtain the multimodal content description information distribution p(′|) based on the latent space information.

[0226] Then, the distribution loss information between the latent space distribution and the content description feature distribution can be calculated. For example, the distribution loss information between the latent space distribution and the content description feature distribution can be calculated based on the relative entropy (Kullback–Leibler divergence, KL).

[0227] In addition, the multimodal content description information distribution can be statistically calculated to obtain statistical information. For example, the content feature distribution q(z|x s ) computes the expectation of the information distribution of multimodal content description.

[0228] Then, the distribution loss information and the statistical information may be subjected to an arithmetic operation to obtain the topic mapping constraint loss information. For example, the distribution loss information and the statistical information may be subtracted to obtain the topic mapping constraint loss information.

[0229] For example, the topic mapping constraint loss information can be as follows:

[0230]

[0231] Among them, L vae Can represent subject mapping constraint information.

[0232] In one embodiment, the information processing model to be trained can be used to adjust the initial video topic features corresponding to each description dimension based on the multimodal description content features to obtain target video topic features for multiple different description dimensions of the video sample. For example, the information processing model to be trained can include a neural network, which is then used to adjust the initial video topic features corresponding to each description dimension based on the multimodal description content features to obtain target video topic features for multiple different description dimensions of the video sample.

[0233] In order to improve the accuracy of the target video's theme features, and thus improve the accuracy of the multimodal content description information, theme feature loss information can be generated, and then the theme feature loss information can be used to constrain the neural network, thereby improving the performance of the neural network.

[0234] Among them, since the embodiment of the present application includes target video theme features on multiple different description dimensions, the self-supervision method of contrastive learning is used to optimize the target video theme features. The core of contrastive learning is to widen the semantic distance between semantically matched positive pairs (i.e., theme features point to the same meaning) and semantically irrelevant negative pairs (i.e., theme features are far apart) to make the learning of features more representative. Specifically, the step of "calculating theme feature loss information based on target video theme features of video samples on multiple different description dimensions" may include:

[0235] Performing a power operation using the target video theme features in multiple different description dimensions as exponents and a preset base to obtain information after the first power operation;

[0236] Performing a power operation with the target video theme feature and the target video theme feature corresponding to the negative sample as the exponent and a preset base to obtain information after the second power operation;

[0237] The information after the first power operation is compared with the information after the second power operation to obtain the topic feature loss information.

[0238] The performing of a comparison operation on the first information after the power operation and the second information after the power operation may refer to performing a division operation or an absolute value operation on the first information after the power operation and the second information after the power operation, and so on.

[0239] For example, the current video sample can be a positive example pair, and other video samples can be negative samples, and it is assumed that there are N negative samples. In addition, it is assumed that the target video theme feature h in the text description dimension is included s and the target video topic feature h in the image description dimension v . Then, the topic feature loss information can be expressed as follows:

[0240]

[0241] Among them, L cl It can represent the topic feature loss information.

[0242] Among them, the positive example believes that the theme of the same video should be consistent, that is, the (h v ,h s ) is a positive pair. For the negative pair set, the theme of other videos will be used to form pairs, such as (h v ,h s′), where s' is other topic features obtained by random sampling that have little to do with the current video, specifically Figure 4 The construction method shown in the figure has a sampling number of 128. Through this self-supervised contrastive learning mechanism, multimodal video theme features can be fully learned.

[0243] In one embodiment, multimodal loss information is further calculated based on the multimodal content description information. For example, the multimodal loss information between the multimodal content description information and the corresponding tags can be calculated. For example, the multimodal loss information can be calculated using cross entropy.

[0244] For example, Y′ can be the multimodal content description information corresponding to the information processing model to be trained. Y can represent the label corresponding to the video sample, that is, the actual annotation. The multimodal loss information can be expressed as follows:

[0245] L ce =CrossEntropy(Y′,Y)

[0246] Among them, L ce Can represent multimodal loss information.

[0247] In one embodiment, the description content feature loss information, the topic mapping constraint loss information, the topic feature loss information and the multimodal loss information may be fused to obtain the target loss information. For example, the target loss information may be as follows:

[0248] L=L ce +λ1L vs +λ2L vae +λ3L cl

[0249] Among them, λ1, λ2, λ3 are parameters that balance multiple loss functions.

[0250] In an embodiment of the present application, the description content features of the video to be processed can be obtained in multiple different description dimensions; the description content features in each description dimension are subjected to topic mapping to obtain the initial video topic features of the video to be processed in multiple different description dimensions; the description content features of multiple different description dimensions are fused to obtain multimodal description content features; based on the multimodal description content features, the initial video topic features corresponding to each description dimension are adjusted to obtain the target video topic features of multiple different description dimensions of the video to be processed; the target video topic features of different description dimensions and the multimodal description content features are decoded to obtain the multimodal content description information of the video to be processed, and accurate subject description information can be generated by the embodiment of the present application. Through the embodiment of the present application, the multimodal content description information of the video to be processed can be generated by combining the description content features of the video to be processed in multiple different description dimensions, which can improve the accuracy of the generated multimodal content description information. In addition, in the embodiment of the application, the multimodal description content features will be used to adjust the initial video topic features corresponding to each description dimension, so that the generation of multimodal content description information can be guided by mining high-level topic concepts, thereby generating a better quality video description.

[0251] The method described in the above embodiment will be further described in detail below with examples.

[0252] The embodiment of the present application will take the video information processing method integrated on the server as an example to introduce the method of the embodiment of the present application.

[0253] In one embodiment, if Figure 5 As shown, a video information processing method, the specific process is as follows:

[0254] 201. The server obtains description contents of the information processing model to be trained and the video samples in multiple different description dimensions.

[0255] In one embodiment, the server may implement the method of the embodiment of the present application using an information processing model. The information processing model may be a model built based on an artificial intelligence algorithm. For example, the information processing module may be a model built based on ASR, MLP, Transformer, DNN, and variational autoencoder. The information processing module may include a topic mapping module, an information encoding module, and an information decoding module.

[0256] The topic mapping module may include ASR, MLP, and variational autoencoder. The topic mapping module may be used to perform topic mapping on the description content features in each description dimension to obtain initial video topic features of the video sample in multiple different description dimensions.

[0257] The information encoding module may include a Transformer and a DNN. The information encoding module can be used to fuse descriptive content features from multiple different descriptive dimensions to generate multimodal descriptive content features. Furthermore, the information encoding module can be used to adjust the initial video theme features corresponding to each descriptive dimension based on the multimodal descriptive content features to generate target video theme features for multiple different descriptive dimensions of the video sample.

[0258] The information decoding module may include a Transformer, which can be used to decode target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video sample.

[0259] The information processing model to be trained may also be a model based on ASR, MLP, Transformer, DNN and variational autoencoder. The information processing module to be trained may also include a topic mapping module, an information encoding module and an information decoding module.

[0260] In one embodiment, after obtaining a video sample, a topic mapping model in the information processing model to be trained can extract audio information from the video sample and convert the audio information into text information. For example, the topic mapping model can use an ASR tool to extract audio information from the video sample and convert the audio information into text information.

[0261] In one embodiment, when the ASR tool is used to convert audio information into text information, the text information may also carry the time information corresponding to the text information in the video sample. Then, the video sample can be cut into multiple video segments based on the time information corresponding to the text information in the video sample.

[0262] For example, a video sample is a one-minute variety show clip. After using an ASR tool to convert the audio information of the video sample into text information, the resulting text information includes "Let's play a game," "The rules of the game are XXX," and "The winners of the game are Xiao Ming and Xiao Hong," etc. The time information for "Let's play a game" is from the 0th to the 3rd second, the time information for "The rules of the game are XXX" is from the 4th to the 15th second, and the time information for "The winners of the game are Xiao Ming and Xiao Hong" is from the 16th to the 30th second, and so on. Then, based on the time information corresponding to the text information in the video sample, the video sample can be divided into several video segments. For example, the video sample from the 0th to the 3rd second is divided into the first video segment, the video sample from the 4th to the 15th second is divided into the second video segment, and the video sample from the 16th to the 30th second is divided into the third video segment, and so on.

[0263] In one embodiment, after the video sample is divided into a plurality of video segments, the plurality of video segments may be processed to obtain multimodal content description information of the video sample.

[0264] In one embodiment, the text information corresponding to each video clip can be cleaned to obtain target text information corresponding to each video clip to optimize the text quality. For example, some spoken words in the text information can be deleted to obtain the target text information.

[0265] 202. The server uses the information processing model to be trained to extract features of the description content of the video sample in multiple different description dimensions to obtain initial description content features of the video sample in multiple different description dimensions.

[0266] In one embodiment, the BERT network can be used to convert the target text information corresponding to each video clip pair into a text representation vector, where the vector dimension of the text representation vector can be 512. Then, a multilayer perceptron (MLP) can be used to map the text representation vector into a feature space to obtain the text content features corresponding to the target text information.

[0267] In one embodiment, each video clip may be converted into an image representation vector using an S3D network, wherein the vector dimension of the image representation vector may also be 512 dimensions.

[0268] 203. The server uses the information processing model to be trained to perform feature normalization mapping on the initial description content features in multiple different description dimensions to obtain the description content features of the video sample in multiple different description dimensions.

[0269] Then, MLP can be used to map the image representation vector into the feature space to obtain the image content features corresponding to the video clip.

[0270] For example, suppose the text representation vector of video clip i of a video sample is s i , the image representation vector is v i In order to solve the semantic gap problem between cross-modal information, the feature projection functions f and g of images and texts can be used to construct a semantic space:

[0271]

[0272]

[0273] Where f and g can be MLPs. It can represent the image content characteristics, It can represent text content features.

[0274] 204. The server uses the information processing model to be trained to perform topic mapping on the description content features in each description dimension to obtain initial video topic features of the video sample in multiple different description dimensions.

[0275] For example, the server can use a topic variational autoencoder to perform topic mapping on the description content features in each description dimension to obtain the initial video topic features of the video sample in multiple different description dimensions.

[0276] Among them, the variational autoencoder is a variational version based on the autoencoder, which is good at modeling semantic space distribution and capturing high-order concepts.

[0277] Specifically, the topic variational autoencoder will first use the encoder to model latent variables through input data. Then, the latent vector will be used to map the description content features into the initial video topic features.

[0278] For example, the first topic variational autoencoder can utilize the encoder Modeling latent variables z through text content features s For example, the second topic variational autoencoder can use the encoder Modeling latent variables z through image content features v .

[0279] Then, the softmax function can be used to map it to the initial video topic features.

[0280] For example, the initial video topic features in the text description dimension can be expressed as:

[0281]

[0282] in, It can represent the initial video theme features corresponding to the i-th video clip in the text description dimension, It can represent the weight matrix.

[0283] It is worth noting that after being projected into a common space, the semantics between image content features and text content features can be considered measurable. Furthermore, considering that the central theme of a segment and the central theme of a sentence are necessarily consistent, a shared weight matrix is ​​used here to obtain the initial video topic features in the image description dimension. For example, the initial video topic features in the image description dimension can be expressed as:

[0284]

[0285] in, It can represent the initial video theme features corresponding to the i-th video clip in the image description dimension.

[0286] 205. The server utilizes the information processing model to be trained to fuse the initial video theme features of multiple different description dimensions to obtain multimodal description content features corresponding to the video sample.

[0287] For example, the server can use Transformer to fuse the description content features of each video clip of the video sample in multiple description dimensions to obtain the multimodal description content features corresponding to each video clip.

[0288] For example, the multimodal description content features can be expressed as follows:

[0289]

[0290] in, It can represent the multimodal description content features of the i-th video clip. Then, the multimodal description content features of each video clip can be spliced ​​to obtain the multimodal description content features x of the video sample c .

[0291] Among them, the obtained multimodal description content feature x c The same is 512-dimensional. Then use x c Each initial video theme feature is adjusted to obtain target video theme features of multiple different description dimensions for the video sample.

[0292] 206. The server adjusts the initial video theme features corresponding to each description dimension based on the multimodal description content features using the information processing model to be trained, and obtains target video theme features of multiple different description dimensions for the video sample.

[0293] For example, the server can use the information processing model to be trained to splice the multimodal description content features and the initial video theme features corresponding to each video clip, and then use a neural network to calculate the similarity between the two, and then obtain the target video theme features after softmax normalization.

[0294] For example, the target video topic features in the image description dimension can be expressed as follows:

[0295]

[0296] Wherein, W2 may represent the parameters of the neural network in the information processing model to be trained.

[0297] For another example, the target video theme features in the text description dimension can be expressed as follows:

[0298]

[0299] Wherein, W3 may represent the parameters of the neural network in the information processing model to be trained.

[0300] 207. The server decodes the target video theme features and the multimodal description content features of different description dimensions to obtain multimodal content description information of the video sample.

[0301] Next, the server may decode the target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video sample.

[0302] For example, the server can use Transformer to decode the target video theme features and multimodal description content features of different description dimensions to obtain multimodal content description information of the video sample.

[0303] For example, the multimodal content description information can be expressed as follows:

[0304] Y′=TransformerDecoder([x c ;h v ;h s ])

[0305] Here, Y′ may represent multimodal content description information.

[0306] 208. Calculate target loss information based on description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample.

[0307] For example, description content feature loss information, topic feature loss information, topic mapping constraint loss information, and multimodal loss information may be calculated, and then the description content feature loss information, the topic mapping constraint loss information, the topic feature loss information, and the multimodal loss information may be fused to obtain the target loss information.

[0308] For example, the description of content feature loss information can be expressed as follows:

[0309]

[0310] For example, the topic mapping constraint loss information can be as follows:

[0311]

[0312] For example, the topic feature loss information can be expressed as follows:

[0313]

[0314] For example, the multimodal loss information can be expressed as follows:

[0315] L ce =CrossEntropy(Y′,Y)

[0316] For example, the target loss information can be as follows:

[0317] L=L ce +λ1L vs +λ2L vae +λ3L cl

[0318] 209. The server trains the information processing model to be trained using the target loss information to obtain an information processing model.

[0319] For example, the parameters in the information processing model to be trained can be adjusted according to the target loss information to obtain the information processing model.

[0320] In an embodiment of the present application, the server can obtain the description content of the information processing model to be trained and the video sample in multiple different description dimensions; use the information processing model to be trained to perform feature extraction on the description content of the video sample in multiple different description dimensions to obtain the initial description content features of the video sample in multiple different description dimensions; use the information processing model to be trained to perform feature normalization mapping on the initial description content features in multiple different description dimensions to obtain the description content features of the video sample in multiple different description dimensions; use the information processing model to be trained to perform topic mapping on the description content features in each description dimension to obtain the initial video topic features of the video sample in multiple different description dimensions; use the information processing model to be trained to map the initial video content features in multiple different description dimensions to obtain the initial video topic features of the video sample in multiple different description dimensions The information processing model to be trained is used to fuse the initial video topic features corresponding to each description dimension based on the multimodal description content features to obtain the target video topic features for multiple different description dimensions of the video sample. The target video topic features and the multimodal description content features of the different description dimensions are decoded to obtain the multimodal content description information of the video sample. The target loss information is calculated based on the description content features of the video sample in multiple different description dimensions, the target video topic features of the video sample in multiple different description dimensions, and the multimodal content description information of the video sample. The target loss information is used to train the information processing model to be trained to obtain the information processing model. The information processing model generates high-quality topic description information, thereby improving the accuracy of video recommendations for users.

[0321] To better implement the video information processing method provided in the embodiments of this application, a video information processing device is also provided in one embodiment. The video information processing device can be integrated into a computer device. The meanings of the terms herein are the same as those in the aforementioned video information processing method. For specific implementation details, please refer to the description in the method embodiment.

[0322] In one embodiment, a video information processing device is provided. The video information processing device can be integrated into a computer device, such as Figure 6 As shown, the video information processing device includes: an acquisition unit 301, a theme mapping unit 302, a fusion unit 303, an adjustment unit 304 and a decoding unit 305, which are specifically as follows:

[0323] An acquisition unit 301 is configured to acquire description content features of a video to be processed in multiple different description dimensions;

[0324] The topic mapping unit 302 is used to perform topic mapping on the description content features in each description dimension to obtain the initial video topic features of the video to be processed in multiple different description dimensions;

[0325] A fusion unit 303 is configured to fuse the description content features of the multiple different description dimensions to obtain a multimodal description content feature;

[0326] An adjustment unit 304 is configured to adjust the initial video theme feature corresponding to each description dimension based on the multimodal description content feature to obtain target video theme features for multiple different description dimensions of the video to be processed;

[0327] The decoding unit 305 is configured to decode the target video theme features of the different description dimensions and the multimodal description content features to obtain the multimodal content description information of the video to be processed.

[0328] In one embodiment, the topic mapping unit 302 may include:

[0329] A first mapping subunit is configured to map the description content features on each description dimension to a preset subject feature space, and obtain feature distribution information of the description content features on each description dimension in the preset subject feature space;

[0330] A probability calculation subunit, configured to calculate probability fitting information of each description content feature based on feature distribution information of the description content feature on each description dimension in a preset subject feature space;

[0331] The feature determination subunit is used to determine the initial video theme features of the video to be processed in multiple different description dimensions based on the probability fitting information of each description content feature.

[0332] In one embodiment, the fusion unit 303 may include:

[0333] A convolution operation subunit, configured to perform a convolution operation on the description content features of multiple different description dimensions to obtain a convolution operation result corresponding to the description content feature of each description dimension;

[0334] The cross attention fusion subunit is used to perform cross attention fusion on the convolution operation results of each description dimension to obtain the cross attention features corresponding to each description dimension;

[0335] The first splicing subunit is used to splice the cross-attention features corresponding to each description dimension to obtain the spliced ​​features;

[0336] The fully connected subunit is used to perform fully connected processing on the spliced ​​features to obtain the multimodal description content features.

[0337] In one embodiment, the adjusting unit 304 may include:

[0338] A second splicing subunit is configured to splice the multimodal content feature with the initial video theme feature corresponding to the current description dimension to obtain a spliced ​​feature corresponding to the current description dimension;

[0339] A second mapping subunit is configured to map the concatenated features corresponding to the current description dimension using the topic mapping information corresponding to the current description dimension to obtain mapped features;

[0340] A nonlinear conversion processing subunit, configured to perform nonlinear conversion processing on the mapped features to obtain nonlinear converted features;

[0341] The feature fusion subunit is used to fuse the nonlinearly converted features with the initial video theme features to obtain the target video theme features.

[0342] In one embodiment, the decoding unit 305 may include:

[0343] a third splicing subunit, configured to splice the target video theme features of different description dimensions and the multimodal description content features to obtain a spliced ​​content feature, wherein the spliced ​​content feature is composed of a multidimensional content vector;

[0344] An information prediction subunit, configured to perform information prediction on the content vector of the concatenated content feature to obtain an information unit corresponding to each dimension of the content vector;

[0345] The combining subunit is configured to combine the information units to obtain the multimodal content description information.

[0346] In one embodiment, the information processing device may further include:

[0347] a content acquisition unit, configured to acquire description content of at least one video segment of a video to be processed in multiple different description dimensions;

[0348] a feature extraction unit, configured to extract features of the description content of the video clip in a plurality of different description dimensions, and obtain initial description content features of the video clip in the plurality of different description dimensions;

[0349] The feature normalization mapping unit is used to perform feature normalization mapping on the initial description content features of the video clips in multiple different description dimensions to obtain the description content features of at least one video clip of the video to be processed in multiple different description dimensions.

[0350] In one embodiment, the information processing device may further include:

[0351] A model acquisition unit, used to obtain descriptions of the information processing model to be trained and the video samples in multiple different description dimensions;

[0352] a processing unit, configured to process the description content of the video sample in multiple different description dimensions using the information processing model to be trained, to obtain description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample;

[0353] a loss calculation unit, configured to calculate target loss information based on description content features of the video sample in a plurality of different description dimensions, target video subject features of the video sample in a plurality of different description dimensions, and multimodal content description information of the video sample;

[0354] A training unit is used to train the information processing model to be trained using the target loss information to obtain the information processing model.

[0355] In one embodiment, the loss calculation unit may include:

[0356] A first loss calculation subunit, configured to calculate description content feature loss information based on description content features of the video sample in a plurality of different description dimensions;

[0357] A second loss calculation subunit is configured to calculate topic mapping constraint loss information based on the description content features of the video sample in multiple description dimensions and the multimodal content description information;

[0358] A third loss calculation subunit is configured to calculate topic feature loss information based on target video topic features of the video sample in a plurality of different description dimensions;

[0359] a fourth loss calculation subunit, configured to calculate multimodal loss information based on the multimodal content description information;

[0360] The loss fusion subunit is used to fuse the description content feature loss information, the topic mapping constraint loss information, the topic feature loss information and the multimodal loss information to obtain the target loss information.

[0361] In one embodiment, the second loss calculation subunit may include:

[0362] The information acquisition module is used to obtain the latent space information corresponding to the information processing model to be trained;

[0363] a distribution calculation module, configured to perform distribution calculation on the latent space information, the descriptive content features, and the multimodal content description information to obtain a latent space distribution corresponding to the latent space information, a descriptive content feature distribution corresponding to the descriptive content features, and a multimodal content description information distribution corresponding to the multimodal content description information;

[0364] A loss calculation module, configured to calculate distribution loss information between the latent space distribution and the content description feature distribution;

[0365] A statistical operation module, configured to perform statistical operations on the distribution of the multimodal content description information to obtain statistical information;

[0366] An arithmetic operation module is used to perform an arithmetic operation on the distribution loss information and the statistical information to obtain the topic mapping constraint loss information.

[0367] In one embodiment, the third loss calculation subunit may include:

[0368] A first exponentiation module, configured to perform an exponentiation operation using the target video theme features in the multiple different description dimensions as exponents and a preset base to obtain first exponentiation information;

[0369] A second exponentiation module is configured to perform an exponentiation operation using the target video theme feature and the target video theme feature corresponding to the negative sample as an exponent and a preset base to obtain information after the second exponentiation operation;

[0370] A comparison operation module is used to perform a comparison operation on the first information after the power operation and the second information after the power operation to obtain the topic feature loss information.

[0371] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.

[0372] The above-mentioned video information processing device can generate high-quality topic description information, thereby improving the accuracy of recommending videos to users.

[0373] The embodiment of the present application also provides a computer device, which may include a terminal or a server. For example, the computer device may be used as a video information processing terminal, which may be a mobile phone, a tablet computer, etc.; for another example, the computer device may be a server, such as a video information processing server. Figure 7 As shown, it shows a schematic diagram of the structure of the terminal involved in the embodiment of the present application, specifically:

[0374] The computer device may include one or more processing core processors 401, one or more computer readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will understand that Figure 7 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0375] Processor 401 is the control center of the computer device, connecting the various components of the entire computer device using various interfaces and circuits. It executes the various functions of the computer device and processes data by running or executing software programs and / or modules stored in memory 402 and accessing data stored in memory 402. Optionally, processor 401 may include one or more processing cores; preferably, processor 401 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.

[0376] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0377] The computer device also includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0378] The computer device may further include an input unit 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0379] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to one or more application processes into the memory 402 according to the following instructions, and the processor 401 will run the application stored in the memory 402 to implement various functions as follows:

[0380] Obtaining description content features of the video to be processed in multiple different description dimensions;

[0381] Performing topic mapping on the description content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions;

[0382] fusing the description content features of the multiple different description dimensions to obtain a multimodal description content feature;

[0383] Based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted to obtain target video theme features of multiple different description dimensions for the video to be processed;

[0384] The target video theme features of the different description dimensions and the multimodal description content features are decoded to obtain multimodal content description information of the video to be processed.

[0385] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0386] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.

[0387] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by a computer program, or by controlling related hardware through a computer program. The computer program may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0388] To this end, embodiments of the present application further provide a storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the video information processing methods provided in embodiments of the present application. For example, the computer program can execute the following steps:

[0389] Obtaining description content features of the video to be processed in multiple different description dimensions;

[0390] Performing topic mapping on the description content features in each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions;

[0391] fusing the description content features of the multiple different description dimensions to obtain a multimodal description content feature;

[0392] Based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted to obtain target video theme features of multiple different description dimensions for the video to be processed;

[0393] The target video theme features of the different description dimensions and the multimodal description content features are decoded to obtain multimodal content description information of the video to be processed.

[0394] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0395] Since the computer program stored in the storage medium can execute the steps in any video information processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any video information processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0396] The above is a detailed introduction to a video information processing method, device, computer equipment and storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A video information processing method, characterized in that: include: Cutting the video to be processed into a plurality of video segments according to time information corresponding to the text information in the video to be processed, wherein the text information is obtained by converting the audio information in the video to be processed, and the plurality of video segments correspond to the plurality of sentences included in the text information; Obtaining description content features of a video to be processed in multiple different description dimensions; wherein, the method includes: obtaining description content features of at least one video segment of the video to be processed in multiple different description dimensions; Performing topic mapping on the description content features on each description dimension to obtain initial video topic features of the video to be processed in multiple different description dimensions; wherein, the method includes: performing topic mapping on the description content features on each description dimension of each video clip to obtain initial video topic features of each video clip of the video to be processed in multiple different description dimensions; The description content features of the multiple different description dimensions are fused to obtain a multimodal description content feature; wherein, the method includes: fusing the description content features of the video clips of the video to be processed at the multiple different description dimensions to obtain a multimodal description content feature corresponding to each video clip; Based on the multimodal description content features, the initial video theme features corresponding to each description dimension are adjusted to obtain target video theme features of multiple different description dimensions for the video to be processed; wherein, the method includes: based on the multimodal description content features corresponding to the video clip, the initial video theme features corresponding to each description dimension of the video clip are updated to obtain target video theme features of each video clip in multiple different description dimensions; the target video theme features corresponding to the video clip in each description dimension are fused to obtain target video theme features of the video to be processed in multiple different description dimensions; The target video theme features of the different description dimensions and the multimodal description content features are decoded to obtain multimodal content description information of the video to be processed.

2. The method according to claim 1, characterized in that The topic mapping of the description content features on each description dimension to obtain the initial video topic features of the to-be-processed video on multiple different description dimensions includes: Mapping the description content features on each description dimension to a preset subject feature space to obtain feature distribution information of the description content features on each description dimension in the preset subject feature space; Calculate the probability fitting information of each description content feature based on the feature distribution information of the description content feature on each description dimension in the preset subject feature space; According to the probability fitting information of each descriptive content feature, the initial video theme features of the video to be processed in multiple different description dimensions are determined.

3. The method according to claim 1, characterized in that The fusing of the description content features of the multiple different description dimensions to obtain a multimodal description content feature includes: Performing a convolution operation on the description content features of multiple different description dimensions to obtain a convolution operation result corresponding to the description content feature of each description dimension; Perform cross-attention fusion on the convolution operation results of each description dimension to obtain the cross-attention features corresponding to each description dimension; Concatenate the cross-attention features corresponding to each description dimension to obtain the concatenated features; Fully connect the concatenated features to obtain the multimodal description content features.

4. The method according to claim 1, wherein The adjusting process of the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features of multiple different description dimensions for the video to be processed includes: Splicing the multimodal content feature and the initial video theme feature corresponding to the current description dimension to obtain a spliced ​​feature corresponding to the current description dimension; Using the topic mapping information corresponding to the current description dimension, mapping the concatenated features corresponding to the current description dimension to obtain mapped features; Performing nonlinear transformation on the mapped features to obtain nonlinear transformed features; The nonlinearly converted features and the initial video theme features are fused to obtain the target video theme features.

5. The method according to claim 1, wherein The decoding of the target video theme features of the different description dimensions and the multimodal description content features to obtain the multimodal content description information of the video to be processed includes: Splicing the target video theme features of different description dimensions and the multimodal description content features to obtain a spliced ​​content feature, wherein the spliced ​​content feature is composed of a multidimensional content vector; Performing information prediction on the content vector of the spliced ​​content feature to obtain an information unit corresponding to each dimension of the content vector; The information units are combined to obtain the multimodal content description information.

6. The method according to claim 1, characterized in that The obtaining of description content features of at least one video segment of the video to be processed in multiple different description dimensions includes: Obtaining description content of at least one video segment of a video to be processed in multiple different description dimensions; Extracting features of the description content of the video clip in multiple different description dimensions to obtain initial description content features of the video clip in the multiple different description dimensions; Performing feature normalization mapping on the initial description content features of the video clips in multiple different description dimensions to obtain description content features of at least one video clip of the video to be processed in multiple different description dimensions.

7. The method according to claim 1, characterized in that The obtaining of description content features of the video to be processed in multiple different description dimensions includes: Using information processing models to obtain descriptive content features of the video to be processed in multiple different descriptive dimensions; The topic mapping of the description content features on each description dimension to obtain the initial video topic features of the video to be processed in multiple different description dimensions includes: Using the information processing model to perform topic mapping on the description content features in each description dimension, to obtain initial video topic features of the video to be processed in multiple different description dimensions; The fusing of the description content features of the multiple different description dimensions to obtain a multimodal description content feature includes: Using the information processing model to perform topic mapping on the description content features in each description dimension, to obtain initial video topic features of the video to be processed in multiple different description dimensions; The adjusting process of the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features of multiple different description dimensions for the video to be processed includes: Using the information processing model to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features, to obtain target video theme features for multiple different description dimensions of the video to be processed; The target video theme features of the different description dimensions and the multimodal description content features are decoded to obtain multimodal content description information of the video to be processed.

8. The method according to claim 7, characterized in that The method further comprises: Obtaining descriptions of the information processing model to be trained and the video samples in multiple different description dimensions; Processing the description content of the video sample in multiple different description dimensions using the information processing model to be trained to obtain description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample; Calculating target loss information based on description content features of the video sample in multiple different description dimensions, target video theme features of the video sample in multiple different description dimensions, and multimodal content description information of the video sample; The information processing model to be trained is trained using the target loss information to obtain the information processing model.

9. The method according to claim 8, characterized in that Calculating target loss information based on description content features of the video sample in multiple description dimensions, target video theme features of the video sample in multiple description dimensions, and multimodal content description information of the video sample includes: Calculating description content feature loss information based on description content features of the video sample in multiple different description dimensions; Calculating topic mapping constraint loss information based on the description content features of the video sample in multiple different description dimensions and the multimodal content description information; Calculating topic feature loss information based on target video topic features of the video sample in multiple different description dimensions; Calculating multimodal loss information according to the multimodal content description information; The description content feature loss information, the topic mapping constraint loss information, the topic feature loss information and the multimodal loss information are fused to obtain the target loss information.

10. The method according to claim 9, characterized in that The calculating the topic mapping constraint loss information according to the description content features of the video sample in multiple description dimensions and the multimodal content description information includes: Obtain latent space information corresponding to the information processing model to be trained; Performing distribution calculation on the latent space information, the descriptive content features, and the multimodal content description information to obtain a latent space distribution corresponding to the latent space information, a descriptive content feature distribution corresponding to the descriptive content features, and a multimodal content description information distribution corresponding to the multimodal content description information; Calculating distribution loss information between the latent space distribution and the content description feature distribution; Performing statistical operations on the distribution of the multimodal content description information to obtain statistical information; Performing an arithmetic operation on the distribution loss information and the statistical information to obtain the topic mapping constraint loss information.

11. The method according to claim 9, characterized in that The calculating of the theme feature loss information according to the target video theme features of the video sample in a plurality of different description dimensions includes: Performing a power operation using the target video theme features in the multiple different description dimensions as exponents and a preset base to obtain first power-operated information; Performing a power operation using the target video theme feature and the target video theme feature corresponding to the negative sample as an exponent and a preset base to obtain information after a second power operation; A comparison operation is performed on the first information after the power operation and the second information after the power operation to obtain the topic feature loss information.

12. A video information processing device, characterized in that: include: The device is configured to cut a video to be processed into a plurality of video segments based on time information corresponding to text information in the video to be processed, wherein the text information is obtained by converting audio information in the video to be processed, and the plurality of video segments correspond to a plurality of sentences included in the text information; The acquisition unit is configured to acquire description content features of the video to be processed in multiple different description dimensions, which includes: acquiring description content features of at least one video segment of the video to be processed in multiple different description dimensions; A topic mapping unit is used to perform topic mapping on the description content features on each description dimension to obtain the initial video topic features of the video to be processed in multiple different description dimensions; wherein, the method includes: performing topic mapping on the description content features on each description dimension of each video clip to obtain the initial video topic features of each video clip of the video to be processed in multiple different description dimensions; A fusion unit is configured to fuse the description content features of the plurality of different description dimensions to obtain a multimodal description content feature; wherein the fusion unit includes: fusing the description content features of the video clips of the video to be processed at the plurality of different description dimensions to obtain a multimodal description content feature corresponding to each video clip; An adjustment unit is configured to adjust the initial video theme features corresponding to each description dimension based on the multimodal description content features to obtain target video theme features for multiple different description dimensions of the video to be processed; wherein the adjustment unit includes: updating the initial video theme features corresponding to each description dimension of the video clip based on the multimodal description content features corresponding to the video clip to obtain target video theme features for each video clip in multiple different description dimensions; fusing the target video theme features corresponding to the video clip in each description dimension to obtain target video theme features for the video to be processed in multiple different description dimensions; The decoding unit is used to decode the target video theme features of the different description dimensions and the multimodal description content features to obtain the multimodal content description information of the video to be processed.

13. A computer device, characterized in that: It comprises a memory and a processor; the memory stores an application, and the processor is used to run the application in the memory to perform the operations in the video information processing method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor to execute the steps in the video information processing method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the video information processing method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Video file classification method and device, medium and electronic equipment

    CN111488489A

  • Video description text generation method based on multi-modal fusion

    CN112069361A