Video information processing method and device based on video information processing model
By obtaining video parameters and generating fusion feature vectors, the problem of inaccurate video representation in the prior art is solved, and multimodal representation and fusion of video information is realized, improving the accuracy of user experience and video recommendation.
Patent Information
- Application Number
- CN201911183241.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-11-27
AI Technical Summary
The prior art lacks structured methods in the learning of video information representation, resulting in inaccurate video representation, especially when video tag information is missing or inaccurate, affecting the user experience.
By obtaining video parameters, determining basic features and multimodal features, using the video processing network in the video information processing model to generate fusion feature vectors, and combining the output of the second video processing network, multimodal representation and fusion of video information are realized.
It improves the accuracy of the representation of video information, can better express video features, improve user experience, and improve the accuracy and rationality of video recommendations.
Smart Images

Figure CN112861580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information processing technology, and in particular to a video information processing method, device, electronic device and storage medium based on a video information processing model. Background Art
[0002] The vectorized representation of video information is the basis of many machine learning algorithms, and how to accurately represent video information is the focus of research in this direction. Most existing technologies are relatively one-sided and do not perform structured representation learning on videos.
[0003] Common learning methods include: 1) Directly using video tags, including video classification, video theme, video release source, etc. Videos can be roughly divided into entertainment videos, sports videos, or subdivided into basketball highlights, film and television highlights. However, this type of representation method is relatively extensive, and the classification tag information needs to be set in advance and updated in a timely manner, and its content representation ability is limited. 2) Text-based learning, including text semantic learning of video titles, video description information or video tags. This type of method relies more on the accuracy of text information, but many videos lack text information, which makes the video representation inaccurate. The random recommendation strategy generated in the previous learning process is simple in logic and has a relatively low accuracy rate, which seriously affects the user experience. Summary of the invention
[0004] In view of this, the embodiments of the present invention provide a video information processing method, device, electronic device and storage medium based on a video information processing model. The technical solution of the embodiments of the present invention is implemented as follows:
[0005] An embodiment of the present invention provides a video information processing method based on a video information processing model, the method comprising:
[0006] Acquire a first target video, and parse the first target video to acquire video parameters of the first target video;
[0007] Determining basic features matching the first target video according to the video parameters of the first target video;
[0008] Determining, according to the video parameters of the first target video, a multimodal feature matching the first target video;
[0009] Based on the basic features and the multimodal features, a fusion feature vector matching the first target video is determined through the first video processing network in the video information processing model, wherein the fusion feature vector is used to be combined with the second target video fusion feature vector output by the second video processing network in the video information processing model to implement a process matching the video information processing model.
[0010] The embodiment of the present invention further provides a processing device based on a video information processing model, the device comprising:
[0011] An information transmission module, used for acquiring a first target video, and parsing the first target video to acquire video parameters of the first target video;
[0012] An information processing module, configured to determine basic features matching the first target video according to video parameters of the first target video;
[0013] The information processing module is used to determine the multimodal features matching the first target video according to the video parameters of the first target video;
[0014] The information processing module is used to determine a fused feature vector matching the first target video based on the basic features and the multimodal features through the first video processing network in the video information processing model, wherein the fused feature vector is used to be combined with the second target video fused feature vector output by the second video processing network in the video information processing model to implement a process matching the video information processing model.
[0015] In the above scheme,
[0016] The information processing module is used to parse the first target video and obtain label information of the first target video;
[0017] The information processing module is used to parse the video information corresponding to the first target video according to the label information of the first target video, so as to obtain the video parameters of the first target video in the basic dimension and the multimodal dimension respectively.
[0018] In the above scheme,
[0019] The information processing module is used to, based on the video parameters of the first target video in the basic dimension,
[0020] The information processing module is used to determine the category parameter, video tag parameter and video publishing source parameter corresponding to the first target video;
[0021] The information processing module is used to extract features from the category parameters, video tag parameters and video publishing source parameters corresponding to the first target video, so as to form basic features that match the first target video.
[0022] In the above scheme,
[0023] The information processing module is used to, based on the video parameters of the first target video in the basic dimension,
[0024] The information processing module is used to determine the title text parameters, image information parameters and visual information parameters corresponding to the first target video;
[0025] The information processing module is used to extract and fuse the title text parameters, image information parameters and visual information parameters corresponding to the first target video respectively, so as to form a multimodal feature matching the first target video.
[0026] In the above scheme,
[0027] The information processing module is used to process the basic features through a basic information processing network in the first video processing network to form a corresponding basic feature vector;
[0028] The information processing module is used to process the image features in the multimodal features through the image processing network in the first video processing network to form a corresponding image feature vector;
[0029] The information processing module is used to process the title text features in the multimodal features through the text processing network in the first video processing network to form a corresponding title text feature vector;
[0030] The information processing module is used to process the visual features in the multimodal features through the visual processing network in the first video processing network to form a corresponding visual feature vector;
[0031] The information processing module is used to perform vector fusion through the first video processing network based on the basic feature vector, the image feature vector, the title text feature vector and the visual feature vector to form a fused feature vector matching the first target video.
[0032] In the above scheme,
[0033] The information processing module is used to obtain the image to be processed and the target resolution corresponding to the playback interface of the first target video;
[0034] The information processing module is used to respond to the target resolution, perform resolution enhancement processing on the image to be processed through the image processing network in the first video processing network, and obtain the corresponding image feature vector to achieve the adaptation of the image feature vector to the target resolution corresponding to the playback interface of the first target video.
[0035] In the above scheme,
[0036] The information processing module is used to extract a text feature vector matching the title text feature through a text processing network;
[0037] The information processing module is used to determine at least one word-level latent variable corresponding to the title text feature according to the text feature vector through the text processing network;
[0038] The information processing module is used to generate, through the text processing network, a processing word corresponding to the word-level latent variable and a probability of selection of the processing word according to the at least one word-level latent variable;
[0039] The information processing module is used to select at least one processing word to form a text processing result corresponding to the title text feature according to the selection probability of the processing result.
[0040] In the above scheme,
[0041] The information processing module is used to determine bit rate information matching the playback environment of the first target video;
[0042] The information processing module is used to adjust the bit rate of the first target video by using the visual features in the multimodal features through the visual processing network in the first video processing network, so as to match the bit rate information of the first target video with the bit rate information of the playback environment.
[0043] In the above scheme,
[0044] The information processing module is used to adjust the parameters of the recurrent convolutional neural network based on the attention mechanism in the first video processing network according to the second target video fusion feature vector output by the second video processing network in the video information processing model when the process matching the video information processing model is a video recommendation process, so as to achieve the adaptation of the parameters of the recurrent convolutional neural network based on the attention mechanism to the fusion feature vector.
[0045] In the above scheme,
[0046] The information processing module is used to adjust the parameters of the second video processing network in the video information processing model;
[0047] Determining a new second target video fusion feature vector through a second video processing network in the video information processing model after parameter adjustment;
[0048] The new second target video fusion feature vector and the first target video fusion feature vector are connected through a classification prediction function that matches the video information processing model to determine the correlation between the first target video and the second target video.
[0049] An embodiment of the present invention further provides an electronic device, the electronic device comprising:
[0050] A memory for storing executable instructions;
[0051] The processor is used to implement the preceding video information processing method based on the video information processing model when running the executable instructions stored in the memory.
[0052] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implements the aforementioned video information processing method based on a video information processing model.
[0053] The embodiments of the present invention have the following beneficial effects:
[0054] The method comprises the following steps: obtaining a first target video and parsing the first target video to obtain video parameters of the first target video; determining basic features matching the first target video according to the video parameters of the first target video; determining multimodal features matching the first target video according to the video parameters of the first target video; and determining a fused feature vector matching the first target video through a first video processing network in the video information processing model based on the basic features and the multimodal features, wherein the fused feature vector is used to be combined with a second target video fused feature vector output by a second video processing network in the video information processing model to achieve a process matching the video information processing model, thereby processing the video information of the first target video to form matching video multimodal information, integrating the multimodal features of the first target video, and being able to better express the features of the first target video, which is beneficial to subsequent operations on the first target video. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 A schematic diagram of a use scenario of a video information processing method based on a video information processing model provided by an embodiment of the present invention;
[0056] Figure 2A schematic diagram of the composition structure of a processing device based on a video information processing model provided by an embodiment of the present invention;
[0057] Figure 3 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention;
[0058] Figure 4 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention;
[0059] Figure 5 This is an optional structural diagram of a text processing network in an embodiment of the present invention;
[0060] Figure 6 Schematic diagram of a process for determining an optional word-level latent variable of a text processing network in an embodiment of the present invention;
[0061] Figure 7 This is a schematic diagram of an optional structure of an encoder in a text processing network in an embodiment of the present invention;
[0062] Figure 8 A schematic diagram of vector concatenation of an encoder in a text processing network according to an embodiment of the present invention;
[0063] Fig. 9 Schematic diagram of the encoding process of an encoder in a text processing network in an embodiment of the present invention;
[0064] Fig.10 Schematic diagram of the decoding process of a decoder in a text processing network in an embodiment of the present invention;
[0065] Fig.11 Schematic diagram of the decoding process of a decoder in a text processing network in an embodiment of the present invention;
[0066] Fig.12 Schematic diagram of the decoding process of a decoder in a text processing network in an embodiment of the present invention;
[0067] Fig.13 This is an optional structural diagram of an image processing network in an embodiment of the present invention;
[0068] Fig.14 This is an optional structural diagram of an image visual processing network in an embodiment of the present invention;
[0069] Fig.15 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention;
[0070] Fig.16Schematic diagram of an application environment of a video information processing method based on a video information processing model in an embodiment of the present invention;
[0071] Fig.17 A schematic diagram of the working process of the video information processing method based on the video information processing model provided by an embodiment of the present invention;
[0072] Fig.18 A structural schematic diagram of a video information processing device of a video information processing model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.
[0074] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0075] Before further describing the embodiments of the present invention in detail, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.
[0076] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed may be in real time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0077] 2) The first target video is various forms of video information available on the Internet, such as video files and multimedia information presented in a client or smart device.
[0078] 3) Convolutional Neural Networks (CNN) is a type of feed-forward neural network that includes convolutional calculations and has a deep structure. It is one of the representative algorithms of deep learning. Convolutional neural networks have representation learning capabilities and can perform shift-invariant classification on input information according to their hierarchical structure.
[0079] 4) Model training, multi-classification learning of image datasets. The model can be built using deep learning frameworks such as Tensor Flow and torch, and uses multiple layers of neural network layers such as CNN to form a multi-classification model. The input of the model is a three-channel or original channel matrix formed by reading the image through tools such as openCV. The model output is a multi-classification probability, and the web page category is finally output through algorithms such as softmax. During training, the model approaches the correct trend through objective functions such as cross entropy.
[0080] 5) Neural Network (NN): Artificial Neural Network (ANN), referred to as neural network or quasi-neural network, is a mathematical model or computational model that imitates the structure and function of biological neural networks (the central nervous system of animals, especially the brain) in the field of machine learning and cognitive science, and is used to estimate or approximate functions.
[0081] 6) Speech Recognition (SR Speech Recognition): Also known as Automatic Speech Recognition (ASR Automatic Speech Recognition), Computer Speech Recognition (CSR Computer Speech Recognition) or Speech to Text (STT Speech To Text), its goal is to use computers to automatically convert human speech content into corresponding text.
[0082] 7) Machine Translation (MT): It belongs to the field of computational linguistics. Its research is to translate text or speech from one natural language into another natural language by computer programs. Neural Machine Translation (NMT) is a technology that uses neural network technology for machine translation.
[0083] 8) Encoder-decoder structure: A network structure commonly used in machine translation technology. It consists of two parts: an encoder and a decoder. The encoder converts the input text into a series of context vectors that can express the characteristics of the input text. The decoder receives the output of the encoder as its own input and outputs the corresponding text sequence in another language.
[0084] 9) Bidirectional Attention Neural Network Model (BERT Bidirectional Encoder Representationsfrom Transformers) Bidirectional Attention Neural Network Model proposed by Google.
[0085] 10) Token: Before any actual processing of the input text, it needs to be segmented into language units such as words, punctuation marks, numbers, or pure alphanumeric characters. These units are called tokens.
[0086] 11) Softmax: Normalized exponential function, which is a generalization of the logic function. It can "compress" a K-dimensional vector containing any real number into another K-dimensional real vector, so that each element is between [0,1] and the sum of all elements is 1.
[0087] 12) Word segmentation: Use Chinese word segmentation tools to segment Chinese text and obtain a set of fine-grained words. Stop words: Characters or words that do not contribute to the semantics of the text or whose contribution can be ignored. Cosine similarity: The cosine similarity of two texts after they are represented as vectors.
[0088] 13) Transformers: A new network structure that uses an attention mechanism to replace the traditional encoder-decoder model that must rely on other neural networks. Word vector: A single word is represented by a distribution vector of fixed dimension. Compound word: A coarse-grained keyword composed of fine-grained keywords, whose semantics are richer and more complete than fine-grained keywords.
[0089] Figure 1 A schematic diagram of a use scenario of a video information processing method based on a video information processing model provided in an embodiment of the present invention, see Figure 1 The terminal (including the terminal 10-1 and the terminal 10-2) is provided with a client of software capable of displaying the corresponding first target video, such as a client or plug-in for video playback. The user can obtain the first target video (or the first target video and the corresponding second target video) through the corresponding client and display it; the terminal is connected to the server 200 through the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two, and a wireless link is used to realize data transmission.
[0090] As an example, the server 200 is used to deploy the processing device based on the video information processing model to implement the video information processing method based on the video information processing model provided by the present invention, so as to obtain a first target video and parse the first target video to obtain the video parameters of the first target video; determine the basic features matching the first target video according to the video parameters of the first target video; determine the multimodal features matching the first target video according to the video parameters of the first target video; based on the basic features and the multimodal features, determine the fused feature vector matching the first target video through the first video processing network in the video information processing model, wherein the fused feature vector is used to combine with the second target video fused feature vector output by the second video processing network in the video information processing model to implement a process matching the video information processing model, and display the output through the terminal (terminal 10-1 and / or terminal 10-2) that matches the first target video and any feature matching the first target video. Of course, the processing device based on the video information processing model provided by the present invention can be applied to video playback. In video playback, first target videos from different data sources are usually processed, and finally the corresponding first target video and any feature information matching the first target video are presented on the user interface (UI). The accuracy and timeliness of the features of the first target video directly affect the user experience. The background database of video playback receives a large amount of video data from different sources every day, and the text information matching the first target video can also be called by other applications.
[0091] Of course, the process of processing the first target video by the processing device based on the video information processing model to achieve matching with the video information processing model specifically includes:
[0092] Acquire a first target video, and parse the first target video to obtain video parameters of the first target video; determine basic features that match the first target video according to the video parameters of the first target video; determine multimodal features that match the first target video according to the video parameters of the first target video; based on the basic features and the multimodal features, determine a fused feature vector that matches the first target video through a first video processing network in the video information processing model, wherein the fused feature vector is used to be combined with a second target video fused feature vector output by a second video processing network in the video information processing model to achieve a process that matches the video information processing model.
[0093] The structure of the processing device based on the video information processing model of the embodiment of the present invention is described in detail below. The processing device based on the video information processing model can be implemented in various forms, such as a dedicated terminal with the processing function of the processing device based on the video information processing model, or a server provided with the processing function of the processing device based on the video information processing model, such as the preceding Figure 1 Server 200 in. Figure 2 The schematic diagram of the composition structure of the processing device based on the video information processing model provided by the embodiment of the present invention can be understood as follows: Figure 2 Only an exemplary structure of a processing device based on a video information processing model is shown, not all structures, and it can be implemented as needed. Figure 2 Partial or complete structure shown.
[0094] The processing device based on the video information processing model provided in the embodiment of the present invention includes: at least one processor 201, a memory 202, a user interface 203 and at least one network interface 204. The various components in the processing device based on the video information processing model are coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, in Figure 2 Various buses are labeled as bus system 205 .
[0095] The user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.
[0096] It is understood that the memory 202 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The memory 202 in the embodiment of the present invention can store data to support the operation of the terminal (such as 10-1). Examples of these data include: any computer program for operating on the terminal (such as 10-1), such as an operating system and an application. Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application can include various applications.
[0097] In some embodiments, the processing device based on the video information processing model provided in the embodiment of the present invention can be implemented in a combination of software and hardware. As an example, the processing device based on the video information processing model provided in the embodiment of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the video information processing method based on the video information processing model provided in the embodiment of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0098] As an example of a processing device based on a video information processing model provided in an embodiment of the present invention being implemented by a combination of software and hardware, the processing device based on a video information processing model provided in an embodiment of the present invention can be directly embodied as a combination of software modules executed by a processor 201, and the software module can be located in a storage medium, and the storage medium is located in a memory 202. The processor 201 reads the executable instructions included in the software module in the memory 202, and combines with necessary hardware (for example, including the processor 201 and other components connected to the bus 205) to complete the video information processing method based on the video information processing model provided in an embodiment of the present invention.
[0099] As an example, the processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0100] As an example of hardware implementation of the processing device based on the video information processing model provided in an embodiment of the present invention, the device provided in an embodiment of the present invention can be directly executed by a processor 201 in the form of a hardware decoding processor, for example, one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components to implement the video information processing method based on the video information processing model provided in an embodiment of the present invention.
[0101] The memory 202 in the embodiment of the present invention is used to store various types of data to support the operation of the processing device based on the video information processing model. Examples of these data include: any executable instructions for operating on the processing device based on the video information processing model, such as executable instructions, and the program for implementing the video information processing method based on the video information processing model of the embodiment of the present invention can be included in the executable instructions.
[0102] In other embodiments, the processing device based on the video information processing model provided in the embodiments of the present invention can be implemented in software. Figure 2 The processing device based on the video information processing model stored in the memory 202 is shown, which may be software in the form of a program and a plug-in, and includes a series of modules. As an example of a program stored in the memory 202, a processing device based on the video information processing model may be included. The processing device based on the video information processing model includes the following software modules:
[0103] Information transmission module 2081 and information processing module 2082. When the software modules in the processing device based on the video information processing model are read into the RAM by the processor 201 and executed, the video information processing method based on the video information processing model provided by the embodiment of the present invention will be implemented, wherein the functions of each software module in the processing device based on the video information processing model include:
[0104] The information transmission module 2081 is used to obtain a first target video and parse the first target video to obtain video parameters of the first target video;
[0105] An information processing module 2082, configured to determine basic features matching the first target video according to video parameters of the first target video;
[0106] The information processing module 2082 is used to determine the multimodal features matching the first target video according to the video parameters of the first target video;
[0107] The information processing module 2082 is used to determine a fusion feature vector matching the first target video based on the basic features and the multimodal features through the first video processing network in the video information processing model, wherein the fusion feature vector is used to be combined with the second target video fusion feature vector output by the second video processing network in the video information processing model to implement a process matching the video information processing model.
[0108] Combination Figure 2 The processing device based on the video information processing model shown in the figure illustrates the video information processing method based on the video information processing model provided by the embodiment of the present invention, see Figure 3 , Figure 3 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention is provided. It can be understood that: Figure 3 The steps shown can be executed by various electronic devices running a processing device based on a video information processing model, for example, a dedicated terminal, a server or a server cluster with a processing device based on a video information processing model, wherein the dedicated terminal with a processing device based on a video information processing model can be the preceding Figure 2 The electronic device with a processing device based on the video information processing model in the embodiment shown. Figure 3 The steps shown are explained.
[0109] Step 301: A processing device based on a video information processing model obtains a first target video, and parses the first target video to obtain video parameters of the first target video.
[0110] In some embodiments of the present invention, parsing the first target video to obtain the video parameters of the first target video may be implemented in the following manner:
[0111] The first target video is parsed to obtain the tag information of the first target video; according to the tag information of the first target video, the video information corresponding to the first target video is parsed to obtain the video parameters of the first target video in the basic dimension and the multimodal dimension respectively. Among them, the tag information of the first target video obtained can be used to decompose the video image frame and the corresponding audio file of the first target video. Since the source of the first target video is uncertain (it can be a video resource on the Internet or a local video file saved by an electronic device), by obtaining the video parameters in the basic dimension and the multimodal dimension corresponding to the first target video, the original first target video can be saved in the corresponding blockchain network, and the video parameters in the basic dimension and the multimodal dimension corresponding to the first target video can be saved in the blockchain network at the same time, so as to realize the traceability of the first target video.
[0112] Step 302: The processing device based on the video information processing model determines basic features that match the first target video according to the video parameters of the first target video.
[0113] Continue to combine Figure 2 The video information processing device of the video information processing model shown in the figure illustrates the video information processing method based on the video information processing model provided by the embodiment of the present invention. Figure 4 , Figure 4 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention is provided. It can be understood that: Figure 4 The steps shown can be performed by various electronic devices of a video information processing device running a video information processing model, for example, a dedicated terminal, a server or a server cluster with a video information processing function of a video information processing model is used to determine the basic features and multimodal dimensional features that match the first target video to determine the model parameters adapted to the video information processing model, specifically including the following steps:
[0114] Step 401: determining a category parameter, a video tag parameter, and a video publishing source parameter corresponding to the first target video according to the video parameters of the first target video in the basic dimension;
[0115] Step 402: Feature extraction is performed on the category parameters, video tag parameters and video publishing source parameters corresponding to the first target video to form basic features that match the first target video.
[0116] Step 403: Determine title text parameters, image information parameters and visual information parameters corresponding to the first target video according to the video parameters of the first target video in the basic dimension.
[0117] Step 404: extract and fuse the title text parameters, image information parameters and visual information parameters corresponding to the first target video to form a multimodal feature that matches the first target video.
[0118] Among them, in some embodiments of the present invention, the basic features are mainly used to describe the video in a basic way through definition, including video multi-level classification categories, video tags, video release sources, video length, release time, and event cities. The basic features are qualitative descriptions of the video, and the information about the content of the video itself is relatively lacking.
[0119] In some embodiments of the present invention, multimodal features are features extracted from the title text, picture information and visual information of a video, which are used to describe the content information of the video. The title and cover image can affect the video's playback click rate, and the video's visual frame image information can affect the video's playback completion rate.
[0120] Step 303: The processing device based on the video information processing model determines the multimodal features matching the first target video according to the video parameters of the first target video.
[0121] In some embodiments of the present invention, based on the basic features and the multimodal features, determining the fusion feature vector matching the first target video through the first video processing network in the video information processing model can be done in the following manner:
[0122] The basic features are processed by the basic information processing network in the first video processing network to form a corresponding basic feature vector; the image features in the multimodal features are processed by the image processing network in the first video processing network to form a corresponding image feature vector; the title text features in the multimodal features are processed by the text processing network in the first video processing network to form a corresponding title text feature vector; the visual features in the multimodal features are processed by the visual processing network in the first video processing network to form a corresponding visual feature vector; based on the basic feature vector, the image feature vector, the title text feature vector and the visual feature vector, vector fusion is performed by the first video processing network to form a fused feature vector matching the first target video. Among them, the video information processing model provided by the present invention includes a first video processing network and a second video processing network, wherein the first video processing network is used to process the first target video to form a fusion feature vector matching the first target video, and the second video processing network is used to process the second target video to form a second target video fusion feature vector matching the second target video; further, the first video processing network can be composed of different sub-networks for processing the features in the multimodal features separately.
[0123] The different sub-networks in the first video processing network are described below.
[0124] In some embodiments of the present invention, the method further comprises:
[0125] A text feature vector matching the title text feature is extracted through a text processing network; at least one word-level latent variable corresponding to the title text feature is determined through the text processing network according to the text feature vector; a processing word corresponding to the word-level latent variable and the probability of the processing word being selected are generated through the text processing network according to the at least one word-level latent variable; and at least one processing word is selected according to the probability of the processing result being selected to form a text processing result corresponding to the title text feature. Thus, not only is the title text feature of the target text processed through the text processing network to determine the appropriate title of the first target video to display, but also the title text feature in the multimodal feature is processed to form a corresponding title text feature vector.
[0126] In some embodiments of the present invention, the text processing network may be a bidirectional attention neural network model (BERTBidirectional Encoder Representations from Transformers). Figure 5 , Figure 5 The following is an optional structural diagram of a text processing network in an embodiment of the present invention, wherein the encoder comprises: N = 6 identical layers, each layer contains two sub-layers. The first sub-layer is a multi-head attention layer followed by a simple fully connected layer. Each sub-layer has a residual connection and normalization.
[0127] The decoder consists of N=6 identical layers, where the layers are different from the encoder. The layers here contain three sub-layers, including a self-attention layer, an encoder-decoder attention layer, and finally a fully connected layer. The first two sub-layers are based on a multi-head attention layer.
[0128] Continue to refer Figure 6 , Figure 6 This is a schematic diagram of the process of determining an optional word-level class latent variable of the text processing network in an embodiment of the present invention, wherein the encoder and decoder parts both contain 6 encoders and decoders. The inputs entering the first encoder are combined with embedding and positional embedding. After passing through the 6 encoders, the output is sent to each decoder of the decoder part; the input target is "Journey to the West 86 Edition Episode 35 Daughter Kingdom" and after being processed by the text processing network, the output word-level class latent variable result is: "Journey to the West-Daughter Kingdom".
[0129] Continue to refer Figure 7 , Figure 7 The figure is an optional structural diagram of an encoder in a text processing network in an embodiment of the present invention, wherein its input consists of a query (Q) and a key (K) of dimension d and a value (V) of dimension d, all keys calculate the dot product of the query, and apply the softmax function to obtain the weight of the value.
[0130] Continue to refer Figure 7 , Figure 7The vector diagram of the encoder in the text processing network in the embodiment of the present invention is shown in FIG, where Q, K and V are obtained by multiplying the vector x of the input encoder with W^Q, W^K, W^V. The dimension of W^Q, W^K, W^V in the article is (512, 64), and then assume that the dimension of our inputs is (m, 512), where m represents the number of words. Therefore, the dimension of Q, K and V obtained after multiplying the input vector with W^Q, W^K, W^V is (m, 64).
[0131] Continue to refer Figure 8 , Figure 8 This is a schematic diagram of the vector concatenation of the encoder in the text processing network in the embodiment of the present invention, where Z0 to Z7 correspond to 8 parallel heads (dimension is (m, 64)), and then after concat these 8 heads, the dimension is (m, 512). Finally, after multiplying with W^O, the output matrix with dimension (m, 512) is obtained, and the dimension of this matrix is consistent with the dimension entering the next encoder.
[0132] Continue to refer Fig. 9 , Fig. 9 This is a schematic diagram of the encoding process of the encoder in the text processing network in the embodiment of the present invention, in which x1 is in the state of z1 after self-attention. The tensor that has passed the self-attetion needs to be processed by the residual network and the Later Norm, and then enters the fully connected feedforward network. The feedforward network needs to perform the same operation, residual processing and normalization. Finally, the output tensor can enter the next encoder, and then this operation is iterated 6 times, and the result of the iterative processing enters the decoder.
[0133] Continue to refer Fig.10 , Fig.10 Schematic diagram of the decoding process of the decoder in the text processing network in an embodiment of the present invention, wherein the input, output and decoding process of the decoder are as follows:
[0134] Output: probability distribution of the output word corresponding to position i;
[0135] Input: encoder output & corresponding i-1 position decoder output. So the middle attention is not self-attention, its K, V come from the encoder, and Q comes from the output of the previous position decoder.
[0136] Continue to refer Fig.11 and Fig.12 , Fig.11Schematic diagram of the decoding process of the decoder in the text processing network in an embodiment of the present invention, wherein the vector output by the last decoder of the decoder network passes through the Linear layer and the softmax layer. Fig.12 This is a schematic diagram of the decoding process of the decoder in the text processing network in an embodiment of the present invention. The function of the Linear layer is to map the vector from the decoder part into a logits vector, and then the softmax layer converts it into a probability value based on the logits vector, and finally finds the position of the maximum probability, thus completing the output of the decoder.
[0137] In some embodiments of the present invention, the method further comprises:
[0138] Acquire an image to be processed and a target resolution corresponding to a playback interface of the first target video;
[0139] In response to the target resolution, the image to be processed is subjected to resolution enhancement processing by the image processing network in the first video processing network, and a corresponding image feature vector is obtained to achieve adaptation of the image feature vector to the target resolution corresponding to the playback interface of the first target video. Thus, not only is it achieved that the image to be processed is processed by the image processing network to determine a suitable cover image for the first target video, but it is also achieved that the image features in the multimodal features are processed to form a corresponding title image feature vector.
[0140] refer to Fig.13 , Fig.13 The present invention is an optional structural diagram of an image processing network in an embodiment of the present invention, wherein the encoder may include a convolutional neural network, and after the image feature vector is input into the encoder, a frame-level image feature vector corresponding to the image feature vector is output. Specifically, the image feature vector is input into the encoder, that is, the convolutional neural network in the encoder is input, and the frame-level image feature vector corresponding to the image feature vector is extracted through the convolutional neural network, and the convolutional neural network outputs the extracted frame-level image feature vector as the output of the encoder, and then the image feature vector output by the encoder is used to perform corresponding image semantic recognition. Alternatively, the encoder may include a convolutional neural network and a recurrent neural network, and after the image feature vector is input into the encoder, a frame-level image feature vector carrying time series information corresponding to the image feature vector is output, such as Fig.13 Specifically, the image feature vector is input into the encoder, that is, the convolutional neural network (e.g. Fig.13 The CNN neural network in the encoder extracts the frame-level image feature vector corresponding to the image feature vector through the convolutional neural network. The convolutional neural network outputs the extracted frame-level image feature vector and inputs it into the recurrent neural network in the encoder (corresponding to Fig.13 The hi-1, hi and other structures in the convolutional neural network are extracted and fused with the time series information of the extracted convolutional neural network feature vector through a recurrent neural network. The recurrent neural network outputs an image feature vector carrying time series information and uses it as the output of the encoder. The image feature vector output by the encoder is then used to perform corresponding processing steps.
[0141] In some embodiments of the present invention, the method further comprises:
[0142] Determine the bit rate information that matches the playback environment of the first target video; adjust the bit rate of the first target video by using the visual features in the multimodal features through the visual processing network in the first video processing network, so that the bit rate of the first target video matches the bit rate information of the playback environment. In this way, not only the visual information is processed by the visual processing network to determine the appropriate dynamic bit rate of the first target video, but also the visual features in the multimodal features are processed to form a corresponding title visual feature vector.
[0143] refer to Fig.14 , Fig.14 This is an optional structural diagram of an image visual processing network in an embodiment of the present invention, wherein a dual-stream long short-term memory network may include a bidirectional vector model, an attention model, a fully connected layer and a sigmoid classifier. The bidirectional vector model recursively processes different feature vectors in the input visual feature vector set, and uses the attention model to merge the recursively processed feature vectors together to form a longer vector, for example, merging associated visual feature vectors together to form a longer vector, and merging the two merged vectors together again to form a longer vector (local aggregation vector). Finally, two fully connected layers are used to map the obtained distributed feature representation to the corresponding sample label space to improve the accuracy of the final code rate. Finally, a sigmoid classifier is used to determine the probability value of each label corresponding to the image visual feature, so as to integrate the text processing results and form new text information corresponding to the image visual feature information.
[0144] Among them, the batch processing parameter (batch size) of the convolutional neural network model can be selected as 32 or 64, the initial learning rate of the adaptive optimizer (adam) selected by the convolutional neural network model optimizer can be selected as 0.0001, and the random inactivation (dropout) can be selected as 0.2. After 100,000 iterations of training, the accuracy of the training set and the test set are both stable at more than 90%, indicating that the model matches the task scenario, can achieve a relatively ideal training effect and fix all parameters of the convolutional neural network model in this state, thereby adjusting the bit rate of the first target video to achieve matching of the bit rate information of the first target video with the bit rate information of the playback environment.
[0145] Step 304: The processing device based on the video information processing model determines a fusion feature vector matching the first target video through the first video processing network in the video information processing model based on the basic features and the multimodal features.
[0146] The fused feature vector is used to combine with the second target video fused feature vector output by the second video processing network in the video information processing model to achieve a process matching the video information processing model.
[0147] Continue to combine Figure 2 The video information processing device of the video information processing model shown in the figure illustrates the video information processing method based on the video information processing model provided by the embodiment of the present invention. Fig.15 , Fig.15 An optional flow chart of a video information processing method based on a video information processing model provided in an embodiment of the present invention is provided. It can be understood that: Fig.15 The steps shown can be performed by various electronic devices of a video information processing device running a video information processing model, for example, a dedicated terminal, a server or a server cluster with a video information processing function of a video information processing model is used to determine the basic features and multimodal dimensional features that match the first target video to determine the model parameters adapted to the video information processing model, specifically including the following steps:
[0148] Step 1501: The video information processing device of the video information processing model determines the type of process that matches the video information processing model.
[0149] Step 1502: When the process matched with the video information processing model is a video recommendation process, the video information processing device of the video information processing model adjusts the parameters of the recurrent convolutional neural network based on the attention mechanism in the first video processing network according to the second target video fusion feature vector output by the second video processing network in the video information processing model.
[0150] In this way, the parameters of the attention mechanism-based recurrent convolutional neural network can be adapted to the fused feature vector.
[0151] Step 1503: The video information processing device of the video information processing model adjusts the parameters of the second video processing network in the video information processing model.
[0152] Step 1504: The video information processing device of the video information processing model determines a new second target video fusion feature vector through the second video processing network in the video information processing model with adjusted parameters.
[0153] Step 1505: The video information processing device of the video information processing model connects the new second target video fusion feature vector and the first target video fusion feature vector through a classification prediction function matching the video information processing model.
[0154] In this way, the correlation between the first target video and the second target video can be determined.
[0155] Further, when the relevance exceeds a corresponding relevance threshold, the first target video may be recommended to the corresponding terminal, otherwise, other videos may be recommended to replace the current first target video.
[0156] The video information processing method provided by the embodiment of the present invention is described below by taking the video recommendation scenario in the short video playback interface as an example, wherein: Fig.16 FIG. 1 is a schematic diagram of an application environment of a video information processing method based on a video information processing model in an embodiment of the present invention, wherein Fig.16 As shown, the short video playback interface can be displayed in the corresponding APP, or it can be triggered by the WeChat applet (the video information processing model can be encapsulated in the corresponding APP after training or saved in the WeChat applet in the form of a plug-in). With the continuous development and increase of short video application products, the carrying capacity of video information is far greater than that of text information. Short videos can be continuously recommended to users through the corresponding application. Therefore, when the above-played video (i.e., the second target video involved in the previous embodiment) is known, recommending subsequent related videos (i.e., the first target video in the previous embodiment) is a very important link. Effective recommendation of subsequent related videos can effectively improve the user experience. The above video represented by the second target video can be either a video played before the first target video is displayed, or a video collection of several videos played before the first target video is displayed. In this process, the vectorized representation of video information is the basis of many machine learning algorithms.
[0157] In traditional technology centers, common learning methods include: 1) Directly using video tags, including video classification, video theme, video release source, etc. Videos can be roughly divided into entertainment videos, sports videos, or subdivided into basketball highlights, film and television highlights. However, this type of representation method is relatively extensive, and the classification tag information needs to be set in advance and updated in a timely manner, and its content representation ability is limited. 2) Text-based learning, including text semantic learning of video titles, video description information or video tags. This type of method relies more on the accuracy of text information, but many videos lack text information, which makes the video representation inaccurate. These methods have a serious impact on the user experience.
[0158] Fig.17 The working process diagram of the video information processing method based on the video information processing model provided by the embodiment of the present invention is as follows: Fig.18 The structure diagram of the video information processing device of the video information processing model provided by the embodiment of the present invention is as follows. Fig.18 The structural schematic diagram of the video information processing device of the video information processing model shown in the figure illustrates the working process of the question-answering model of the video information processing method of the present invention, which specifically includes the following steps:
[0159] Step 1701: Obtain relevant video pair annotations in the video data source.
[0160] Among them, the title can be used in the video data source of the video server to simply calculate the text relevance to obtain related video pairs. It should be noted that the related video pairs here are not directional, and testers need to manually annotate part of the seed data. This part of the seed data is a directional data pair, that is, video A can recommend video B, but not vice versa. Furthermore, the video information processing model can be used to learn the prediction results to process the video data, spot check the quality of the video pairs, and gradually expand the annotated data pairs to achieve the expansion of the data pairs. After obtaining the positive sample, you can also randomly select video pairs of the same magnitude as negative examples for training the video information processing model.
[0161] Step 1702: Obtain video multimodal features that match the first target video.
[0162] Among them, video features are summarized into two categories, namely basic features and multimodal features. Specifically, basic features mainly provide basic descriptions of videos through definitions, including: multi-level video classification categories, video tags, video release sources, video length, release time, and event cities. Basic features are qualitative descriptions of videos, but they lack information about the content of the video itself.
[0163] Multimodal features are features extracted from the title text, image information, and visual adjustments of a video. They are used to describe the content information of a video. The title and cover image can affect the video's playback click-through rate, and the video's visual frame image information can affect the video's playback completion rate.
[0164] The title feature is extracted using a pre-trained model of natural language processing. The pre-trained model uses an optional structure as a bidirectional attention neural network model BERT (Bidirectional Encoder Representation from Transformers), which is used to send the video title sentence into the model task to obtain a 64-dimensional (the dimension size can be customized) title feature vector. The BERT model is used to further increase the generalization ability of the word vector model and achieve sentence-level representation capabilities.
[0165] The cover image feature is extracted using a pre-trained convolutional neural network based on deep residual resnet50, and the video cover image information is extracted into a 128-dimensional feature vector. Resnet is currently a widely used extraction network in image feature extraction, which is conducive to the representation of cover image information. The cover image information has a great eye-catching appeal before users watch it, and a reasonable and appropriate cover image can greatly improve the video's playback click-through rate.
[0166] Visual features are extracted using the netvladVector of locally aggregated descriptors processed by video, and the video frame image is converted into a 128-bit feature vector. During video viewing, the video frame information reflects the specific content and quality of the video, which is directly related to the user's viewing time.
[0167] Step 1703: Process the first target video information through a video information processing model to form matching video multimodal information.
[0168] Among them, continue to refer to Fig.18 The structural diagram of the video information processing model of the video information processing model shown in FIG. Fig.18The multimodal information shown in can be divided into four domains, namely basic information, cover image information, title information, and visual information. The ID-type features in the basic information are initialized one-hot features and then learned in the embedding layer (embedding) of the video information processing model. The cover image, title vector, and visual information are embeddings obtained through the pre-training method mentioned in the previous step. Fully connected re-representation as a 128-dimensional vector is performed in each domain, and then the vectors of these four domains are concatenated to generate a 128*4=512-dimensional vector, followed by two layers of fully connected layers. The first layer of full connection reduces the dimension to 256 dimensions, and then the second layer of full connection generates the final 128-dimensional video vector representation.
[0169] Next, continue to combine Fig.18 ,right Fig.18 The attention mechanism shown in the figure is introduced. Through the attention mechanism, when generating each feature domain for predicting related videos, the above attention mechanism is added to learn the weights of each vector. The attention mechanism (Attention) process is as follows:
[0170] 1) First, calculate the weight of the second target video vector to the relevant video in this domain as
[0171]
[0172] 2) Then use softmax to normalize the weights
[0173]
[0174] 3) Finally, weight the weights and corresponding key values to obtain the vector representation under the above attention mechanism
[0175] Where Q abv represents the second target video vector, K rel The vector learned by the embedding network of the related video. After the attention mechanism is processed for each domain, it is equivalent to associating the learning of the second target video vector and the related video vector by weighting.
[0176] Furthermore, for example, the video output is re-expressed in different dimensions to realize the difference in information content between the second target video and the related video. The related video is re-expressed as a 128-dimensional new vector high embedding based on the 128-dimensional video, while the second target video only generates a 64-dimensional new vector low embedding, thereby reducing the amount of information in the second target video and relatively increasing the amount of information in the related video. After that, the two embeddings are spliced together and represented separately as two 64-dimensional intermediate vectors. Finally, the two intermediate vectors are spliced together and the sigmoid function classification is performed to predict whether the second target video can point to the related video.
[0177] Classification prediction sigmod function:
[0178]
[0179] The loss function uses weighted cross entropy loss:
[0180]
[0181] where θ k represents the input of the kth sample, p k represents the estimated classification of the kth sample, y k Indicates the actual classification of the kth sample. k Represents the sample weight, which is proportional to the number of times the second target video and the related video appear.
[0182] Beneficial technical effects:
[0183] 1) Compared with the traditional technology, the technical solution provided by the present application processes the video information of the first target video to form matching video multimodal information, which integrates the multimodal features of the first target video and can better express the features of the first target video, which is beneficial to the subsequent operations on the first target video.
[0184] 2) When video multimodal information is used to realize video push or prediction, the second target video and related videos can be effectively distinguished, and the different information representations contained in the second target video and related videos in the information representation dimension can be distinguished. At the same time, the lack of scalability of the simple video tag method in the traditional method can be corrected, and targeted recommendations can be made based on time-related videos that meet the relevance requirements, thereby improving the rationality of the recommended content and enhancing the user experience.
[0185] The above descriptions are merely embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A video information processing method based on a video information processing model, characterized in that: The method comprises: Acquire a first target video, and parse the first target video to acquire video parameters of the first target video; Determining basic features matching the first target video according to the video parameters of the first target video; Determining, according to the video parameters of the first target video, a multimodal feature matching the first target video; Based on the basic features and the multimodal features, a fusion feature vector matching the first target video is determined through the first video processing network in the video information processing model, wherein the fusion feature vector is used to combine with the second target video fusion feature vector output by the second video processing network in the video information processing model to achieve a process matching the video information processing model; the video data source used in training the video information processing model includes related video pairs, and the related video pairs are non-directional; the related video pairs include part of seed data, and the seed data is a data pair with directionality; the prediction result of the seed data is learned through the video information processing model, the labeled data pairs are gradually expanded to obtain positive samples, and video pairs of the same magnitude are randomly selected as negative examples; the video information processing model is trained through the positive samples and the negative examples.
2. The method according to claim 1, characterized in that The parsing the first target video to obtain the video parameters of the first target video includes: Parsing the first target video to obtain tag information of the first target video; According to the tag information of the first target video, the video information corresponding to the first target video is parsed to respectively obtain video parameters of the first target video in a basic dimension and a multimodal dimension.
3. The method according to claim 2, characterized in that The determining, according to the video parameters of the first target video, basic features that match the first target video includes: According to the video parameters of the first target video in the basic dimension, Determining a category parameter, a video tag parameter, and a video publishing source parameter corresponding to the first target video; Feature extraction is performed on the category parameters, video tag parameters and video publishing source parameters corresponding to the first target video to form basic features that match the first target video.
4. The method according to claim 2, characterized in that: The determining, according to the video parameters of the first target video, a multimodal feature matching the first target video includes: According to the video parameters of the first target video in the basic dimension, Determining title text parameters, image information parameters, and visual information parameters corresponding to the first target video; Feature extraction and fusion are performed on the title text parameters, image information parameters and visual information parameters corresponding to the first target video to form a multimodal feature that matches the first target video.
5. The method according to any one of claims 1 to 4, characterized in that: The determining, based on the basic features and the multimodal features, a fusion feature vector matching the first target video through a first video processing network in the video information processing model includes: Processing the basic features through a basic information processing network in the first video processing network to form a corresponding basic feature vector; Processing the image features in the multimodal features through the image processing network in the first video processing network to form a corresponding image feature vector; Processing the title text features in the multimodal features through the text processing network in the first video processing network to form a corresponding title text feature vector; Processing the visual features in the multimodal features through a visual processing network in the first video processing network to form a corresponding visual feature vector; Based on the basic feature vector, the image feature vector, the title text feature vector and the visual feature vector, vector fusion is performed through the first video processing network to form a fused feature vector that matches the first target video.
6. The method according to claim 5, characterized in that The method further comprises: Acquire an image to be processed and a target resolution corresponding to a playback interface of the first target video; In response to the target resolution, the image to be processed is subjected to resolution enhancement processing by an image processing network in the first video processing network, and a corresponding image feature vector is obtained to achieve adaptation of the image feature vector to the target resolution corresponding to the playback interface of the first target video.
7. The method according to claim 5, characterized in that The method further comprises: Extracting a text feature vector matching the title text feature through a text processing network; Determining at least one word-level latent variable corresponding to the title text feature according to the text feature vector through the text processing network; Generate, by means of the text processing network, a processing word corresponding to the word-level latent variable and a probability of selection of the processing word according to the at least one word-level latent variable; According to the probability of the processing words being selected, at least one processing word is selected to form a text processing result corresponding to the title text feature.
8. The method according to claim 5, characterized in that The method further comprises: Determining bit rate information that matches the playback environment of the first target video; The bit rate of the first target video is adjusted by using the visual features in the multimodal features through the visual processing network in the first video processing network, so that the bit rate of the first target video matches the bit rate information of the playback environment.
9. The method according to claim 1, characterized in that: The method further comprises: When the process matching the video information processing model is a video recommendation process, According to the second target video fusion feature vector output by the second video processing network in the video information processing model, the parameters of the attention-based recurrent convolutional neural network in the first video processing network are adjusted to achieve adaptation of the parameters of the attention-based recurrent convolutional neural network to the fusion feature vector.
10. The method according to claim 9, characterized in that The method further comprises: Adjusting parameters of a second video processing network in the video information processing model; Determining a new second target video fusion feature vector through a second video processing network in the video information processing model after parameter adjustment; The new second target video fusion feature vector and the first target video fusion feature vector are connected through a classification prediction function that matches the video information processing model to determine the correlation between the first target video and the second target video.
11. A processing device based on a video information processing model, characterized in that: The device comprises: An information transmission module, used for acquiring a first target video, and parsing the first target video to acquire video parameters of the first target video; An information processing module, configured to determine basic features matching the first target video according to video parameters of the first target video; The information processing module is used to determine the multimodal features matching the first target video according to the video parameters of the first target video; The information processing module is used to determine a fused feature vector matching the first target video based on the basic features and the multimodal features through the first video processing network in the video information processing model, wherein the fused feature vector is used to combine with the second target video fused feature vector output by the second video processing network in the video information processing model to achieve a process matching the video information processing model; the video data source used when training the video information processing model includes related video pairs, and the related video pairs are non-directional; the related video pairs include part of seed data, and the seed data is a data pair with directionality; the prediction result of the seed data is learned through the video information processing model, the labeled data pairs are gradually expanded to obtain positive samples, and video pairs of the same magnitude are randomly selected as negative examples; the video information processing model is trained through the positive samples and the negative examples.
12. The device according to claim 11, characterized in that The information processing module is used to parse the first target video and obtain label information of the first target video; The information processing module is used to parse the video information corresponding to the first target video according to the label information of the first target video, so as to obtain the video parameters of the first target video in the basic dimension and the multimodal dimension respectively.
13. The device according to claim 12, characterized in that The information processing module is used to, based on the video parameters of the first target video in the basic dimension, The information processing module is used to determine a category parameter, a video tag parameter and a video publishing source parameter corresponding to the first target video; The information processing module is used to extract features from the category parameters, video tag parameters and video publishing source parameters corresponding to the first target video, so as to form basic features that match the first target video.
14. An electronic device, characterized in that: The electronic device comprises: A memory for storing executable instructions; The processor is used to implement the video information processing method based on the video information processing model described in any one of claims 1 to 10 when running the executable instructions stored in the memory.
15. A computer-readable storage medium storing executable instructions, characterized in that: When the executable instructions are executed by the processor, the video information processing method based on the video information processing model described in any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Video recall method and device and storage medium
CN110446065A