Video information recommendation method, device, electronic device and storage medium

Video data is obtained through a video information recommendation model, and dynamic adjustments are made using the content category and timeliness category processing network. This solves the problems of unstable quality and low efficiency caused by manual labeling, achieves efficient and accurate video recommendations, and improves user experience.

CN115146107BActive Publication Date: 2025-09-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110346151.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-09-19
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

The existing video recommendation methods rely on manual labeling, which results in unstable labeling quality and low efficiency, making it difficult to achieve accurate and efficient video recommendation.

Method used

Video data is obtained through the video information recommendation model, the attribute feature vector of the video information is determined, and the content category and timeliness category processing network are used to dynamically adjust the recommendation strategy, including feature vector extraction and model training to adapt to different video playback environments.

Benefits of technology

It achieves accurate and timely recommendation of video information, improves user experience, and improves the quality and efficiency of video recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115146107B_ABST
    Figure CN115146107B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, and electronic device for recommending video information. The method comprises: processing the video information to be recommended through a text processing network in a video information recommendation model to determine the title feature vector and tag feature vector corresponding to the video information to be recommended; determining the content category of the video information to be recommended through a content category processing network of the video information recommendation model based on feature vectors of multiple attributes; and determining the timeliness category of the video information to be recommended through a timeliness category processing network of the video information recommendation model based on feature vectors of multiple attributes. This method not only enhances the accuracy and timeliness of video information recommendations, but also effectively improves the quality of video information recommendations and enhances the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information processing technology, and in particular to a video information recommendation method, device, and electronic equipment. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive field in computer science. By studying the design principles and implementation methods of various intelligent machines, AI enables them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, including natural language processing and machine learning / deep learning. With technological advancements, AI will be applied in even more areas and play an increasingly important role.

[0003] In traditional technology, on the one hand, the timeliness of the video to be pushed is determined by manual annotation. Manual methods often have subjective timeliness determination standards, which will lead to unstable annotation quality. At the same time, manual annotation is inefficient and costly. On the one hand, by extracting information of various types of text in the video to be pushed (such as subtitle text information), the timeliness corresponding to each type of text information is determined, and then the timeliness of the video to be pushed is determined by combining the timeliness corresponding to each type of text information. Since the process of extracting various types of text information is time-consuming, it will affect the efficiency of video timeliness determination, especially for a large number of videos to be pushed. Therefore, it is necessary to provide an accurate and efficient video recommendation method so that users can get a better video recommendation experience. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method, device, electronic device, and storage medium for recommending video information. The technical solution of the embodiments of the present invention is implemented as follows:

[0005] An embodiment of the present invention provides a video information recommendation method including:

[0006] Obtain information about videos to be recommended from a video data source;

[0007] Processing the video information to be recommended through a processing network in a video information recommendation model to determine feature vectors of multiple attributes corresponding to the video information to be recommended;

[0008] Based on the feature vectors of the multiple attributes, determining the content category of the video information to be recommended through the content category processing network of the video information recommendation model;

[0009] Based on the feature vectors of the multiple attributes, determining the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model;

[0010] Determining the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended;

[0011] Based on the type of the video information to be recommended, the recommendation strategy of the video to be recommended is dynamically adjusted.

[0012] An embodiment of the present invention further provides a video information recommendation device, comprising:

[0013] An information transmission module is used to obtain information of videos to be recommended from a video data source;

[0014] An information processing module, configured to process the video information to be recommended through a processing network in a video information recommendation model, and determine feature vectors of multiple attributes corresponding to the video information to be recommended;

[0015] The information processing module is configured to determine the content category of the video information to be recommended through the content category processing network of the video information recommendation model based on the feature vectors of the multiple attributes;

[0016] The information processing module is configured to determine the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model based on the feature vectors of the multiple attributes;

[0017] The information processing module is used to determine the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended;

[0018] The information processing module is used to dynamically adjust the recommendation strategy of the video to be recommended based on the type of the video information to be recommended.

[0019] In the above scheme,

[0020] The information processing module is used to perform data screening processing on the video information to be recommended, and parse and obtain the title and tag of the video information to be recommended;

[0021] The information processing module is used to trigger the target word segmentation library and perform word segmentation processing on the title and label of the video information to be recommended respectively through the target word segmentation library to obtain the video information to be recommended at the word level;

[0022] The information processing module is used to perform vectorization processing on the word-level video information to be recommended through the text processing network in the video information recommendation model to form a multi-dimensional word-level title feature vector and a multi-dimensional word-level label feature vector of the video information to be recommended.

[0023] In the above scheme,

[0024] The information processing module is configured to determine historical parameters of the video information to be recommended based on the type of the video playback environment in which the video information to be recommended is located;

[0025] The information processing module is configured to determine a training sample set that matches the video information recommendation model based on historical parameters of the video information to be recommended, wherein the training sample set includes at least one group of training samples;

[0026] The information processing module is used to extract a training sample set that matches the training sample through a noise threshold that matches the video information recommendation model;

[0027] The information processing module is used to train the video information recommendation model according to a training sample set that matches the training sample.

[0028] In the above scheme,

[0029] The information processing module is used for, when the network structures of the content category processing network and the timeliness category processing network are different,

[0030] The information processing module is used to determine a multi-task loss function that matches the video information recommendation model;

[0031] The information processing module is used to adjust the parameters of the content category processing network and the network parameters of the timeliness category processing network in the video information recommendation model based on the multi-task loss function until the loss functions of different dimensions corresponding to the video information recommendation model reach corresponding convergence conditions; so as to achieve the adaptation of the parameters of the video information recommendation model to the video playback environment.

[0032] In the above scheme,

[0033] The information processing module is configured to determine a dynamic noise threshold that matches the usage environment of the video information recommendation model when the video playback environment of the video information to be recommended is a short video recommendation;

[0034] The information processing module is configured to perform noise removal processing on the first training sample set according to the dynamic noise threshold to form a second training sample set that matches the dynamic noise threshold;

[0035] The information processing module is used to determine a fixed noise threshold corresponding to the video information recommendation model when the video playback environment of the video information to be recommended is a long video playback, and to remove noise from the first training sample set according to the fixed noise threshold to form a second training sample set that matches the fixed noise threshold.

[0036] In the above scheme,

[0037] The information processing module is used to, when the network structures of the content category processing network and the timeliness category processing network are the same,

[0038] The information processing module is used to determine a loss function that matches the video information recommendation model;

[0039] The information processing module is used to adjust the parameters of the content category processing network and the network parameters of the timeliness category processing network in the video information recommendation model based on the loss function until the loss functions of different dimensions corresponding to the video information recommendation model reach corresponding convergence conditions; so as to achieve the adaptation of the parameters of the video information recommendation model to the video playback environment.

[0040] In the above scheme,

[0041] The information processing module is configured to dynamically adjust the exposure rate corresponding to the video information to be recommended; or

[0042] Adjusting the exposure channel corresponding to the video information to be recommended; or

[0043] The exposure position corresponding to the video information to be recommended is adjusted.

[0044] In the above scheme,

[0045] The information processing module is configured to monitor exposure parameters of a short video when the video information to be recommended is embedded advertising information in the short video;

[0046] The information processing module is used to determine the Manrong visual effect index corresponding to the embedded advertising information according to the exposure parameters of the short video;

[0047] The information processing module is used to adjust the playback configuration information corresponding to the embedded advertising information in the short video through the topology file according to the Manrong visual effect index corresponding to the embedded advertising information.

[0048] In the above scheme,

[0049] The information processing module is used to send the exposure parameters of the short video during playback to the monitoring server, so that the monitoring server can monitor the exposure of the short video;

[0050] The information processing module is used to use the exposure parameters of the short video saved by the monitoring server during playback as the data source of the playback effect parameters of the video information to be recommended.

[0051] In the above scheme,

[0052] The information processing module is used to obtain the historical browsing information of the target user;

[0053] The information processing module is configured to determine, based on the historical browsing information of the target user, an exposure history of the video information to be recommended corresponding to the historical browsing information;

[0054] The information processing module is used to dynamically adjust the playback strategy of the video information to be recommended based on the exposure history of the video information to be recommended corresponding to the historical browsing information.

[0055] In the above scheme,

[0056] The information processing module is used to determine the usage environment of the terminal display interface;

[0057] The information processing module is used to determine the category of the video information to be played or recommended according to the usage environment of the terminal display interface;

[0058] The information processing module is used to trigger a matching data source of the video information to be recommended in response to the category of the video information to be played and recommended, so as to adjust the video information to be played and recommended through the data source of the video information to be recommended that matches the category of the video information to be played and recommended.

[0059] An embodiment of the present invention further provides an electronic device, comprising:

[0060] a memory for storing executable instructions;

[0061] The processor is configured to implement the aforementioned video information recommendation method when running the executable instructions stored in the memory.

[0062] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, which implement the aforementioned video information recommendation method when executed by a processor.

[0063] The present invention obtains the video information to be recommended from a video data source; processes the video information to be recommended through a text processing network in a video information recommendation model to determine the title feature vector and label feature vector corresponding to the video information to be recommended; determines the content category of the video information to be recommended through a content category processing network of the video information recommendation model based on the feature vectors of multiple attributes; determines the timeliness category of the video information to be recommended through a timeliness category processing network of the video information recommendation model based on the feature vectors of multiple attributes; determines the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended; and dynamically adjusts the recommendation strategy of the video to be recommended based on the type of the video information to be recommended. In this way, the video information recommendation model can classify and recommend video information in the usage environment, while enhancing the accuracy and timeliness of video information recommendations, effectively improving the quality of video information recommendations, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of a usage scenario of the video information recommendation method provided by an embodiment of the present invention;

[0065] Figure 2 A schematic diagram of the structure of a video information recommendation device provided by an embodiment of the present invention;

[0066] Figure 3 An optional flowchart of a video information recommendation method provided by an embodiment of the present invention;

[0067] Figure 4 This is a schematic diagram of an optional time-sensitive short video playback in an embodiment of the present invention;

[0068] Figure 5 This is a schematic diagram of an optional time-sensitive short video recommendation process in an embodiment of the present invention;

[0069] Figure 6 An optional flowchart of a video information recommendation method provided by an embodiment of the present invention;

[0070] Figure 7 Schematic diagram of an application environment of a training method for a video information recommendation model according to an embodiment of the present invention;

[0071] Figure 8 A schematic diagram of an optional data processing architecture of the video information recommendation method provided in an embodiment of the present invention;

[0072] Figure 9 A schematic diagram of the working process of the video information recommendation method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0074] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0075] Before further explaining the embodiments of the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.

[0076] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0077] 2) Based on, used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0078] 3) Model training: Multi-classification learning is performed on image datasets. This model can be built using deep learning frameworks such as TensorFlow and Torch, using multiple layers of neural network layers such as CNN to form a multi-classification model. The model input is a three-channel or raw channel matrix of images read using tools such as OpenCV. The model output is multi-class probabilities, and finally the webpage category is output using algorithms such as softmax. During training, the model approaches the correct trend using objective functions such as cross entropy.

[0079] 4) Neural Network (NN): Artificial Neural Network (ANN), also known as neural network or quasi-neural network, is a mathematical model or computational model that imitates the structure and function of biological neural networks (the central nervous system of animals, especially the brain) in the fields of machine learning and cognitive science. It is used to estimate or approximate functions.

[0080] 5) Multi-task Learning: In the field of machine learning, by simultaneously jointly learning and optimizing multiple related tasks, better model accuracy can be achieved than that of a single task. Multiple tasks help each other by sharing a representation layer. This training method is called multi-task learning, also known as joint learning.

[0081] 6) Video Timeliness: Video content has a certain effect over a period of time, and this effect is measured by user interest in the video content. Generally speaking, timeliness means that videos pushed to users within a certain timeframe do not expire. Timeliness plays a crucial role in end-to-end online user retention, clickthrough rates, and CTR. Pushing content to users within the content's timeliness period has a positive impact; otherwise, it can cause user dissatisfaction.

[0082] The embodiments of the present invention may be implemented in conjunction with cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network or local area network to enable data computing, storage, processing, and sharing. It can also be understood as a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. Backend services of technical network systems, such as video websites, image websites, and more portal websites, require a large amount of computing and storage resources. Therefore, cloud technology needs to be supported by cloud computing.

[0083] It should be noted that cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, the resources in the "cloud" appear to be infinitely scalable and can be accessed at any time, used on demand, and expanded at any time, with a pay-per-use fee. As a provider of cloud computing's basic capabilities, a cloud computing resource pool platform, often referred to as Infrastructure as a Service (IaaS), is established. Various types of virtual resources are deployed within the resource pool for external customers to choose from. The cloud computing resource pool primarily includes computing devices (which can be virtualized machines, including operating systems), storage devices, and network devices.

[0084] 7) Video information: various forms of information available on the Internet, such as video files presented on the client or smart devices, recommended video information, news information, etc.

[0085] Figure 1 Schematic diagram of the use scenario of the video information recommendation method provided by the embodiment of the present invention, see Figure 1The terminals (including terminal 10-1 and terminal 10-2) are equipped with software clients capable of displaying different video information, such as video playback clients or plug-ins. Users can obtain and display different video information (such as different short video information or news information) through the corresponding clients. The terminals are connected to server 200 via network 300. Network 300 can be a wide area network or a local area network, or a combination of the two, using a wireless link to achieve data transmission. Server 200 can also be a node in the blockchain.

[0086] As an example, the server 200 is used to deploy a corresponding video information recommendation model to implement the video information recommendation method provided by the present invention, or to deploy a video information recommendation device to implement the video information recommendation method. Specifically, the video information recommendation processing includes: obtaining the video information to be recommended from the video data source; processing the video information to be recommended through the text processing network in the video information recommendation model to determine the title feature vector and label feature vector corresponding to the video information to be recommended; based on the feature vectors of multiple attributes, determining the content category of the video information to be recommended through the content category processing network of the video information recommendation model; based on the feature vectors of multiple attributes, determining the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model; determining the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended; based on the type of the video information to be recommended, dynamically adjusting the recommendation strategy of the video to be recommended, and displaying and outputting the video information to be recommended that matches the target user through the terminal (terminal 10-1 and / or terminal 10-2). Taking short video information as an example, the video information recommendation model provided by the present invention can be applied to short video playback. In short video playback, different short video information from different data sources are usually processed, and finally the corresponding different video information and the corresponding short video recommendation process are presented on the user interface UI (User Interface). The accuracy and timeliness of the characteristics of different video information directly affect the user experience. The background database of video playback receives a large amount of video data from different sources every day, and the different video information obtained for recommending video information to target users can also be called by other applications (for example, the recommendation results of the short video recommendation process are migrated to the long video recommendation process or the news recommendation process). Of course, the video information recommendation model that matches the corresponding target user can also be migrated to different video recommendation processes (for example, a web video recommendation process, a mini-program video recommendation process, or a video recommendation process of a long video client).

[0087] Among them, the video information recommendation method provided in the embodiment of the present application is based on artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0088] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0089] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned speech processing technology and machine learning. For example, it may involve the speech recognition technology (Automatic Speech Recognition, ASR) in speech technology, including speech signal preprocessing, speech signal frequency analysis, speech signal feature extraction, speech signal feature matching / recognition, speech training, etc.

[0090] For example, it can involve machine learning (ML), which is a multidisciplinary interdisciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory, and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning generally includes technologies such as deep learning. Deep learning includes artificial neural networks (ANNs), such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep neural networks (DNNs).

[0091] It is understandable that the video information recommendation method and voice processing provided in this application can be applied to smart devices (Intelligent device). The smart device can be any device with information display function, such as smart terminals, smart home devices (such as smart speakers, smart washing machines, etc.), smart wearable devices (such as smart watches), in-vehicle intelligent central control systems (displaying video information to users through mini-programs that perform different tasks) or AI smart medical devices (displaying treatment cases by displaying video information), etc.

[0092] The structure of the video information recommendation device according to the embodiment of the present invention is described in detail below. The video information recommendation device can be implemented in various forms, such as a dedicated terminal with a video information recommendation processing function, or a server with a video information recommendation processing function, such as the preceding embodiment. Figure 1 Server 200 in. Figure 2 This is a schematic diagram of the structure of the video information recommendation device provided by an embodiment of the present invention. It can be understood that Figure 2 Only the exemplary structure of the video information recommendation device is shown, not the complete structure. Figure 2 Partial or complete structure shown.

[0093] The video information recommendation device provided in the embodiment of the present invention includes: at least one processor 201, a memory 202, a user interface 203 and at least one network interface 204. The various components in the video information recommendation device are coupled together via a bus system 205. It is understood that the bus system 205 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 205 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 205 is not described in detail. Figure 2 Various buses are labeled as bus system 205 .

[0094] The user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a camera, a touch pad or a touch screen.

[0095] It is understood that the memory 202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The memory 202 in the embodiment of the present invention can store data to support the operation of the terminal (such as 10-1). Examples of such data include: any computer program used to operate on the terminal (such as 10-1), such as an operating system and an application program. Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application program can include various application programs.

[0096] In some embodiments, the video information recommendation device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the video information recommendation device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the video information recommendation model provided by the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0097] As an example of a video information recommendation device provided by an embodiment of the present invention being implemented by a combination of software and hardware, the video information recommendation device provided by an embodiment of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software module can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software module in the memory 202, and combines with the necessary hardware (for example, including the processor 201 and other components connected to the bus 205) to complete the training method of the video information recommendation model provided by the embodiment of the present invention.

[0098] As an example, the processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0099] As an example of a hardware implementation of the video information recommendation device provided in an embodiment of the present invention, the device provided in an embodiment of the present invention can be directly executed using a processor 201 in the form of a hardware decoding processor. For example, the training method of the video information recommendation model provided in an embodiment of the present invention can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0100] The memory 202 in this embodiment of the present invention is used to store various types of data to support the operation of the video information recommendation device. Examples of such data include any executable instructions for operating on the video information recommendation device, such as executable instructions. A program implementing the training method for a video information recommendation model in this embodiment of the present invention may be included in the executable instructions.

[0101] In other embodiments, the video information recommendation device provided by the embodiment of the present invention can be implemented in software. Figure 2 The video information recommendation device stored in the memory 202 is shown. The device may be software in the form of a program or plug-in, and may include a series of modules. As an example of a program stored in the memory 202, a video information recommendation device may be included. The video information recommendation device includes the following software modules:

[0102] Information transmission module 2081 and information processing module 2082. When the software modules in the video information recommendation device are read into RAM and executed by the processor 201, the training method of the video information recommendation model provided by the embodiment of the present invention will be implemented. The functions of each software module in the video information recommendation device include: an information transmission module for obtaining the target user's behavior parameter information in response to a video information recommendation request;

[0103] The information processing module 2081 is used to obtain the video information to be recommended from the video data source.

[0104] An information processing module 2082 is configured to process the video information to be recommended through a processing network in a video information recommendation model to determine feature vectors of multiple attributes corresponding to the video information to be recommended;

[0105] The information processing module 2082 is used for the information processing module to determine the content category of the video information to be recommended through the content category processing network of the video information recommendation model based on the feature vectors of the multiple attributes.

[0106] The information processing module 2082 is configured to determine the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model based on the feature vectors of the multiple attributes.

[0107] The information processing module 2082 is configured to determine the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended.

[0108] The information processing module 2082 is configured to dynamically adjust the recommendation strategy of the video to be recommended based on the type of the video information to be recommended.

[0109] Combine Figure 2 The video information recommendation device shown in the figure illustrates the video information recommendation method provided by the embodiment of the present invention. Figure 3 , Figure 3 This is an optional flow chart of the video information recommendation method provided by an embodiment of the present invention. It can be understood that: Figure 3 The steps shown can be executed by various electronic devices running the video information recommendation device, for example, a dedicated terminal, a server or a server cluster with a video information recommendation device, or an advertising information server in video playback, wherein the dedicated terminal with the video information recommendation device can be the preamble Figure 2 The electronic device with the video information recommendation device in the embodiment shown is as follows. Figure 3 The steps shown are explained.

[0110] Step 301: The video information recommendation apparatus receives a video information recommendation request sent by a terminal.

[0111] Among them, in related technologies, when judging the timeliness of specific video content through machine learning or deep learning, it can only be processed in a relatively coarse-grained manner. For example, an intelligent classification model for the timeliness of video content is used to classify the timeliness according to the video content category. The specific number of timeliness categories is generally given by humans subjectively and based on experience. If the number of categories is set to a large number, it will not be beneficial to the training of the entire model, and the relative accuracy rate cannot be too high. Therefore, when the number of categories is relatively reasonable, the timeliness classification is relatively coarse-grained.

[0112] Using rules to match video timeliness is simpler. If a time-related keyword is matched, the corresponding timeliness is assigned. For example, if the keywords "today," "tomorrow," or "yesterday" are matched, the timeliness is assigned to a value of less than 24 hours. However, rules have a relatively narrow coverage, and keywords cannot cover all scenarios. This can lead to high accuracy but low recall in real-world scenarios, making them useful only as a supplementary or more granular method for timeliness assessment.

[0113] Step 302: The video information recommendation apparatus obtains the video information to be recommended from the video data source in response to the video information recommendation request.

[0114] In some embodiments of the present invention, the various user behaviors matched by the corresponding client can be collected through different program components, and the original logs of the user behavior data can be effectively extracted, such as the user's device number (user account), video information type, video information browsing time, and video information browsing completeness parameters. Among them, the user's historical click behavior and the browsing time of the corresponding information will be recorded through the subscription service and stored in Redis. The online recommendation system will pull the corresponding user's historical click behavior when the user request arrives. Furthermore, based on the use environment of the video information recommendation model, an exposure threshold that matches the use environment of the video information recommendation model can be determined; the exposure parameters carried by different video information in the video information source can be obtained; the exposure parameters carried by the different video information can be traversed through the exposure threshold to determine the video information to be promoted in the use environment of the video information recommendation model.

[0115] Step 303: The video information recommendation device processes the video information to be recommended through the processing network in the video information recommendation model to determine feature vectors of multiple attributes corresponding to the video information to be recommended.

[0116] Taking advertising video recommendation as an example, advertisers place video advertisements as videos to be recommended, which are inserted into the video information to be played for users to watch. In order to achieve a more accurate recommendation effect, it is necessary to identify the content of the advertising video. In this process, the processing network of the video information recommendation model can be a text processing network for text information recognition. The text processing network in the video information recommendation model performs feature extraction on the video information to be recommended to obtain feature extraction results; based on the feature extraction results of the text processing network, the title feature vector and label feature vector corresponding to the video information to be recommended are determined. In this way, the label content and title content in the video to be recommended can be accurately determined.

[0117] In some embodiments of the present invention, processing the video information to be recommended by a text processing network in a video information recommendation model to determine the title feature vector and label feature vector corresponding to the video information to be recommended can be achieved by:

[0118] The video information to be recommended is subjected to data screening and parsing to obtain the title and tags of the video information to be recommended. A target word segmentation library is triggered and the title and tags of the video information to be recommended are segmented separately using the target word segmentation library to obtain the video information to be recommended at the word level. The video information to be recommended is vectorized using the text processing network in the video information recommendation model to form a multi-dimensional word-level title feature vector and a multi-dimensional word-level tag feature vector for the video information to be recommended. The videos to be recommended in the video data source typically carry basic features to facilitate recommendation to different video clients. Basic features primarily provide a basic description of the video through definitions, including multi-level video classification categories, video tags, video release source, video length, release time, and event city. Basic features provide a qualitative description of the video but lack information about the video's content. Feature extraction is performed on the video's title text and tag information to describe the video's content information and temporal characteristics, allowing for more accurate identification of the video's type.

[0119] Step 304: The video information recommendation apparatus determines the content category of the video information to be recommended through the content category processing network of the video information recommendation model based on the feature vectors of the multiple attributes.

[0120] Step 305: The video information recommendation device determines the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model based on the feature vectors of multiple attributes.

[0121] Step 306: The video information recommendation device determines the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended.

[0122] Step 307: The video information recommendation device dynamically adjusts the recommendation strategy of the video to be recommended based on the type of the video information to be recommended.

[0123] In some embodiments of the present invention, dynamically adjusting the recommendation strategy of the video to be recommended based on the type of the video information to be recommended can be achieved in the following ways:

[0124] Dynamically adjust the exposure rate corresponding to the video information to be recommended; or adjust the exposure channel corresponding to the video information to be recommended; or adjust the exposure position corresponding to the video information to be recommended.

[0125] See also Figure 4 , Figure 4 This is a schematic diagram of an optional time-sensitive short video playback in an embodiment of the present invention. Figure 3 The video recommendation method provided by the embodiment of the present invention can replace the advertising information played in the advertising position in the time-sensitive short video playback environment, replacing Advertisement A with Advertisement B, so as to configure more playback traffic for Advertisement B and enable users to have a better viewing experience. Specifically, based on the traffic parameters matched by the playback strategy of the advertising information and the iterative experimental parameters, when the playback strategy of the advertising information is dynamically adjusted, the advertising exposure rate can be increased. In some embodiments of the present invention, the exposure channel of Advertisement A can also be adjusted from exposure in the current short video client to advertising in the contact status information of the instant messaging client. Of course, when the exposure position of Advertisement A is adjusted, it can be adjusted from the friend circle advertisement of the instant messaging client to the opening screen advertisement to comply with different dynamically adjusted playback strategies, so that the time-sensitive short video can be recommended to different users in a short time to obtain a better video recommendation effect.

[0126] When dynamically adjusting the playback strategy of the time-sensitive short video, the target user's historical browsing information can be obtained; based on the target user's historical browsing information, the exposure history of the time-sensitive short video corresponding to the historical browsing information is determined; based on the exposure history of the time-sensitive short video corresponding to the historical browsing information, the playback strategy of the time-sensitive short video is dynamically adjusted to Figure 4 For example, when it is determined that Ad B has been blocked in the historical browsing information of target user 1, Ad A can be replaced with other advertising information (such as Ad C) by dynamically adjusting the playback strategy. When it is determined that Ad C has been blocked in the historical browsing information of target user 2, Ad A can be replaced with other advertising information (such as Ad D) by dynamically adjusting the playback strategy to conform to the usage habits of the target users, so that the users can have a better usage experience.

[0127] See also Figure 5 , Figure 5 This is a schematic diagram of an optional time-sensitive short video recommendation process in an embodiment of the present invention. All tasks share a common network structure, but features extracted at different network layers correspond to different tasks. Generally speaking, the bottom layer of the model structure corresponds to less complex NLP tasks. During use, when the target resources include different ads from the same advertiser, the time-sensitive short video ads from different resource groups can be played sequentially in the time-sensitive short video playback window. When all time-sensitive short video playback areas on the display interface are contracted by advertisers, after the ad playback ends, the ads from the same advertiser can be cyclically presented in the time-sensitive short video playback window on the ad information display interface. Furthermore, when the advertiser's time-sensitive short videos are video ads, the same advertiser's video ads can be cyclically presented, and the audio volume of the videos can be adjusted to maximum to prompt the user to watch the video ads. Thus, by dynamically adjusting the ad playback strategy based on traffic parameters matched to the time-sensitive short video playback strategy and iterative experiment parameters, the advertiser's effective cost per thousand impressions (ECPM) can be reduced, achieving better ad playback performance.

[0128] Combine Figure 2 The video information recommendation device shown in the figure illustrates the video information recommendation method provided by the embodiment of the present invention. Figure 6 , Figure 6 This is an optional flow chart of the video information recommendation method provided by an embodiment of the present invention. It can be understood that: Figure 6 The steps shown can be executed by various electronic devices running the video information recommendation device, for example, a dedicated terminal, a server or a server cluster with a video information recommendation device, wherein the dedicated terminal with the video information recommendation device can be the preamble Figure 2 The electronic device with the video information recommendation device in the embodiment shown is as follows. Figure 6 The steps shown are explained.

[0129] Step 601: Determine historical parameters of the video information to be recommended based on the type of the video playback environment in which the video information to be recommended is located.

[0130] Step 602: Based on the historical parameters of the video information to be recommended, determine a training sample set that matches the video information recommendation model.

[0131] Step 603: extracting a training sample set that matches the training sample based on the noise threshold that matches the video information recommendation model.

[0132] Step 604: Determine a multi-task loss function that matches the video information recommendation model.

[0133] Step 605: Based on the multi-task loss function, adjust the parameters of the content category processing network and the network parameters of the timeliness category processing network in the video information recommendation model.

[0134] Thus, during the training process, until the loss functions of different dimensions corresponding to the video information recommendation model reach the corresponding convergence conditions, the parameters of the video information recommendation model can be adapted to the video playback environment. For example, when the use environment of the video information recommendation model is short video recommendation, in the process of recommending different short videos to users in the short video process, the short video playback interface can be displayed in the corresponding APP or triggered by the WeChat applet (the video information recommendation model can be encapsulated in the corresponding APP after training or saved in the WeChat applet as a plug-in). With the continuous development and increase of short video application products, the carrying capacity of video information is far greater than that of text information. Different types of short videos in the short video server can be continuously recommended to users through the corresponding application. In this training process, in the use environment of triggering short video recommendation through the WeChat applet, the dynamic noise threshold that matches the use environment of the video information recommendation model needs to be smaller than the dynamic noise threshold of recommending short videos to users directly in the short video client.

[0135] In some embodiments of the present invention, when the video information recommendation model is applied to a news information recommendation process, a fixed noise threshold corresponding to the news information recommendation process is determined, and a first training sample set is denoised based on the fixed noise threshold to form a second training sample set that matches the fixed noise threshold. When the video information recommendation model is embedded in a corresponding hardware device (e.g., a news reading terminal, an e-book terminal, or a financial news terminal), and the usage environment is to push different news information to users via the news reading terminal or e-book terminal, fixing the fixed noise threshold corresponding to the video information recommendation model can effectively improve the training speed of the video information recommendation model and reduce user waiting time. In a fixed noise environment, the training sample set can be derived from the target user's historical data. The historical recommended video information browsing data can be the recommended video information viewing behavior data generated when recommended videos were recommended to the target user, and can be extracted from historical browsing logs. The historical recommended video information browsing data can include all historical recommended video information browsing data; alternatively, considering the timeliness of the behavior data, it can include only historical recommended video information browsing data within a preset time period, such as historical recommended video information browsing data within a week. Taking the financial news terminal as an example, the corresponding financial news user cluster can be further subdivided into local financial news user cluster, stock financial news user cluster, and futures financial news user cluster, and can be marked according to the user cluster classification set by the user.

[0136] The following describes the training method of the video information recommendation model provided by the embodiment of the present invention by taking the video news information recommendation scenario in the short video playback interface as an example, wherein: Figure 7 FIG. 1 is a schematic diagram of an application environment of a training method for a video information recommendation model according to an embodiment of the present invention, wherein Figure 7As shown, the video news information playback interface can be displayed in the corresponding APP or triggered by the WeChat applet (the video information recommendation model can be encapsulated in the corresponding APP after training or saved in the WeChat applet as a plug-in, and the usage environment is the recommendation of news information). With the continuous development and increase of short video application products, the carrying capacity of video news video information is far greater than that of text information. Video news information can be continuously recommended to users through the corresponding application. For example, the "Take a look" entrance included in the discovery page of the WeChat application, or the audio recommendation entrance of the audio application, or the video recommendation entrance of the video application, or the live broadcast recommendation entrance of the live broadcast application, etc. When the target terminal runs the target application according to the user operation and controls the target application to display the application page including the trigger entrance for triggering the opening of the recommended content display page, it can detect the trigger operation of the trigger entrance. When the trigger operation corresponding to the trigger entrance is generated, a recommendation request is sent to the server, and after receiving the recommended content fed back by the server in response to the recommendation request, the recommended content is displayed in the recommended content display page in the recommended order.

[0137] refer to Figure 8 , Figure 8 The present invention provides a schematic diagram of an optional data processing architecture for the video information recommendation method provided in an embodiment of the present invention. The personalized news recommendation in the video information recommendation method provided in this application can be divided into two stages: recall and sorting. The two stages each perform their duties and complete different tasks, and each has a different focus. Specifically: In the recall stage, the main task is to filter out important content. The focus is on how to quickly and effectively extract content that a large number of users may be interested in from a large amount of news. The difficulty lies in how to accurately recommend timely video information to users in the face of a large amount of news and a large number of users. Therefore, the focus of the sorting process is to comprehensively and accurately classify the video information in the video data source in order to make accurate recommendations. Figure 8 In the data processing architecture of the video information recommendation method shown, in personalized news recommendation, it is necessary to first achieve accurate classification of individual video content. Specifically, in the personalized video recommendation process, it is first necessary to perform feature extraction processing on the recommended video information through the text processing network in the video information recommendation model to obtain feature extraction results; based on the feature extraction results of the text processing network, determine the title feature vector and label feature vector corresponding to the video information to be recommended, and fuse the title feature vector and label feature vector. Using a multi-task learning model, perform content category training and effectiveness training respectively to obtain the classification results of the video, so that recommendations can be made in accordance with the usage needs of different users.

[0138] refer to Figure 9 , Figure 9The working process diagram of the video information recommendation method provided by the embodiment of the present invention is as follows. Figure 9 The video information recommendation method shown in the figure illustrates the working process of the video information recommendation model provided by the present invention, which specifically includes the following steps:

[0139] Step 901: Obtain video information and determine corresponding video recommendation training samples.

[0140] Step 902: Determine initial model parameters of the video information recommendation model.

[0141] Step 903: Train different networks in the video information recommendation model to determine update parameters of the video information recommendation model.

[0142] Among them, according to the different functions of the video information recommendation model, a text processing network, a content category processing network, and a timeliness category processing network can be configured in the video information recommendation model, and feature fusion can be performed based on the processing results of different self-networks to obtain more accurate recommendation results.

[0143] Step 904: According to the updated parameters of the video information recommendation model, the initial parameters of the video information recommendation model are iteratively updated through training samples.

[0144] Step 905: deploying a video information recommendation model, and determining the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended.

[0145] When the video information to be recommended played in the short video's video information playback window is a video advertisement, a trigger operation for the video information to be recommended playback window is received; in response to the trigger operation, the business processing interface is presented to jump to the product display interface indicated by the video information to be recommended, thereby allowing users to more conveniently purchase the products presented in the video information playback window.

[0146] Step 906: Recommend the video to different video clients according to the type of the video information to be recommended.

[0147] When dynamically adjusting the playback strategy of the video information to be recommended, the historical browsing information of the target user can be obtained; based on the historical browsing information of the target user, the exposure history of the video information to be recommended corresponding to the historical browsing information is determined; based on the exposure history of the video information to be recommended corresponding to the historical browsing information, the playback strategy of the video information to be recommended is dynamically adjusted to Figure 4For example, in any ad slot from ad slot 1 to ad slot 3, an ad that serves as a recommended video can be displayed. When it is determined that ad B has been blocked in the historical browsing information of target user 1, other ad information (such as ad C) can be used to replace ad A by dynamically adjusting the playback strategy. When it is determined that ad C has been blocked in the historical browsing information of target user 2, other ad information (such as ad D) can be used to replace ad A by dynamically adjusting the playback strategy to conform to the usage habits of the target user, so that the user has a better usage experience.

[0148] Beneficial technical effects:

[0149] The present invention obtains the video information to be recommended from a video data source; processes the video information to be recommended through a text processing network in a video information recommendation model to determine the title feature vector and label feature vector corresponding to the video information to be recommended; determines the content category of the video information to be recommended through a content category processing network of the video information recommendation model based on the feature vectors of multiple attributes; determines the timeliness category of the video information to be recommended through a timeliness category processing network of the video information recommendation model based on the feature vectors of multiple attributes; determines the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended; and dynamically adjusts the recommendation strategy of the video to be recommended based on the type of the video information to be recommended. In this way, the video information recommendation model can classify and recommend video information in the usage environment, while enhancing the accuracy and timeliness of video information recommendations, effectively improving the quality of video information recommendations, and enhancing the user experience.

[0150] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A video information recommendation method, characterized in that: The method comprises: Obtain information about videos to be recommended from a video data source; Processing the video information to be recommended through a processing network in a video information recommendation model to determine feature vectors of multiple attributes corresponding to the video information to be recommended; Based on the feature vectors of the multiple attributes, determining the content category of the video information to be recommended through the content category processing network of the video information recommendation model; Based on the feature vectors of the multiple attributes, determining the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model; Determining the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended; Based on the type of the video information to be recommended, the recommendation strategy of the video to be recommended is dynamically adjusted.

2. The method according to claim 1, characterized in that The processing of the video information to be recommended by the processing network in the video information recommendation model to determine the feature vectors of multiple attributes corresponding to the video information to be recommended includes: When the video information recommendation model is used to process advertising videos, Performing feature extraction processing on the video information to be recommended through a text processing network in a video information recommendation model to obtain a feature extraction result; Based on the feature extraction result of the text processing network, a title feature vector and a tag feature vector corresponding to the video information to be recommended are determined.

3. The method according to claim 2, characterized in that The determining of the title feature vector and the label feature vector corresponding to the video information to be recommended based on the feature extraction result of the text processing network includes: Performing data screening on the video information to be recommended, parsing and obtaining the title and tag of the video information to be recommended; Triggering a target word segmentation library, and performing word segmentation processing on the title and tag of the video information to be recommended respectively through the target word segmentation library to obtain the video information to be recommended at the word level; The word-level video information to be recommended is vectorized through the text processing network in the video information recommendation model to form a multi-dimensional word-level title feature vector and a multi-dimensional word-level label feature vector of the video information to be recommended.

4. The method according to claim 1, wherein The method further comprises: determining historical parameters of the video information to be recommended according to the type of the video playback environment in which the video information to be recommended is located; Determining a training sample set that matches the video information recommendation model based on historical parameters of the video information to be recommended, wherein the training sample set includes at least one group of training samples; Extracting a training sample set that matches the training sample through a noise threshold that matches the video information recommendation model; The video information recommendation model is trained according to a training sample set that matches the training sample.

5. The method according to claim 4, characterized in that The method further comprises: When the network structures of the content category processing network and the timeliness category processing network are different, Determining a multi-task loss function that matches the video information recommendation model; Based on the multi-task loss function, the parameters of the content category processing network and the network parameters of the timeliness category processing network in the video information recommendation model are adjusted until the loss functions of different dimensions corresponding to the video information recommendation model reach the corresponding convergence conditions; so as to achieve the adaptation of the parameters of the video information recommendation model to the video playback environment.

6. The method according to claim 4, characterized in that The method further comprises: When the video playback environment of the video information to be recommended is a short video recommendation, determining a dynamic noise threshold that matches the usage environment of the video information recommendation model; performing noise removal processing on the first training sample set according to the dynamic noise threshold to form a second training sample set matching the dynamic noise threshold; When the video playback environment of the video information to be recommended is a long video playback, a fixed noise threshold corresponding to the video information recommendation model is determined, and the first training sample set is subjected to noise removal processing according to the fixed noise threshold to form a second training sample set that matches the fixed noise threshold.

7. The method according to claim 4, characterized in that The method further comprises: When the network structures of the content category processing network and the timeliness category processing network are the same, Determining a loss function that matches the video information recommendation model; Based on the loss function, the parameters of the content category processing network and the network parameters of the timeliness category processing network in the video information recommendation model are adjusted until the loss functions of different dimensions corresponding to the video information recommendation model reach the corresponding convergence conditions; so as to achieve the adaptation of the parameters of the video information recommendation model to the video playback environment.

8. The method according to claim 1, characterized in that The dynamically adjusting the recommendation strategy of the video to be recommended based on the type of the video information to be recommended includes: Dynamically adjusting the exposure rate corresponding to the video information to be recommended; or Adjusting the exposure channel corresponding to the video information to be recommended; or The exposure position corresponding to the video information to be recommended is adjusted.

9. The method according to claim 1, characterized in that The method further comprises: When the video information to be recommended is embedded advertising information in a short video, monitoring exposure parameters of the short video; Determining a Manrong visual effect index corresponding to the embedded advertising information according to the exposure parameters of the short video; According to the Manrong visual effect index corresponding to the embedded advertising information, the playback configuration information corresponding to the embedded advertising information in the short video is adjusted through the topology file.

10. The method according to claim 9, characterized in that The method further comprises: Sending the exposure parameters of the short video during playback to a monitoring server, so that the monitoring server can monitor the exposure of the short video; The exposure parameters of the short video during playback stored by the monitoring server are used as a data source for the playback effect parameters of the video information to be recommended.

11. The method according to claim 1, wherein The method further comprises: Obtain the target user's historical browsing information; Based on the historical browsing information of the target user, determining the exposure history of the video information to be recommended corresponding to the historical browsing information; Based on the exposure history of the video information to be recommended corresponding to the historical browsing information, the playback strategy of the video information to be recommended is dynamically adjusted.

12. The method according to claim 1, characterized in that The method further comprises: Determine the usage environment of the terminal display interface; Determining the category of the video information to be played or recommended based on the usage environment of the terminal display interface; In response to the category of the video information to be played and recommended, a matching data source of the video information to be recommended is triggered to adjust the video information to be played and recommended through the data source of the video information to be recommended that matches the category of the video information to be played and recommended.

13. A video information recommendation device, characterized in that: The device comprises: An information transmission module is used to obtain information of videos to be recommended from a video data source; An information processing module, configured to process the video information to be recommended through a processing network in a video information recommendation model, and determine feature vectors of multiple attributes corresponding to the video information to be recommended; The information processing module is configured to determine the content category of the video information to be recommended through the content category processing network of the video information recommendation model based on the feature vectors of the multiple attributes; The information processing module is configured to determine the timeliness category of the video information to be recommended through the timeliness category processing network of the video information recommendation model based on the feature vectors of the multiple attributes; The information processing module is used to determine the type of the video information to be recommended according to the content category of the video information to be recommended and the timeliness category of the video information to be recommended; The information processing module is used to dynamically adjust the recommendation strategy of the video to be recommended based on the type of the video information to be recommended.

14. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; The processor is configured to implement the video information recommendation method according to any one of claims 1 to 12 when running the executable instructions stored in the memory.

15. A computer-readable storage medium storing executable instructions, characterized in that: When the executable instructions are executed by a processor, the video information recommendation method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Video recommendation method and device

    CN106686414A

  • Video timeliness determination method and device, electronic equipment and medium

    CN112399201A