Video pushing method and device, electronic equipment and computer readable medium

By extracting keywords, performing correlation analysis and semantic matching of video tag index information sets, and determining video popularity information, the video push process is optimized, solving the problems of excessive computing resource consumption and difficulty in timely previewing of new video materials, and achieving efficient video material push and circulation.

CN119052584BActive Publication Date: 2026-04-28CHENGDU GUANGCHANG CREATIVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU GUANGCHANG CREATIVE TECH CO LTD
Filing Date
2024-08-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as excessive consumption of computing resources, difficulty in timely delivery of video materials that users are interested in, and difficulty in timely previewing of new video materials during the video push process.

Method used

By extracting keywords, performing correlation analysis on video tag index information sets, semantic matching, and determining video popularity information, the video push process is optimized, reducing computing resource consumption and improving the efficiency of video material circulation.

Benefits of technology

It enables timely delivery of video materials that users are interested in while reducing the consumption of computing resources, and ensures that new video materials can be previewed in a timely manner, thereby improving the efficiency of video material circulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052584B_ABST
    Figure CN119052584B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video pushing method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises: performing keyword extraction processing on video retrieval text information to obtain a retrieval keyword sequence; performing association analysis processing on a video label index information set and the retrieval keyword sequence to obtain an associated video information set; performing semantic matching processing on each video material segment corresponding to the video retrieval text information and the associated video information set to obtain a to-be-displayed video information set; determining video heat information corresponding to each to-be-displayed video information to obtain a video heat information set; performing sorting processing on each sample video material segment based on the video heat information set to obtain a to-be-previewed video sequence; and generating a video pushing result based on a video preview image information set and the to-be-previewed video sequence. The embodiment can push video material of interest to a user in a timely manner, thereby improving the flow efficiency of the video material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the field of computer technology, specifically to video push methods, apparatuses, electronic devices, and computer-readable media. Background Technology

[0002] Digital content distribution platforms are dedicated to providing users with abundant video resources. Therefore, how to promptly push video resources that users are interested in to improve the efficiency of video resource distribution has become a core issue in the field. Currently, the common approach to video push is as follows: First, based on an image-language pre-trained model transferred to video-text matching, various video clips that semantically match the searched text are directly selected from the video resource library. Then, sample video clips of the selected video resources are sorted according to their popularity and pushed to the user's terminal for preview.

[0003] However, in practice, it has been found that the following technical problems often arise when using the above method for video push:

[0004] First, if we use the transferred image-language pre-trained model to directly perform semantic matching between the search text and all the video materials in the video material library for each user's input search text, it will consume a lot of computing power and time, resulting in a large consumption of computing resources and difficulty in pushing video materials that users are interested in to users in a timely manner.

[0005] Second, because the transferred image-language pre-trained models usually have difficulty capturing the temporal features of video clips and ignore the semantics of the search text when extracting video features, there are many clips with low user interest in the selected video materials. As a result, it is difficult for users to quickly identify the video materials to be transferred and the efficiency of video material transfer is reduced.

[0006] Third, since new video materials are unlikely to accumulate high popularity in a very short period of time when they are first launched, if they are sorted and pushed according to the popularity of the video materials, the new video materials will easily be displayed at a relatively low position, thus making it difficult for users to preview the newer video materials that they are interested in in a timely manner.

[0007] The information disclosed in this background section is only intended to enhance the understanding of the background of the concept of this application, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0008] The summary section of this application is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0009] Some embodiments of this application provide video push methods, apparatuses, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the background section above.

[0010] In a first aspect, some embodiments of this application provide a video push method, which includes: in response to receiving video retrieval text information, performing keyword extraction processing on the video retrieval text information to obtain a retrieval keyword sequence; performing association analysis processing on a preset video tag index information set and the retrieval keyword sequence to obtain an associated video information set; performing semantic matching processing on each video material segment corresponding to the video retrieval text information and the associated video information set to obtain a video information set to be displayed; determining the video popularity information corresponding to each video information to be displayed in the video information set to be displayed to obtain a video popularity information set; sorting each sample video material segment corresponding to the video information set to be displayed based on the video popularity information set to obtain a video sequence to be previewed; and generating a video push result based on a pre-generated video preview image information set and the video sequence to be previewed, for push to a target user for preview.

[0011] Secondly, some embodiments of this application provide a video push device, comprising: a keyword extraction processing unit configured to, in response to receiving video retrieval text information, perform keyword extraction processing on the video retrieval text information to obtain a retrieval keyword sequence; an association analysis processing unit configured to perform association analysis processing on a preset video tag index information set and the retrieval keyword sequence to obtain an associated video information set; a semantic matching processing unit configured to perform semantic matching processing on the video retrieval text information and each video material segment corresponding to the associated video information set to obtain a video information set to be displayed; a determination unit configured to determine the video popularity information corresponding to each video information set to be displayed in the video information set to be displayed, to obtain a video popularity information set; a sorting processing unit configured to sort each sample video material segment corresponding to the video information set to be displayed based on the video popularity information set to obtain a video sequence to be previewed; and a generation unit configured to generate a video push result based on a pre-generated video preview image information set and the video sequence to be previewed, for push to a target user for preview.

[0012] Thirdly, some embodiments of this application provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0013] Fourthly, some embodiments of this application provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0014] The above embodiments of this application have the following beneficial effects: the video push method of some embodiments of this application can consume less computing resources and push video materials of interest to users in a timely manner. Specifically, the reason for consuming more resources and making it difficult to push video materials of interest to users in a timely manner is that if a transferred image-language pre-trained model is used to directly perform semantic matching between the search text input by each user and all video materials in the video material library, it will consume a lot of computing power and time, resulting in a large consumption of computing resources and difficulty in pushing video materials of interest to users in a timely manner. Based on this, the video push method of some embodiments of this application firstly, in response to receiving video search text information, performs keyword extraction processing on the video search text information to obtain a search keyword sequence. Thus, each keyword in the video search text can be obtained. Secondly, a correlation analysis processing is performed on the preset video tag index information set and the above search keyword sequence to obtain a related video information set. Thus, based on the correlation relationship between keywords and the tags marked on the video materials, potential video materials of interest to users can be initially screened. Then, semantic matching processing is performed on the video retrieval text information and the corresponding video clips in the associated video information set to obtain the video information set to be displayed. Thus, based on the semantic matching relationship between the retrieval text and the video content, the video clips that the user is actually interested in can be accurately selected from the various video clips that the potential user is interested in, so as to push them to the user. Next, the video popularity information corresponding to each video clip in the video information set to be displayed is determined to obtain the video popularity information set. Thus, the attention popularity of each video clip to be pushed to the user can be obtained. Then, based on the video popularity information set, the sample video clips corresponding to the video information set to be displayed are sorted to obtain the video sequence to be previewed. Thus, high-quality video clips with higher attention popularity can be placed at the top, making it easier for users to preview them in a timely manner. Finally, based on the pre-generated video preview image information set and the video sequence to be previewed, a video push result is generated for push to the target user for preview. Therefore, the video push method of some embodiments of this application, which matches video clip tags with the user's retrieval text first and then performs semantic matching with the video clips, can reduce the number of video clips involved in semantic matching and reduce the computational power consumption and time cost of filtering videos of interest. This reduces the consumption of computing resources and allows for the timely delivery of video materials that users are interested in, thereby improving the efficiency of video streaming. Attached Figure Description

[0015] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0016] Figure 1 This is a flowchart of some embodiments of the video push method according to this application;

[0017] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the video push device according to this application;

[0018] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of this application. Detailed Implementation

[0019] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0020] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0021] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0022] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0023] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0024] The present application will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] Figure 1A flow 100 of some embodiments of the video push method according to this application is shown. The video push method includes the following steps:

[0026] Step 101: In response to receiving video retrieval text information, perform keyword extraction processing on the video retrieval text information to obtain a retrieval keyword sequence.

[0027] In some embodiments, the execution entity of the video push method (e.g., a computing device) may, in response to receiving video retrieval text information, perform keyword extraction processing on the video retrieval text information to obtain a retrieval keyword sequence. The video retrieval text information may be text information input by a user for retrieving video materials from a video material library. The video material library may be a database for storing various video material segments. The video material segments may be various video segments used in the video production process. Each video material segment may be associated with a video material identifier. The video material identifier may be a unique identifier for the corresponding video material segment. The retrieval keyword sequence may be a set arranged according to the order in which the retrieval keywords appear in the video retrieval text. The retrieval keywords may be keywords used to retrieve video material segments. The retrieval keyword sequence can be obtained by performing keyword extraction processing on the video retrieval text information using a preset keyword extraction method. For example, the keyword extraction method may be, but is not limited to, one of the following: a keyword extraction method based on word embedding, or a keyword extraction method based on an attention mechanism.

[0028] Step 102: Perform correlation analysis on the preset video tag index information set and the search keyword sequence to obtain the associated video information set.

[0029] In some embodiments, the aforementioned executing entity can perform correlation analysis on the preset video tag index information set and the aforementioned search keyword sequence in various ways to obtain a correlated video information set. Each correlated video information piece in the correlated video information set can correspond one-to-one with a video clip. The video tag index information in the aforementioned video tag index information set can be an index composed of tag words and the corresponding video clips. Tag words can be obtained by manually or machine-annotating the video clips. Tag words can characterize the attribute features of the video clips. Attribute features can include, but are not limited to: objects in the video, video format, and video style. The correlated video information in the aforementioned correlated video information set can be information about video clips related to any search keyword.

[0030] In some optional implementations of certain embodiments, each video tag index information in the aforementioned video tag index information set may include tag words and a set of labeled video identifiers. The labeled video identifiers in the aforementioned set of labeled video identifiers may be video clip identifiers for video clips labeled with the corresponding tag words. The aforementioned execution entity can perform correlation analysis processing on the preset video tag index information set and the aforementioned search keyword sequence through the following steps to obtain a correlated video information set:

[0031] The first step is to select, for each search keyword in the above search keyword sequence, video tag index information whose tag words match the search keyword from the above video tag index information set, as target video tag index information, thus obtaining a target video tag index information group. Matching the search keyword means that the tag words included in the video tag index information contain characters or strings that are exactly the same as the search keyword. For example, when the search keyword is "flower", the tag words that match the search keyword can include "flower", "peony", and "spring flowers".

[0032] The second step is to identify each labeled video identifier in each labeled video identifier set corresponding to the target video tag index information group as associated video information, thus obtaining the associated video information set.

[0033] Step 103: Perform semantic matching processing on the video retrieval text information and the corresponding video material segments of the associated video information set to obtain the video information set to be displayed.

[0034] In some embodiments, the aforementioned executing entity may perform semantic matching processing on the aforementioned video retrieval text information and the various video clips corresponding to the aforementioned associated video information set in various ways to obtain a video information set to be displayed. The aforementioned video information set to be displayed may be the video retrieval results after precise retrieval based on video semantics.

[0035] In some optional implementations of certain embodiments, the execution entity may perform semantic matching processing on the video retrieval text information and the video material segments corresponding to the associated video information set through the following steps to obtain the video information set to be displayed:

[0036] The first step is to perform feature extraction processing on the aforementioned video retrieval text information to obtain text features. These text features can be vectors representing the video retrieval text information. Specifically, the following steps can be performed:

[0037] The first sub-step involves determining the text representation vector corresponding to the aforementioned video retrieval text information using a word embedding model.

[0038] The second sub-step involves projecting the aforementioned text representation vector onto a preset feature space using a feedforward neural network to obtain the projected text vector. This preset feature space can be a feature space of a preset dimension. The preset dimension can be a pre-defined dimension.

[0039] The third sub-step involves learning contextual feature representations from the projected text vectors using a multi-attention-based Transformer model to obtain text features.

[0040] The second step is to perform the following steps for each video clip corresponding to the above-mentioned associated video information set:

[0041] The first sub-step involves inputting the aforementioned text features and video clips into a pre-trained text-video semantic matching model to obtain video-text matching values. This model may include a video spatial feature extraction module, a video temporal feature extraction module, a video spatiotemporal feature extraction module, a video text encoding module, a video text decoding module, and a classification module. The video spatial feature extraction module extracts spatial features from the video clips. The video temporal feature extraction module extracts temporal features from the video clips. The video spatiotemporal feature extraction module extracts spatiotemporal features from the video clips. The video text encoding module fuses and encodes the video spatiotemporal features and text features to obtain associated features. The video text decoding module decodes the original video spatiotemporal features based on the associated features to obtain decoded video features. The classification module outputs a label value representing whether the text and video semantically match based on the decoded video features. The video-text matching value can be a probability value representing whether the corresponding video clip is semantically similar to the retrieved text.

[0042] In some optional implementations of certain embodiments, the execution entity can input the aforementioned text features and video clips into a pre-trained text-video semantic matching model, and obtain video-text matching values ​​through the following steps:

[0043] Step 1: Input the aforementioned video clips into the video spatial feature extraction module of the pre-trained text-video semantic matching model to obtain a hierarchical extracted video feature sequence and video spatial features. The video feature extraction module may include various video feature extraction layers. For example, it may include four video feature extraction layers. These layers can be executed sequentially. Each video feature extraction layer can consist of a single-layer Transformer encoder. The Transformer encoder may include a hierarchical normalization layer, a feedforward network layer, and a multi-head self-attention layer. Except for the first layer, the output of the previous layer can be the input of the next layer. The hierarchical extracted video features in the aforementioned hierarchical extracted video feature sequence can correspond one-to-one with the video feature extraction layers. The hierarchical extracted video feature sequence can be an ordered set of hierarchical extracted video features arranged in the execution order of the layers. The hierarchical extracted video features can be the output of a layer other than the last layer. The video spatial features can be the final output of the aforementioned video feature extraction module. These video spatial features can characterize the spatial information of the video clips. Spatial information may include, but is not limited to, the layout of the scene in the video, the shape of the objects, the texture, and the color distribution.

[0044] Step two: Input the aforementioned text features and the aforementioned layered extracted video feature sequence into the video temporal feature extraction module included in the aforementioned text-video semantic matching model to obtain video temporal features. The aforementioned video temporal feature extraction module may include various temporal feature extraction layers. These temporal feature extraction layers can be executed sequentially. Each temporal feature extraction layer may consist of a single-layer Transformer encoder. First, input the aforementioned text features and the layered extracted video feature with index 1 from the aforementioned layered extracted video feature sequence into the first-layer temporal feature extraction layer of the aforementioned video temporal feature extraction module to obtain the temporal feature output result corresponding to the first-layer temporal feature extraction layer. Then, for each temporal feature extraction layer in the aforementioned video temporal feature extraction module that meets the preset number of layers condition, the temporal feature output result of the previous-layer temporal feature extraction layer and the corresponding layered extracted video feature of the aforementioned temporal feature extraction layer can be added together, and the result of the addition can be input into the aforementioned temporal feature extraction layer. The aforementioned meeting the preset number of layers condition may mean that the temporal feature extraction layer is not the first-layer temporal feature extraction layer in the aforementioned video temporal feature extraction module. The aforementioned video temporal features can characterize the temporal information of video clips.

[0045] Step 3: Input the aforementioned video temporal features and video spatial features into the video spatiotemporal feature extraction module included in the text-video semantic matching model to obtain video spatiotemporal features. These video spatiotemporal features characterize the spatiotemporal information resulting from the fusion of temporal and spatial information of a video segment. The video spatiotemporal feature extraction module may include a feature fusion layer, a fusion feature dimensionality reduction layer, and a fusion feature extraction layer. The feature fusion layer can be used to concatenate the aforementioned video temporal features and video spatial features to obtain fused features. The fusion feature dimensionality reduction layer can be used to reduce the dimensionality of the fused features through linear mapping to obtain dimensionality-reduced fused features. The fusion feature extraction layer can extract features from the dimensionality-reduced fused features using a single-layer Transformer encoder to obtain video spatiotemporal features.

[0046] Step four involves inputting the aforementioned video spatiotemporal features and text features into the video text encoding module included in the text-video semantic matching model to obtain video text association features. These video text association features characterize the relationship between video clips and the retrieved text. The video text encoding module uses a single-layer Transformer encoder to encode the aforementioned video spatiotemporal features based on the text features, thereby obtaining the video text association features.

[0047] Step five involves inputting the aforementioned video-text association features and video spatiotemporal features into the video-text decoding module included in the text-video semantic matching model to obtain a decoded video feature vector. This video-text decoding module can be composed of a multi-layer Transformer encoder. The video-text decoding module can decode the aforementioned video spatiotemporal features based on the video-text association features to obtain a decoded video feature vector. This decoded video feature vector characterizes whether the corresponding video clip is semantically similar to the video retrieval text.

[0048] Step six: Input the decoded video feature vector into the classification module of the text-video semantic matching model to obtain the video-text matching value. The classification module may include a linear layer and an activation function layer. The activation function layer may include the Sigmoid activation function.

[0049] The second sub-step involves determining, in response to the determination that the video text matching value is greater than a preset matching threshold, the associated video information corresponding to the video clip is identified as the video information to be displayed. The preset matching threshold can be a pre-set lower limit for the video text matching value.

[0050] Optionally, the above text-video semantic matching model can be pre-trained through the following steps:

[0051] The first step is to obtain a sample text-video matching information set. Each piece of sample text-video matching information in this set may include sample video retrieval text information and a sample video text matching information set. The sample video text matching information in this set may include sample video clips and sample video text matching tags.

[0052] The second step involves selecting sample text-video matching information from the aforementioned set of sample text-video matching information and performing the following training steps:

[0053] The first sub-step involves performing feature extraction processing on the sample video retrieval text information included in the selected sample text video matching information to obtain sample text features.

[0054] The second sub-step involves performing the following steps for each sample video clip in at least one sample video clip included in the selected sample text-video matching information:

[0055] In sub-step one, the aforementioned sample text features and sample video clips are input into the initial text-video semantic matching model to obtain the video-text matching value. The initial text-video semantic matching model can be an untrained or incompletely trained text-video semantic matching model.

[0056] Sub-step two: In response to determining that the video text matching value is greater than a preset matching threshold, a preset matching tag is determined as the video text matching tag. The preset matching tag can be a pre-set tag representing the semantic similarity between the sample video retrieval text information and the sample video clip. For example, when the preset matching tag is 0, it indicates that the sample video retrieval text information and the sample video clip are semantically dissimilar. When the preset matching tag is 1, it indicates that the sample video retrieval text information and the sample video clip are semantically similar.

[0057] The third sub-step involves comparing the video text matching tag corresponding to each sample video clip in the at least one sample video clip with the corresponding sample video text matching tag.

[0058] The fourth sub-step involves determining, based on the comparison results, whether the initial text-video semantic matching model has achieved a preset optimization objective. This optimization objective could be that the accuracy of the video text matching tags generated by the initial text-video semantic matching model is greater than a preset accuracy threshold.

[0059] The fifth sub-step is to determine the initial text-video semantic matching model as the text-video semantic matching model in response to the determination that the initial text-video semantic matching model has reached the preset optimization target.

[0060] Optionally, in response to determining that the initial text-video semantic matching model has not achieved the aforementioned optimization objective, the aforementioned execution entity may adjust the network parameters of the initial text-video semantic matching model, and use the unselected sample text-video matching information, using the adjusted initial text-video semantic matching model as the initial text-video semantic matching model, and then execute the aforementioned training steps again. Specifically, backpropagation and gradient descent algorithms can be used to adjust the network parameters of the initial text-video semantic matching model.

[0061] Step 103 and its related content, as an inventive point of this application, solves the second technical problem mentioned in the background art: "Users have difficulty quickly identifying the video material to be transferred, and the efficiency of video material transfer is reduced." The reasons why users have difficulty quickly identifying the video material to be transferred and why the efficiency of video material transfer is reduced are often as follows: because the transferred image-language pre-trained model usually has difficulty capturing the temporal features of video segments, and ignores the semantics of the retrieval text when extracting video features, many of the selected video materials contain materials with low user interest. Solving the above problems can achieve the effect of enabling users to quickly identify the video material to be transferred and improving the efficiency of video material transfer. To achieve this effect, firstly, the text features corresponding to the video retrieval text information are determined. Then, the text features and the video material segments to be matched are input into the above-mentioned text-video semantic matching model to determine whether the video material segments and the video retrieval text are semantically matched. Specifically, after extracting video spatial features, the above-mentioned text-video semantic matching model uses the hierarchical structure of the Transformer encoder to extract video temporal features, and performs feature fusion extraction of video spatial features and video temporal features to obtain video spatiotemporal features. Furthermore, textual features are incorporated into the video temporal feature extraction process. This not only improves the extraction capability of video temporal features but also yields more comprehensive video spatiotemporal features aligned with textual features. After extracting the video spatiotemporal features, the aforementioned text-video semantic matching model obtains relatively accurate video-text association features through a video text encoding module. Then, a video text decoding module decodes the video spatiotemporal features based on these association features, resulting in a more accurate decoded video feature vector that characterizes the semantic similarity between video clips and the video retrieval text. Finally, video clips that semantically match the video retrieval text are identified as the video clips of interest to be displayed to the user. This reduces the push of video clips with low user interest, making it easier and faster for users to identify the video clips to be streamed, thereby improving the efficiency of video clip streaming.

[0062] Step 104: Determine the video popularity information corresponding to each video information in the video information set to be displayed, and obtain the video popularity information set.

[0063] In some embodiments, the aforementioned executing entity may determine the video popularity information corresponding to each video information to be displayed in the aforementioned video information set through various methods, thereby obtaining a video popularity information set. The video popularity information in the aforementioned video popularity information set may be information regarding the degree of user attention received by the corresponding video clip.

[0064] In some optional implementations of certain embodiments, the aforementioned execution entity can determine the video popularity information corresponding to each video information in the set of video information to be displayed through the following steps, thereby obtaining the video popularity information set:

[0065] For each video in the above set of video information to be displayed, perform the following steps:

[0066] The first step is to select user browsing behavior statistics that match the video information to be displayed from a pre-defined set of user browsing behavior statistics. Each user browsing behavior statistic may include a viewed video identifier and a behavior category information group. The viewed video identifier may be an identifier of a sample video clip that the user has viewed. The sample video clip may be a video clip composed of some video frames included in the corresponding video clip. The sample video clip and the corresponding video clip have the same identifier. Each behavior category information in the behavior category information group may include a behavior type identifier and a behavior count value. The behavior type identifier may be a unique identifier for the user behavior type. The user behavior type may be, but is not limited to, one of the following: like, favorite, comment, place an order. The behavior count value may be the number of times the corresponding user behavior type has occurred since the video clip went live. Matching the video information to be displayed means that the viewed video identifier included in the user browsing behavior statistics may be the same as the video clip identifier corresponding to the video information to be displayed.

[0067] The second step involves generating user interest scores based on a preset set of behavioral weight information and a selected set of behavioral category information including user browsing behavior statistics. Each behavioral weight information in the aforementioned set can include a behavior type identifier and a behavior weight value. The behavior weight value represents the proportion of contribution of the corresponding user behavior type to the user interest score. The user interest score can be obtained by weighted summing of the counts of each behavior included in the aforementioned behavioral category information set, based on the aforementioned behavioral weight information set.

[0068] The third step is to obtain the video upload date corresponding to the aforementioned video information to be displayed. This video upload date can be the upload date of the video clip corresponding to the video information to be displayed. The upload date can be the date on which the video clip can be retrieved and displayed on a digital content distribution platform. This digital content distribution platform can be a website used for digital content distribution (e.g., trading). The aforementioned digital content can include, but is not limited to, videos and music. The video upload date corresponding to the aforementioned video information to be displayed can be obtained from a pre-set video attribute information database via wired or wireless connection. This video attribute information database can be a database used to store video attribute information. Video attribute information can include, but is not limited to, video clip identifier, video provider identifier, video upload date, and video format. The video provider identifier can be a unique identifier for the video provider.

[0069] It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future wireless connection methods.

[0070] The fourth step involves generating video popularity information corresponding to the aforementioned video information to be displayed, based on the user interest and video release date. The executing entity can generate this video popularity information using various methods, based on the aforementioned user interest and video release date.

[0071] In some optional implementations of certain embodiments, each video provider attribute in the aforementioned video provider attribute information set may include a provider reputation score and the number of video works. The provider reputation score may be the video provider's reputation rating. The video provider may be the copyright owner of the corresponding video clip. The number of video works may be the total number of video clips uploaded by the video provider. Furthermore, the executing entity may also perform the following steps to generate video popularity information corresponding to the aforementioned video information to be displayed, based on the aforementioned user interest and the aforementioned video upload date:

[0072] Step 1: Determine the video provider attribute information corresponding to the video information to be displayed above as the target video provider attribute information.

[0073] Step two: Based on the preset single-work score and the number of video works included in the target video provider attribute information, generate the video provider's work score. The single-work score can be the score corresponding to a single video clip. The product of the single-work score and the number of video works can be used to determine the video provider's work score.

[0074] Step 3: Based on the video provider's work score and the provider's reputation score (including the target video provider's attribute information), the video provider's popularity score is determined. This popularity score represents the degree of attention the video provider receives. The popularity score is obtained by weighting and summing the video provider's work score and reputation score according to the work weight and reputation weight. The work weight represents the percentage contribution of the number of videos provided by the video provider to the popularity score. The reputation weight represents the percentage contribution of the video provider's reputation score to the popularity score.

[0075] Step four: The sum of the above user interest score and the above video provider popularity score is determined as the initial video popularity score.

[0076] Step 5: Based on the video's release date and the current date, generate the number of days the video has been online. This number can be the number of days the video clip has been searchable by users since it was uploaded to the digital content streaming platform, excluding the current date. The difference between the current date and the video's release date can be used to determine the number of days online.

[0077] Step six: Determine the time decay coefficient by summing the number of days already uploaded and the preset initial number of days. The preset initial number of days can be a pre-set initial value for the number of days. This ensures that the time decay coefficient of newly uploaded video clips is not too small. For example, the preset initial number of days could be 7 days.

[0078] Step seven: Based on a preset adjustable attenuation coefficient parameter, update the aforementioned time attenuation coefficient to obtain an updated time attenuation coefficient. This updated time attenuation coefficient characterizes the rate at which video popularity decays over time. The adjustable attenuation coefficient parameter can be used to adjust the magnitude of the time attenuation coefficient. The updated time attenuation coefficient can be determined by using the time attenuation coefficient as the base and the adjustable attenuation coefficient parameter as the exponent.

[0079] Step 8: Update the initial video popularity value according to the aforementioned update time decay coefficient to obtain the updated video popularity value. The updated video popularity value can be a video popularity value that decays with the number of days the video has been online. The ratio of the initial video popularity value to the aforementioned update time decay coefficient can be determined as the updated video popularity value.

[0080] Step nine: Based on the aforementioned video information to be displayed and the aforementioned updated video popularity value, generate video popularity information. Specifically, the video material identifier corresponding to the aforementioned video information to be displayed and the aforementioned updated video popularity value can be identified as the video popularity information.

[0081] Step 104 and its related content, as an inventive point of this application, solves the third technical problem mentioned in the background art: "Users have difficulty previewing newer video materials of interest in a timely manner." The reasons why users have difficulty previewing newer video materials of interest in a timely manner are often as follows: because new video materials are difficult to accumulate high attention in a very short time when they are first launched, if they are sorted and pushed according to their attention level, the display position of new video materials is likely to be relatively low. If the above problem is solved, it is possible to ensure that high-quality video materials with high attention are ranked higher, while users can also preview newer video materials of interest in a timely manner. To achieve this effect, firstly, considering that video materials from reputable video providers with abundant works are more likely to attract attention, and that the historical browsing behavior of users with video needs directly reflects the degree of attention received by video materials, the attention level of video materials is determined from two dimensions: the historical browsing behavior of video providers and users with video needs. Secondly, considering that the attention level of video materials will decay over time and newly launched video materials need a buffer period to accumulate attention, a time decay coefficient can be determined based on the number of days since launch and the preset initial number of days. The preset initial number of days serves as a buffer period for newly released video materials to accumulate popularity. The time decay coefficient can then be adjusted using adjustable parameters. This mitigates the impact of time on the popularity of video materials when there are few available materials, while accelerating the impact when there are many. Finally, the adjusted update time decay coefficient further adjusts the popularity of video materials. This ensures that when sorting and pushing videos based on quality and popularity, high-quality video materials with high popularity rank higher, while newly released video materials also receive a relatively high ranking. This allows users to preview newer video materials of interest more quickly, thereby improving the efficiency of subsequent video material distribution.

[0082] Step 105: Based on the video popularity information set, sort the sample video clips corresponding to the video information set to be displayed to obtain the video sequence to be previewed.

[0083] In some embodiments, the execution entity may sort the sample video clips corresponding to the video information set to be displayed based on the video popularity information set to obtain a video sequence to be previewed. Each sample video clip can be a video segment composed of a portion of video frames included in the corresponding video clip. First, the video popularity information set is sorted in descending order according to the updated video popularity values ​​using a preset sorting algorithm to obtain a video popularity information sequence. For example, the sorting algorithm may be, but is not limited to, one of the following: bubble sort or quick sort. Then, the video clip identifier sequence included in the video popularity information sequence is determined as the target video clip identifier sequence. Finally, the sample video clips corresponding to the video information set to be displayed are sorted according to the order of the video clip identifiers corresponding to the sample video clips in the target video clip identifier sequence using the sorting algorithm, and each sorted sample video clip is used as a video to be previewed to obtain the video sequence to be previewed.

[0084] Step 106: Based on the pre-generated set of video preview image information and the sequence of videos to be previewed, generate video push results for push to the target user for preview.

[0085] In some embodiments, the aforementioned execution entity can generate video push results based on a pre-generated set of video preview image information and the aforementioned video sequence to be previewed, in various ways, for push to the target user for preview. The target user is the user who sends the aforementioned video retrieval text information through a user terminal. The video preview image information in the aforementioned video preview image information set can correspond one-to-one with video clips. The video preview image information in the aforementioned video preview image information set can include video clip identifiers and video frame image sequences. The video frame images in the aforementioned video frame image sequence can be non-adjacent video frame images extracted sequentially from the corresponding video clips. The aforementioned video push results can be information about various video clips of interest retrieved by the target user.

[0086] In some optional implementations of certain embodiments, the aforementioned execution entity can generate a video push result based on a pre-generated set of video preview image information and the aforementioned video sequence to be previewed through the following steps:

[0087] The first step, for each video in the above video sequence to be previewed, is to perform the following steps:

[0088] The first sub-step involves selecting video preview images that match the video to be previewed from the aforementioned set of video preview images as target video preview images. Matching the video to be previewed can mean that the video material identifier included in the video preview image information is the same as the video material identifier corresponding to the video to be previewed.

[0089] The second sub-step involves determining the aforementioned video to be previewed and the aforementioned target video preview image information as the video content information to be previewed.

[0090] The second step is to sort the determined video content information to obtain a sequence of video content information. This can be achieved using the sorting algorithm described above, where the determined video content information is sorted according to the order of the videos within the overall video sequence.

[0091] The third step is to obtain the historical preview mode information corresponding to the target user. This historical preview mode information can be obtained from a user information database, showing the preview modes used by the user when previewing video push results in historical periods. The user information database can be a database used to store user information. User information can include both static and dynamic characteristics of the user. Static characteristics can be characteristics that do not easily change over time. For example, static characteristics may include, but are not limited to, user identifiers and registration time. Dynamic characteristics can be characteristics that change over time. For example, dynamic characteristics may include, but are not limited to, user behavior log information and historical preview mode information. The preview mode can characterize the display style of the video push results. The preview mode can be one of the following: classic mode, thumbnail mode, and grid mode. The classic mode can be a display style with the video on top and the grid-style image below. The thumbnail mode can be a display style that only includes the video. The grid mode can be a display style that only includes the grid-style image.

[0092] Optionally, the aforementioned historical preview mode information may be the preview mode used by the target user when they last searched for and viewed the video push results.

[0093] The fourth step is to determine the sequence of the above-mentioned historical preview mode information and the above-mentioned video content information to be previewed as the video push result.

[0094] Optionally, before generating the video push result based on the pre-generated video preview image information set and the aforementioned video sequence to be previewed, the execution entity may also perform the following steps for each video clip corresponding to the aforementioned video sequence to be previewed, in order to generate video preview image information in the video preview image information set:

[0095] Step 1: Obtain the video keyframe sequence and configured grid number corresponding to the aforementioned video clips. The video keyframes in the sequence can be obtained by sequentially extracting the video clips using a preset video frame extraction interface. This interface can include video classification methods, time-interval-based uniform frame extraction methods, and video scene extraction methods. The video classification method can be used to segment and identify the background of each video frame to determine if the video has a relatively small background variation. The time-interval-based uniform frame extraction method can be used to extract frames from video types with small background variations to obtain a video keyframe sequence. The video scene extraction method can be used to extract frames from video types with large background fluctuations to obtain a video keyframe sequence. The configured grid number can be the number of grids pre-set by the video provider for displaying the video keyframes.

[0096] Step 2: Determine the number of each video keyframe in the above video keyframe sequence as the number of keyframes.

[0097] Step 3: Based on the number of keyframes and the number of configured grids mentioned above, generate the frame extraction step size. The frame extraction step size can be determined by rounding down the quotient of the number of keyframes and the number of configured grids.

[0098] Step four: Determine the initial frame-sampling video image sequence number. This initial frame-sampling video image sequence number can be the sequence number of the starting image when extracting frames from the video keyframe sequence. A random number generator can be used to randomly select a number from 0 to the frame extraction step size as the initial frame-sampling video image sequence number.

[0099] Step 5: Based on the aforementioned frame extraction step size and the aforementioned initial frame extraction video image sequence number, perform frame extraction processing on the aforementioned video keyframe sequence to obtain a video frame image sequence. Specifically, starting from the video keyframe corresponding to the aforementioned initial frame extraction video image sequence number, sequential frame extraction processing can be performed on the aforementioned video keyframe sequence according to the aforementioned frame extraction step size to obtain the video frame image sequence.

[0100] Step six: Based on the aforementioned video clips and video frame sequence, generate video preview image information. Specifically, the video clip identifier corresponding to the aforementioned video clips and the aforementioned video frame sequence can be identified as the video preview image information.

[0101] Optionally, the aforementioned executing entity may send the generated video push results to the user terminal corresponding to the target user. The user terminal may display the video push results for the target user to preview and select video clips to be streamed (e.g., for purchase).

[0102] The above embodiments of this application have the following beneficial effects: the video push method of some embodiments of this application can consume less computing resources and push video materials of interest to users in a timely manner. Specifically, the reason for consuming more resources and making it difficult to push video materials of interest to users in a timely manner is that if a transferred image-language pre-trained model is used to directly perform semantic matching between the search text input by each user and all video materials in the video material library, it will consume a lot of computing power and time, resulting in a large consumption of computing resources and difficulty in pushing video materials of interest to users in a timely manner. Based on this, the video push method of some embodiments of this application firstly, in response to receiving video search text information, performs keyword extraction processing on the video search text information to obtain a search keyword sequence. Thus, each keyword in the video search text can be obtained. Secondly, a correlation analysis processing is performed on the preset video tag index information set and the above search keyword sequence to obtain a related video information set. Thus, based on the correlation relationship between keywords and the tags marked on the video materials, potential video materials of interest to users can be initially screened. Then, semantic matching processing is performed on the video retrieval text information and the corresponding video clips in the associated video information set to obtain the video information set to be displayed. Thus, based on the semantic matching relationship between the retrieval text and the video content, the video clips that the user is actually interested in can be accurately selected from the various video clips that the potential user is interested in, so as to push them to the user. Next, the video popularity information corresponding to each video clip in the video information set to be displayed is determined to obtain the video popularity information set. Thus, the attention popularity of each video clip to be pushed to the user can be obtained. Then, based on the video popularity information set, the sample video clips corresponding to the video information set to be displayed are sorted to obtain the video sequence to be previewed. Thus, high-quality video clips with higher attention popularity can be placed at the top, making it easier for users to preview them in a timely manner. Finally, based on the pre-generated video preview image information set and the video sequence to be previewed, a video push result is generated for push to the target user for preview. Therefore, the video push method of some embodiments of this application, which matches video clip tags with the user's retrieval text first and then performs semantic matching with the video clips, can reduce the number of video clips involved in semantic matching and reduce the computational power consumption and time cost of filtering videos of interest. This reduces the consumption of computing resources and allows for the timely delivery of video materials that users are interested in, thereby improving the efficiency of video streaming.

[0103] Further reference Figure 2 As an implementation of the methods shown in the above figures, this application provides some embodiments of a video push device, which are similar to... Figure 1Corresponding to the method embodiments shown, the video push device 200 can be specifically applied to various electronic devices.

[0104] like Figure 2 As shown, the video push device 200 in some embodiments includes: a keyword extraction processing unit 201, an association analysis processing unit 202, a semantic matching processing unit 203, a determination unit 204, a sorting processing unit 205, and a generation unit 206. The keyword extraction processing unit 201 is configured to extract keywords from the received video retrieval text information to obtain a retrieval keyword sequence in response to receiving the video retrieval text information; the association analysis processing unit 202 is configured to perform association analysis on a preset video tag index information set and the retrieval keyword sequence to obtain an associated video information set; the semantic matching processing unit 203 is configured to perform semantic matching on the video retrieval text information and the associated video information set to obtain a video information set to be displayed; the determination unit 204 is configured to determine the video popularity information corresponding to each video information set to be displayed in the video information set to be displayed, to obtain a video popularity information set; the sorting processing unit 205 is configured to sort the sample video material segments corresponding to the video information set to be displayed based on the video popularity information set, to obtain a video sequence to be previewed; and the generation unit 206 is configured to generate video push results based on a pre-generated video preview image information set and the video sequence to be previewed, for push to target users for preview.

[0105] It is understandable that the units described in the video push device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the video push device 200 and the units contained therein, and will not be repeated here.

[0106] Further reference Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of this application. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.

[0107] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0108] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0109] In particular, according to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this application.

[0110] It should be noted that, in some embodiments of this application, the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0111] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0112] The aforementioned computer-readable medium may be included in the aforementioned device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: , in response to receiving video retrieval text information, perform keyword extraction processing on the video retrieval text information to obtain a retrieval keyword sequence; perform association analysis processing on a preset video tag index information set and the aforementioned retrieval keyword sequence to obtain an associated video information set; perform semantic matching processing on each video clip corresponding to the aforementioned video retrieval text information and the aforementioned associated video information set to obtain a video information set to be displayed; determine the video popularity information corresponding to each video information in the aforementioned video information set to be displayed to obtain a video popularity information set; based on the aforementioned video popularity information set, sort each sample video clip corresponding to the aforementioned video information set to be displayed to obtain a video sequence to be previewed; and generate video push results based on a pre-generated video preview image information set and the aforementioned video sequence to be previewed, for push to target users for preview.

[0113] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0115] The units described in some embodiments of this application can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor can be described as including: a keyword extraction processing unit, an association analysis processing unit, a semantic matching processing unit, a determination unit, a sorting processing unit, and a generation unit. The names of these units do not necessarily limit the specific unit; for example, the keyword extraction processing unit can also be described as "a unit that performs keyword extraction processing on video retrieval text information to obtain a sequence of retrieval keywords."

[0116] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0117] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.

Claims

1. A video push method, comprising: In response to receiving video retrieval text information, keyword extraction processing is performed on the video retrieval text information to obtain a retrieval keyword sequence; A correlation analysis is performed on the preset video tag index information set and the search keyword sequence to obtain a correlated video information set, wherein each correlated video information set corresponds to a video clip. Semantic matching processing is performed on the video retrieval text information and the corresponding video clips in the associated video information set to obtain the video information set to be displayed, including: The text information retrieved from the video is processed by feature extraction to obtain text features; The text representation vector corresponding to the video retrieval text information is determined by using a word embedding model; The text representation vector is projected onto a preset feature space using a feedforward neural network to obtain a projected text vector. By using a multi-attention-based Transformer model, contextual feature representations are learned from the projected text vectors to obtain text features. For each video clip corresponding to the associated video information set, perform the following steps: The process of inputting the text features and the video clips into a pre-trained text-video semantic matching model to obtain video-text matching values ​​includes: inputting the video clips into a video spatial feature extraction module included in the pre-trained text-video semantic matching model to obtain a hierarchically extracted video feature sequence and video spatial features; inputting the text features and the hierarchically extracted video feature sequence into a video temporal feature extraction module included in the text-video semantic matching model to obtain video temporal features; inputting the video temporal features and video spatial features into a video spatiotemporal feature extraction module included in the text-video semantic matching model to obtain video spatiotemporal features; inputting the video spatiotemporal features and text features into a video text encoding module included in the text-video semantic matching model to obtain video text association features; inputting the video text association features and video spatiotemporal features into a video text decoding module included in the text-video semantic matching model to obtain a decoded video feature vector; and inputting the decoded video feature vector into a classification module included in the text-video semantic matching model to obtain video-text matching values. In response to determining that the video text matching value is greater than a preset matching threshold, the associated video information corresponding to the video material segment is determined as the video information to be displayed. Determine the video popularity information corresponding to each video information in the set of video information to be displayed to obtain the video popularity information set; Based on the video popularity information set, the sample video material segments corresponding to the video information set to be displayed are sorted to obtain the video sequence to be previewed. Each sample video material segment is a video segment composed of some video frames included in the corresponding video material segment. Based on the pre-generated set of video preview image information and the video sequence to be previewed, a video push result is generated for push to the target user for preview, wherein the target user is the user who sends the video retrieval text information through the user terminal.

2. The method according to claim 1, wherein, Each video tag index in the video tag index information set includes tag words and a set of labeled video identifiers; and the association analysis processing of the preset video tag index information set and the search keyword sequence to obtain the associated video information set includes: For each search keyword in the search keyword sequence, video tag index information whose tag words match the search keyword is selected from the video tag index information set as target video tag index information, thus obtaining a target video tag index information group; Each labeled video identifier in each labeled video identifier set corresponding to the target video tag index information group is determined as associated video information, thus obtaining an associated video information set.

3. The method according to claim 1, wherein, The step of determining the video popularity information corresponding to each video information in the video information set to be displayed, to obtain the video popularity information set, includes: For each video information to be displayed in the set of video information to be displayed, perform the following steps: User browsing behavior statistics that match the video information to be displayed are selected from a preset set of user browsing behavior statistics. Each user browsing behavior statistics includes a browsing video identifier and a behavior classification information group. Each behavior classification information in the behavior classification information group includes a behavior type identifier and a behavior count value. Based on the preset behavioral weight information group and the selected user browsing behavior statistics including the behavioral classification information group, a user interest score is generated; Obtain the video release date corresponding to the video information to be displayed; Based on the user's interest level and the video's release date, video popularity information corresponding to the video information to be displayed is generated.

4. The method according to any one of claims 1-3, wherein, The process of generating video push results based on a pre-generated set of video preview image information and the sequence of videos to be previewed includes: For each video in the video sequence to be previewed, perform the following steps: Select the video preview image information that matches the video to be previewed from the video preview image information set as the target video preview image information; The video to be previewed and the target video preview image information are determined as the video content information to be previewed; The determined video content information to be previewed is sorted to obtain a sequence of video content information to be previewed; Obtain the historical preview mode information corresponding to the target user; The sequence of historical preview mode information and the video content information to be previewed is determined as the video push result.

5. The method according to claim 4, wherein, The historical preview mode information refers to the preview mode used by the target user when they last searched for and viewed the video push results.

6. The method according to claim 1, wherein, The step of performing semantic matching processing on the video retrieval text information and the corresponding video material segments of the associated video information set to obtain the video information set to be displayed includes: The text information retrieved from the video is processed by feature extraction to obtain text features; For each video clip corresponding to the associated video information set, perform the following steps: The text features and the video clips are input into a pre-trained text-video semantic matching model to obtain video-text matching values. The text-video semantic matching model includes a video spatial feature extraction module, a video temporal feature extraction module, a video spatiotemporal feature extraction module, a video text encoding module, a video text decoding module, and a classification module. In response to determining that the video text matching value is greater than a preset matching threshold, the associated video information corresponding to the video material segment is determined as the video information to be displayed.

7. The method according to claim 4, wherein, Before generating the video push result based on the pre-generated video preview image information set and the video sequence to be previewed, the method further includes: For each video clip corresponding to the video sequence to be previewed, perform the following steps to generate video preview image information in the video preview image information set: Obtain the video keyframe sequence and the number of grid cells corresponding to the video clip; The number of each video keyframe in the video keyframe sequence is determined as the number of keyframes. Based on the number of keyframes and the number of configured grids, a frame extraction step size is generated; Determine the initial sequence number of the extracted video frames; Based on the frame extraction step size and the initial frame extraction video image sequence number, the video keyframe sequence is subjected to frame extraction processing to obtain a video frame image sequence. Based on the video footage clips and the video frame image sequence, video preview image information is generated.

8. A video push device, comprising: The keyword extraction processing unit is configured to perform keyword extraction processing on the video retrieval text information in response to receiving video retrieval text information, and obtain a retrieval keyword sequence. The association analysis and processing unit is configured to perform association analysis and processing on a preset video tag index information set and the search keyword sequence to obtain an association video information set; The semantic matching processing unit is configured to perform semantic matching processing on the video retrieval text information and the corresponding video material segments of the associated video information set to obtain a video information set to be displayed, including: The text information retrieved from the video is processed by feature extraction to obtain text features; The text representation vector corresponding to the video retrieval text information is determined by using a word embedding model; The text representation vector is projected onto a preset feature space using a feedforward neural network to obtain a projected text vector. By using a multi-attention-based Transformer model, contextual feature representations are learned from the projected text vectors to obtain text features. For each video clip corresponding to the associated video information set, perform the following steps: The process of inputting the text features and the video clips into a pre-trained text-video semantic matching model to obtain video-text matching values ​​includes: inputting the video clips into a video spatial feature extraction module included in the pre-trained text-video semantic matching model to obtain a hierarchically extracted video feature sequence and video spatial features; inputting the text features and the hierarchically extracted video feature sequence into a video temporal feature extraction module included in the text-video semantic matching model to obtain video temporal features; inputting the video temporal features and video spatial features into a video spatiotemporal feature extraction module included in the text-video semantic matching model to obtain video spatiotemporal features; inputting the video spatiotemporal features and text features into a video text encoding module included in the text-video semantic matching model to obtain video text association features; inputting the video text association features and video spatiotemporal features into a video text decoding module included in the text-video semantic matching model to obtain a decoded video feature vector; and inputting the decoded video feature vector into a classification module included in the text-video semantic matching model to obtain video-text matching values. In response to determining that the video text matching value is greater than a preset matching threshold, the associated video information corresponding to the video material segment is determined as the video information to be displayed. The determining unit is configured to determine the video popularity information corresponding to each video information to be displayed in the video information set to be displayed, thereby obtaining the video popularity information set; The sorting processing unit is configured to sort the sample video clips corresponding to the video information set to be displayed based on the video popularity information set, so as to obtain the video sequence to be previewed. The generation unit is configured to generate video push results based on a pre-generated set of video preview image information and the video sequence to be previewed, for push to the target user for preview.

9. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Data retrieval method and cross-modal data matching model processing method and device

    CN113987119A

  • Video display method and device, electronic equipment and medium

    CN115687693A