Video recommendation method and device, electronic equipment, storage medium and program product

By extracting and perceptual enhancement of the user's query video and candidate videos, and using the hash similarity recommendation method, the problem of poor video recommendation results when the user's behavior data is small, achieving more accurate user interest capture and personalized recommendation.

CN120256677AInactive Publication Date: 2025-07-04CHINA MOBILE M2M +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380405.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing video recommendation methods are difficult to accurately capture content that users are interested in when there is little user behavior data, resulting in poor recommendation results.

Method used

By extracting the video semantic feature of the user's query video and candidate video, using the video twin detection network to obtain the semantic feature map, perform semantic perception enhancement and hash to binary code, calculate the Hamming distance to determine the hash similarity, and recommend candidate videos based on the similarity.

Benefits of technology

With less user behavior data, it can more accurately capture content that users are interested in, improve the accuracy and personalization of video recommendations, and improve user participation and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256677A_ABST
    Figure CN120256677A_ABST
Patent Text Reader

Abstract

The invention provides a video recommendation method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of video recommendation, and the method comprises the steps: carrying out the video semantic feature extraction of a query video of a user and a candidate video corresponding to the query video, and obtaining a plurality of video semantic feature maps; performing semantic perception enhancement on each video semantic feature map to obtain a plurality of enhanced video semantic feature maps; hashing each enhanced video semantic feature map to a binary code, and determining the Hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary code; and based on the Hash similarity, determining recommendation information of the candidate video. According to the method, the complexity and diversity of the video content can be understood, so that the essence of the video content can be captured more accurately, and the matching degree of video recommendation and user interest is improved. And under the condition of less user behavior data, the content interested by the user is accurately captured, personalized video recommendation is generated, and the participation degree and the satisfaction degree of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recommendation, and particularly to a video recommendation method, device, electronic device, storage medium, and program product. Background Art

[0002] Currently, video platforms have become an important channel for users to obtain information and entertainment. To improve the user experience, video recommendation systems have emerged. The core of video recommendation lies in recommending the most relevant and interesting video content to users based on the types or themes that users are interested in, so as to improve user participation and satisfaction. However, current video recommendation methods are difficult to accurately capture the content that users are interested in when there is less user behavior data, resulting in poor recommendation effects. Summary of the Invention

[0003] The present invention provides a video recommendation method, device, electronic device, storage medium, and program product to solve the defect of poor video recommendation effect in the prior art, and to accurately capture the content that users are interested in and generate personalized video recommendations when there is less user behavior data, thereby improving user participation and satisfaction.

[0004] The present invention provides a video recommendation method, including the following steps: Extract video semantic features of the user's query video and candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; Perform semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; Hash each of the enhanced video semantic feature maps into binary codes, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; Determine the recommendation information of the candidate videos based on the hash similarity.

[0005] According to a video recommendation method provided by the present invention, the enhanced video semantic feature maps include an enhanced query video semantic feature map and an enhanced candidate video semantic feature map; The step of performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps includes: Input the video semantic feature map into the first branch of the feature multi-scale perception enhancement module to obtain a first-channel video semantic local activation feature vector output by the first branch; Input the video semantic feature map into the second branch of the feature multi-scale perception enhancement module to obtain a second-channel compressed video semantic feature map output by the second branch; Input the video semantic feature map into the third branch of the feature multi-scale perception enhancement module to obtain the video semantic receptive field expansion global activation feature matrix output by the third branch; Perform feature fusion on the second-channel compressed video semantic feature map with the first-channel video semantic local activation feature vector and the video semantic receptive field expansion global activation feature matrix respectively to obtain the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map; Add the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map by position to obtain the second-channel compressed video semantic multi-scale fusion activation feature map; Perform dilated convolution encoding on the second-channel compressed video semantic multi-scale fusion activation feature map to obtain the enhanced video semantic feature map.

[0006] According to a video recommendation method provided by the present invention, the first branch is used for: Perform point convolution processing on the video semantic feature map to obtain the first-channel video semantic compressed feature map; Perform global average pooling on each feature matrix along the channel dimension in the first-channel video semantic compressed feature map to obtain the first-channel video semantic compressed feature vector; Perform non-linear activation on the first-channel video semantic compressed feature vector to obtain the first-channel video semantic local activation feature vector.

[0007] According to a video recommendation method provided by the present invention, the second branch is used for: Perform point convolution processing on the video semantic feature map to obtain the second-channel compressed video semantic feature map.

[0008] According to a video recommendation method provided by the present invention, the third branch is used for: Perform dilated convolution encoding on the video semantic feature map to obtain the video semantic receptive field expansion feature map; Perform point convolution processing on the video semantic receptive field expansion feature map to obtain the video semantic receptive field expansion global feature matrix; Perform non-linear activation on the video semantic receptive field expansion global feature matrix to obtain the video semantic receptive field expansion global activation feature matrix.

[0009] According to a video recommendation method provided by the present invention, the video semantic feature map includes a query video semantic feature map and a candidate video semantic feature map; Performing video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps, including: Input the query video and the candidate video into a video twin detection network including a first video encoder and a second video encoder to obtain the query video semantic feature map and the candidate video semantic feature map; Among them, the first video encoder and the second video encoder have the same network structure.

[0010] According to a video recommendation method provided by the present invention, determining the recommendation information of the candidate video based on the hash similarity includes: When the hash similarity is greater than a preset threshold, recommend the candidate video to the user.

[0011] The present invention also provides a video recommendation device, including the following modules: A video semantic feature extraction module, configured to perform video semantic feature extraction on the user's query video and the candidate video corresponding to the query video to obtain multiple video semantic feature maps; A semantic perception enhancement module, configured to perform semantic perception enhancement on each of the video semantic feature maps to obtain multiple enhanced video semantic feature maps; A hash similarity determination module, configured to hash each of the enhanced video semantic feature maps into binary codes, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; A video recommendation module, configured to determine the recommendation information of the candidate video based on the hash similarity.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the video recommendation method described in any one of the above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the video recommendation method described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the video recommendation method described in any one of the above is implemented.

[0015] The video recommendation method, device, electronic device, storage medium and program product provided by the present invention extract video semantic features from a user's query video and candidate videos corresponding to the query video to obtain multiple video semantic feature maps; perform semantic perception enhancement on each video semantic feature map to obtain multiple enhanced video semantic feature maps; hash each enhanced video semantic feature map into a binary code, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; determine the recommendation information of the candidate videos based on the hash similarity. This method can understand the complexity and diversity of video content, more accurately capture the essence of video content, and improve the matching degree between video recommendations and user interests. In the case of less user behavior data, it can accurately capture the content that users are interested in and generate personalized video recommendations, enhancing user participation and satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0017] Figure 1 is one of the flow diagrams of the video recommendation method provided by the present invention.

[0018] Figure 2 is the second flow diagram of the video recommendation method provided by the present invention.

[0019] Figure 3 is the architecture diagram of the video recommendation method provided by the present invention.

[0020] Figure 4 is the flow diagram of determining the enhanced video semantic feature map provided by the present invention.

[0021] Figure 5 is the processing flow diagram of the first branch provided by the present invention.

[0022] Figure 6 is the processing flow diagram of the third branch provided by the present invention.

[0023] Figure 7 is the flow diagram of feature fusion provided by the present invention.

[0024] Figure 8 is the structural diagram of the video recommendation device provided by the present invention.

[0025] Figure 9 is the structural diagram of the electronic device provided by the present invention. Detailed Implementation Manner

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0027] The following combines Figures 1-9 to describe the video recommendation method, device, electronic device, storage medium, and program product of the present invention.

[0028] Figure 1 is one of the flow diagrams of the video recommendation method provided by the present invention. As Figure 1 shown, the method includes the following:[[]]END]] Step 101: Extract video semantic features from the user's query video and the candidate videos corresponding to the query video to obtain multiple video semantic feature maps.

[0029] First, obtain the user's query video, and then extract the candidate videos corresponding to the query video from the candidate video library. Herein, the candidate video may be a first candidate video. It should be understood that the query video refers to the video that the user is currently watching or has shown interest in. In the recommendation system, this video is used as the starting point for analyzing the user's interests and preferences; the first candidate video is the video selected from the candidate video library according to a preset criterion or algorithm, and it is considered to be the most relevant to the query video or the video that the user may be most interested in; the candidate video library is a set containing multiple videos available for recommendation. Based on this, by analyzing the query video that the user is watching, the recommendation system can understand the user's current interest points and extract videos from the candidate video library to provide personalized recommendations. By recommending content related to the video that the user is currently watching, the viewing satisfaction of the user and the platform stickiness can be improved.

[0030] Considering that both the query video and the candidate videos contain a large amount of content information and semantic features of video frames. Based on this, in order to capture and extract the video semantic feature information contained in the query video and the candidate videos respectively. In the embodiments of the present invention, the query video and the candidate videos are input into a video twin detection network including a first video encoder and a second video encoder to obtain video semantic feature maps. Among them, the video semantic feature maps include a query video semantic feature map and a candidate video semantic feature map.

[0031] Considering that both the query video and the candidate video contain a large amount of content information and semantic features of video frames. Based on this, in order to capture and extract the video semantic feature information contained in the query video and the candidate video respectively, in the embodiments of the present invention, the query video and the candidate video are input into a video twin detection network including a first video encoder and a second video encoder to obtain a query video semantic feature map and a candidate video semantic feature map.

[0032] Among them, the first video encoder and the second video encoder have the same network structure. That is to say, the video twin detection network including the first video encoder and the second video encoder can capture the implicit semantic feature information in the video more deeply, such as the scene, action, object, etc. described in the video, which provides help and support for subsequent evaluation of the similarity between the query video and the candidate video, because it determines whether the candidate video matches the user's query intention.

[0033] It can be understood that the query video semantic feature map represents the semantic information of the query video, such as the scene, action, object, etc.; the candidate video semantic feature map represents the semantic information of the candidate video. These feature maps are high-dimensional representations of the video content and can capture the deep semantic information hidden in the video.

[0034] It can be understood that the video twin detection network is a symmetric two-branch neural network for processing paired input videos (such as the query video and the candidate video). It includes two video encoders with the same structure (the first video encoder and the second video encoder), which are respectively used to extract the semantic features of the query video and the candidate video. Among them, the first video encoder is used to process the query video and extract its semantic features; the second video encoder is used to process the candidate video and extract its semantic features; the two video encoders have the same network structure and parameters to ensure the same feature extraction method for the query video and the candidate video; each encoder will generate a semantic feature map, which respectively represents the semantic information of the query video and the candidate video.

[0035] It can be understood that the main functions of the video encoder include: 1) Video frame extraction: Extract key frames or continuous frame sequences from the input video. 2) Feature extraction: Extract visual features from the video frames through a convolutional neural network or other deep learning models. 3) Temporal modeling: Use a recurrent neural network or a Transformer model, etc., to capture the temporal information in the video (such as actions, scene changes, etc.). 4) Semantic feature generation: Fuse the visual features and the temporal information to generate a high-dimensional semantic feature map, which represents the deep semantic information of the video.

[0036] By using a video twin detection network to extract the semantic feature maps of the query video and the candidate video, it is possible to capture the semantic information in the video (such as scenes, actions, objects, etc.) at a deeper level, thereby achieving more accurate video recommendation or matching tasks. The core advantages of this method lie in its symmetry and deep feature extraction ability, which can effectively improve the performance of video analysis tasks.

[0037] Step 102: Perform semantic perception enhancement on each of the video semantic feature maps to obtain multiple enhanced video semantic feature maps.

[0038] It can be understood that the main purpose of semantic perception enhancement is to enhance the expressive ability of the video semantic feature maps so that they can more accurately capture the semantic information in the video (such as scenes, actions, objects, etc.). Through the enhancement process, the feature maps can better reflect the core content of the video, thereby improving the accuracy of video recommendation or matching.

[0039] Specifically, perform semantic perception enhancement on the query video semantic feature map and the candidate video semantic feature map respectively to obtain an enhanced query video semantic feature map and an enhanced candidate video semantic feature map. Among them, the enhanced video semantic feature maps include the enhanced query video semantic feature map and the enhanced candidate video semantic feature map.

[0040] Step 103: Hash each of the enhanced video semantic feature maps into binary codes, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes.

[0041] In order to quickly compare the similarity between the enhanced query video semantic feature map and the enhanced candidate video semantic feature map, so as to more accurately quantify the semantic similarity between the candidate video and the query video, in the embodiments of the present invention, the hash similarity between the enhanced query video semantic feature map and the enhanced candidate video semantic feature map is calculated. It should be understood that the hash similarity can be calculated by hashing the feature maps into binary codes and then comparing the Hamming distances of these binary codes. This calculation process is very fast and efficient, and can provide a quantitative evaluation of the similarity between two features. It should be understood that the purpose of hash coding is to convert the high-dimensional enhanced video semantic feature maps into binary codes in order to reduce storage and computational costs and support fast matching. Specifically, binary codes occupy less storage space and have higher computational efficiency; by calculating the Hamming distance between binary codes, the similarity between video semantic feature maps can be quickly judged.

[0042] For example, the hash similarity can be calculated in the following way: 1) Hash each enhanced video semantic feature map into a binary code so that each enhanced video semantic feature map is converted into a binary code of a fixed length (such as 64 bits, 128 bits, etc.). Among them, it can be achieved through hash coding methods such as locality-sensitive hashing and Deep Hashing. These methods map high-dimensional features into binary codes through specific hash functions.

[0043] 2) Calculate the Hamming distance between binary codes. The Hamming distance refers to the number of different values at the same position of two binary codes. For example, the Hamming distance between binary codes 1010 and 1100 is 2. It should be understood that the smaller the Hamming distance, the more similar the two binary codes are; the larger the Hamming distance, the less similar the two binary codes are.

[0044] 3) Based on the Hamming distance between binary codes, calculate the hash similarity between enhanced video semantic feature maps. For example, the hash similarity can be defined as: hash similarity = 1 - (Hamming distance / binary code length). It should be understood that the hash similarity can be a value between 0 and 1, and the larger the value, the more similar the two video semantic feature maps are.

[0045] By calculating the hash similarity between the enhanced query video semantic feature map and the enhanced candidate video semantic feature map, the semantic difference information between the query video and the candidate video can be detected in a timely manner, making the obtained similarity value more accurate, thereby improving the accuracy of video recommendation. In addition, through hash coding, the system can quickly judge the similarity between video semantic feature maps, thereby supporting tasks such as large-scale video recommendation, retrieval, and deduplication.

[0046] Step 104, determine the recommendation information of the candidate video based on the hash similarity.

[0047] Based on the comparison between the hash similarity and a preset threshold, determine whether to recommend the candidate video. In one embodiment, when the hash similarity is greater than the preset threshold, recommend the candidate video to the user. It should be understood that the preset threshold is a pre-set similarity threshold used to determine whether the candidate video is similar enough to the query video to decide whether to recommend. If the hash similarity is greater than the preset threshold, it is considered that the candidate video is similar enough to the query video, and the candidate video is recommended to the user; if the hash similarity is less than or equal to the preset threshold, it is considered that the candidate video is not similar enough to the query video, and the candidate video is not recommended.

[0048] By comparing the hash similarity between the enhanced query video semantic feature map and the enhanced candidate video semantic feature map with the preset threshold, thereby automatically determining whether to recommend the candidate video. Through this method, the complexity and diversity of video content can be understood to more accurately capture the essence of video content and improve the matching degree between video recommendation and user interests.

[0049] The video recommendation method provided by the embodiments of the present invention extracts video semantic feature maps for the user's query video and candidate videos corresponding to the query video, obtaining multiple video semantic feature maps; performs semantic perception enhancement on each video semantic feature map to obtain multiple enhanced video semantic feature maps; hashes each enhanced video semantic feature map into a binary code, and determines the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary code; determines the recommendation information of the candidate videos based on the hash similarity. Based on the above method, the complexity and diversity of video content can be understood to more accurately capture the essence of video content and improve the matching degree between video recommendation and user interests. At the same time, in the case of less user behavior data, the content that the user is interested in can be accurately captured and personalized video recommendations can be generated, thereby enhancing the user's participation and satisfaction.

[0050] In one embodiment, the performing semantic perception enhancement on each of the video semantic feature maps to obtain multiple enhanced video semantic feature maps includes: Step 1020: Input the video semantic feature map into the first branch of the feature multi-scale perception enhancement module to obtain a first-channel video semantic local activation feature vector output by the first branch; Step 1021: Input the video semantic feature map into the second branch of the feature multi-scale perception enhancement module to obtain a second-channel compressed video semantic feature map output by the second branch; Step 1022: Input the video semantic feature map into the third branch of the feature multi-scale perception enhancement module to obtain a video semantic receptive field expansion global activation feature matrix output by the third branch; Step 1023: Perform feature fusion on the second-channel compressed video semantic feature map with the first-channel video semantic local activation feature vector and the video semantic receptive field expansion global activation feature matrix respectively to obtain a second-channel compressed video semantic local activation feature map and a second-channel compressed video semantic global activation feature map; Step 1024: Add the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map by position to obtain a second-channel compressed video semantic multi-scale fusion activation feature map; Step 1025: Perform dilated convolution encoding on the second-channel compressed video semantic multi-scale fusion activation feature map to obtain the enhanced video semantic feature map.

[0051] Considering that the semantic feature maps of the query video and the candidate video exhibit different semantic feature information at different time scales. That is to say, both the semantic feature map of the query video and the semantic feature map of the candidate video have manifestations of semantic feature information at different scales. For example, some semantic features may change rapidly within a few frames, while other features may remain stable over a longer time period. Therefore, in order to more finely understand and analyze the local and global semantic information expressed by the query video semantic feature map and the candidate video semantic feature map at different time scales, so as to obtain a comprehensive and all-round video and feature representation. In the embodiments of the present invention, the semantic feature map of the query video and the semantic feature map of the candidate video are respectively input into the feature multi-scale perception reinforcement module to obtain the enhanced query video semantic feature map and the enhanced candidate video semantic feature map output by the feature multi-scale perception reinforcement module.

[0052] It can be understood that the feature multi-scale perception reinforcement module performs feature processing on the input feature map through three different branches to respectively capture and extract the local information and global information in the input feature map for semantic enhancement processing of the input feature map. Specifically, first, in the first branch, the channel compression and non-linear activation processing of the video semantic feature map capture and mine the local detailed semantic feature information hidden in the video semantic feature map. Then, in the second branch, point convolution processing is performed on the video semantic feature map to reduce the consumption of computing resources and improve the running efficiency of the model. Then, in the third branch, dilated convolution encoding is used for the video semantic feature map to expand the global perception feature information of the video semantic feature map. Finally, the feature information obtained in the second branch is respectively fused and added with the information obtained in the first branch and the third branch, and then the receptive field is expanded again to enhance the local information and global semantic feature information in the video semantic feature map, so as to obtain the enhanced query video semantic feature map and the enhanced candidate video semantic feature map containing rich semantic information.

[0053] Specifically, input the video semantic feature map into the first branch of the feature multi-scale perception enhancement module to obtain the first-channel video semantic local activation feature vector output by the first branch; input the video semantic feature map into the second branch of the feature multi-scale perception enhancement module to obtain the second-channel compressed video semantic feature map output by the second branch; input the video semantic feature map into the third branch of the feature multi-scale perception enhancement module to obtain the video semantic receptive field expansion global activation feature matrix output by the third branch; perform feature fusion on the second-channel compressed video semantic feature map with the first-channel video semantic local activation feature vector and the video semantic receptive field expansion global activation feature matrix respectively to obtain the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map; perform position-wise addition on the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map to obtain the second-channel compressed video semantic multi-scale fusion activation feature map; perform dilated convolution encoding on the second-channel compressed video semantic multi-scale fusion activation feature map to obtain the enhanced video semantic feature map.

[0054] By performing multi-scale processing on the video semantic feature map, it is possible to capture the multi-scale information of the video, enhance the semantic expression ability, and thus improve the accuracy and personalization ability of video recommendations.

[0055] In one embodiment, the first branch is used to: perform point convolution processing on the video semantic feature map to obtain the first-channel video semantic compressed feature map; perform global average pooling on each feature matrix along the channel dimension in the first-channel video semantic compressed feature map to obtain the first-channel video semantic compressed feature vector; perform non-linear activation on the first-channel video semantic compressed feature vector to obtain the first-channel video semantic local activation feature vector.

[0056] It can be understood that the first branch is a branch in the feature multi-scale perception enhancement module, focusing on local information extraction and compression of the video semantic feature map.

[0057] Specifically, the video semantic feature map is subjected to point convolution processing, wherein point convolution is a 1x1 convolution, which is mainly used for the transformation of channel dimension. It can compress or expand the number of channels of the input feature map while retaining the information of the spatial dimension. The first channel video semantic compression feature map obtained by point convolution retains the key spatial information. Then, the first channel video semantic compression feature map is subjected to global average pooling along the channel dimension, wherein global average pooling compresses the feature matrix of each channel into a scalar, that is, taking the average value of all spatial positions of each channel; the first channel video semantic compression feature vector obtained by global average pooling is a one-dimensional vector, which represents the global information of each channel. Finally, the first channel video semantic compression feature vector is subjected to nonlinear activation (such as ReLU, Sigmoid, etc.), wherein nonlinear activation introduces nonlinear transformation to enhance the expression ability of the feature vector. The first channel video semantic local activation feature vector obtained by nonlinear activation is a vector that has undergone nonlinear transformation and can better represent local semantic information.

[0058] The first branch extracts and compresses local information by performing point convolution, global mean pooling and nonlinear activation on the video semantic feature map to generate the first channel video semantic local activation feature vector. This branch focuses on the extraction of local information and is an important part of the feature multi-scale perception enhancement module. It can improve the expression ability of video semantic features, thereby supporting more accurate video recommendation or matching tasks.

[0059] In one embodiment, the second branch is used to perform point convolution processing on the video semantic feature map to obtain a second channel compressed video semantic feature map.

[0060] It can be understood that the second branch is another branch in the feature multi-scale perceptual enhancement module, focusing on global information extraction and compression of video semantic feature maps.

[0061] Specifically, the video semantic feature map is subjected to point convolution, where point convolution is a 1x1 convolution, which is mainly used for channel dimension transformation. It can compress or expand the number of channels of the input feature map while retaining the information of the spatial dimension. The second channel compressed video semantic feature map obtained by point convolution retains the key spatial information.

[0062] Through point convolution processing, the second branch focuses on extracting global information in the video semantic feature map. Point convolution compresses the high-dimensional feature map into a low-dimensional feature map, reducing the amount of calculation while retaining key information.

[0063] In one embodiment, the third branch is used to: perform hole convolution encoding on the video semantic feature map to obtain a video semantic receptive field expansion feature map; perform point convolution processing on the video semantic receptive field expansion feature map to obtain a video semantic receptive field expansion global feature matrix; perform nonlinear activation on the video semantic receptive field expansion global feature matrix to obtain a video semantic receptive field expansion global activation feature matrix.

[0064] It can be understood that the third branch is another branch in the feature multi-scale perception enhancement module, focusing on multi-scale information extraction and fusion of video semantic feature maps.

[0065] Specifically, the video semantic feature map is encoded by dilated convolution, where dilated convolution can expand the receptive field without increasing the number of parameters by introducing holes (i.e., intervals) in the convolution kernel, thereby capturing a wider range of contextual information. The video semantic receptive field expansion feature map obtained by dilated convolution contains richer multi-scale information. Then, the video semantic receptive field expansion feature map is processed by point convolution. For example, the number of channels of the input feature map can be compressed or expanded while retaining the information of the spatial dimension. The video semantic receptive field expansion global feature matrix obtained by point convolution is a more compact feature representation. Finally, the video semantic receptive field expansion global feature matrix is ​​nonlinearly activated, where nonlinear activation introduces nonlinear transformation to enhance the expression ability of the feature matrix. The video semantic receptive field expansion global activation feature matrix obtained by nonlinear activation is a feature matrix that has undergone nonlinear transformation and can better represent multi-scale semantic information.

[0066] The third branch extracts and fuses multi-scale information by performing hole convolution coding, point convolution processing and non-linear activation on the video semantic feature map, generating a global activation feature matrix for expanding the video semantic receptive field. This branch focuses on the extraction of multi-scale information and is an important part of the feature multi-scale perception enhancement module. It can improve the expression ability of video semantic features, thereby supporting more accurate video recommendation or matching tasks.

[0067] In one embodiment, step 1023 includes: positionally multiplying the first channel video semantic local activation feature vector with each feature matrix of the second channel compressed video semantic feature map along the channel dimension to obtain the second channel compressed video semantic local activation feature map; and positionally multiplying the video semantic receptive field expansion global activation feature matrix with each corresponding feature matrix of the second channel compressed video semantic feature map along the channel dimension to obtain the second channel compressed video semantic global activation feature map.

[0068] It can be understood that by multiplying by position, the first-channel video semantic local activation feature vector (local information) is fused with the second-channel compressed video semantic feature map (global information) to generate the second-channel compressed video semantic local activation feature map, which contains both local details and global context information and can represent the semantic features of the video more comprehensively. Moreover, by multiplying by position, the video semantic receptive field expansion global activation feature matrix (multi-scale information) is fused with the second-channel compressed video semantic feature map (global information) to generate the second-channel compressed video semantic global activation feature map, which contains both multi-scale and global context information and can represent the semantic features of the video more accurately.

[0069] By fusing local information, global information, and multi-scale information, the generated feature map can represent the semantic features of the video more comprehensively. The fused feature map can more accurately reflect the core content of the video, thereby improving the matching accuracy between the query video and the candidate videos. At the same time, the fused feature map can better reflect the user's interests and preferences, thus supporting more personalized video recommendations.

[0070] In one embodiment, the video semantic feature map is input into the feature multi-scale perception reinforcement module and processed according to the following reinforcement formula to obtain the reinforced video semantic feature map. Wherein, the reinforcement formula is: ; Wherein, represents the video semantic feature map, represents point convolution processing on the feature map, represents global average pooling processing on each feature matrix along the channel dimension in the feature map, represents non-linear activation processing, represents the first-channel video semantic local activation feature vector, represents the second-channel compressed video semantic feature map, represents dilated convolution encoding on the feature map, represents the video semantic receptive field expansion global activation feature matrix, represents addition by position, represents multiplication by position, represents the reinforced video semantic feature map.

[0071] To further analyze and explain the video recommendation method proposed by the present invention, refer to the following embodiments.

[0072] As Figures 2-3 shown, the video recommendation method provided by the embodiment of the present invention mainly includes the following steps: S110, obtain the query video; S120, extract the first candidate video from the candidate video library; S130, perform video semantic feature extraction on the query video and the first candidate video to obtain a query video semantic feature map and a first candidate video semantic feature map; S140, perform semantic perception enhancement on the query video semantic feature map and the first candidate video semantic feature map to obtain an enhanced query video semantic feature map and an enhanced first candidate video semantic feature map; S150, calculate the hash similarity between the enhanced query video semantic feature map and the enhanced first candidate video semantic feature map; S160, based on the hash similarity, determine whether to recommend the first candidate video.

[0073] In S110 - S120, it should be understood that by obtaining the query video, extracting the first candidate video from the candidate video library, and analyzing the query video that the user is currently watching, the recommendation system can understand the user's current interest points and extract videos from the candidate library to provide personalized recommendations. In this way, by recommending content related to the video that the user is currently watching, the user's viewing satisfaction and platform stickiness can be improved.

[0074] In S130, it should be understood that the query video and the first candidate video are input into a video twin detection network including a first video encoder and a second video encoder to obtain a query video semantic feature map and a first candidate video semantic feature map. Correspondingly, considering that both the query video and the first candidate video contain a large amount of content information and semantic features of video frames. Based on this, in order to capture and extract the video semantic feature information contained in the query video and the first candidate video respectively, in the technical solution of the present invention, the query video and the first candidate video are input into a video twin detection network including a first video encoder and a second video encoder to obtain the query video semantic feature map and the first candidate video semantic feature map output by the video twin detection network. It should be understood that through the video twin detection network including the first video encoder and the second video encoder, the semantic feature information hidden in the video can be captured more deeply, such as the scenes, actions, objects, etc. described in the video, which provides help and support for subsequent evaluation of the similarity between the query video and the first candidate video, because it determines whether the first candidate video matches the user's query intention.

[0075] In S140, it should be understood that the query video semantic feature map and the first candidate video semantic feature map are input into the feature multi-scale perception reinforcement module to obtain the reinforced query video semantic feature map and the reinforced first candidate video semantic feature map output by the feature multi-scale perception reinforcement module. Considering that the semantic feature information shown by the query video semantic feature map and the first candidate video semantic feature map at different time scales is different, that is to say, both the query video semantic feature map and the first candidate video semantic feature map have the manifestation of semantic feature information at different scales. For example, some semantic features may change rapidly within a few frames, while other features may remain stable over a longer period of time. Therefore, in order to more finely understand and analyze the local and global semantic information expressed by the query video semantic feature map and the first candidate video semantic feature map at different time scales to obtain a comprehensive and overall video and feature representation, in the technical solution of the present invention, the query video semantic feature map and the first candidate video semantic feature map are input into the feature multi-scale perception reinforcement module to obtain the reinforced query video semantic feature map and the reinforced first candidate video semantic feature map. It should be understood that the feature multi-scale perception reinforcement module performs feature processing on the input feature map through three different branches to respectively capture and extract the local information and global information in the input feature map for semantic reinforcement processing of the input feature map. Specifically, first, in the first branch, channel compression and non-linear activation processing on the video semantic feature map capture and mine the local detailed semantic feature information hidden in the video semantic feature map. Then, in the second branch, point convolution processing is performed on the video semantic feature map to reduce the consumption of computing resources and improve the running efficiency of the model. Then, in the third branch, dilated convolution encoding is used on the video semantic feature map to expand the global perception feature information of the video semantic feature map. Finally, the feature information obtained in the second branch is respectively fused and added to the information obtained in the first branch and the third branch, and then the receptive field is expanded again to strengthen the local information and global semantic feature information in the video semantic feature map, thereby obtaining the reinforced query video semantic feature map and the reinforced first candidate video semantic feature map containing rich semantic information.

[0076] As Figure 4 shown, taking the determination of the reinforced query video semantic feature map as an example, S140 may include: S210, in the first branch, perform first-channel feature extraction on the query video semantic feature map to obtain a first-channel query video semantic local activation feature vector; S220, in the second branch, perform second-channel feature extraction on the query video semantic feature map to obtain a second-channel compressed query video semantic feature map; S230, in the third branch, perform receptive field expansion on the query video semantic feature map to obtain a query video semantic receptive field expansion global activation feature matrix; S240. Feature fusion is performed on the compressed query video semantic feature map of the second channel with the query video semantic receptive field expansion global activation feature matrix and the first channel query video semantic local activation feature vector respectively to obtain the second channel compressed query video semantic global activation feature map and the second channel compressed query video semantic local activation feature map; S250. The second channel compressed query video semantic global activation feature map and the second channel compressed query video semantic local activation feature map are added by position to obtain the second channel compressed query video semantic multi-scale fusion activation feature map; S260. Dilated convolution encoding is performed on the second channel compressed query video semantic multi-scale fusion activation feature map to obtain the enhanced query video semantic feature map.

[0077] It should be understood that the enhancement processing method of the first candidate video semantic feature map is consistent with that of the query video semantic feature map.

[0078] As Figure 5 shown, S210 may include: S310. In the first branch, point convolution processing is performed on the query video semantic feature map to obtain the first channel query video semantic compressed feature map; S320. Global average pooling is performed on each feature matrix along the channel dimension in the first channel query video semantic compressed feature map to obtain the first channel query video semantic compressed feature vector; S330. Nonlinear activation is performed on the first channel query video semantic compressed feature vector to obtain the first channel query video semantic local activation feature vector.

[0079] In an embodiment of the present invention, S220 may include: in the second branch, point convolution processing is performed on the query video semantic feature map to obtain the second channel compressed query video semantic feature map.

[0080] As Figure 6 shown, S230 may include: S410. In the third branch, dilated convolution encoding is performed on the query video semantic feature map to obtain the query video semantic receptive field expansion feature map; S420. Point convolution processing is performed on the query video semantic receptive field expansion feature map to obtain the query video semantic receptive field expansion global feature matrix; S430. Nonlinear activation is performed on the query video semantic receptive field expansion global feature matrix to obtain the query video semantic receptive field expansion global activation feature matrix.

[0081] As Figure 7As shown, S240 may include: S510, multiplying the expanded global activation feature matrix of the query video semantic receptive field and the respective corresponding feature matrices of the second-channel compressed query video semantic feature map along the channel dimension to obtain the second-channel compressed query video semantic global activation feature map; S520, multiplying the first-channel query video semantic local activation feature vector and the respective feature matrices of the second-channel compressed query video semantic feature map along the channel dimension to obtain the second-channel compressed query video semantic local activation feature map.

[0082] In S150, it should be understood that in order to quickly compare the similarity between the enhanced query video semantic feature map and the enhanced first candidate video semantic feature map, so as to more accurately quantify the semantic similarity between the first candidate video and the query video, in the technical solution of the present invention, the hash similarity between the enhanced query video semantic feature map and the enhanced first candidate video semantic feature map is calculated. It is worth mentioning that the hash similarity can be calculated by hashing the feature map into binary codes and then comparing the Hamming distances of these binary codes. This calculation process is very fast and efficient and can provide a quantitative evaluation of the similarity between two features. That is to say, by calculating the hash similarity between the enhanced query video semantic feature map and the enhanced first candidate video semantic feature map, the semantic difference information between the query video and the first candidate video can be detected in a timely manner, making the obtained similarity value more accurate, thereby improving the accuracy of video recommendation.

[0083] In S160, it should be understood that based on the comparison between the hash similarity and a preset threshold, it is determined whether to recommend the first candidate video. That is, by comparing the hash similarity between the enhanced query video semantic feature map and the enhanced first candidate video semantic feature map with the preset threshold, it is automatically determined whether to recommend the first candidate video. Through this method, the complexity and diversity of video content can be understood to more accurately capture the essence of video content and improve the matching degree between video recommendation and user interests.

[0084] The video recommendation method provided by the embodiments of the present invention obtains a query video, extracts a first candidate video from a candidate video library, and uses video processing and analysis techniques based on a deep learning neural network to perform video semantic feature extraction and multi-scale enhancement on the query video and the first candidate video, and thus automatically determines whether to recommend the first candidate video according to the comparison between the semantic similarity between the query video and the first candidate video and a preset threshold. Through this method, the complexity and diversity of video content can be understood to more accurately capture the essence of video content and improve the matching degree between video recommendation and user interests.

[0085] The video recommendation device provided by the present invention will be described below. The video recommendation device described below can be mutually referred to the video recommendation method described above.

[0086] Reference Figure 8 , Figure 8 FIG. is a schematic structural diagram of the video recommendation device provided by the present invention. The video recommendation device provided by the present invention includes a video semantic feature extraction module 801, a semantic perception enhancement module 802, a hash similarity determination module 803, and a video recommendation module 804.

[0087] The video semantic feature extraction module 801 is configured to perform video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; The semantic perception enhancement module 802 is configured to perform semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; The hash similarity determination module 803 is configured to hash each of the enhanced video semantic feature maps into a binary code, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary code; The video recommendation module 804 is configured to determine the recommendation information of the candidate videos based on the hash similarity.

[0088] The video recommendation device provided by the embodiments of the present invention performs video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; performs semantic perception enhancement on each video semantic feature map to obtain a plurality of enhanced video semantic feature maps; hashes each enhanced video semantic feature map into a binary code, and determines the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary code; and determines the recommendation information of the candidate videos based on the hash similarity. This device can understand the complexity and diversity of video content, more accurately capture the essence of video content, and improve the matching degree between video recommendations and user interests. At the same time, in the case of less user behavior data, it can accurately capture the content that the user is interested in and generate personalized video recommendations, thereby enhancing user participation and satisfaction.

[0089] In one embodiment, the semantic perception enhancement module 802 is specifically configured to: Input the video semantic feature map into the first branch of the feature multi-scale perception enhancement module to obtain a first-channel video semantic local activation feature vector output by the first branch; Input the video semantic feature map into the second branch of the feature multi-scale perception enhancement module to obtain a second-channel compressed video semantic feature map output by the second branch; Input the video semantic feature map into the third branch of the feature multi-scale perception enhancement module to obtain the video semantic receptive field expansion global activation feature matrix output by the third branch; Perform feature fusion on the second-channel compressed video semantic feature map with the first-channel video semantic local activation feature vector and the video semantic receptive field expansion global activation feature matrix respectively to obtain the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map; Perform position-wise addition on the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map to obtain the second-channel compressed video semantic multi-scale fusion activation feature map; Perform dilated convolution encoding on the second-channel compressed video semantic multi-scale fusion activation feature map to obtain the enhanced video semantic feature map.

[0090] In one embodiment, the first branch is used for: Perform point convolution processing on the video semantic feature map to obtain the first-channel video semantic compressed feature map; Perform global average pooling on each feature matrix along the channel dimension in the first-channel video semantic compressed feature map to obtain the first-channel video semantic compressed feature vector; Perform non-linear activation on the first-channel video semantic compressed feature vector to obtain the first-channel video semantic local activation feature vector.

[0091] In one embodiment, the second branch is used for: Perform point convolution processing on the video semantic feature map to obtain the second-channel compressed video semantic feature map.

[0092] In one embodiment, the third branch is used for: Perform dilated convolution encoding on the video semantic feature map to obtain the video semantic receptive field expansion feature map; Perform point convolution processing on the video semantic receptive field expansion feature map to obtain the video semantic receptive field expansion global feature matrix; Perform non-linear activation on the video semantic receptive field expansion global feature matrix to obtain the video semantic receptive field expansion global activation feature matrix.

[0093] In one embodiment, the video semantic feature extraction module 801 is specifically used for: Input the query video and the candidate video into the video siamese detection network including the first video encoder and the second video encoder to obtain the query video semantic feature map and the candidate video semantic feature map; Among them, the first video encoder and the second video encoder have the same network structure.

[0094] In one embodiment, the video recommendation module 804 is specifically configured to: When the hash similarity is greater than a preset threshold, recommend the candidate video to the user.

[0095] Figure 9 Illustrates a schematic physical structure diagram of an electronic device, as Figure 9 shown. The electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call the logical instructions in the memory 930 to execute the video recommendation method, which includes: extracting video semantic feature maps for the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; hashing each of the enhanced video semantic feature maps into binary codes, and determining the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; and determining the recommendation information of the candidate videos based on the hash similarity.

[0096] In addition, when the logical instructions in the foregoing memory 930 can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0097] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video recommendation method provided by each of the above methods. The method includes: extracting video semantic features from the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; hashing each of the enhanced video semantic feature maps into binary codes, and determining the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; and determining the recommendation information of the candidate videos based on the hash similarity.

[0098] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the video recommendation method provided by each of the above methods. The method includes: extracting video semantic features from the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; hashing each of the enhanced video semantic feature maps into binary codes, and determining the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; and determining the recommendation information of the candidate videos based on the hash similarity.

[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A video recommendation method, characterized in that, Including: Performing video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; Performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; Hashing each of the enhanced video semantic feature maps into binary codes, and determining the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; Determining the recommendation information of the candidate videos based on the hash similarity.

2. The video recommendation method according to claim 1, wherein The enhanced video semantic feature maps include an enhanced query video semantic feature map and an enhanced candidate video semantic feature map; The performing semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps includes: Inputting the video semantic feature map into the first branch of the feature multi-scale perception enhancement module to obtain a first-channel video semantic local activation feature vector output by the first branch; Inputting the video semantic feature map into the second branch of the feature multi-scale perception enhancement module to obtain a second-channel compressed video semantic feature map output by the second branch; Inputting the video semantic feature map into the third branch of the feature multi-scale perception enhancement module to obtain a video semantic receptive field expansion global activation feature matrix output by the third branch; Performing feature fusion on the second-channel compressed video semantic feature map with the first-channel video semantic local activation feature vector and the video semantic receptive field expansion global activation feature matrix respectively to obtain a second-channel compressed video semantic local activation feature map and a second-channel compressed video semantic global activation feature map; Performing position-wise addition on the second-channel compressed video semantic local activation feature map and the second-channel compressed video semantic global activation feature map to obtain a second-channel compressed video semantic multi-scale fusion activation feature map; Performing dilated convolution encoding on the second-channel compressed video semantic multi-scale fusion activation feature map to obtain the enhanced video semantic feature map.

3. The video recommendation method according to claim 2, wherein The first branch is used for: Performing point convolution processing on the video semantic feature map to obtain a first-channel video semantic compressed feature map; Performing global average pooling on each feature matrix along the channel dimension in the first-channel video semantic compressed feature map to obtain a first-channel video semantic compressed feature vector; Performing non-linear activation on the first-channel video semantic compressed feature vector to obtain the first-channel video semantic local activation feature vector.

4. The video recommendation method according to claim 2, wherein The second branch is used for: Performing point convolution processing on the video semantic feature map to obtain the second-channel compressed video semantic feature map.

5. The video recommendation method according to claim 2, wherein The third branch is used for: Performing dilated convolution encoding on the video semantic feature map to obtain a video semantic receptive field expansion feature map; Performing point convolution processing on the video semantic receptive field expansion feature map to obtain a video semantic receptive field expansion global feature matrix; Performing non-linear activation on the video semantic receptive field expansion global feature matrix to obtain the video semantic receptive field expansion global activation feature matrix.

6. The video recommendation method according to claim 1, wherein The video semantic feature maps include a query video semantic feature map and a candidate video semantic feature map; Performing video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps, including: Inputting the query video and the candidate videos into a video twin detection network including a first video encoder and a second video encoder to obtain the query video semantic feature map and the candidate video semantic feature maps; Wherein, the first video encoder and the second video encoder have the same network structure.

7. The video recommendation method according to claim 1, wherein Determining the recommendation information of the candidate videos based on the hash similarity, including: When the hash similarity is greater than a preset threshold, recommending the candidate video to the user.

8. A video recommendation device, characterized in that, Including: A video semantic feature extraction module, configured to perform video semantic feature extraction on the user's query video and the candidate videos corresponding to the query video to obtain a plurality of video semantic feature maps; A semantic perception enhancement module, configured to perform semantic perception enhancement on each of the video semantic feature maps to obtain a plurality of enhanced video semantic feature maps; A hash similarity determination module, configured to hash each of the enhanced video semantic feature maps into binary codes, and determine the hash similarity between the enhanced video semantic feature maps based on the Hamming distance of the binary codes; A video recommendation module, configured to determine the recommendation information of the candidate videos based on the hash similarity.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the video recommendation method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the video recommendation method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the video recommendation method according to any one of claims 1 to 7.