Explanation recommendation method

By calculating the similarity between the video stream and the client's screen data, the video stream with the highest similarity is recommended, solving the problem of low commentary recommendation accuracy and achieving an immersive viewing experience for users.

CN120640074APending Publication Date: 2025-09-12MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510814020.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When watching a game, users cannot determine the commentary perspective due to the low accuracy of the commentary recommendations, resulting in a poor viewing experience.

Method used

By calculating the target similarity between the picture data of multiple video streams and the client's picture data, N video streams with the highest similarity are recommended to the client, improving the accuracy of commentary recommendations and enabling users to quickly enter the commentary perspective.

Benefits of technology

By improving the accuracy of commentary recommendations, users can quickly switch to video streams with high similarity, achieving an immersive experience and improving the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640074A_ABST
    Figure CN120640074A_ABST
Patent Text Reader

Abstract

The invention provides a commentary recommendation method, and relates to the technical field of video processing, and the method comprises the steps: receiving first picture data sent by a client, the first picture data being picture data of a target event played by the client in a preset time period; acquiring second picture data of each video stream in a plurality of video streams in the preset time period, wherein the plurality of video streams are explained video streams corresponding to the target event; calculating a target similarity between the second picture data of each video stream and the first picture data; determining N video streams from the plurality of video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the plurality of video streams; and sending the push information corresponding to the N video streams to the client. According to the invention, the accuracy of speaking recommendation is improved, so that the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a commentary recommendation method. Background Art

[0002] Free-view technology is a key technology for improving the user viewing experience in live streaming. It allows users to freely view scene videos from multiple angles. In related technologies, commentary video streams are present during the playback of videos or events such as matches. A commentator explains the development of the video or match, and the client recommends different commentary video streams to users based on the number of viewers. However, because the user's perspective differs from the commentator's, the user cannot determine the commentary's perspective or perceive the content of the commentary. After selecting a recommended commentary video stream, the user loses the immersive experience, resulting in a poor viewing experience.

[0003] It can be seen that the related art has the problem of low accuracy of commentary recommendation, which leads to poor experience for users watching the game. Summary of the Invention

[0004] The embodiment of the present invention provides a commentary recommendation method to solve the problem in the related art that the commentary recommendation accuracy is low, resulting in a poor experience for users watching the game.

[0005] To solve the above problems, the present invention is achieved as follows:

[0006] In a first aspect, an embodiment of the present invention provides a commentary recommendation method, applied to a server, comprising:

[0007] Receiving first picture data sent by a client, where the first picture data is picture data of a target event played by the client within a preset time period;

[0008] Obtaining second frame data of each video stream in the preset time period from a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event;

[0009] Calculating a target similarity between the second picture data and the first picture data of each video stream;

[0010] Determining N video streams from the multiple video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams;

[0011] Send push information corresponding to the N video streams to the client.

[0012] In a second aspect, an embodiment of the present invention provides a commentary recommendation method, applied to a client, comprising:

[0013] Collecting first frame data of a target event played by the client within a preset time period;

[0014] Sending the first screen data to the server;

[0015] Receive push information of N video streams, where target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0016] Display the push information corresponding to the N video streams.

[0017] In a third aspect, an embodiment of the present invention further provides a commentary recommendation device, comprising:

[0018] A receiving module, configured to receive first picture data sent by a client, where the first picture data is picture data of a target event played by the client within a preset time period;

[0019] An acquisition module, configured to acquire second frame data of each video stream in the preset time period from among a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event;

[0020] a calculation module, configured to calculate a target similarity between the second picture data and the first picture data of each video stream;

[0021] a determining module, configured to determine N video streams from the multiple video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams;

[0022] The sending module is used to send push information corresponding to the N video streams to the client.

[0023] In a fourth aspect, an embodiment of the present invention further provides a commentary recommendation device, comprising:

[0024] An acquisition module, configured to acquire first frame data of a target event played by a client within a preset time period;

[0025] A sending module, configured to send the first screen data to a server;

[0026] a first receiving module, configured to receive push information of N video streams, wherein target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0027] The display module is used to display the push information corresponding to the N video streams.

[0028] In a fifth aspect, an embodiment of the present invention further provides an electronic device, comprising a transceiver and a processor,

[0029] The transceiver is configured to receive first image data sent by a client, wherein the first image data is image data of a target event played by the client within a preset time period;

[0030] The processor is configured to obtain second frame data of each video stream in a preset time period from a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event;

[0031] The processor is further configured to calculate a target similarity between the second picture data and the first picture data of each video stream;

[0032] The processor is further configured to determine N video streams from the multiple video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams;

[0033] The transceiver is further configured to send push information corresponding to the N video streams to the client.

[0034] In a sixth aspect, an embodiment of the present invention further provides an electronic device, including a transceiver and a processor,

[0035] The processor is configured to collect first frame data of a target event played by a client within a preset time period;

[0036] The transceiver is used to send the first picture data to the server;

[0037] The transceiver is further configured to receive push information of N video streams, wherein target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0038] The processor is further configured to display push information corresponding to the N video streams.

[0039] In the seventh aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, and a program stored on the memory and runnable on the processor, wherein the program, when executed by the processor, implements the steps of the explanation recommendation method described in the first or second aspect above.

[0040] In an eighth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the commentary recommendation method described in the first or second aspect are implemented.

[0041] In a ninth aspect, the present invention further provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the interpretation recommendation method as described in the first or second aspect above.

[0042] In an embodiment of the present invention, a first screen data sent by a client is received, wherein the first screen data is screen data of a target event played by the client within a preset time period; second screen data of each video stream in the preset time period is obtained, wherein the multiple video streams are video streams with commentary corresponding to the target event; target similarity between the second screen data of each video stream and the first screen data is calculated; based on the target similarity, N video streams are determined from the multiple video streams, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams; and push information corresponding to the N video streams is sent to the client. In this way, by calculating the target similarity between the second screen data of each video stream in the multiple video streams with commentary and the first screen data, and then determining N video streams, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams, the recommendation accuracy is improved, so that after the client selects a video stream from the N video streams, the client switches from the screen corresponding to the first screen data to the screen corresponding to the second screen data of the N video streams, which can quickly bring the user into the commentary perspective, achieve an immersive experience, and thus improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0044] Figure 1This is a flow chart of a commentary recommendation method applied to a server provided by an embodiment of the present invention;

[0045] Figure 2 This is a flowchart of a method for recommending explanations applied to a client provided by an embodiment of the present invention;

[0046] Figure 3 This is a flow chart of recommending a video stream with commentary provided by an embodiment of the present invention;

[0047] Figure 4 This is a schematic diagram of switching the display interface of the client provided by an embodiment of the present invention;

[0048] Figure 5 is a structural diagram of an explanation recommendation device provided by an embodiment of the present invention;

[0049] Figure 6 is a structural diagram of an explanation recommendation device provided by an embodiment of the present invention;

[0050] Figure 7 is a structural diagram of an electronic device provided by an embodiment of the present invention;

[0051] Figure 8 This is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0053] See Figure 1 , Figure 1 This is a flow chart of a method for recommending explanations applied to a server provided by an embodiment of the present invention. Figure 1 As shown, the following steps are included:

[0054] Step 101: Receive first picture data sent by a client, where the first picture data is picture data of a target event played by the client within a preset time period.

[0055] The first screen data is collected when the client is in a free viewing angle. If the client is watching a video stream with commentary, it is not necessary to collect the client's screen data.

[0056] The target event is an event played by the client. The target event can be a video or live broadcast, such as a game video or a live basketball game. When user-recommended commentary is required, the screen data of the target event played by the client within a preset time period is collected to obtain the first screen data. The preset time period can be set as needed, for example, to collect screen data within 3 seconds before the time of collection.

[0057] Step 102: Acquire second frame data of each video stream in the preset time period from a plurality of video streams, wherein the plurality of video streams are video streams corresponding to the target event and having explanations.

[0058] The above-mentioned multiple video streams are video streams with commentary. To obtain multiple video streams, a commentator or host may comment on the target event, and then upload the video stream including the commentary content to a server and store it. When commentary recommendation is needed, the video stream including the commentary is obtained from the storage location.

[0059] In some embodiments, the commentator or anchor pushes the collected audio and video data (i.e., commentary content) to the live streaming server through streaming software or live streaming software development kit (SDK). The server is responsible for receiving the video stream pushed by the anchor and storing or temporarily caching the video stream. The streaming protocol can be selected from the Real-Time Messaging Protocol (RTMP), Web Real-Time Communication (WebRTC), and a private protocol based on the User Data Protocol (UDP). When commentary recommendation is required, the server obtains each video stream from the storage location and extracts the second picture data corresponding to each video stream based on a preset time period.

[0060] Furthermore, the video images in the first image data and the second image data are synthesized by collecting multiple streams of images from multiple cameras and performing three-dimensional modeling, and the video images can be displayed through two-dimensional images.

[0061] The second picture data is picture data of a video stream with commentary. The video stream includes commentary data and picture data. When commentary recommendation is required, the second picture data of a preset time period is first extracted from the video stream, and then the video stream to be recommended is determined.

[0062] Step 103: Calculate the target similarity between the second picture data and the first picture data of each video stream.

[0063] The target similarity can reflect the degree of similarity between the first frame data played by the client and the second frame data of each video stream. It should be noted that the higher the degree of similarity between the first frame data and the second frame data, the faster the user can enter the commentary perspective after the client switches video streams, achieving an immersive experience.

[0064] Among them, the target similarity can be the color features, motion features, shape features and / or texture features between the first picture data and the second picture data, and the degree of similarity between the first picture data and the second picture data can be determined through the color features, motion features, shape features and / or texture features.

[0065] Step 104: Determine N video streams from the multiple video streams based on the target similarity, where the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams.

[0066] The target similarity of the above N video streams is greater than the target similarity of other video streams. Through the target similarity, N video streams are screened from multiple video streams, making the N video streams the N video streams most similar to the first picture data of the target event played by the client, which can better improve the user experience after switching the video stream.

[0067] Step 105: Send push information corresponding to the N video streams to the client.

[0068] In an embodiment of the present invention, a first screen data sent by a client is received, wherein the first screen data is screen data of a target event played by the client within a preset time period; second screen data of each video stream in the preset time period is obtained, wherein the multiple video streams are video streams with commentary corresponding to the target event; target similarity between the second screen data of each video stream and the first screen data is calculated; based on the target similarity, N video streams are determined from the multiple video streams, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams; and push information corresponding to the N video streams is sent to the client. In this way, by calculating the target similarity between the second screen data of each video stream in the multiple video streams with commentary and the first screen data, and then determining N video streams, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams, the recommendation accuracy is improved, so that after the client selects a video stream from the N video streams, the client switches from the screen corresponding to the first screen data to the screen corresponding to the second screen data of the N video streams, which can quickly bring the user into the commentary perspective, achieve an immersive experience, and thus improve the user experience.

[0069] In one embodiment, calculating the target similarity between the second picture data and the first picture data of each video stream includes:

[0070] Calculating at least one of color feature similarity, motion feature similarity, shape feature similarity, and texture feature similarity between second picture data of a target video stream and the first picture data, wherein the target video stream is one of the multiple video streams;

[0071] The target similarity between the second picture data of the target video stream and the first picture data is obtained based on weighting of at least one of the color feature similarity, the motion feature similarity, the shape feature similarity, and the texture feature similarity.

[0072] In an embodiment of the present invention, at least one of color feature similarity, motion feature similarity, shape feature similarity, and texture feature similarity is calculated between the second image data of a target video stream and the first image data, wherein the target video stream is one of the multiple video streams; and the target similarity between the second image data of the target video stream and the first image data is obtained based on a weighted combination of at least one of the color feature similarity, the motion feature similarity, the shape feature similarity, and the texture feature similarity. In this way, the target similarity is determined based on at least one of the color feature similarity, the motion feature similarity, the shape feature similarity, and the texture feature similarity, thereby improving the accuracy of the target similarity.

[0073] In one embodiment, the color feature similarity is calculated as follows:

[0074] generating a first color histogram at each moment within the preset time period based on the video frame of the first picture data;

[0075] generating a second color histogram at each moment based on a video frame of the second picture data of the target video stream;

[0076] Calculating the correlation coefficient corresponding to each moment based on the first color histogram and the second color histogram;

[0077] Based on the correlation coefficients corresponding to all moments in the preset time period, the color feature similarity between the second picture data of the target video stream and the first picture data is calculated.

[0078] The first color histogram and the second color histogram can be obtained by converting the video frame from the red, green, blue (RGB) color space to the hue, saturation, and brightness (HSV) space. The first color histogram and the second color histogram represent the distribution of each color in the video frame. The color distribution of the video frame can be determined using the first color histogram and the second color histogram, thereby determining the color feature similarity.

[0079] The correlation coefficient may be a Pearson correlation coefficient. The closer the value of the Pearson correlation coefficient is to 1, the more similar the two histograms are. The closer the value of the Pearson correlation coefficient is to -1, the less similar the two histograms are. In an embodiment of the present invention, the color feature similarity is obtained by calculating the Pearson correlation coefficient corresponding to each moment and then calculating the Pearson correlation coefficient for all moments.

[0080] In some embodiments, calculating the color feature similarity between the second frame data and the first frame data of the target video stream based on the correlation coefficients corresponding to all moments within the preset time period may include calculating the variance of the correlation coefficients corresponding to all moments and then converting the variance into color feature similarity. A smaller variance indicates a higher similarity.

[0081] In some embodiments, based on the correlation coefficients corresponding to all moments within the preset time period, the color feature similarity between the second picture data of the target video stream and the first picture data is calculated. The average value of the correlation coefficients corresponding to all moments can be calculated through a chi-square test algorithm, and then the average value is converted into color feature similarity.

[0082] In an embodiment of the present invention, a first color histogram is generated at each moment within a preset time period based on a video frame of the first picture data; a second color histogram is generated at each moment based on a video frame of the second picture data of the target video stream; a correlation coefficient corresponding to each moment is calculated based on the first color histogram and the second color histogram; and the color feature similarity between the second picture data and the first picture data of the target video stream is calculated based on the correlation coefficients corresponding to all moments within the preset time period.

[0083] In one embodiment, the motion feature similarity is calculated as follows:

[0084] Calculating an entropy value corresponding to each of a plurality of first video frames included in the first picture data, and an entropy value corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0085] Extracting scale-invariant feature transform (SIFT) feature points corresponding to each of the first video frames and SIFT feature points corresponding to each of the second video frames;

[0086] Determine a first key frame based on the entropy value corresponding to each first video frame and the SIFT feature points corresponding to each first video frame, where the first key frame is one of the multiple first video frames;

[0087] Determining a second key frame based on the entropy value corresponding to each second video frame and the SIFT feature points corresponding to each second video frame, where the second key frame is one of the plurality of second video frames;

[0088] The motion feature similarity between the second picture data and the first picture data is calculated based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame.

[0089] The entropy value corresponding to the first video frame is used to represent the degree of detail change in the first video frame, and the entropy value corresponding to the second video frame is used to represent the degree of detail change in the second video frame. A higher entropy value indicates a greater degree of detail change in the image, and a lower entropy value indicates a lesser degree of detail change in the image. It should be noted that the more obvious the motion features in the video frame, the higher the entropy value. Determining key frames based on entropy values ​​can improve the accuracy of key frames.

[0090] In some embodiments, multiple first intermediate frames can be determined from multiple first video frames using the entropy value corresponding to the first video frame, so that the first key frame can be subsequently screened from the multiple first intermediate frames; and multiple second intermediate frames can be determined from multiple second video frames using the entropy value corresponding to the second video frame, so that the second key frame can be subsequently screened from the multiple second intermediate frames.

[0091] The above-mentioned SIFT feature points are used to detect and describe local features in the picture of the video frame. The movement of the video frame can be determined by the distribution of SIFT feature points. The key frame is determined by the SIFT feature points corresponding to the first video frame and the SIFT feature points corresponding to the second video frame, which can improve the accuracy of the key frame.

[0092] In some embodiments, after obtaining multiple first intermediate frames using entropy values ​​corresponding to the first video frame, the first key frame can be determined from the multiple first intermediate frames using entropy values ​​corresponding to the first intermediate frames and SIFT feature points corresponding to the first intermediate frames.

[0093] In some embodiments, after obtaining multiple second intermediate frames using entropy values ​​corresponding to the second video frame, a second key frame can be determined from the multiple second intermediate frames using entropy values ​​corresponding to the second intermediate frames and SIFT feature points corresponding to the second intermediate frames.

[0094] The calculation of the motion feature similarity between the second picture data and the first picture data based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame can be performed by using a Fast Library for Approximate Nearest Neighbors (FLANN) or a Brute-Force matcher to find the correspondence between the SIFT feature points of the first key frame and the SIFT feature points of the second key frame, using an optical flow method to estimate pixel motion between consecutive frames, and then using a cosine similarity method to calculate the motion feature similarity between the SIFT feature points of the first key frame and the SIFT feature points of the second key frame. The closer the value of the motion feature similarity is to 1, the higher the similarity.

[0095] In an embodiment of the present invention, an entropy value corresponding to each of a plurality of first video frames included in the first picture data and an entropy value corresponding to each of a plurality of second video frames included in the second picture data of the target video stream are calculated; scale-invariant feature transform (SIFT) feature points corresponding to each first video frame and SIFT feature points corresponding to each second video frame are extracted; a first key frame and a second key frame are determined by using the entropy value and the SIFT feature points, and the motion feature similarity between the second picture data and the first picture data is calculated based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame.

[0096] In one embodiment, the shape feature similarity is calculated as follows:

[0097] Extracting first edge information corresponding to each of the plurality of first video frames included in the first picture data, and second edge information corresponding to each of the plurality of second video frames included in the second picture data of the target video stream;

[0098] Calculating a first key point corresponding to each first video frame based on the first edge information, and calculating a second key point corresponding to each second video frame based on the second edge information;

[0099] Calculating a first shape context descriptor value for each of the first video frames based on the first key point, and calculating a second shape context descriptor value for each of the second video frames based on the second key point;

[0100] Setting the first video frame having the largest value of the first shape context descriptor among the plurality of first video frames as a third key frame, and setting the second video frame having the largest value of the second shape context descriptor among the plurality of second video frames as a fourth key frame;

[0101] The shape feature similarity between the second picture data and the first picture data is calculated based on the first shape context descriptor value of the third key frame and the second shape context descriptor value of the fourth key frame.

[0102] The first edge information and the second edge information can determine the edge of the shape. The edge of the video frame can be determined using the first edge information and the second edge information, thereby extracting the shape features. The first edge information can be extracted from each first video frame using an edge detection algorithm (Canny), and the second edge information can be extracted from each second video frame using an edge detection algorithm.

[0103] The first key point is a point where the shape of the first video frame changes significantly, and the shape characteristics of the first video frame can be determined by the first key point. The second key point is a point where the shape of the second video frame changes significantly, and the shape characteristics of the second video frame can be determined by the second key point. The first key point and the second key point can be obtained by corner detection using the Harris key point detection algorithm.

[0104] The above-mentioned first shape context descriptor value is used to describe the shape elements in the picture of the first video frame, and the above-mentioned second shape context descriptor value is used to describe the shape elements in the picture of the second video frame. It should be noted that the larger the values ​​of the first shape context descriptor value and the second shape context descriptor value are, the more significant the shape features of the first video frame and the second video frame are. By setting the first video frame with the largest first shape context descriptor value among the multiple first video frames as the third key frame, and setting the second video frame with the largest second shape context descriptor value among the multiple second video frames as the fourth key frame, visually similar redundant video frames are removed, so that the third key frame is the video frame with the most obvious shape features among the multiple first video frames, and the fourth key frame is the video frame with the most obvious shape features among the multiple second video frames.

[0105] In some embodiments, scene changes can be first detected by an inter-frame difference method, multiple third intermediate video frames can be selected from multiple first video frames, multiple fourth intermediate video frames can be selected from multiple second video frames, and then a third key frame can be determined from the multiple third intermediate video frames, and a fourth key frame can be determined from the multiple fourth intermediate video frames.

[0106] In an embodiment of the present invention, first edge information corresponding to each first video frame in a plurality of first video frames included in the first picture data and second edge information corresponding to each second video frame in a plurality of second video frames included in the second picture data of the target video stream are extracted; a first key point corresponding to each first video frame is calculated based on the first edge information, and a second key point corresponding to each second video frame is calculated based on the second edge information; a first shape context descriptor value of each first video frame is calculated based on the first key point, and a second shape context descriptor value of each second video frame is calculated based on the second key point; the first video frame having the largest first shape context descriptor value among the plurality of first video frames is set as a third key frame, and the second video frame having the largest second shape context descriptor value among the plurality of second video frames is set as a fourth key frame; based on the first shape context descriptor value of the third key frame and the second shape context descriptor value of the fourth key frame, the shape feature similarity between the second picture data and the first picture data is calculated.

[0107] In one embodiment, the texture feature similarity is calculated as follows:

[0108] Extracting a first texture feature corresponding to each of a plurality of first video frames included in the first picture data, and a second texture feature corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0109] Clustering the first texture features corresponding to each first video frame to obtain a first intermediate texture feature;

[0110] Clustering the second texture features corresponding to each second video frame to obtain second intermediate texture features;

[0111] The texture feature similarity between the second picture data and the first picture data is calculated based on the first intermediate texture feature and the second intermediate texture feature.

[0112] The first texture feature and the second texture feature can be extracted by algorithms such as Local Binary Patterns (LBP).

[0113] In an embodiment of the present invention, a first texture feature corresponding to each of a plurality of first video frames included in the first picture data and a second texture feature corresponding to each of a plurality of second video frames included in the second picture data of the target video stream are extracted; the first texture feature corresponding to each of the first video frames is clustered to obtain a first intermediate texture feature; the second texture feature corresponding to each of the second video frames is clustered to obtain a second intermediate texture feature; and the texture feature similarity between the second picture data and the first picture data is calculated using the first intermediate texture feature and the second intermediate texture feature.

[0114] In some embodiments, in addition to calculating the texture feature similarity between the second picture data and the first picture data based on the first intermediate texture feature and the second intermediate texture feature, the texture feature similarity may also be calculated in the following manner.

[0115] Specifically, after obtaining the first intermediate texture feature and the second intermediate texture feature, calculating a first difference between the first texture feature and the first intermediate texture feature corresponding to each first video frame, and setting the first video frame corresponding to the smallest first difference as the fifth key frame;

[0116] calculating a second difference between the second texture feature and the second intermediate texture feature corresponding to each second video frame, and setting the second video frame corresponding to the smallest second difference as a sixth key frame;

[0117] The texture feature similarity between the second picture data and the first picture data is calculated based on the first texture feature corresponding to the fifth key frame and the second key frame corresponding to the sixth key frame.

[0118] In one embodiment, before sending the push information corresponding to the N video streams to the client, the method further includes:

[0119] The N video streams are distributed to a content delivery network (CDN), and edge nodes of the CDN are used to cache the N video streams.

[0120] In an embodiment of the present invention, the N video streams are distributed to a content delivery network (CDN), so that the client can obtain the required video streams from the CDN, thereby effectively improving the transmission efficiency of the video streams.

[0121] See Figure 2 , Figure 2 This is a flow chart of a method for recommending explanations applied to a client provided by an embodiment of the present invention. Figure 2 As shown, the following steps are included:

[0122] Step 201: collecting first frame data of a target event played by the client within a preset time period;

[0123] Step 202: Send the first screen data to the server;

[0124] Step 203: Receive push information of N video streams, where target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being the target similarity between the second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0125] Step 204: Display push information corresponding to the N video streams.

[0126] In an embodiment of the present invention, the first screen data of the target event played by the client within a preset time period is collected; the first screen data is sent to the server; push information of N video streams is received, the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams, the target similarity is the target similarity between the second screen data of each video stream in the multiple video streams and the first screen data, the multiple video streams include the N video streams and the other video streams, N is a positive integer greater than 1; and the push information corresponding to the N video streams is displayed. In this way, by collecting the first screen data and sending the first screen data to the server, the server can determine N video streams based on the first screen data, and the client displays the N video streams to recommend video streams with explanations to the user.

[0127] The collecting of the first screen data of the target event played by the client within a preset time period includes:

[0128] The first picture data is obtained by collecting the pictures of the target event played by the client within a preset time period based on a preset cycle through a preset sliding window.

[0129] In an embodiment of the present invention, a preset sliding window is used to capture images of the target event played by the client within a preset time period based on a preset period to obtain the first image data, thereby achieving image capture of the target time played by the client. The preset sliding window uses each video frame as a unit, and the different video frames collected within the preset time period based on a preset period are used to obtain the video frames included in the first image data.

[0130] In one embodiment, after displaying the push information corresponding to the N video streams, the method further includes:

[0131] receiving a target operation, wherein the target operation is selecting a target video stream based on the push information, and the target video stream is a video stream among the N video streams;

[0132] Obtain the target video stream from the edge node of the content delivery network CDN;

[0133] Play the target video stream.

[0134] In this embodiment of the present invention, a target operation is received, wherein the target operation is to select a target video stream from the N video streams based on the push information; the target video stream is obtained from an edge node of a content delivery network (CDN); and the target video stream is played. In this way, the user selects the target video stream from the N video streams displayed on the client through the target operation, and then obtains the target video stream from the CDN edge node, thereby achieving the goal of recommending and playing the video stream with commentary to the user.

[0135] The overall process of recommending video streams with commentary is as follows: Figure 3 and Figure 4 As shown, the perspective in the client is first determined to be the free perspective A, and the first picture data of the client is collected; N video streams are obtained by matching based on the first picture data through the server, and the N video streams are displayed through the client, for example Figure 4 The commentary perspective video stream A1, commentary perspective video stream A2, commentary perspective video stream A3 and commentary perspective video stream A4 are included in the video stream; the user selects a target video stream to play from N video streams according to demand, for example, selects commentary perspective video stream A1 to play.

[0136] Furthermore, after the client plays the target video stream, if the user no longer needs to watch the target video stream or needs to watch other video streams, the user can choose to exit the target video stream by clicking the "Exit" button in the interface. At this time, the client switches from the target video stream to the free view angle A. After the client switches to the free view angle, the video stream with commentary can continue to be recommended based on the client's screen, for example Figure 4 The commentary perspective video stream B1, commentary perspective video stream B2, commentary perspective video stream B3 and commentary perspective video stream B4.

[0137] See Figure 5 , Figure 5 This is a structural diagram of an explanation recommendation device provided by an embodiment of the present invention. Figure 5 As shown, the explanation recommendation device 500 includes:

[0138] The receiving module 501 is configured to receive first image data sent by a client, where the first image data is image data of a target event played by the client within a preset time period;

[0139] An acquisition module 502 is configured to acquire second frame data of each video stream in a preset time period from a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event;

[0140] A calculation module 503 is configured to calculate a target similarity between the second picture data and the first picture data of each video stream;

[0141] A determination module 504 is configured to determine N video streams from the multiple video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams;

[0142] The sending module 505 is configured to send push information corresponding to the N video streams to the client.

[0143] In one embodiment, the calculation module 503 includes:

[0144] a first calculating unit, configured to calculate at least one of color feature similarity, motion feature similarity, shape feature similarity, and texture feature similarity between second picture data of a target video stream and the first picture data, wherein the target video stream is one of the multiple video streams;

[0145] The second calculation unit is used to obtain the target similarity between the second picture data of the target video stream and the first picture data based on the weighting of at least one of the color feature similarity, the motion feature similarity, the shape feature similarity and the texture feature similarity.

[0146] In one embodiment, the color feature similarity is calculated as follows:

[0147] generating a first color histogram at each moment within the preset time period based on the video frame of the first picture data;

[0148] generating a second color histogram at each moment based on a video frame of the second picture data of the target video stream;

[0149] Calculating the correlation coefficient corresponding to each moment based on the first color histogram and the second color histogram;

[0150] Based on the correlation coefficients corresponding to all moments in the preset time period, the color feature similarity between the second picture data of the target video stream and the first picture data is calculated.

[0151] In one embodiment, the motion feature similarity is calculated as follows:

[0152] Calculating an entropy value corresponding to each of a plurality of first video frames included in the first picture data, and an entropy value corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0153] Extracting scale-invariant feature transform (SIFT) feature points corresponding to each first video frame and SIFT feature points corresponding to each second video frame;

[0154] Determine a first key frame based on the entropy value corresponding to each first video frame and the SIFT feature points corresponding to each first video frame, where the first key frame is one of the multiple first video frames;

[0155] Determining a second key frame based on the entropy value corresponding to each second video frame and the SIFT feature points corresponding to each second video frame, where the second key frame is one of the plurality of second video frames;

[0156] The motion feature similarity between the second picture data and the first picture data is calculated based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame.

[0157] In one embodiment, the shape feature similarity is calculated as follows:

[0158] Extracting first edge information corresponding to each of the plurality of first video frames included in the first picture data, and second edge information corresponding to each of the plurality of second video frames included in the second picture data of the target video stream;

[0159] Calculating a first key point corresponding to each first video frame based on the first edge information, and calculating a second key point corresponding to each second video frame based on the second edge information;

[0160] Calculating a first shape context descriptor value for each of the first video frames based on the first key point, and calculating a second shape context descriptor value for each of the second video frames based on the second key point;

[0161] Setting the first video frame having the largest value of the first shape context descriptor among the plurality of first video frames as a third key frame, and setting the second video frame having the largest value of the second shape context descriptor among the plurality of second video frames as a fourth key frame;

[0162] The shape feature similarity between the second picture data and the first picture data is calculated based on the first shape context descriptor value of the third key frame and the second shape context descriptor value of the fourth key frame.

[0163] In one embodiment, the texture feature similarity is calculated as follows:

[0164] Extracting a first texture feature corresponding to each of a plurality of first video frames included in the first picture data, and a second texture feature corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0165] Clustering the first texture features corresponding to each first video frame to obtain a first intermediate texture feature;

[0166] Clustering the second texture features corresponding to each second video frame to obtain second intermediate texture features;

[0167] The texture feature similarity between the second picture data and the first picture data is calculated based on the first intermediate texture feature and the second intermediate texture feature.

[0168] In one embodiment, the commentary recommendation device 500 further includes:

[0169] The distribution module is used to distribute the N video streams to a content distribution network (CDN), and the edge nodes of the CDN are used to cache the N video streams.

[0170] The explanation recommendation device provided in the embodiment of the present invention is capable of implementing each process of each embodiment of the above-mentioned explanation recommendation method applied to the server. The technical features correspond one to one and can achieve the same technical effect. To avoid repetition, they will not be described here.

[0171] It should be noted that the commentary recommendation device in the embodiment of the present invention may be a device, or a component, integrated circuit, or chip in an electronic device.

[0172] See Figure 6 , Figure 6 is a structural diagram of an explanation recommendation device provided by an embodiment of the present invention, such as Figure 6 As shown, the explanation recommendation device 600 includes:

[0173] The acquisition module 601 is used to acquire the first frame data of the target event played by the client within a preset time period;

[0174] A sending module 602 is configured to send the first screen data to a server;

[0175] A first receiving module 603 is configured to receive push information of N video streams, wherein target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0176] The display module 604 is used to display the push information corresponding to the N video streams.

[0177] In one embodiment, the acquisition module 601 includes:

[0178] The collecting unit is configured to collect the images of the target event played by the client within a preset time period based on a preset cycle through a preset sliding window to obtain the first image data.

[0179] In one embodiment, the commentary recommendation device 600 further includes:

[0180] A second receiving module is configured to receive a target operation, wherein the target operation is selecting a target video stream based on the push information, and the target video stream is a video stream among the N video streams;

[0181] An acquisition module is used to obtain a target video stream from an edge node of a content delivery network (CDN);

[0182] A playing module is used to play the target video stream.

[0183] The commentary recommendation device provided in the embodiment of the present invention is capable of implementing the various processes of the various embodiments of the above-mentioned commentary recommendation method applied to the client. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0184] It should be noted that the commentary recommendation device in the embodiment of the present invention may be a device, or a component, integrated circuit, or chip in an electronic device.

[0185] An embodiment of the present invention also provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the various processes of the above-mentioned explanation recommendation method embodiment applied to the server are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0186] For details, see Figure 7 As shown, an embodiment of the present invention further provides an electronic device, including a bus 701 , a transceiver 702 , an antenna 703 , a bus interface 704 , a processor 705 and a memory 707 .

[0187] The transceiver 702 is configured to receive first image data sent by a client, where the first image data is image data of a target event played by the client within a preset time period;

[0188] The processor 705 is configured to obtain second frame data of each video stream in a preset time period from among a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event;

[0189] The processor 705 is further configured to calculate a target similarity between the second picture data and the first picture data of each video stream;

[0190] The processor 705 is further configured to determine N video streams from the multiple video streams based on the target similarity, where the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams;

[0191] The transceiver 702 is further configured to send push information corresponding to the N video streams to the client.

[0192] In one embodiment, calculating the target similarity between the second picture data and the first picture data of each video stream includes:

[0193] Calculating at least one of color feature similarity, motion feature similarity, shape feature similarity, and texture feature similarity between second picture data of a target video stream and the first picture data, wherein the target video stream is one of the multiple video streams;

[0194] The target similarity between the second picture data of the target video stream and the first picture data is obtained based on weighting of at least one of the color feature similarity, the motion feature similarity, the shape feature similarity, and the texture feature similarity.

[0195] In one embodiment, the color feature similarity is calculated as follows:

[0196] generating a first color histogram at each moment within the preset time period based on the video frame of the first picture data;

[0197] generating a second color histogram at each moment based on a video frame of the second picture data of the target video stream;

[0198] Calculating the correlation coefficient corresponding to each moment based on the first color histogram and the second color histogram;

[0199] Based on the correlation coefficients corresponding to all moments in the preset time period, the color feature similarity between the second picture data of the target video stream and the first picture data is calculated.

[0200] In one embodiment, the motion feature similarity is calculated as follows:

[0201] Calculating an entropy value corresponding to each of a plurality of first video frames included in the first picture data, and an entropy value corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0202] Extracting scale-invariant feature transform (SIFT) feature points corresponding to each first video frame and SIFT feature points corresponding to each second video frame;

[0203] Determine a first key frame based on the entropy value corresponding to each first video frame and the SIFT feature points corresponding to each first video frame, where the first key frame is one of the multiple first video frames;

[0204] Determining a second key frame based on the entropy value corresponding to each second video frame and the SIFT feature points corresponding to each second video frame, where the second key frame is one of the plurality of second video frames;

[0205] The motion feature similarity between the second picture data and the first picture data is calculated based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame.

[0206] In one embodiment, the shape feature similarity is calculated as follows:

[0207] Extracting first edge information corresponding to each of the plurality of first video frames included in the first picture data, and second edge information corresponding to each of the plurality of second video frames included in the second picture data of the target video stream;

[0208] Calculating a first key point corresponding to each first video frame based on the first edge information, and calculating a second key point corresponding to each second video frame based on the second edge information;

[0209] Calculating a first shape context descriptor value for each of the first video frames based on the first key point, and calculating a second shape context descriptor value for each of the second video frames based on the second key point;

[0210] Setting the first video frame having the largest value of the first shape context descriptor among the plurality of first video frames as a third key frame, and setting the second video frame having the largest value of the second shape context descriptor among the plurality of second video frames as a fourth key frame;

[0211] The shape feature similarity between the second picture data and the first picture data is calculated based on the first shape context descriptor value of the third key frame and the second shape context descriptor value of the fourth key frame.

[0212] In one embodiment, the texture feature similarity is calculated as follows:

[0213] Extracting a first texture feature corresponding to each of a plurality of first video frames included in the first picture data, and a second texture feature corresponding to each of a plurality of second video frames included in the second picture data of the target video stream;

[0214] Clustering the first texture features corresponding to each first video frame to obtain a first intermediate texture feature;

[0215] Clustering the second texture features corresponding to each second video frame to obtain second intermediate texture features;

[0216] The texture feature similarity between the second picture data and the first picture data is calculated based on the first intermediate texture feature and the second intermediate texture feature.

[0217] In one embodiment, the transceiver 702 is further configured to distribute the N video streams to a content delivery network (CDN), and edge nodes of the CDN are configured to cache the N video streams.

[0218] exist Figure 7 In the embodiment, the bus architecture (represented by bus 701) is shown. Bus 701 may include any number of interconnected buses and bridges. Bus 701 links various circuits including one or more processors represented by processor 705 and memory represented by memory 707. Bus 701 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and are therefore not described further herein. Bus interface 704 provides an interface between bus 701 and transceiver 702. Transceiver 702 may be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. Data processed by processor 705 is transmitted on a wireless medium via antenna 703. Furthermore, antenna 703 receives data and transmits the data to processor 705.

[0219] The processor 705 is responsible for managing the bus 701 and general processing, and may also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 707 may be used to store data used by the processor 705 when performing operations.

[0220] Optionally, the processor 705 may be a CPU, an ASIC, an FPGA, or a CPLD.

[0221] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the aforementioned embodiment of the explanation recommendation method applied to the server and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0222] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The corresponding processes of the explanation recommendation method embodiment applied to the server can achieve the same technical effect, so they will not be repeated here to avoid repetition.

[0223] An embodiment of the present invention also provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the various processes of the above-mentioned explanation recommendation method embodiment applied to the client are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0224] For details, see Figure 8 As shown, an embodiment of the present invention further provides an electronic device, including a bus 801 , a transceiver 802 , an antenna 803 , a bus interface 804 , a processor 805 and a memory 808 .

[0225] The processor 805 is configured to collect first frame data of a target event played by a client within a preset time period;

[0226] The transceiver 802 is configured to send the first image data to the server;

[0227] The transceiver 802 is further configured to receive push information of N video streams, wherein target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1;

[0228] The processor 805 is further configured to display push information corresponding to the N video streams.

[0229] In one embodiment, collecting the first screen data of the target event played by the client within a preset time period includes:

[0230] The first picture data is obtained by collecting the pictures of the target event played by the client within a preset time period based on a preset cycle through a preset sliding window.

[0231] In one embodiment, the transceiver 802 is further configured to receive a target operation, wherein the target operation is selecting a target video stream based on the push information, and the target video stream is a video stream among the N video streams;

[0232] The transceiver 802 is further configured to obtain a target video stream from an edge node of a content delivery network CDN;

[0233] The processor 805 is further configured to play the target video stream.

[0234] exist Figure 8 In the embodiment, the bus architecture (represented by bus 801) is shown. Bus 801 may include any number of interconnected buses and bridges. Bus 801 links various circuits including one or more processors represented by processor 805 and memory represented by memory 808. Bus 801 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and are therefore not described further herein. Bus interface 804 provides an interface between bus 801 and transceiver 802. Transceiver 802 may be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. Data processed by processor 805 is transmitted on a wireless medium via antenna 803. Furthermore, antenna 803 receives data and transmits the data to processor 805.

[0235] The processor 805 is responsible for managing the bus 801 and general processing, and may also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 808 may be used to store data used by the processor 805 when performing operations.

[0236] Optionally, the processor 805 may be a CPU, an ASIC, an FPGA, or a CPLD.

[0237] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the aforementioned embodiment of the client-side commentary recommendation method and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0238] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 2 The corresponding processes of the explanation recommendation method embodiment applied to the server can achieve the same technical effect, so they will not be repeated here to avoid repetition.

[0239] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0240] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0241] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A commentary recommendation method, applied to a server, characterized in that: include: Receiving first picture data sent by a client, where the first picture data is picture data of a target event played by the client within a preset time period; Obtaining second frame data of each video stream in the preset time period from a plurality of video streams, wherein the plurality of video streams are video streams with commentary corresponding to the target event; Calculating a target similarity between the second picture data and the first picture data of each video stream; Determining N video streams from the multiple video streams based on the target similarity, wherein the target similarity of the N video streams is greater than the target similarity of other video streams in the multiple video streams; Send push information corresponding to the N video streams to the client.

2. The method according to claim 1, wherein The calculating the target similarity between the second picture data and the first picture data of each video stream includes: Calculating at least one of color feature similarity, motion feature similarity, shape feature similarity, and texture feature similarity between second picture data of a target video stream and the first picture data, wherein the target video stream is one of the multiple video streams; The target similarity between the second picture data of the target video stream and the first picture data is obtained based on weighting of at least one of the color feature similarity, the motion feature similarity, the shape feature similarity, and the texture feature similarity.

3. The method according to claim 2, wherein The color feature similarity is calculated as follows: generating a first color histogram at each moment within the preset time period based on the video frame of the first picture data; generating a second color histogram at each moment based on a video frame of the second picture data of the target video stream; Calculating the correlation coefficient corresponding to each moment based on the first color histogram and the second color histogram; Based on the correlation coefficients corresponding to all moments in the preset time period, the color feature similarity between the second picture data of the target video stream and the first picture data is calculated.

4. The method according to claim 2, wherein The motion feature similarity is calculated as follows Calculating an entropy value corresponding to each of a plurality of first video frames included in the first picture data, and an entropy value corresponding to each of a plurality of second video frames included in the second picture data of the target video stream; Extracting scale-invariant feature transform (SIFT) feature points corresponding to each first video frame and SIFT feature points corresponding to each second video frame; Determining a first key frame based on the entropy value corresponding to each first video frame and the SIFT feature points corresponding to each first video frame, where the first key frame is one of the multiple first video frames; Determining a second key frame based on the entropy value corresponding to each second video frame and the SIFT feature points corresponding to each second video frame, where the second key frame is one of the plurality of second video frames; The motion feature similarity between the second picture data and the first picture data is calculated based on the SIFT feature points of the first key frame and the SIFT feature points of the second key frame.

5. The method according to claim 2, wherein The shape feature similarity is calculated as follows: Extracting first edge information corresponding to each of the plurality of first video frames included in the first picture data, and second edge information corresponding to each of the plurality of second video frames included in the second picture data of the target video stream; Calculating a first key point corresponding to each first video frame based on the first edge information, and calculating a second key point corresponding to each second video frame based on the second edge information; Calculating a first shape context descriptor value for each of the first video frames based on the first key point, and calculating a second shape context descriptor value for each of the second video frames based on the second key point; Setting the first video frame having the largest value of the first shape context descriptor among the plurality of first video frames as a third key frame, and setting the second video frame having the largest value of the second shape context descriptor among the plurality of second video frames as a fourth key frame; The shape feature similarity between the second picture data and the first picture data is calculated based on the first shape context descriptor value of the third key frame and the second shape context descriptor value of the fourth key frame.

6. The method according to claim 2, wherein The texture feature similarity is calculated as follows: Extracting a first texture feature corresponding to each of a plurality of first video frames included in the first picture data, and a second texture feature corresponding to each of a plurality of second video frames included in the second picture data of the target video stream; Clustering the first texture features corresponding to each first video frame to obtain a first intermediate texture feature; Clustering the second texture features corresponding to each second video frame to obtain second intermediate texture features; The texture feature similarity between the second picture data and the first picture data is calculated based on the first intermediate texture feature and the second intermediate texture feature.

7. The method according to any one of claims 1 to 6, characterized in that Before sending the push information corresponding to the N video streams to the client, the method further includes: The N video streams are distributed to a content delivery network (CDN), and edge nodes of the CDN are used to cache the N video streams.

8. A commentary recommendation method, applied to a client, characterized in that: include: Collecting first frame data of a target event played by the client within a preset time period; Sending the first screen data to the server; Receive push information of N video streams, where target similarities of the N video streams are greater than target similarities of other video streams in a plurality of video streams, the target similarity being a target similarity between second picture data and the first picture data of each video stream in the plurality of video streams, the plurality of video streams including the N video streams and the other video streams, and N being a positive integer greater than 1; Display the push information corresponding to the N video streams.

9. The method according to claim 8, wherein The collecting of first screen data of a target event played by the client within a preset time period includes: The first picture data is obtained by collecting the pictures of the target event played by the client within a preset time period based on a preset cycle through a preset sliding window.

10. The method according to claim 8, wherein After displaying the push information corresponding to the N video streams, the method further includes: receiving a target operation, wherein the target operation is selecting a target video stream based on the push information, and the target video stream is a video stream among the N video streams; Obtain the target video stream from the edge node of the content delivery network CDN; Play the target video stream.