Video display method and device, storage medium and program product

By combining user viewing needs and historical profiles, personalized video clip recommendations are made, solving the problems of long user confirmation time and low click-through conversion rate in existing technologies, and achieving a more efficient video recommendation experience.

CN121644903APending Publication Date: 2026-03-10ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

When recommending videos, existing AI dialogue assistants require users to spend a lot of time confirming whether they are interested, resulting in a poor user experience. Furthermore, the monotonous preview videos fail to accurately match user interests, reducing video click-through rates.

Method used

By acquiring users' natural language viewing needs and historical viewing profiles, a video recommendation model is used to identify target video segments that users are interested in from the video library, and personalized preview videos are created.

Benefits of technology

It improved video click-through rates by accurately recommending video clips that users are interested in, reducing user confirmation time and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644903A_ABST
    Figure CN121644903A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video display method and device, a storage medium and a program product. In the embodiment of the invention, during video recommendation, in combination with the watching demand of the user and the historical watching portrait of the watching interest of the user, the target video clip in which the user is interested is determined from the target videos matched with the watching demand of the user as the preview video of the video; personalized preview videos are made for users with different watching requirements and watching interests. Therefore, the preview video can better hit the interest and the watching demand of the user, the user is helped to further click and watch the complete video, and the video click conversion rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet, and particularly relates to a video display method and device, a storage medium and a program product. BACKGROUND

[0002] With the continuous development of artificial intelligence (AI) technology, AI conversation assistants are applied to various user scenarios. In a video application (APP), an AI conversation assistant can help a user quickly find desired video content.

[0003] When recommending a video to a user, an existing AI conversation assistant mainly displays the name, brief information and cover information of the video. If the user wants to confirm whether the video is interesting, the user needs to play the video from the beginning to more accurately determine whether the video is interesting. This way causes the user to spend a lot of time to determine whether the recommended video is really interesting, wastes the user's time, causes poor user experience, and cannot encourage the user to use the video application.

[0004] In order for a user to quickly understand the content of a video, such as a movie, some film companies make promotional videos for movies, and a user can understand the video content through the promotional video. However, promotional videos of videos are all the same, which leads to a low video click-through rate. SUMMARY

[0005] Aspects of the present application provide a video display method, device, storage medium and program product to provide a preview video of interest to a user, which helps to improve the video click-through rate.

[0006] In a first aspect, an embodiment of the present application provides a video display method, comprising:

[0007] obtaining viewing demand information described in natural language and a historical viewing portrait of a target user providing the viewing demand information; wherein the historical viewing portrait is obtained by profiling the viewing behavior of the target user.

[0008] determining a target video adapted to the viewing demand information from a video library by a video recommendation model according to the viewing demand information and video classification tags of videos in the video library;

[0009] determining a target video segment from the target video by the video recommendation model according to the historical viewing portrait;

[0010] sending video data of the target video segment to a terminal device providing the viewing demand information, so that the terminal device displays the video data of the target video segment.

[0011] In a second aspect, the embodiments of the present application further provide a video display method, applicable to a cloud server, and the method comprises:

[0012] In response to a request for calling a target service, determining a processing resource corresponding to the target service; the target service provides a video recommendation service;

[0013] Using the processing resource corresponding to the target service to execute the steps in the video display method provided in the first aspect.

[0014] In a third aspect, the embodiments of the present application further provide an electronic device, comprising a memory, a processor and a communication component; wherein the memory is used to store a computer program;

[0015] The processor is coupled to the memory and the communication component, and is used to execute the computer program to execute the steps in the video display method provided in the first aspect and / or the second aspect.

[0016] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium storing computer instructions, when the computer instructions are executed by one or more processors, causing the one or more processors to execute the steps in the video display method provided in the first aspect and / or the second aspect.

[0017] In a fifth aspect, the embodiments of the present application further provide a computer program product, comprising a computer program, when the computer program is executed by one or more processors, causing the one or more processors to execute the steps in the video display method provided in the first aspect and / or the second aspect.

[0018] In the embodiments of the present application, when recommending a video, a target video segment of interest of a user is determined as a preview video of the video from a target video adapted to the viewing demand of the user, in combination with a historical viewing profile of the user's viewing demand and viewing interest, and a personalized preview video is made for users with different viewing demands and viewing interests. Therefore, the preview video can better hit the interests and viewing demands of the user, helping the user to further click and watch the complete video, and improving the video click conversion rate. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 A video recommendation effect diagram of an AI dialogue assistant provided by a traditional scheme;

[0021] Figure 2A structural schematic diagram of a video display system provided by an embodiment of the present application is shown in the figure;

[0022] Figure 3 A flowchart of a video display method provided by an embodiment of the present application is shown in the figure;

[0023] Figure 4 and Figure 5 A video processing process schematic diagram of an embodiment of the present application is shown in the figure;

[0024] Figure 6 A video display effect schematic diagram provided by an embodiment of the present application is shown in the figure;

[0025] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0028] As shown in the figure, Figure 1 In a traditional AI dialogue assistant, a user can describe his viewing demand through a human-computer interaction interface as shown in the figure, Figure 1 such as "help me recommend an action movie". The terminal device of the user can send the viewing demand information described in natural language to a server device. The server device can filter out a target video that matches the viewing demand information from a video library according to the viewing demand information, and send target information of the target video such as the name, synopsis and cover information of the target video to the terminal device. The terminal device displays the target information of the target video on the human-computer interaction interface, such as the name of the movie XXXX, the synopsis and cover of XXXX. If you want to confirm in detail whether you are interested in the video, you still need to play the video from the beginning to judge more accurately. This way leads to the user spending a lot of time to determine whether he is really interested, and the user experience is poor.

[0029] To quickly understand the video content, such as movie content, some film and television companies make promotional videos for movie production, and users can understand the video content summary through the promotional videos. However, the promotional videos of the videos are all the same. Because different users are interested in different contents of the videos, the reasons for recommending the same video to different users may be different. For example, some users are recommended because of the love segments in the movie, and some users are recommended because of the disaster segments in the movie. The preview video (such as the promotional video) cannot necessarily hit the interests of the user who currently provides the viewing demand information. If the preview video cannot hit the interests of the user, the video click conversion rate will be affected.

[0030] The video click conversion rate refers to the proportion of the actual click of the user in a series of videos displayed to the user. The video click conversion rate can be expressed as the ratio of the number of times the video is clicked to play to the number of times the video is displayed.

[0031] To solve the above problems, in some embodiments of the present application, when a video is recommended, the target video segment that interests the user is determined as a preview video of the video from the target video that matches the user's viewing demand, and personalized preview videos are made for users with different viewing demands and viewing interests. Therefore, the preview video can better hit the interests and viewing demands of the user, which helps the user to further click and watch the complete video, and improves the video click conversion rate.

[0032] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0033] It should be noted that the same reference numerals in the following drawings and embodiments represent the same objects, and therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in subsequent drawings and embodiments.

[0034] The video display method provided by the embodiments of the present application can be applied to any video-related application scenario, such as a video recommendation scenario, a video search scenario, etc. In order to better understand and illustrate the solutions of the embodiments of the present application, the solutions of the present application will be described in detail below in combination with a specific embodiment. In this embodiment, the solutions provided by the embodiments of the present application can be applied to a video recommendation scenario. Specifically, the solutions provided by the embodiments of the present application can be implemented as a video application program or a functional plug-in in a video application program. The user can install the application program in any terminal device, and can realize the display of the video through the application program installed on the terminal device. For example, the user can add an AI dialogue assistant in the video application program, and the user can input the viewing demand information through the man-machine interface of the AI dialogue assistant (such as the voice input interface of the AI dialogue assistant). Figure 1The video recommendation result is displayed in the instant communication interface for human-computer interaction. The video application can be a separate application installed in the terminal device, or can be a video application opened by a browser or the like.

[0035] As an optional implementation manner, as shown in Figure 2 Figure 2 An architecture schematic diagram of a video display system is provided in the embodiment. A user can perform information interaction with a terminal device 10, and the terminal device 10 is in communication connection with a server device 20. The terminal device 10 is installed with a video application, and the server device 20 in communication connection with the terminal device 10 can be a server device corresponding to the video application, or can be a server device called by the server device corresponding to the video application to provide a video recommendation service.

[0036] The server device 20 can be a single server device, or a clouded server array, or a virtual machine (VM) running in the clouded server array. In addition, the server device can also refer to other computing devices with corresponding service capabilities, such as a computer terminal device (running a service program) and the like.

[0037] The video application on the terminal device 10 can obtain video data from the server device 20 and display the video data. When the user performs information interaction with the terminal device, the user can input viewing demand information into the terminal device. For example, the terminal device can display a human-computer interaction interface (such as an instant communication interface for human-computer dialogue), and the user can input the viewing demand information through the human-computer interaction interface. The viewing demand information is described in natural language. The natural language can be text or voice and the like. The server device 20 can receive the viewing demand information, and determine a preview video according to the viewing demand information. The implementation manner of determining the preview video according to the viewing demand information by the server device is described in detail below from the perspective of the server device.

[0038] Figure 3 A flow schematic diagram of a video display method is provided in the embodiment. As shown in Figure 3 The video display method mainly includes:

[0039] 301. Obtain viewing demand information described in natural language.

[0040] 302. Determine a target video adapted to the viewing demand information from a video library by a video recommendation model according to the viewing demand information and video classification tags of videos in the video library pre-stored.

[0041] 303. Obtain a historical viewing portrait of a target user providing the viewing demand information. ​

[0042] 304. determining, by the video recommendation model, a target video segment from the target video according to the viewing demand information and the historical viewing portrait.

[0043] 305. sending video data of the target video segment to the terminal device providing the viewing demand information, so as to display the video data of the target video segment by the terminal device.

[0044] In some embodiments, the terminal device can provide a human-computer interaction interface, such as Figure 1 The instant communication interface of the human-computer dialogue is shown, and the user can input the viewing demand information in the human-computer interaction interface. The user can input the viewing demand information by using text or voice. The terminal device can obtain the viewing demand information submitted by the user in the human-computer interaction interface, and send the viewing demand information to the server device. Among them, the scheme of describing the viewing demand information in the form of voice can improve the user expression efficiency compared with describing the viewing demand information in the form of text, and then help to improve the subsequent video recommendation efficiency. The user can describe the viewing demand information by text or voice in the embodiment of the application, and the flexibility of human-computer interaction is higher, especially the voice interaction is more convenient than the text interaction.

[0045] For the server device, in step 301, the viewing demand information described in natural language can be obtained. Further, the target video to which the viewing demand information is adapted can be determined according to the viewing demand information.

[0046] In order to improve the video recommendation efficiency, the video in the video library can be pre-labeled with a classification tag, that is, a video classification tag. The video classification tag of the video in the video library is pre-labeled. The video classification tag can be labeled according to the classification of the attributes of the video. The attributes of the video are used to reflect the information of different dimensions of the video, which can include part or all of the attributes such as the field to which the video belongs, the genre, the theme, the content type, the online time, the language, the shooting unit and the main actors, but not limited to this.

[0047] In some embodiments, the video classification tag can label a video classification tag for a complete video in the video library. In the embodiment of the application, the specific implementation form of the video in the video library is not limited. For example, a complete video can be a movie, a television episode or a television series, a play or a news video, etc., but not limited to this. For any video A in the video library, the video A can be classified by a video analysis model to determine the video classification tag of the video A.

[0048] The video analysis model can be a neural network model. In some embodiments, when the number of parameters in the video analysis model is large, for example, when the number of parameters in the video analysis model is in the millions, hundreds of millions, billions, or even more, the video analysis model can also be called a large model. In the embodiments of this application, a large model is defined as a neural network model whose number of parameters conforms to a preset parameter number range. The preset parameter number range corresponds to a very large number of model parameters, which can be in the millions, hundreds of millions, billions, or even more, and the specific value can be determined by AI field standards.

[0049] In this application, the specific implementation architecture of the video analysis model is not limited. The video analysis model can adopt a large-scale model architecture with usage rights and open source, or it can adopt a self-developed large-scale model architecture. Generative large-scale models do not require retraining when used by users.

[0050] In some embodiments, the video analytics model can be a model that needs to be trained, such as a large model that needs to be trained in two phases: pre-training and fine-tuning. In the pre-training phase, the model is trained on a large-scale, general video dataset. Then, in the fine-tuning phase, the model is further trained on a smaller, more domain-specific dataset. Fine-tuning allows the model to better understand and generate language specific to that domain, thereby better performing the specific task.

[0051] During the training phase of the video analysis model, video samples and their corresponding pre-labeled video classification tags are obtained. Then, with the goal of minimizing the loss function, the initial model of the video analysis model is trained using the video samples to obtain the final video analysis model. The loss function can be represented by the difference between the video classification tags predicted by the video analysis model during training and the pre-labeled video classification tags.

[0052] After obtaining the video analysis model, such as Figure 4 As shown, for any video A in the video library, a video analysis model can be used to classify video A to determine its video category label. Using the same method, all videos in the video library can be classified to obtain their video category labels. Furthermore, the video category labels for each video in the video library can be stored. For example... Figure 4 As shown, video category tags for each video in the video library can be stored in the tag management system. The stored video category tags can include the correspondence between video identifiers and video category identifiers. A video identifier is information that uniquely identifies a video, and can be a video's identity number (ID) and / or name, etc. This video identifier can uniquely pinpoint a video.

[0053] To improve the accuracy of video classification, and thus improve the accuracy of subsequent video classification, the video can be divided into segments, and each video segment can be classified. Specifically, as shown in Figure 5 For any video A in the video library, the video analysis model can classify the video frames included in the video A to obtain the video classification labels of the video frames included in the video A.

[0054] Specifically, the video analysis model can perform semantic analysis on the video A to determine the attribute information of the video frames included in the video A. Further, the video classification labels of the video frames included in the video A can be determined according to the attribute information of the video frames included in the video A.

[0055] Further, from the video A, multiple video frames with the same video classification label and in sequence can be determined as a video segment to obtain multiple video segments included in the video A. Then, as shown in Figure 5 The playing positions of the multiple video segments included in the video A in the video A can also be determined. The playing positions of the video segments included in the video A in the video A can be represented by the starting playing time of the video segment in the video A, for example, the playing positions of the video segments included in the video A in the video A can be represented as (identification of the video A, 50:32). 50:32 means 50 minutes and 32 seconds in the video A. Alternatively, the playing positions of the video segments included in the video A in the video A can be represented by the starting playing time and the ending playing time of the video segment in the video A. The starting playing time and the ending playing time can be accurate to seconds or milliseconds, etc. In the same way, the video classification labels and the playing positions of the video segments included in each video in the video library can be determined; and the video classification labels and the playing positions of the video segments included in each video in the video library can be stored. In this embodiment, the video classification labels of the videos in the video library can include the video classification labels of the video segments included in each video in the video library. The stored video classification labels of the video segments can include the correspondence between the identification of the video segment and the video classification label. The identification of the video segment refers to information that uniquely identifies a video segment, which can be the identification of the video to which the video segment belongs and the playing position of the video segment in the video, etc. Through the identification of the video segment, a video segment can be uniquely locked.

[0056] After obtaining the video classification labels of the videos in the video library, the video classification labels and the playing positions of the video segments included in each video in the video library can be stored. As shown in Figure 5As shown, the video classification tags and playing positions of the video segments contained in each video in the video library can be stored in the tag management system. When online video recommendation or search is performed, based on the viewing requirement information provided by the user and the pre-stored video classification tags of the videos in the video library, the target video that matches the viewing requirement information can be determined from the video library.

[0057] Based on the pre-labeled video classification tags of the videos in the video library, in step 302, the video recommendation model can be used to determine the target video that matches the viewing requirement information from the video library based on the viewing requirement information and the pre-stored video classification tags of the videos in the video library.

[0058] In some embodiments, when the number of model parameters of the video recommendation model is large, for example, the number of parameters of the video recommendation model is millions, hundreds of millions, billions or even more, the video recommendation model can also be referred to as a large model. For the description of the large model, please refer to the foregoing related content, which will not be described here.

[0059] Specifically, the video recommendation model can be used to perform semantic analysis on the viewing requirement information to determine the keywords reflecting the video classification tags contained in the viewing requirement information. Further, the keywords can be matched with the video classification tags of the videos in the video library, and the target video classification tags of the videos that match the keywords can be determined from the pre-stored video classification tags of the videos in the video library. For example, the target video classification tags containing the keywords can be determined from the pre-stored video classification tags of the videos in the video library. The videos corresponding to the target video classification tags are the target videos that match the viewing requirement information.

[0060] For the embodiment in which the video classification tags of a video contain the video classification tags of the multiple video segments contained in the video, the above-mentioned step 302 can be implemented as follows: the video recommendation model is used to determine the candidate video segments that match the viewing requirement information from the video library based on the viewing requirement information and the pre-stored video classification tags of the video segments, as the target videos. The foregoing video segments are the video segments included in the videos in the video library.

[0061] Specifically, the video recommendation model can be used to perform semantic analysis on the viewing requirement information to determine the keywords reflecting the video classification tags contained in the viewing requirement information. Further, the keywords can be matched with the video classification tags of the videos in the video library, and the target video classification tags of the videos that match the keywords can be determined from the pre-stored video classification tags of the video segments. For example, the target video classification tags containing the keywords can be determined from the pre-stored video classification tags of the video segments. The videos corresponding to the target video classification tags are the candidate video segments that match the viewing requirement information. The number of candidate video segments is one or more. More than two means.

[0062] Since different users have different interests in the content of a video, the reason why the same video is recommended to different users can be different. For example, some users are recommended because of the love segment in a movie, and some users are recommended because of the disaster segment in the movie. A preview video (such as a promotional video) that is the same for all users may not hit the interests of the current user providing the viewing demand information. If the preview video does not hit the interests of the user, the video click conversion efficiency will be affected.

[0063] Based on this, in the embodiments of the present application, in order to recommend a preview video that a user is really interested in to different users, the viewing behavior of the user can also be profiled to obtain the historical viewing profile of the user. Specifically, in step 303, the historical viewing profile of the target user providing the viewing demand information can be obtained.

[0064] Specifically, the viewing behavior data of the target user can be obtained. The viewing behavior data refers to various interactions and activity data generated by the user when watching a video, which can help analyze the viewing habits, preferences of the user, and the performance of the video content. The viewing behavior data can include, but is not limited to, video classification tags of videos watched by the target user in the past, the number of times of watching videos by the target user, the watching time, the completion rate, the repeated playback rate, the video classification tags of fast-forward videos, the video classification tags of rewinding videos, comment data on videos, and sharing or recommending data of videos, etc. Further, the viewing behavior data of the target user can be used to profile the viewing interests of the target user to obtain the historical viewing profile of the target user.

[0065] In some embodiments, the viewing behavior data of the target user can be input into a user profiling model, in which the video classification tags of interest to the user are determined according to the viewing behavior data of the target user, as the historical viewing profile of the target user.

[0066] The user profiling model can be a deep learning model, such as a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a long short-term memory (LSTM) model, or a self-attention model, but is not limited thereto.

[0067] The historical viewing profile of the user obtained by profiling the viewing interests of the user can provide reference information for subsequent video recommendation. When recommending videos, the videos recommended in combination with the viewing interests of the user can more accurately match the interests of the user, which helps the user to further watch the complete video and improves the video click conversion rate.

[0068] Based on the target user's historical viewing profile, in step 304, the target video segment can be determined from the target video using a video recommendation model based on the historical viewing profile.

[0069] In some embodiments, such as Figure 4 As shown, the target video is a complete video. The video recommendation model may include a video analysis module. The implementation of the video analysis module can be found in the aforementioned video analysis model and will not be repeated here. Accordingly, step 304 can be implemented as follows: The video analysis module performs semantic analysis on the video frames contained in the target video to determine the video classification labels of the video frames contained in the target video; from the target video, multiple consecutive video frames with the same video classification label are identified as a video segment, thus identifying multiple video segments contained in the target video. Then, the video recommendation model can determine the matching degree between the multiple video segments contained in the target video and the target user's interests based on historical viewing profiles and the video classification labels of the multiple video segments contained in the target video; and based on the matching degree between the multiple video segments contained in the target video and the target user's interests, a target video segment is determined from the multiple video segments contained in the target video. For example, the video segment with the highest matching degree with the target user's interests can be determined from the multiple video segments contained in the target video as the target video segment. Alternatively, at least one video segment with a matching degree greater than or equal to a set matching degree threshold with the target user's interests can be determined from the multiple video segments contained in the target video. Then, according to the order in which these at least one video segment appears in the target video, the at least one video segment can be spliced ​​together to form the target video segment, etc.

[0070] Alternatively, after determining the video category tags of the video frames contained in the target video, a video recommendation model can be used to determine the matching degree between the video frames contained in the target video and the target user's interests based on historical viewing profiles and the video category tags of the video frames contained in the target video. Specifically, the video recommendation model can determine the number of identical sub-tags between the video category tags of the video frames contained in the target video and ...

[0071] For example, the video classification tags of a video frame A contained in the target video are {business, love, movie, Zhao Li, after 2010, Li Si director}, and the video classification tags of the target user are {business, action, movie, Zhang San, after 2010, Wang Wu director}. The video classification tags of the target user and the video classification tags of the video frame A have 3 same sub-tags, and thus the matching degree between the video frame A and the interest of the target user can be 3 / 6*100% = 50%. The video classification tags of another video frame B contained in the target video are {business, action, movie, Zhang San, after 2010, Li Si director}, and the video classification tags of the target user have 5 same sub-tags, and thus the matching degree between the video frame A and the interest of the target user can be 5 / 6*100% ≈ 83.33%.

[0072] Further, the target video segment is determined from the target video according to the matching degree between the video frame contained in the target video and the interest of the target user. Specifically, the target video frame can be determined from the target video according to the matching degree between the video frame contained in the target video and the interest of the target user, and the target video frame is spliced into the target video segment.

[0073] Alternatively, the target video frame with the matching degree greater than or equal to the set matching degree threshold can be determined from the multiple target video segments contained in the target video. Then, the target video frame can be spliced into the target video segment according to the order of the multiple target video segments in the target video, and the like.

[0074] In some other embodiments, the target video is implemented as Figure 5 The multiple candidate video segments are shown. Accordingly, the step 304 can be implemented as: determining the matching degree between the multiple candidate video segments and the interest of the target user according to the historical viewing portrait and the video classification tags of the multiple candidate video segments by the video recommendation model; and determining the target video segment from the multiple candidate video segments according to the matching degree between the multiple candidate video segments and the interest of the target user. For example, the candidate video segment with the highest matching degree between the target user and the interest can be determined from the multiple candidate video segments as the target video segment. Or, at least one candidate video segment with the matching degree greater than or equal to the set matching degree threshold between the target user and the interest can be determined from the multiple candidate video segments. When the at least one candidate video segment with the matching degree greater than or equal to the set matching degree threshold between the target user and the interest is one, the one candidate video segment is the target video segment. When the at least one candidate video segment with the matching degree greater than or equal to the set matching degree threshold between the target user and the interest is multiple, the multiple video segments can be spliced into the target video segment according to the order of the multiple video segments in the target video, and the like.

[0075] In some embodiments, in order to improve the quality of the identified target video segment, such as Figure 5 As shown, the quality scores of the videos to which multiple candidate video segments belong can also be obtained. In this embodiment, the specific scoring method for video quality scores is not limited. In some embodiments, for any video A in the video library, information on multiple quality influencing factors of video A can be obtained. Here, quality influencing factors can be understood as quality influencing parameters, and the information on quality influencing factors can be understood as the parameter values ​​of the quality influencing parameters. Quality influencing factors refer to factors that affect the quality of a video. Different types of videos may have different quality influencing factors. For example, for videos with weak timeliness, such as movies, TV series, plays, and documentaries, quality influencing factors may include: the video's website rating, the popularity of the people associated with the video (director, screenwriter, and / or actors, etc.), and the video's awards, etc. For videos with strong timeliness, such as news, quality influencing factors may include: timeliness, content source, and the accuracy of news descriptions, etc.

[0076] Furthermore, the quality score of video A can be obtained by weighting the information of multiple quality influencing factors according to their pre-set weights. The quality score of each video in the video library can be determined using the same method. Furthermore, the quality scores of the videos in the video library can be stored.

[0077] Based on the quality scores of videos in a pre-stored video library, the quality scores of multiple candidate video segments can be obtained. These multiple candidate video segments can belong to the same video or multiple different videos. The number of different videos is less than or equal to the number of candidate video segments.

[0078] Based on the quality scores of the videos to which multiple candidate video segments belong, the target video segment can be determined from multiple candidate video segments according to the matching degree between the multiple candidate video segments and the target user's interests, as well as the quality scores of the videos to which the multiple candidate video segments belong.

[0079] In some embodiments, at least one candidate video segment can be determined from a plurality of candidate video segments whose matching degree with the target user's interests is greater than or equal to a set matching degree threshold. Then, based on the quality scores of these at least one candidate video segment, the candidate video segment with the highest quality score can be determined as the target video segment.

[0080] Alternatively, the matching degree between multiple candidate video segments and the target user's interests, as well as the quality score of the videos to which the multiple candidate video segments belong, can be weighted according to pre-set interest matching weights and video quality score weights to obtain the recommendation probability of multiple candidate video segments. Then, the target video segment can be determined from the multiple candidate video segments based on the recommendation probability. Optionally, the candidate video segment with the highest recommendation probability can be determined as the target video segment from the multiple candidate video segments.

[0081] In this embodiment, the target video segment to be recommended is determined by combining the target user's viewing interests with the video quality. Combining the target user's viewing interests ensures that the recommended videos more accurately match the user's interests, encouraging the user to watch the full video and improving the video click-through rate. Combining video quality allows for the recommendation of higher-quality videos. High-quality videos are generally more popular with users; therefore, recommending high-quality target video segments helps guide users to watch the full video, further improving the video click-through rate.

[0082] Furthermore, the target video segment can be extracted from its parent video based on its playback position within the parent video.

[0083] The method for determining the target video segment shown in the foregoing embodiments is merely illustrative and does not constitute a limitation. After determining the target video segment from the target video, the video data of the target video segment can be sent to the terminal device in step 305. This allows the terminal device to display the video data of the target video segment. For example, as... Figure 6 As shown, the terminal device can display the video data of the target video segment on a human-computer interaction interface (such as an instant messaging interface).

[0084] The target video segment is a preview video, and its video data includes, but is not limited to, the target video segment itself and / or its cover image. Optionally, in embodiments where the target video segment's video data includes the target video segment itself, the terminal device can automatically play the target video segment when the video data is displayed on the human-computer interaction interface. Or, as... Figure 6As shown, the terminal device can display a preview control corresponding to the target video clip on the human-computer interaction interface. The user can click the preview control to view the target video clip. The terminal device can respond to the triggering operation of the preview control to play the target video clip, etc. In an embodiment where the video data of the target video clip is the cover image of the target video clip, the terminal device can respond to the triggering operation of the preview control to request the target video clip from the server device. The server device, in response to the request from the terminal device, sends the target video clip to the terminal device. The terminal device receives the target video clip and automatically plays it.

[0085] In this embodiment, during video recommendation, the system combines the user's historical viewing profile, including their viewing needs and interests, to identify target video segments that match their viewing needs and serve as preview videos. Personalized preview videos are then created for users with different viewing needs and interests. Therefore, these preview videos better target the user's interests and viewing needs, encouraging them to click and watch the full video, thus improving video click-through rates.

[0086] In this embodiment of the application, target information of the video to which the target video segment belongs can also be determined. This target information, which allows users to quickly understand the content of the video to which the target video segment belongs, may include one or more of the following: the name of the video to which the target video segment belongs, a content summary, cover information, and information of the participants. "Multiple" refers to two or more of these.

[0087] Furthermore, the target information of the video to which the target video segment belongs can be sent to the terminal device. For example... Figure 6 As shown, the terminal device can display target information of the video to which the target video clip belongs. For example, the terminal device can display target information in a human-computer interaction interface (such as an instant messaging interface). In this way, after previewing the target video clip, the user can directly play the video to which the target video clip belongs based on the target information. For example, the user can play the video to which the target video clip belongs by clicking the playback control corresponding to the target information; or, the target information of the video to which the target video clip belongs can be associated with a video network address; the user can click the target information to play the video to which the target video clip belongs, and so on.

[0088] The server-side device can send the target video segment and the video to which it belongs together to the terminal device. For example, the target video segment can be spliced ​​before the video to which it belongs, resulting in a spliced ​​video, which is then sent to the terminal device. Alternatively, the server-side device can encapsulate the target video segment and the video to which it belongs into a single data packet and send the data packet to the terminal device. Of course, the server-side device can also send the target video segment and the video to which it belongs to the terminal device separately.

[0089] Similarly, the video display method provided in this application embodiment can be deployed on any computing device. Optionally, the video display method provided in this application embodiment can also be deployed on a cloud server as a Software as a Service (SaaS) application. For a cloud server with this SaaS application deployed, the steps in the above video display method can be executed in response to a request to call the target service.

[0090] In this embodiment, the target service refers to the video recommendation service. The processing resources corresponding to the target service refer to the processing resources required to execute the above video display method, including but not limited to: processor resources, memory resources, and input / output (IO) resources.

[0091] The video display method provided in this embodiment can be deployed on a cloud server to provide users with video recommendation services for target applications, i.e., target services. Users can be service providers for the target applications, or users or clients of the target applications. Optionally, the cloud server can provide an Application Programming Interface (API) to the user. The service requester (i.e., the user) can call the API to invoke the target service. Accordingly, the request to invoke the target service is implemented as a call event generated by calling the API. The service requester (i.e., the user) can also invoke the target service through Remote Procedure Call (RPC) or Remote Direct Memory Access (RDMA) technologies.

[0092] For cloud servers, in response to a request to call the target service, the processing resources corresponding to the target service can be determined; and the processing resources corresponding to the target service can be used to execute the aforementioned steps 301-305 to recommend the target video clip to the target user.

[0093] In this embodiment, during video recommendation, the system combines the user's historical viewing profile, including their viewing needs and interests, to identify target video segments that match their viewing needs and serve as preview videos. Personalized preview videos are then created for users with different viewing needs and interests. Therefore, these preview videos better target the user's interests and viewing needs, encouraging them to click and watch the full video, thus improving video click-through rates.

[0094] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 301 and 302 can be device A; or the execution subject of step 301 can be device A, and the execution subject of step 302 can be device B; and so on.

[0095] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 301, 302, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0096] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps in the aforementioned video display methods.

[0097] This application also provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the aforementioned video display methods. In this application, the specific implementation of the computer program product is not limited. In some embodiments, the computer program product may be implemented as an application (APP), a mini-program, a PC client, a program module, a plug-in, an installation package, a software development kit (SDK), an optical disc image file (such as an ISO file), a plug-in, or software in the form of Software as a Service (SaaS), etc., but is not limited thereto.

[0098] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7As shown, the electronic device includes a memory 70a and a processor 70b. The memory 70a is used to store computer programs.

[0099] The processor 70b is coupled to the memory 70a and is used to execute a computer program to perform the steps in the video display methods provided in the foregoing embodiments. Specific implementation details of each step can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.

[0100] In some alternative implementations, such as Figure 7 As shown, the electronic device may also include optional components such as a communication component 70c, a power supply component 70d, a display component 70e, and an audio component 70f. Figure 7 The diagram only shows some components and does not mean that the electronic device must contain them. Figure 7 The inclusion of all components does not imply that an electronic device can only include... Figure 7 The components shown.

[0101] in addition, Figure 7 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the form factor of the electronic device. The electronic device in this embodiment can be a desktop computer, laptop computer, mobile phone, or IoT device; it can also be a traditional server, cloud server, or server cluster, or other server equipment.

[0102] In this embodiment, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Electrically Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0103] In the embodiments of this application, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), a complex programmable logic device (CPLD), or other programmable devices; or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited thereto.

[0104] In this embodiment, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device housing the communication component can access wireless networks based on communication standards, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In another exemplary embodiment, the communication component may also be implemented based on Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), or other technologies.

[0105] In embodiments of this application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0106] In this embodiment, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.

[0107] In embodiments of this application, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. For example, in devices with voice interaction capabilities, voice interaction with the user can be achieved through the audio component.

[0108] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.

[0110] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0113] In a typical configuration, a computing device includes one or more processors (CPU, etc.), input / output interfaces, network interfaces, and memory.

[0114] Memory may include non-persistent storage in computer-readable media, such as random-access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0115] Computer storage media are readable storage media, also known as removable media. Removable and non-removable media can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient media, such as modulated data signals and carrier waves.

[0116] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.

[0117] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A video display method characterized by, The method comprises: obtaining viewing demand information described in natural language and a historical viewing profile of a target user providing the viewing demand information, wherein the historical viewing profile is obtained by profiling viewing behavior of the target user; determining, by a video recommendation model, a target video from a video library according to the viewing demand information and video classification tags of videos in the video library; determining, by the video recommendation model, a target video segment from the target video according to the historical viewing profile; sending video data of the target video segment to a terminal device providing the viewing demand information, so that the terminal device displays the video data of the target video segment.

2. The method of claim 1, wherein, For any video in the video library, the video classification tags of the any video include video classification tags of a plurality of video segments contained in the any video. The method comprises: determining, by a video recommendation model, candidate video segments from the video library according to the viewing demand information and video classification tags of each video segment, as the target video, wherein the each video segment is a video segment included in each video in the video library.

3. The method of claim 2, wherein, The candidate video segments are a plurality of; the method comprises: determining, by the video recommendation model, matching degrees between the plurality of candidate video segments and interests of the target user according to the historical viewing profile and the video classification tags of the plurality of candidate video segments; determining the target video segment from the plurality of candidate video segments according to the matching degrees between the plurality of candidate video segments and the interests of the target user.

4. The method of claim 3, wherein, The method further comprises: obtaining quality scores of videos to which the plurality of candidate video segments belong; The method comprises: determining the target video segment from the plurality of candidate video segments according to the matching degrees between the plurality of candidate video segments and the interests of the target user, and the quality scores of the videos to which the plurality of candidate video segments belong.

5. The method of claim 4, wherein, The method comprises: weighting the matching degrees between the plurality of candidate video segments and the interests of the target user, and the quality scores of the videos to which the plurality of candidate video segments belong, according to pre-set interest matching degree weights and video quality score weights, to obtain recommended probabilities of the plurality of candidate video segments; determining the target video segment from the plurality of candidate video segments according to the recommended probabilities of the plurality of candidate video segments.

6. The method according to any one of claims 2-5, characterized in that, The method further comprises, before sending the video data of the target video clip to the terminal device sending the viewing demand information: determining the playing position of the target video clip in the corresponding video from the playing positions of the pre-stored video clips in the corresponding video; extracting the target video clip from the video to which the target video clip belongs according to the playing position of the target video clip in the corresponding video.

7. The method of claim 6, wherein, Further comprising: performing semantic analysis on any video in the video library by a video analysis model to determine attribute information of video frames contained in the any video; determining video classification labels of the video frames contained in the any video according to the attribute information of the video frames contained in the any video; determining a plurality of video frames with the same video classification label and in succession in the any video as a video clip to determine a plurality of video clips contained in the any video; determining the playing positions of the plurality of video clips contained in the any video in the any video; storing the video classification labels of the plurality of video clips contained in the any video and the playing positions of the plurality of video clips contained in the any video in the any video.

8. The method according to any one of claims 1 to 5, characterized in that, Obtaining a historical viewing profile of a target user providing the viewing demand information comprises: obtaining viewing behavior data of the target user; performing a viewing interest profiling on the target user according to the viewing behavior data to obtain the historical viewing profile.

9. The method of claim 4, wherein, Before obtaining the quality scores of the videos to which the plurality of candidate video clips belong, further comprising: obtaining information of a plurality of quality influencing factors of any video in the video library; weighting the information of the plurality of quality influencing factors according to the weights of the plurality of quality influencing factors to obtain the quality score of the any video; storing the quality score of the any video; The obtaining of the quality scores of the videos to which the plurality of candidate video clips belong comprises: obtaining the quality scores of the videos to which the plurality of candidate video clips belong from the pre-stored quality scores of the videos in the video library.

10. The method of claim 1, wherein, The determining of a target video clip from a target video by the video recommendation model according to the historical viewing profile comprises: performing semantic analysis on video frames contained in the target video by the video recommendation model to determine attribute information of the video frames contained in the target video; determining video classification labels of the video frames contained in the target video according to the attribute information of the target video frames; determining a matching degree between the video frames contained in the target video and the interests of the target user according to the historical viewing profile and the video classification labels of the video frames contained in the target video; determining target video frames from the video frames contained in the target video as the target video clip according to the matching degree between the video frames contained in the target video and the interests of the target user.

11. The method of claim 1, wherein, The obtaining of the viewing demand information described in natural language comprises: obtaining the viewing demand information described in natural language sent by the terminal device; the viewing demand information is obtained by the terminal device through a human-computer interaction interface. The sending the video data of the target video clip to the terminal device sending the viewing demand information comprises: sending the video data of the target video clip to the terminal device for displaying the video data of the target video clip on the human-computer interaction interface by the terminal device; The method further comprises: determining target information of a video to which the target video clip belongs; and sending the target information to the terminal device for displaying the target information on the human-computer interaction interface by the terminal device.

12. A video display method, applicable to a cloud server, characterized in that, The method comprises: in response to a request for invoking a target service, determining a processing resource corresponding to the target service; the target service providing a video recommendation service; using the processing resource corresponding to the target service to execute the steps in the method of any one of claims 1-11.

13. An electronic device, comprising: comprise: a memory, a processor and a communication component; wherein the memory is configured to store a computer program; the processor is coupled to the memory and the communication component, and is configured to execute the computer program to execute the steps in the method of any one of claims 1-12.

14. A computer readable storage medium having stored thereon computer instructions, wherein, when the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method of any one of claims 1-12.

15. A computer program product, characterised in that, comprise a computer program, when the computer program is executed by one or more processors, the one or more processors are caused to execute the steps in the method of any one of claims 1-12.