Video recommendation method and device, electronic device, computer readable storage medium and computer program product
By displaying a video collection control in the media stream interface and loading video collection information on demand, the problem of video collection creation relying on a single publishing object is solved, achieving efficient and accurate video recommendation and reducing network interaction frequency and resource consumption.
Patent Information
- Application Number
- CN202610771306.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, the creation of video collections relies on a single publishing object, which is inflexible and the frequent network interactions lead to excessive network bandwidth and memory consumption, affecting the efficiency and accuracy of video recommendations.
Display video collection controls in the media stream interface, load video collection information on demand, reduce graphics rendering calculations, batch retrieve video information from different publishers on the same topic, and avoid multiple retrieval requests and network handshakes.
It effectively saves network bandwidth and memory consumption, reduces processor power consumption, improves the accuracy and efficiency of video recommendations, and ensures smooth video playback.
Smart Images

Figure CN122640593A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a video recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of mobile internet and multimedia technology, media streaming has become the main way for users to obtain information. In the current media streaming playback interface, the person who publishes the video can create a video collection for the published video, and the person who watches the video can view the video collections published by different people by searching for the name of the video collection. In related technologies, multiple independent retrieval requests and network handshakes need to be initiated, resulting in frequent network input and output interactions. Summary of the Invention
[0003] This application provides a video recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can reduce the frequency of network input and output interactions and improve the accuracy of video recommendations.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a video recommendation method, the method comprising: Display the first video in the media stream interface, and also display the video collection control; The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same theme as the first video. In response to a video viewing command triggered by the video collection control, video information of each of the second videos is displayed; The video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to multiple pieces of video information is different.
[0005] This application embodiment provides a video recommendation device, the device comprising: The display module is used to display the first video in the media stream interface and to display the video collection control; The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same theme as the first video. The response module is used to respond to a video viewing command triggered based on the video collection control and display the video information of each of the second videos; The video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to multiple pieces of video information is different.
[0006] In the above scheme, the response module is further configured to, before displaying the video information of each of the second videos, in response to a trigger operation on the video collection control, display the collection information of the video collection and display a first viewing control for viewing the video information of the second videos; wherein, the collection information includes at least one of the following: the name of the video collection, the cover of the video collection, and a brief description of the video collection; and in response to a trigger operation on the first viewing control, trigger the video viewing instruction.
[0007] In the above scheme, the first video belongs to different candidate video sets, and different candidate video sets correspond to different types. The response module is also used to respond to the trigger operation of the video set control by displaying multiple candidate types, wherein the candidate type is the type of the candidate video set to which the first video belongs, and the multiple candidate types include the target type of the video set; and to respond to the trigger operation of the target type among the multiple candidate types by displaying the set information of the video set of the target type.
[0008] In the above scheme, the display module is further configured to display at least one of the following information for each candidate type: the number of objects watching videos in the video collection of the candidate type, the recommendation level of the video collection of the candidate type, and the update time of the video collection of the candidate type.
[0009] In the above scheme, the collection information of the video collection is displayed on the information display interface. The display module is also used to display the video collection control using a first display style. The first display style is used to guide and trigger the video collection control. The response module is further configured to, in response to a trigger operation on the video collection control for the first display style, display a playback control for playing the first video in the information display interface; in response to a trigger operation on the playback control, return to the media stream interface to play the first video, and display the video collection control using a second display style; wherein the second display style is different from the first display style, and the second display style is used to indicate that the video collection control has been triggered.
[0010] In the above scheme, the response module is also used to respond to the trigger operation of the video collection control, display the viewing progress of the current object watching the video collection; when the viewing progress reaches the preset viewing progress, display the virtual reward resources corresponding to the preset viewing progress, and display the claim control for claiming the virtual reward resources.
[0011] In the above scheme, the response module is further configured to, in response to a video viewing instruction triggered based on the video collection control, display the first video in the first display area of the media stream interface; if the number of the second videos is less than or equal to a first quantity threshold, display the video information of each of the second videos in the second display area of the media stream interface; if the number of the second videos is greater than the first quantity threshold, display multiple filter tags in the second display area of the media stream interface, the multiple filter tags including a target filter tag; in response to a trigger operation for the target filter tag, display the video information of each of the second videos filtered according to the target filter tag in the second display area.
[0012] In the above scheme, the display module is also used to display a setting control before displaying the video information of each of the second videos in response to the video viewing instruction triggered by the video collection control. The setting control is used to set the filtering criteria for filtering the second videos. The response module is also configured to respond to a setting instruction triggered based on the setting control. If the number of the second videos is greater than the first quantity threshold, the module displays the filtering criteria set by the setting instruction and displays multiple filtering labels after filtering the second videos according to the filtering criteria in the second display area of the media stream interface.
[0013] In the above scheme, the display module is further configured to display the video collection control when at least one of the following conditions is met: the duration of playing the first video is greater than a duration threshold; or an interactive operation is performed on the first video.
[0014] In the above scheme, the display module is further configured to, when displaying the first video, if there is a video collection including the first video, display a collection switch; wherein, the collection switch is used to enable the video collection mode; and when the video collection mode is enabled based on the collection switch, a video collection control is displayed.
[0015] In the above scheme, the response module is further configured to, after displaying the video information of each of the second videos, respond to a cancellation display instruction for the video information of a target second video among the plurality of second videos, cancel the display of the video information of the target second video; when the number of video information of the second videos that have not been canceled decreases to a second quantity threshold, display an add control; wherein, the add control is used to add videos to be displayed based on the content of the second videos that have not been canceled.
[0016] In the above scheme, the response module is further configured to, after the display of the add control, in response to a trigger operation on the add control, display at least one add condition; wherein the add condition is a condition satisfied by the added video; in response to a selection operation on a target add condition among the at least one add condition, display an added third video; wherein the third video is determined based on the fact that the second video was not canceled from display and meets the target add condition.
[0017] In the above scheme, each second video has a corresponding serial number. The display module is also used to display the video information of each second video in ascending order of the serial number. The response module is further configured to, in response to a cancellation display instruction for the video information of a target second video among the plurality of second videos, update the sequence number of the second videos arranged after the target second video, and re-display the video information of the remaining second videos according to the updated sequence number.
[0018] In the above scheme, the content of the first video is the first sub-event of the target event occurring at the first time point. The display module is also used to display a progress axis with at least one node, wherein each node has a corresponding sub-event, and the sub-event belongs to the target event. In the associated area of each node in the progress axis, video information of the second video corresponding to the node is displayed. The content of the second video is at least one of the following: the second sub-event of the target event occurring at the first time point, and the third sub-event of the target event occurring at the second time point.
[0019] In the above scheme, the response module is further configured to display a shared viewing control for the video collection after displaying the video information of each of the second videos; wherein the shared viewing control is configured to invite other objects to watch the videos in the video collection together; based on the shared viewing control, in response to the shared viewing instruction for the other objects, an invitation message for inviting them to watch the video collection together is sent to the other objects.
[0020] In the above scheme, the response module is further configured to, upon receiving consent information, respond to a selection operation triggered by the video information of the second video. If the selected video is consistent with the video selected by the other object, a shared viewing interface is displayed, and the selected video is displayed in the shared viewing interface. The consent information is used to indicate that the other object agrees to share the video collection based on the invitation information. Upon receiving the consent information, in response to the selection operation triggered by the video information of the second video, if the selected video is inconsistent with the video selected by the other object, the selected video is displayed, and the object identifier and viewing progress of the other object are displayed in the associated area of the video information of the video selected by the other object.
[0021] In the above scheme, the response module is further configured to, after displaying the video information of each of the second videos, control the selected second video to be in a selected state in response to the selection operation of the plurality of second videos; wherein, the number of second videos in the selected state is less than the number of second videos in the video collection; in response to the sharing instruction for the second video in the selected state, create a new video collection of the second video in the selected state, and share the new video collection according to the sharing instruction.
[0022] In the above scheme, the response module is further configured to respond to a video viewing instruction triggered based on the video collection control. If the second video included in the video collection is a subset of videos selected by other objects from the target video collection, a second viewing control is displayed. The second viewing control is used to view the target video collection. The response module is also configured to, after displaying the video information of each of the second videos, in response to a trigger operation on the second viewing control, switch the display of the video information of the second videos to the video information of the videos in the target video collection.
[0023] In the above scheme, the response module is further configured to, after displaying the video information of each of the second videos, play the second video in response to a playback command for the second video; and when the number of unplayed second videos decreases to a third quantity threshold, display at least one recommended video collection based on the video collection.
[0024] In the above scheme, the second video includes a target video that has an arrangement order with the first video, and other videos that supplement the content of the first video; the response module is further configured to display a first playback control and a second playback control after the first video is displayed in the media stream interface and the first video finishes playing; wherein, the first playback control is used to play the target video, and the second playback control is used to play the other videos; in response to a trigger operation on the first playback control, the target video that is adjacent to and follows the first video in the arrangement order is played; in response to a trigger operation on the second playback control, the other videos are played; when the other videos finish playing, the target video that is adjacent to and follows the first video in the arrangement order is played.
[0025] In the above scheme, the plurality of second videos includes a target second video. The response module is further configured to, after displaying the video information of each second video, in response to a trigger operation on the video information of the target second video, switch the display of the first video to display the target second video; if the target second video also belongs to other video collections, display a jump control; in response to a trigger operation on the jump control, switch the displayed video information of each second video to the video information of videos in the other video collections.
[0026] This application provides an electronic device, the electronic device comprising: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the video recommendation method provided in the embodiments of this application.
[0027] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the video recommendation method provided in this application when executed by a processor.
[0028] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the video recommendation method provided in this application.
[0029] The embodiments of this application have the following beneficial effects: In the video recommendation method provided in this application embodiment, a first video is displayed in the media stream interface, and a video collection control is also displayed. The video collection control is used to view a video collection recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same topic as the first video. In response to a video viewing instruction triggered based on the video collection control, video information of each second video is displayed. The video information includes the video publishing object that published the second video, and at least one of the video publishing objects corresponding to the multiple video information is different.
[0030] The video recommendation method provided in this application displays a video collection control instead of preloading the media stream data or video information of all second videos in the video collection when displaying the first video. This effectively saves network bandwidth and memory consumption during the operation of the electronic device, ensuring the smooth playback of the first video. Responding to video viewing instructions triggered by the video collection control, the video information of each second video is displayed, which is equivalent to on-demand loading. This reduces graphics rendering calculations, thereby reducing processor power consumption and computational load. Furthermore, responding to video viewing instructions triggered by the video collection control, video information of second videos belonging to the same topic but from different video publishers can be retrieved in batches. This avoids the electronic device initiating multiple independent search requests or multiple network handshakes to obtain videos published by different video publishers, significantly reducing the frequency of network input and output interactions of the electronic device and improving data acquisition efficiency. Moreover, since the first video and the second videos in the recommended video collection belong to the same topic, the accuracy of video recommendations is improved. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the structure of the video recommendation system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the video recommendation method provided in the embodiments of this application. Figure 1 ; Figure 4 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 1 ; Figure 5 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 2 ; Figure 6 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 3 ; Figure 7 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 4 ; Figure 8 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 5 ; Figure 9 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 6 ; Figure 10 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 7 ; Figure 11 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 8 ; Figure 12 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 9 ; Figure 13 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 ; Figure 14 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 one; Figure 15 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 two; Figure 16 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 three; Figure 17 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 Four; Figure 18 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 five; Figure 19 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 six; Figure 20 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 seven. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0034] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0035] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0036] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0037] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0038] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0039] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0040] 2) The human-computer interaction interface can be any display interface involved in the embodiments of this application, used to provide human-computer interaction functions, and to display videos. For example, graphical user interface (GUI) displays, such as augmented reality (AR) interfaces, virtual reality (VR) interfaces, voice user interfaces (VUI), interactive projection interfaces (using projection technology to display information on a plane), eye-tracking interfaces (interfaces controlled by detecting the user's gaze), holographic interfaces (three-dimensional holograms formed by projecting images using holographic projection technology, allowing the user to see stereoscopic images without wearing special glasses), multimodal interfaces (interfaces that combine multiple interaction methods, such as tactile, visual, and auditory interaction), brain-machine interface (BMI) interfaces, etc.
[0041] 3) Client, also known as user terminal, refers to the program that provides local services to users in contrast to the server. Except for some applications that can only run locally, it is generally installed on ordinary client machines and needs to cooperate with the server to run. That is, there needs to be a corresponding server and service program in the network to provide the corresponding services. Thus, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application.
[0042] 4) A media streaming interface refers to an interface used to receive and continuously present streaming media data. Streaming media data includes, but is not limited to, at least one of the following: continuous short videos (videos with a duration less than a duration threshold), long videos (videos with a duration greater than or equal to a duration threshold), live streams, and text and image information streams that are downloaded and played simultaneously over a network. The presentation form of a media streaming interface is not limited to two-dimensional physical screens (such as graphical interfaces displayed on smartphones and personal computer screens), but also includes three-dimensional spatial display interfaces (such as virtual windows in virtual reality and augmented reality) or holographic projection interfaces. Media streaming interfaces rely on hardware rendering calculations and memory data caching, and can trigger sequential loading, automatic playback, or seamless switching of streaming media data based on user input commands (such as swiping, clicking) or preset logic.
[0043] 5) A video publishing object refers to a logical entity or identity credential that performs video data uploading, authorization, or distribution operations within a multimedia distribution network. Specific forms of video publishing objects include, but are not limited to, personal user accounts or corporate accounts registered on multimedia platforms, or automated program entities (such as virtual digital humans or content distribution robots) that automatically generate and deliver digital content based on artificial intelligence algorithms. Each video publishing object is bound to a unique identifier and possesses specific control permissions over the published video data (such as setting access levels and managing statuses like content removal), as well as associated digital attributes (such as account nicknames, qualification certification identifiers, and social relationship chains).
[0044] The applicant discovered the following technical problems in the relevant technology: In related technologies, video collections are created by the object that publishes the video. For example, if object 1 publishes video 1, video 2, and video 3, object 1 can build a video collection from video 1 and video 2. The video collection only includes the videos published by the single object that published the video. This method is highly dependent on manual intervention and has extremely poor flexibility.
[0045] This application provides a video recommendation method, a video recommendation device, an electronic device, a computer-readable storage medium, and a computer program product, which can reduce the frequency of network input and output interactions and improve the accuracy of video recommendations.
[0046] See Figure 1 , Figure 1 This is a schematic diagram of the structure of the video recommendation system provided in the embodiments of this application. Figure 1 The video recommendation system 100 shown supports a video recommendation application. The terminal 400 connects to the server 200 through the network 300, which can be a wide area network, a local area network, or a combination of both.
[0047] Terminal 400 can display a media stream interface, which displays a first video and a video collection control. The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, which belong to the same theme as the first video. The data corresponding to the media stream interface can be sent from server 200 to terminal 400. In response to a video viewing command triggered by the video collection control, the video information of each second video is displayed. The video information includes the video publishing object that published the second video, and at least one of the video publishing objects corresponding to the multiple video information is different.
[0048] In the video recommendation method provided in this application embodiment, a first video is displayed in the media stream interface, and a video collection control is also displayed. The video collection control is used to view a video collection recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same topic as the first video. In response to a video viewing instruction triggered based on the video collection control, video information of each second video is displayed. The video information includes the video publishing object that published the second video, and at least one of the video publishing objects corresponding to the multiple video information is different.
[0049] The video recommendation method provided in this application displays a video collection control instead of preloading the media stream data or video information of all second videos in the video collection when displaying the first video. This effectively saves network bandwidth and memory consumption during the operation of the electronic device, ensuring the smooth playback of the first video. Responding to video viewing instructions triggered by the video collection control, the video information of each second video is displayed, which is equivalent to on-demand loading. This reduces graphics rendering calculations, thereby reducing processor power consumption and computational load. Furthermore, responding to video viewing instructions triggered by the video collection control, video information of second videos belonging to the same topic but from different video publishers can be retrieved in batches. This avoids the electronic device initiating multiple independent search requests or multiple network handshakes to obtain videos published by different video publishers, significantly reducing the frequency of network input and output interactions of the electronic device and improving data acquisition efficiency. Moreover, since the first video and the second videos in the recommended video collection belong to the same topic, the accuracy of video recommendations is improved.
[0050] The following describes an electronic device that performs the video recommendation method provided in the embodiments of this application. The electronic device implementing the video recommendation method in the embodiments of this application can be a terminal, a server, or a combination of both. Therefore, the executing entity of each step will not be repeated below. In some embodiments, the terminal can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals.
[0051] In some embodiments, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and server can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.
[0052] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The illustrated electronic device includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components of the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.
[0053] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0054] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0055] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0056] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0057] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0058] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with user interface 430 (e.g., a display screen, a speaker, etc.). The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0059] In some embodiments, the video recommendation device provided in this application can be implemented in software. Figure 2 A video recommendation device 455 stored in memory 450 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a display module 4551 and a response module 4552. These modules are logically linked and can therefore be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0060] In some embodiments, the terminal or server can implement the video recommendation method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as news applications, video applications, social applications, instant messaging applications, and online education applications; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0061] The video recommendation method provided in the embodiments of this application will be described below. As mentioned above, the electronic device implementing the video recommendation method in the embodiments of this application can be a terminal, a server, or a combination of both. See [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating the video recommendation method provided in the embodiments of this application. Figure 1 The following is combined with Figure 3 The steps shown are illustrated using an electronic device as a terminal to illustrate the video recommendation method provided in this application embodiment.
[0062] In step 101, the first video is displayed in the media stream interface, and a video collection control is also displayed.
[0063] In practical applications, the terminal is equipped with an application that can view videos. This application can be any of the following: news application, video application, social application, instant messaging application, or online education application.
[0064] In some embodiments, a media streaming interface may be displayed in response to a triggering operation on the application. In some embodiments, an application display interface may be displayed in response to a triggering operation on the application, the application having video functionality and the application display interface displaying video functionality controls, and a media streaming interface may be displayed in response to a triggering operation on the video functionality controls.
[0065] Among them, the triggering operation refers to the behavior of the user to trigger a certain function or event by interacting with the display interface of the terminal. The triggering operation can include one or more of the following: single click operation, double click operation, long press operation, drag operation, swipe operation, hover operation, shortcut key, voice control, and gesture operation. The triggering operations provided in the embodiments of this application can be referred to the above description, and will not be repeated hereafter.
[0066] In some embodiments, a first video can be displayed in the media stream interface, which is the audio and video currently being presented in the media stream interface. In addition, a video collection control can be displayed in the media stream interface. The video collection control is used to view a video collection recommended based on the first video. The video collection includes multiple second videos.
[0067] Multiple second videos belong to the same theme as the first video. The following explains why multiple second videos belong to the same theme as the first video. In the news field, the theme can be news events and special columns, etc. In the film and television field, the theme can be movies and documentaries, etc. In the sports field, the theme can be a game or a specific round of the game, etc.
[0068] The timing of displaying the video collection control is explained below. In some embodiments, when displaying the first video, if the first video belongs to a video collection, the video collection control can be displayed. In related technologies, when recommending video collections in a media stream interface, there is a lack of a precise control mechanism for the timing of displaying the video collection control. If the video collection control is displayed directly when the terminal starts playing the first video, it is easy to obscure the main visual image of the first video, interfere with the normal viewing process of the first video content, and lead to low human-computer interaction efficiency.
[0069] In some embodiments, the video collection control may be displayed if at least one of the following conditions is met: the duration of the first video is greater than a duration threshold; an interactive operation is performed on the first video; a friend of the current object performs an interactive operation on the first video; or the current popularity of the first video is greater than a popularity threshold.
[0070] In some embodiments, when displaying the first video, if the first video belongs to a video collection, the video collection control may not be displayed initially. When it is detected that the duration of the first video exceeds a duration threshold, the video collection control is overlaid at the bottom of the first video's display area. See [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 1 When the duration of the first video is detected to be greater than the duration threshold, a video collection control 401 is overlaid at the bottom of the display area of the first video.
[0071] In some embodiments, upon receiving an instruction to perform an interactive operation on the first video (corresponding to performing an interactive operation on the first video), the electronic device overlays a video collection control at the bottom of the display area of the first video. The interactive operation instruction for the first video may be triggered by a like control, comment control, or share control based on the first video.
[0072] In other words, responding to a "like" action on the first video can display a video collection control. Responding to a comment command on the first video can result in a comment being posted on the first video, and the video collection control will also be displayed. Responding to a share command on the first video can share the first video to the object corresponding to the share command, and the video collection control will also be displayed. In other words, the interactive actions can be liking, commenting, and sharing the first video; of course, other settings can be configured according to actual usage needs, which will not be elaborated upon here.
[0073] In some embodiments, the friend objects of the current object perform interactive operations on the first video. That is, the friend objects corresponding to the friend accounts added by the current account used by the current object perform interactive operations on the first video during the viewing process. The interactive operations can be referred to the foregoing description and will not be repeated here.
[0074] In some embodiments, determining the current popularity of a first video when its current popularity exceeds a popularity threshold includes: acquiring behavioral data of the first video within the current statistical period, the behavioral data including at least one of the following: play counts, exposure counts, like counts, comment counts, share counts, favorite counts, completed playback counts, click counts, and follower conversion counts; determining at least one popularity indicator based on the behavioral data; and determining the current popularity of the first video based on the at least one popularity indicator. For example, the current popularity of the first video can be obtained by weighted summation of at least one popularity indicator.
[0075] In some embodiments, the terminal loads and plays the first video in the media streaming interface. At the same time, the terminal sends the unique identifier information of the first video to the server. The server receives the unique identifier information of the first video and checks whether the first video belongs to a completed video collection. If so, the server sends a video collection existence identifier to the terminal. If not, the server sends a no-collection identifier to the terminal. The terminal does not load the video collection control during the subsequent playback of the first video to ensure an uninterrupted user experience.
[0076] After receiving the collection's existence identifier, the terminal starts a local timer to record the playback duration of the first video. The terminal then determines whether the playback duration of the first video exceeds a pre-configured duration threshold. If so, the terminal renders and displays the video collection control in the preset coordinate area of the media stream interface. If not, the terminal determines whether it has received an interactive operation to execute for the first video.
[0077] If an interactive operation is received, the terminal renders and displays the video collection control in the preset coordinate area of the media stream interface; if no interactive operation is received, the terminal determines whether the playback progress of the first video has reached 100%. If it has, the terminal renders and displays the video collection control in the preset coordinate area of the media stream interface; if it has not, the terminal continues to determine whether the playback duration of the first video exceeds a pre-configured duration threshold.
[0078] The video recommendation method provided in this application limits the triggering time of displaying the video collection control. Specifically, the terminal displays the video collection control only when the duration of the first video exceeds a duration threshold or when an interactive operation is performed on the first video. After confirming that the current object has sufficient viewing interest or interaction intention for the first video, the video collection control is accurately provided. This reduces visual interference from invalid information, improves the accuracy of the terminal in displaying related data, and thus significantly improves the efficiency of human-computer interaction.
[0079] In related technologies, when electronic devices continuously request, parse, and render relevant video collection data and corresponding controls in the background, they consume the computing resources and memory space of the electronic device's central processing unit. This lack of state isolation mechanism in the interaction flow logic can easily cause the electronic device to experience screen stuttering or frame rate drops when displaying the first video, reducing the system resource utilization and human-computer interaction efficiency of the electronic device.
[0080] To address the aforementioned technical issues, when the first video is displayed, if a video collection including the first video exists, a collection switch is displayed. This collection switch is used to enable a video collection mode. Video collection mode refers to a specific operating state where the electronic device allows the underlying video collection data flow, parsing, and front-end user interface control rendering. Displaying a video collection control includes: displaying a video collection control when the video collection mode is enabled based on the collection switch.
[0081] In some embodiments, a collection switch is displayed. If the collection switch indicates that the video collection mode is enabled, the video collection mode can be disabled in response to a trigger operation on the collection switch. The electronic device hides the video collection control in the media streaming interface and no longer displays pixel rendering of the video collection control in the layer. If the collection switch indicates that the video collection mode is not enabled, the video collection mode can be enabled in response to a trigger operation on the collection switch.
[0082] See Figure 5 , Figure 5 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 2The video collection switch 501 is displayed in the media stream interface. When the video collection mode is turned off based on the video collection switch 501, the video collection control 502 is not displayed. When the video collection mode is turned on based on the video collection switch 501, the video collection control 502 is displayed.
[0083] In some embodiments, the server receives a request for a first video from the terminal. The server extracts the title features of the first video, uses a multi-channel parallel recall strategy to obtain a set of candidate videos, and inputs the feature data of the candidate video set into a large language model. The server uses the large language model to determine whether the videos in the candidate video set are consistent with the first video. If yes, the server determines that a video set including the first video exists and sends the set data to the terminal, carrying a set existence identifier and a three-dimensional sorting tuple (including a type isolation bit, an integer value for the number of episodes, and a publication timestamp). If no, the server determines that no video set including the first video exists and sends empty identifier data to the terminal. The terminal then prevents the drawing of the set switch in the media stream interface.
[0084] The terminal parses the collection's identifier and draws and displays the collection switch in a preset layer of the screen coordinate system. The terminal continuously determines the control variable value of the collection switch and checks if it matches the preset enable value. If so, the terminal determines that the video collection mode is enabled, parses the received collection data, and renders and displays the video collection control on the front-end interface. If not, the terminal determines that the video collection mode is not enabled, releases the memory space reserved for the video collection control, and skips the rendering step of the video collection control.
[0085] The video recommendation method provided in this application displays a collection switch when it is determined that a video collection including the first video exists, and displays a video collection control when the video collection mode is enabled based on the collection switch; however, the video collection control is not displayed when the collection switch is disabled. A state isolation mechanism is constructed through the collection switch, allocating memory resources and central processing unit rendering threads only when the video collection mode is explicitly enabled, avoiding invalid and redundant control data loading, reducing the system hardware resource utilization of the electronic device, thereby ensuring the smoothness of the first video playback and improving the underlying operating performance of the electronic device.
[0086] See also Figure 3 In step 102, in response to a video viewing command triggered by the video collection control, video information of each second video is displayed.
[0087] In some embodiments, step 102 can be implemented in the following way: in response to a trigger operation on the video collection control, a video viewing instruction can be triggered; in response to the video viewing instruction, video information of each second video can be displayed, wherein the video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to the multiple video information is different.
[0088] Video information includes at least one of the following: basic video information, video content information, video association information, video status information, and video interaction information. Basic video information is used to identify the second video and display its basic attributes, such as at least one of the following: video identifier, video title, video cover, video description, video duration, video publication time, video publishing account (i.e., the account used by the video publisher), the video's collection identifier, and the video's sorting information within the video collection.
[0089] Video content information is used to characterize the video content of the second video, and includes at least one of the following: video tags, video categories, video themes, video keywords, video summaries, video clip descriptions, video subtitle information, video audio information, and video frame description information. Video association information is used to characterize the association between the second video and other videos or video collections, and includes at least one of the following: the identifier of the first video corresponding to the second video, the association between the second video and the first video, information about the video collection to which the second video belongs, information about the preceding videos, subsequent videos, similar videos, and related recommended videos.
[0090] Video status information is used to characterize the playback status and availability status of the second video, and includes at least one of the following: playback status, viewing progress, whether it has been watched, whether it is playable, whether it is expired, whether it has been removed from the platform, whether it is cached, whether it supports skipping to playback, and whether it supports continuous playback. Video interaction information is used to characterize the user interaction or dissemination status of the second video, and includes at least one of the following: number of plays, number of likes, number of comments, number of shares, number of favorites, number of reposts, number of bullet comments, number of completed plays, popularity value, popularity ranking, user like status, user favorite status, and user follow status.
[0091] In related technologies, in response to a trigger operation targeting the collection entry point, a network request for all associated video data is initiated, and pixel rendering logic for a large-scale list structure is immediately executed on the front end. This flawed interactive logic, which directly loads a large amount of video information with a single click, causes electronic devices to face extremely high data parsing pressure and image rendering load instantly. This leads to a surge in CPU usage and significant invalid memory occupation, easily causing interface lag or crashes, severely limiting the hardware performance and data processing efficiency of electronic devices.
[0092] To address the aforementioned technical issues, before displaying the video information of each second video, in response to a trigger operation on the video collection control, collection information of the video collection is displayed, and a first viewing control for viewing the video information of the second videos is also displayed. The collection information includes at least one of the following: the name of the video collection, the cover image of the video collection, and a brief description of the video collection.
[0093] The collection information of a video collection may include at least one of the following: basic collection information, collection content information, collection ownership information, collection organization information, collection status information, collection interaction information, and collection recommendation information.
[0094] Basic collection information is used to identify video collections and display their basic attributes, such as collection identifier, collection name (i.e., the name of the video collection), collection cover (i.e., the cover of the video collection), collection description (i.e., the description of the video collection), collection creation time, collection update time, collection release time, collection release account, collection creation account, collection source, collection type, collection tags, and collection category.
[0095] Collection content information is used to characterize the video content contained in the video collection. For example, it includes at least one of the following: number of videos in the video collection, video list, information on the first video, information on the latest video, information on representative videos, information on highlights videos, video summary, collection theme, collection keywords, collection content description, collection subtitle information, collection audio information, and collection screen description information.
[0096] Collection attribution information is used to characterize the attribution relationship between videos and video collections. For example, it includes at least one of the following: collection identifier to which the video belongs, collection name to which the video belongs, sorting information of the video in the video collection, time when the video was added to the video collection, association relationship between the video and the video collection, whether the video is the first video in the video collection, whether the video is the latest video in the video collection, and whether the video is the representative video in the video collection.
[0097] Collection organization information is used to characterize the organization of videos within a video collection. For example, it includes at least one of the following: video collection sorting method, video collection grouping method, video collection chapter information, video collection directory information, video collection update rules, video collection playback order, video collection continuous playback rules, and video collection jump rules.
[0098] Collection status information is used to characterize the availability, update status, or current viewing status of a video collection. For example, it includes at least one of the following: whether the video collection is accessible, playable, expired, removed, completed, continuously updated, subscribed, favorited, watched, viewing progress, recently watched videos, recent viewing time, and number of unwatched videos.
[0099] The collection interaction information is used to characterize the user interaction or dissemination of the video collection, such as at least one of the following: number of plays, number of exposures, number of clicks, number of likes, number of comments, number of shares, number of favorites, number of subscriptions, number of reposts, number of completed plays, number of consecutive plays, popularity value, popularity ranking, and recommendation index.
[0100] The collection recommendation information is used to characterize the basis for displaying video collections in recommendation or display scenarios. For example, it includes at least one of the following: the recommendation reason corresponding to the video collection, recommendation tag, recommendation scenario, recommendation source, recommendation weight, tag matching the user's (i.e., object's) interests, the degree of relevance between the collection and the currently playing video, and the degree of matching between the collection and the current user's (i.e., the current object's) historical behavior.
[0101] After displaying the first viewing control, in response to the triggering operation of the first viewing control, a video viewing command is triggered, which in turn can respond to the video viewing command triggered based on the video collection control to display the video information of each second video.
[0102] See Figure 6 , Figure 6 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 3 In response to a trigger operation on the video collection control 601, the collection information of the video collection is displayed, and a first viewing control 602 for viewing the video information of the second video is displayed. The cover of the video collection is a certain image frame in the first video. The cover of the video collection can be preset for the current object, or it can be captured from the image frame corresponding to the video in the video collection, which is reasonable.
[0103] See Figure 7 , Figure 7 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 4 , undertake Figure 6 In response to a trigger operation on the first viewing control 602, a video viewing instruction is triggered. In response to a video viewing instruction triggered based on the video collection control, video information of each second video is displayed in the area indicated by area 701.
[0104] In some embodiments, in response to a trigger operation on the video collection control, playback of the first video is paused, and collection information of the video collection in a full-screen overlay style is displayed; simultaneously, a first viewing control for viewing video information of the second videos is displayed. The collection information includes the name of the video collection, the cover image of the video collection, and a description of the video collection. The cover image of the video collection is displayed as the interface background layer, the name of the video collection is displayed in a highlighted font size overlaid on the cover image, and the description text of the video collection is displayed below the name. In response to a trigger operation on the first viewing control, the electronic device triggers a video viewing command. In response to the video viewing command, the display of the cover image and the description of the video collection is canceled, and a list view pops up in the media stream interface. The video information of each second video is rendered and displayed in the list view.
[0105] In some embodiments, in response to a trigger operation on the video collection control, the terminal generates a collection metadata request and sends it to the server. The server receives the collection metadata request, searches the database, and determines whether a video collection to which the first video belongs exists in the database. If so, the server extracts the name of the video collection, the cover image of the video collection, and a description of the video collection and sends them to the terminal; otherwise, the server calls a large language model to generate a description of the video collection in real time using the multimodal features of the first video, and sends the real-time generated description of the video collection to the terminal.
[0106] The terminal receives data from the server, allocates initial video memory, and renders and displays the collection information and the first viewing control within that memory. During the display of the collection information and the first viewing control, the terminal suspends the batch loading thread for the video information of each second video, blocking the underlying list data request logic. The terminal determines whether it has received a trigger operation for the first viewing control. If so, the terminal triggers a video viewing instruction and requests the second video data set sorted by multi-dimensional tuple mapping from the server; otherwise, the terminal maintains its current rendering state and does not perform any additional data loading tasks.
[0107] The video recommendation method provided in this application, before displaying the video information of each second video, responds to the trigger operation of the video collection control, prioritizes displaying collection information including the name of the video collection, the cover of the video collection, and the description of the video collection, and displays the first viewing control. Finally, it responds to the trigger operation of the first viewing control to trigger the video viewing instruction. This is equivalent to introducing the first viewing control as a buffer node for data loading, strictly separating the presentation of coarse-grained information of the video collection from the request for fine-grained underlying list data. The first viewing control constructs a state isolation mechanism, avoiding the execution of invalid underlying video information data parsing and pixel rendering when the current object has no clear intention to watch in depth. This significantly reduces the unnecessary consumption rate of system hardware resources of electronic devices, effectively reduces memory load, and improves the overall operating stability and underlying rendering efficiency of electronic devices.
[0108] In related technologies, when recommending video content in a media stream interface, if the first video simultaneously possesses multi-dimensional content attributes, resulting in the first video being associated with multiple video sets of different types, the terminal often adopts either the logic of directly loading and displaying a single type of data by default, or the logic of forcibly loading all related data to the front-end cache. The logic of directly loading a single type by default lacks a mechanism to judge the true intent of the triggering entity, easily leading to the displayed data deviating from actual needs; while the flawed interaction flow logic of full loading causes electronic devices to face extremely high data request and parsing pressure instantly, not only needlessly consuming the electronic device's running memory and processor computing resources, easily causing interface lag, but also severely reducing the human-computer interaction efficiency of the electronic device.
[0109] To address the aforementioned technical issues, when the first video belongs to different candidate video collections, and these different candidate video collections correspond to different types, in response to a trigger operation on the video collection control, multiple candidate types are displayed. The candidate type is the type of the candidate video collection to which the first video belongs. Among the multiple candidate types is the target type of the video collection. In response to a trigger operation on the target type among the multiple candidate types, the collection information of the video collection of the target type is displayed.
[0110] In some embodiments, the first video belongs to different candidate video sets, and different candidate video sets correspond to different types. A candidate video set can be a collection of videos that includes the first video and aggregates the first video and other videos according to a preset organizational dimension. The type of a candidate video set can be used to characterize the organizational dimension, content attributes, or display scenario of the candidate video set. For example, the type of a candidate video set can include at least one of the following: work type, role type, theme type, course type, knowledge point type, learning stage type, question type, destination type, itinerary type, product type, brand type, event type, topic type, training plan type, training part type, and account series type.
[0111] For example, the first video is a film and television video, and its content is "explanation of a movie clip." This first video can be categorized into four candidate video collections: works, characters, themes, and commentary series. Specifically, a works-related candidate video collection could be a "complete movie commentary collection," used to aggregate videos related to the same film or television work; a character-related candidate video collection could be a "highlight clips of a character collection," used to aggregate videos related to the same character; a theme-related candidate video collection could be a "collection of emotional movie commentaries," used to aggregate videos with the same narrative theme or content style; and a commentary series candidate video collection could be a "classic movie commentary series," used to aggregate videos from the same account or program series.
[0112] See Figure 8 , Figure 8 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 5 When the first video is "The process of making tomato beef brisket", in response to the trigger operation of the video collection control, multiple candidate types are displayed, including home cooking tutorial 801, beef dishes 802 and group meal dishes 803. Then, in response to the trigger operation of the target type among the multiple candidate types, the collection information of the video collection of the target type is displayed.
[0113] In some embodiments, the server receives a network request from the terminal containing the identifier information of a first video, initiates a multi-channel recall strategy using the multimodal features of the first video, obtains an associated set of video features, and determines whether the recalled set of video features contains heterogeneous feature data. If so, the server inputs the multimodal features into a large language model, and through the semantic understanding mechanism of the large language model, determines whether the first video belongs to different candidate video sets. The server extracts different types corresponding to different candidate video sets and sends them to the terminal as multiple candidate types; otherwise, the server returns an empty type identifier to the terminal. The terminal receives the data sent by the server.
[0114] In response to a trigger operation on the video collection control, the terminal instantiates and displays multiple candidate types on the front-end interface. The terminal continuously determines the touch events acting on the interface coordinate system and checks whether the coordinates of the touch events are within the response area of the target type. If so, the terminal packages the parameter information of the target type and sends it to the server. The server sorts the data using a multidimensional tuple mapping algorithm and returns the collection information of the video collection of the target type. The terminal then displays the collection information of the video collection of the target type on the interface. If not, the terminal maintains the display state of multiple candidate types.
[0115] The video recommendation method provided in this application addresses the situation where a first video belongs to different candidate video sets and different candidate video sets correspond to different types. In response to a trigger operation on a video set control, it displays multiple candidate types of the candidate video set to which the first video belongs. Then, in response to a trigger operation on a target type among the multiple candidate types, it accurately displays the set information of the video set of the target type. This introduces an intermediate state buffer mechanism for candidate type selection, strictly separating coarse-grained type guidance from fine-grained data rendering. The terminal only loads the set information of the video set of the target type after capturing a clear intent instruction. This effectively avoids redundant network requests and invalid underlying pixel rendering, significantly reducing the instantaneous data processing load of electronic devices and improving the accuracy of information distribution and the underlying operating performance of electronic devices.
[0116] In some embodiments, while displaying collection information of a video collection of the target type, in response to a type switching instruction, the collection information of the video collection of the target type is de-displayed in the display layer, and collection information of video collections of other candidate types besides the target type is displayed instead.
[0117] In related technologies, when presenting multiple candidate types in a media streaming interface, only plain text identifiers for the candidate types are typically provided. There is a lack of a pre-display mechanism for key parameters such as underlying popularity and timeliness characteristics of each candidate type. This leads to the terminal performing numerous invalid blind clicks and repeated attempts after the interface is displayed due to a lack of information for reference. Consequently, the terminal is forced to frequently respond to invalid touch signals and repeatedly execute network requests for underlying heterogeneous video list data and interface redrawing logic. This significantly increases the unproductive computational burden on the electronic device's central processing unit, causing the device's RAM to be occupied by a large amount of useless cached data, resulting in a substantial decrease in the system resource utilization of the electronic device.
[0118] To address the aforementioned technical issues, for each candidate type, at least one of the following information is displayed: the number of people watching videos in the video collection for the candidate type, the recommendation level of the video collection for the candidate type, the update time of the video collection for the candidate type, the type name of the candidate type, the type description of the candidate type, the collection name of the video collection for the candidate type, the collection cover, the collection introduction, the number of videos, the number of videos watched, the number of videos not watched, the viewing progress, the sorting position of the first video in the video collection, the associated position of the first video in the video collection, and the total duration of the video collection. The following information is considered to assist users in determining their target video category from multiple candidate categories: length, estimated viewing duration, latest video title, latest video cover, latest video release time, video collection update frequency, video collection completion status, video collection popularity score, video collection popularity ranking, number of plays, number of likes, number of comments, number of shares, number of favorites, number of subscriptions, number of completed plays, number of recently viewed videos, relevance to the first video, match with the current user's interests, reason for recommending the video collection, content tags of the video collection, applicable scenarios for the video collection, quality rating of the video collection, and the account that published the video collection. This information is used to assist users in determining their target category from multiple candidate categories.
[0119] In some embodiments, the recommendation level for a video collection of candidate types is determined as follows: Feature information of the video collection of candidate types, feature information of the first video, and object feature information of the current object are obtained; based on the feature information of the video collection of candidate types, the feature information of the first video, and the object feature information of the current object, at least one of the following is determined: the correlation between the video collection of candidate types and the first video; the matching degree between the video collection of candidate types and the current user's interests; the popularity of the video collection of candidate types; the freshness of the video collection of candidate types; the viewing cost of the video collection of candidate types; the continuous viewing benefit of the video collection of candidate types; and the historical selection rate of the video collection of candidate types; the recommendation level is determined based on at least one of the following: correlation, matching degree, popularity, freshness, viewing cost, continuous viewing benefit, and historical selection rate.
[0120] In other words, the degree of recommendation can be determined based on at least one of the following dimensions: the relevance between the candidate video collection and the first video, the matching degree between the candidate video collection and the current user's interests, the popularity of the candidate video collection, the freshness of the candidate video collection, the quality score of the candidate video collection, the viewing cost of the candidate video collection, the continuous viewing benefit of the candidate video collection, and the historical selection rate of the candidate video collection.
[0121] The relevance between the candidate video collection and the first video can be determined based on at least one of the following: tag similarity, theme similarity, category consistency, publishing account consistency, the first video's ranking position within the video collection, and content similarity between the first video and other videos within the video collection. The matching degree between the candidate video collection and the current user's interests can be determined based on the matching degree between the video collection's tags, themes, categories, or publishing accounts and the user's historical viewing tags, historical viewing types, followed accounts, saved videos, and liked videos. The freshness of the video collection can be determined based on the time interval between the video collection's update time and the current time; the shorter the time interval, the higher the freshness. The viewing cost of the video collection can be determined based on the number of videos in the collection, the number of unwatched videos, the total duration, or the estimated viewing duration; the lower the viewing cost, the higher the recommendation level; or, if the user has a preference for in-depth viewing, the higher the viewing cost, the higher the recommendation level. Continuous viewing benefits can be determined based on the continuity of the video sorting within the video collection, the completeness of the chapters, the number of unwatched videos, the user's viewing progress, and whether the video collection supports continuous playback. The historical selection rate can be determined based on the percentage of times multiple users or the current user selects a particular candidate type after multiple candidate types have been displayed.
[0122] In some embodiments, the content score, user score, timeliness score, and interaction score of the video collection of candidate types can be determined first, and then the recommendation level can be determined based on the content score, user score, timeliness score, and interaction score. The content score can be determined based on the tag similarity, topic similarity, and classification consistency between the video collection of candidate types and the first video; the user score can be determined based on the degree of matching between the video collection of candidate types and the current user's historical viewing behavior; the timeliness score can be determined based on the update time, update frequency, or completion status of the video collection of candidate types; and the interaction score can be determined based on at least one of the following: number of viewers, number of plays, number of likes, number of comments, number of shares, number of favorites, number of subscriptions, and number of completed plays of the video collection of candidate types.
[0123] Recommendation level can be displayed in the form of numerical values, ratings, tags, or sorting indicators. For example, recommendation level can be displayed as a recommendation score, recommendation star rating, recommendation level, recommendation tags such as "highly recommended," "relatively recommended," and "generally recommended," or as text information such as "Reason for recommendation: related to the video you are watching," "Reason for recommendation: recently updated," "Reason for recommendation: chosen by most users," and "Reason for recommendation: you often watch this type of video."
[0124] See Figure 9 , Figure 9 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 6 , undertake Figure 8For Home-Style Cooking Tutorial 801, it can display the number of people who have watched the videos in the video collection of Home-Style Cooking Tutorial 801. Figure 9 The text in the middle indicates "X objects are watching". For beef dish 802, it shows that the recommendation level of the video collection for beef dish 802 is high. For group meal dish 803, it shows that the video collection for group meal dish 803 was updated 1 minute ago.
[0125] The video recommendation method provided in this application provides that, for each candidate type, at least one of the following is displayed in advance: the number of objects watching videos in the video collection of the candidate type, the recommendation degree of the video collection of the candidate type, and the update time of the video collection of the candidate type. Decision-aiding parameters are output synchronously during the display of the type identifier, effectively guiding the accurate capture of trigger operations for the target type. This avoids invalid network communication concurrent requests and redundant graphics rendering pipeline calls from electronic devices, significantly reducing the unnecessary consumption of underlying hardware computing resources of electronic devices and improving the resource utilization of electronic devices.
[0126] In related technologies, when displaying video collection controls in a media stream interface, there is a lack of a visual identification mechanism for the historical state of control interactions and a quick return path across interface levels. Because the current object cannot intuitively identify whether the control is in a triggered state, it is very easy to generate repeated clicks on the same entry point. As a result, the terminal frequently responds to redundant touch signals and repeatedly executes network requests for underlying video list data, as well as re-decoding the video image and redrawing the interface logic. This leads to a large amount of ineffective occupation of the central processing unit computing resources of electronic devices, serious consumption of running memory, and a significant reduction in the underlying processing efficiency of electronic devices.
[0127] To address the aforementioned technical issues, the video collection information is displayed on the information display interface. In other words, in response to a trigger operation on the video collection control, the information flow interface can be switched to the information display interface, where the video collection information is displayed. Figure 3 The display of the video collection control in step 101 shown can be implemented in the following way: the video collection control is displayed using a first display style. The first display style is used to guide and trigger the video collection control. The first display style is used to indicate that the display state of the video collection control in the media stream interface is the guided and triggered state.
[0128] In some embodiments, a first video that is playing can be displayed in the media stream interface. In response to a trigger operation on a video collection control for a first display style, the playback of the first video is paused, and a playback control (i.e., a continue playback control) for playing the first video is displayed in the information display interface. In response to a trigger operation on the playback control, the media stream interface is returned to continue playing the first video, and the video collection control is displayed using a second display style.
[0129] The second display style is different from the first display style. The second display style is used to indicate that the video collection control has been triggered. The second display style is used to indicate that the display state of the video collection control in the media stream interface has been switched from the guided trigger state to the normal display state. The different display styles can be characterized by one or more of the following: whether it is bold, display color, whether an underline is added, and whether there are display effects. These will not be elaborated here.
[0130] See Figure 10 , Figure 10 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 7 The video collection control 1001 can be displayed using the first display style (correspondingly bold). See also... Figure 11 , Figure 11 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 8 The second display style (corresponding to the video collection not being bolded and displayed in italics) can be used to display the video collection control 1001.
[0131] In some embodiments, the terminal initializes local state parameters for the video collection control in the main rendering thread of the media stream interface. The terminal determines whether a trigger flag for the corresponding video collection control is recorded in the local cache. If yes, the terminal calls the rendering code block corresponding to the second display style in the interface component library and displays the video collection control using the second display style. If no, the terminal calls the rendering code block corresponding to the first display style and displays the video collection control using the first display style. The terminal captures the trigger operation for the video collection control using the first display style, suspends the underlying decoding process of the first video in the media stream interface, allocates new independent video memory space for loading the graphics data stream of the information display interface, and renders and displays the playback control in the specified coordinate area of the information display interface, continuously determining the touch events acting on the playback control area. In response to the trigger operation for the playback control, the terminal writes the trigger flag to the local state parameters. The terminal destroys the independent video memory space occupied by the information display interface, resumes the underlying decoding process of the first video to return to the media stream interface to play the first video; at the same time, based on the updated local state parameters, the terminal triggers the partial redraw mechanism of the media stream interface and displays the video collection control using the second display style.
[0132] This application's embodiments employ a first display style to display the video collection control, and display a playback control for playing the first video in the information display interface. In response to a trigger operation on the playback control, the system returns to the media stream interface to play the first video. A second display style, different from the first, is also used to display the video collection control. This establishes a clear control state transition mechanism and a memory view switching channel in the front-end graphical user interface. The second display style provides a clear basis for judging the interaction history of the current object, effectively preventing invalid repeated touch signals for the same video collection control. Simultaneously, the playback control enables rapid destruction of the interface and restoration of the first video decoding thread, significantly reducing the invalid computational overhead caused by re-requesting underlying data and repeatedly initializing the video decoder, lowering the memory consumption rate of the electronic device, and improving the working efficiency of the underlying computing hardware.
[0133] In related technologies, it is impossible to provide layered front-end interaction guidance based on the actual viewing depth of the current object, resulting in invalid swiping and search operations. The terminal frequently responds to redundant touch commands, repeatedly loading data from the underlying heterogeneous video list and redrawing the entire interface pixel by pixel. This increases the undue computational burden on the electronic device's central processing unit, causing RAM to be occupied by redundant cached data for extended periods, thus reducing the scheduling efficiency of the electronic device's underlying hardware resources and the efficiency of human-computer interaction.
[0134] In some embodiments, in response to a trigger operation on the video collection control, an information display interface is displayed. In the information display interface, the viewing progress of the current object in watching the video collection is displayed. When the viewing progress reaches a preset viewing progress, the virtual reward resources corresponding to the preset viewing progress are displayed, and a claim control for claiming the virtual reward resources is displayed. In response to a trigger operation on the claim control, the current object's account can be controlled to obtain virtual reward resources.
[0135] See Figure 12 , Figure 12 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 9 In the information display interface, the current object's viewing progress of the video collection can be displayed as 45%. When the viewing progress reaches 40% (corresponding to the preset viewing progress), the virtual reward resources corresponding to 40% can be displayed, and the claiming control 1201 for claiming the virtual reward resources can be displayed.
[0136] Virtual reward resources are digital rights certificates allocated to the current user based on established rules. These resources manifest as virtual trial cards for unlocking specific videos, platform digital points, or specific digital badges. Virtual reward resources can be displayed when showing the user's account information and can also be used to redeem access to different video collections; the specific settings can be customized to meet actual usage needs.
[0137] The visual representation of the viewing progress includes a horizontally filled progress bar and a percentage text that updates dynamically with the progress bar. The viewing progress of a video collection can be determined based on the viewing history of the current object for multiple videos in the video collection. The viewing history can include at least one of the following: the video identifier that the current object has watched, the viewing duration of each video, the total playback duration of each video, the viewing completion status of each video, the identifier of the most recently watched video, the most recently watched time, and the breakpoint playback position.
[0138] Determining the viewing progress of a video collection includes: retrieving multiple videos from the video collection; retrieving the viewing history of the current object for each video within the collection; determining the viewing progress of each individual video based on the viewing history; and determining the current object's overall viewing progress of the video collection based on the individual video viewing progress for each video. The individual video viewing progress can be determined based on the ratio of the current object's already watched time for that video to the total video duration of that video.
[0139] The viewing progress of a video collection can be determined based on the average viewing progress of each individual video, or the ratio of the current user's cumulative viewing time within the collection to the total duration of the collection, or the ratio of the number of videos already watched to the total number of videos in the collection, or the viewing progress of each individual video and its corresponding weight. Therefore, based on the current user's viewing history across multiple videos within the collection, the viewing progress of the current user within the collection can be accurately displayed.
[0140] In some embodiments, the terminal sends a trigger signal for a video collection control and the object identifier of the current object to the server. The server receives the trigger signal and uses a large language model to perform deep semantic extraction on the video collection nodes retrieved from multiple channels, obtaining a set of main episode sequences and a set of peripheral bonus content that constitute the video collection. The server queries the backend database based on the object identifier to extract the historical playback node data of the current object in the main episode sequence set. The server divides the historical playback node data by the total number of nodes in the main episode sequence set to calculate the viewing progress and sends it to the terminal. The terminal parses the instruction and renders and displays the viewing progress on the front-end interface. At the same time, the server determines whether the calculated viewing progress is greater than or equal to the preset viewing progress. If so, the server generates a distribution password for virtual reward resources corresponding to the preset viewing progress in the backend memory pool and sends an interface rendering instruction package containing the distribution password and the claim control to the terminal. After receiving the interface rendering instruction package, the terminal calls the graphics component library to display the virtual reward resources and the claim control on the interface. If not, the server sends a progress not met status code to the terminal, the terminal suspends the resource loading thread, and blocks the graphics rendering pipeline of the virtual reward resources and the claim control.
[0141] The video recommendation method provided in this application displays the viewing progress of the current object in the video collection when the video collection control is triggered. When the viewing progress reaches the preset viewing progress, the virtual reward resources corresponding to the preset viewing progress are displayed, and a control for claiming the virtual reward resources is displayed. By quantitatively presenting the viewing progress and triggering the virtual reward resources on demand in a node-based manner, a precise interactive guidance path is built on the front end, which effectively curbs the generation of invalid interactive operations, significantly reduces the redundant underlying data calculation overhead and graphics rendering pipeline call frequency generated by electronic devices to handle invalid operations, and improves the allocation efficiency and operational stability of the underlying hardware computing resources of electronic devices.
[0142] In related technologies, a fixed interaction logic is used to load and render all underlying video list data at once. When the amount of associated video data is large, it can cause electronic devices to instantiate too many interface elements in the same display layer. This not only causes interface information overload and low human-computer interaction efficiency, but also forces electronic devices to maintain extremely high concurrent network requests and graphics rendering load for a long time, needlessly consuming the computing resources of the electronic device's central processing unit, and causing memory overflow or interface response lag.
[0143] In some embodiments, Figure 3 Step 102 shown can be implemented in the following way: in response to a video viewing command triggered by a video collection control, a first video is displayed in the first display area of the media stream interface; if the number of second videos is less than or equal to a first quantity threshold, video information of each of the second videos is displayed in the second display area of the media stream interface.
[0144] If the number of second videos exceeds the first quantity threshold, multiple filter tags are displayed in the second display area of the media stream interface. Among the multiple filter tags is the target filter tag. In response to the trigger operation for the target filter tag, the video information of each second video after being filtered according to the target filter tag is displayed in the second display area.
[0145] In some embodiments, if the number of second videos exceeds a first quantity threshold, multiple filter tags can be generated according to at least one filtering criterion and displayed in the second display area of the media stream interface. The filtering criterion may include at least one of video sequence number, publication time, publication target, video theme, video duration, viewing status, update batch, video popularity, video type, video source, video chapter, and video relevance. Each filter tag corresponds to a subset of second videos in the second video set. In response to a trigger operation targeting a specific filter tag, video information of the second video corresponding to the target filter tag is displayed in the second display area.
[0146] When the filtering criterion is video sequence number, the filter tags can be different sequence number ranges. The video sequence number can be the ranking number of the second video within the video collection; for example, see [link to relevant documentation]. Figure 13 , Figure 13 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 When the video collection includes 50 second videos, multiple filter tags can include "1-10" (corresponding to the content indicated by 1301), "11-20", "21-30", "31-40", and "41-50".
[0147] In some embodiments, the filtering is based on publication time, with multiple filter tags corresponding to different publication time intervals. The publication time intervals can be divided by day, week, month, quarter, year, or a preset time period. For example, multiple filter tags may include at least two of "Today," "Last 7 Days," "Last 30 Days," "May 2026," "April 2026," and "2025." In response to a trigger operation targeting a specific filter tag, video information for each second video whose publication time falls within the publication time interval corresponding to the target filter tag is displayed.
[0148] In some embodiments, the filtering criterion is the publishing object, with multiple filtering tags corresponding to different publishing objects. Publishing objects may include at least one of the following: publishing account, publishing user, publishing organization, publishing media, publishing broadcaster, publishing lecturer, and publishing merchant. For example, multiple filtering tags may include at least two of "publishing account A," "publishing account B," "official account," "lecturer A," and "organization A." In response to a triggering operation targeting a specific filtering tag, video information for each second video published by the publishing object corresponding to the target filtering tag is displayed.
[0149] In some embodiments, the filtering is based on video themes, with multiple filter tags corresponding to different video themes. Video themes can be determined based on the video title, video tags, video category, video summary, video keywords, or video content recognition results of the second video. For example, multiple filter tags may include at least two of "knowledge explanation," "example problem analysis," "food," "tourist attractions," "strength training," and "stretching and relaxation." In response to a triggering operation targeting a specific filter tag, video information for each second video whose video theme matches the target filter tag is displayed. In some embodiments, the filtering is based on viewing status, with multiple filter tags corresponding to different viewing statuses of the current object for the second video. Viewing status can include at least one of "not watched," "watching," "completed," "recently watched," "favorited," and "liked." For example, multiple filter tags include at least two of "all," "not watched," "watching," "completed," and "recently watched." In response to a triggering operation targeting a specific filter tag, video information for each second video whose viewing status matches the target filter tag is displayed.
[0150] In some embodiments, the filtering is based on video duration, with multiple filter tags corresponding to different video duration ranges. For example, the multiple filter tags include at least two of "less than 1 minute", "1-3 minutes", "3-5 minutes", "5-10 minutes", and "more than 10 minutes". In response to a trigger operation targeting a specific filter tag, video information for each second video whose video duration falls within the video duration range corresponding to the target filter tag is displayed.
[0151] In some embodiments, the filtering is based on video popularity, with multiple filtering tags corresponding to different video popularity ranges or popularity sorting methods. Video popularity can be determined based on at least one of the following: number of plays, number of likes, number of comments, number of shares, number of favorites, number of completed plays, and popularity value. In response to a trigger operation targeting a specific filtering tag, video information for each second video whose popularity meets the corresponding condition of the target filtering tag is displayed.
[0152] In some embodiments, the filtering is based on video chapters, with multiple filter tags corresponding to different chapters, units, stages, or directory nodes in the video collection. For example, the multiple filter tags include at least two of "Chapter 1", "Chapter 2", "Chapter 3", "Week 1", "Week 2", "Day 1", and "Day 2". In response to a triggering operation targeting a specific filter tag, video information for each second video belonging to the chapter, unit, stage, or directory node corresponding to the target filter tag is displayed.
[0153] In some embodiments, the filtering is based on video type, with multiple filter tags corresponding to different video types. The video type can be determined based on the content format, playback purpose, or business attributes of the second video. In response to a trigger operation targeting a specific filter tag, video information for each second video whose video type matches the target filter tag is displayed.
[0154] In some embodiments, the filtering is based on the relevance to the first video, with multiple filtering tags corresponding to different relevance levels or relationship types. For example, the multiple filtering tags include at least two of "strongly relevant," "related recommendations," "same topic," "same author," and "same series." In response to a trigger operation targeting a specific filtering tag, video information for each second video whose relevance to the first video meets the conditions corresponding to the target filtering tag is displayed.
[0155] The video recommendation method provided in this application directly displays video information when the number of second videos is less than or equal to a first quantity threshold. When the number exceeds the first quantity threshold, multiple filter tags are introduced, and the filtered video information is displayed in response to triggering operations on the target filter tags. This constructs an adaptive dynamic rendering mechanism and view state isolation barrier based on data volume. When the data volume is large, the multiple filter tags act as a rendering buffer layer, breaking down the massive data rendering task into smaller parts, avoiding the instantiation of too many node controls in the same display layer by the electronic device. This significantly reduces the peak instantaneous memory usage and underlying graphics computation overhead of the electronic device, ensures the smoothness of the first video decoding and playback, and significantly improves the utilization rate of the device's underlying hardware performance and the interactive efficiency of information retrieval.
[0156] In related technologies, when processing large amounts of video data and generating associated navigation tags in media streaming interfaces, there is a lack of pre-defined constraints for external interaction to actively define the underlying data loading and classification dimensions. This causes electronic devices to perform data classification and tag rendering calculations on the entire multi-dimensional underlying video list data by default in the background. Electronic devices build a large number of redundant and ineffective cache tree structures in their local memory, needlessly consuming the computing resources of the electronic device's central processing unit and the bandwidth of the graphics rendering pipeline, severely reducing the effectiveness and efficiency of the underlying hardware data processing.
[0157] To address the aforementioned technical issues, in response to a video viewing command triggered by the video collection control, a settings control is displayed before showing the video information of each second video. This setting control is used to configure the filtering criteria for the second videos. The filtering criteria can also be called filtering dimensions. Filtering dimensions are rule parameters for attribute isolation or dimension filtering of videos in the video collection. Filtering dimensions can be found in the aforementioned explanation. Filtering criteria can also include video type isolation parameters (such as including only the main feature or only the extras), release time interval parameters, etc., which can be set according to actual usage requirements.
[0158] In some embodiments, in response to a setting instruction triggered based on a setting control, if the number of second videos is greater than the first quantity threshold, the filtering criteria set by the setting instruction are displayed, and multiple filtering labels after filtering the second videos according to the filtering criteria are displayed in the second display area of the media stream interface.
[0159] See Figure 14 , Figure 14 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 First, the setting control 1401 can be displayed. In response to the setting command triggered by the setting control, if the number of second videos is greater than the first quantity threshold, the filtering criteria (video publishing object) set by the setting command will be displayed. In the second display area of the media stream interface, multiple filter labels after filtering the second videos according to the filtering criteria will be displayed, corresponding to object 1, object 2 and object 3.
[0160] In some embodiments, in response to a triggering operation on a setting control, at least one candidate filtering criterion can be displayed, and in response to a triggering operation on a target filtering criterion among the at least one candidate filtering criterion, a setting instruction can be triggered, wherein the filtering criterion set by the setting instruction is the target filtering criterion.
[0161] In some embodiments, the terminal captures setting instructions triggered by setting controls, extracts the filtering criteria set by the setting instructions, and reports the data packet containing the filtering criteria to the server. The server utilizes the semantic understanding mechanism of a large language model to pre-extract the chain attributes of each second video within the video collection. Based on the filtering criteria, it performs a data filtering algorithm on the second videos in the background database, removing heterogeneous second video data that does not match the filtering criteria. It then calls a counter to count the number of filtered second videos and determines whether the number of filtered second videos exceeds a first quantity threshold. If so, the server calls multidimensional tuple mapping logic, reads the integer set number or publication timestamp of the remaining second videos, performs bucket aggregation, generates structured data corresponding to multiple filtering tags that match the filtering criteria, and sends a data stream containing the filtering criteria string and structured data to the terminal. The terminal parses the data stream and displays the filtering criteria on the front end, as well as multiple filtering tags after filtering the second videos according to the filtering criteria in the second display area. If not, the server directly extracts the video information of the filtered second videos and sends it to the terminal. The terminal blocks the rendering threads of multiple filtering tags and directly renders the video information of each second video.
[0162] The video recommendation method provided in this application embodiment can display a setting control before displaying the video information of each second video. In response to a setting instruction triggered by the setting control, if the number of second videos is greater than a first quantity threshold, the filtering criteria set by the setting instruction are displayed, and multiple filtering tags after filtering the second videos according to the filtering criteria are displayed in the second display area. By advancing the isolation operation of data processing through the filtering criteria, redundant heterogeneous video data is eliminated from the source, effectively avoiding the generation of underlying filtering calculations and invalid tag nodes that do not conform to the interaction intent. This greatly reduces the length of the rendering buffer queue in the system memory, reduces the underlying hardware resource load of electronic devices, and significantly improves the memory scheduling efficiency of electronic devices.
[0163] In related technologies, after video information is removed, the terminal either directly displays a blank area or forces the terminal to trigger a full data refresh request in the background for each removal operation. Frequent triggering of redundant network concurrent requests and full redraw logic of the front-end interface needlessly consumes the computing resources and running memory of the electronic device's central processing unit, significantly reducing the effective flow efficiency of the electronic device's underlying resources.
[0164] In some embodiments, after displaying the video information of the second video, in response to a cancellation display instruction for the video information of the target second video among a plurality of second videos, the display of the video information of the target second video is cancelled. When the number of video information of the second videos that have not been cancelled (i.e., the remaining video information of the second videos) decreases to a second quantity threshold, an increase control is displayed. The remaining video information of the second videos is the video information of the currently displayed second videos excluding the video information of the target second video.
[0165] See Figure 15 , Figure 15 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 Second, when the number of video information of the second video that has not been canceled decreases to a second quantity threshold, the add control 1501 is displayed. The add control 1501 is used to add the video to be displayed based on the content of the second video that has not been canceled.
[0166] In some embodiments, a corresponding cancel display control can be displayed for each second video's video information. In response to the triggering operation of the cancel display control corresponding to the target second video, a cancel display instruction is triggered for the video information of the target second video among multiple second videos.
[0167] In some embodiments, a cancel display control is displayed for video information of multiple second videos. In response to a trigger operation on the cancel display control, the video information of the second video can be controlled to be in a candidate state. In response to a trigger operation on the video information of a target second video, the video information of the target second video can be controlled to be in a selected state. In response to a confirmation instruction on the selected target second video, the confirmation instruction can be used as a cancel display instruction on the video information of the target second video among the multiple second videos.
[0168] In some embodiments, the terminal maintains a global counter variable locally. In response to a cancellation command for the video information of the target second video, the global counter variable is decremented by one. The terminal then determines whether the number of uncancelled second video information items recorded in the global counter variable is less than or equal to a second quantity threshold. If so, the terminal calls the underlying graphical user interface library to instantiate and display an add control at a preset coordinate on the screen. In response to a trigger operation on the add control, the terminal extracts the remaining multimodal feature data of the second video and packages it for transmission to the server.
[0169] The server uses a large language model to perform deep semantic extraction on the multimodal feature data of the remaining second video, constructs a dynamic preference association chain, initiates multi-path recall based on the dynamic preference association chain, generates a new heterogeneous video data stream through a multidimensional tuple mapping algorithm and sends it to the terminal. The terminal parses the heterogeneous video data stream to increase the number of videos in the video collection; otherwise, the terminal maintains the current value of the global counter variable, blocks the graphics processing thread that instantiates and adds controls, and suspends the network communication port that requests new data from the server.
[0170] In some embodiments, the content features of the second video that has not been canceled from display are obtained. These content features include at least one of the following: video title, video tags, video category, video theme, video keywords, video summary, subtitle text, audio recognition text, image recognition result, publishing account, publishing time, video duration, video popularity, and user interaction data. Based on the content features of the second video that has not been canceled from display, candidate videos matching the second video that has not been canceled from display are determined from a candidate video set. Based on at least one of the following: content similarity between the candidate video and the second video that has not been canceled from display, popularity of the candidate video, publishing time of the candidate video, quality score of the candidate video, and matching degree between the candidate video and the current object's interests, a video to be added is determined from the candidate videos. The video information of the video to be added is then added and displayed in the second display area. The candidate video set includes at least one of the following: videos not displayed in the current video set, videos canceled from display in the current video set, videos in other video sets related to the current video set, videos related to the first video, videos related to the second video that has not been canceled from display, videos not watched by the current object, and videos matching the current object's historical viewing behavior.
[0171] In some embodiments, content similarity can be determined based on at least one of the following: title similarity, tag similarity, category consistency, theme similarity, keyword similarity, subtitle text similarity, audio text similarity, screen content similarity, and publishing account consistency between the videos to be compared.
[0172] The video recommendation method provided in this application, after canceling the display of the target second video in response to a cancellation display command, determines the number of second video information that has not been canceled. An add control is displayed only when the number drops to a second threshold. This add control is used to add more videos to the display, effectively avoiding redundant network communication requests and repetitive interface pixel redrawing caused by a single removal operation. Simultaneously, targeted data supplementation calculations are performed based on the remaining second videos, improving the accuracy of extracting heterogeneous video data on the same theme in the background, significantly reducing the invalid information processing load on the underlying hardware, and significantly improving the resource scheduling efficiency of the electronic device's local memory and graphics rendering pipeline.
[0173] In related technologies, when triggering the addition of associated video data in a media stream interface, an unconditional, direct retrieval logic is used. This flawed interactive process logic, lacking a pre-filtering dimension, forces electronic devices to concurrently request and fully load massive amounts of heterogeneous underlying video list data in the background after receiving an append request. A large amount of redundant video data, which does not conform to the actual preferences of the triggering entity, is forcibly loaded into the running memory. This not only needlessly consumes the network bandwidth and CPU computing resources of the electronic device, but also easily leads to memory overflow and interface stuttering when the graphics processor renders a large object tree, significantly reducing the underlying operating performance and hardware resource utilization of the electronic device.
[0174] In some embodiments, after displaying the add control, in response to a triggering operation on the add control, at least one add condition is displayed, wherein the add condition is a condition satisfied by the added video, and in response to a selection operation on a target add condition among the at least one add condition, a third video is displayed, wherein the third video is determined based on the second video that was not canceled and meets the target add condition.
[0175] Adding conditions are used to limit the filtering, recommendation, or sorting criteria that the video to be added must meet. Adding conditions can include at least one of the following: The added condition can be similarity to the content of the second video that is not displayed. Correspondingly, the third video can be a video that matches the second video that is not displayed in at least one of the following: title, tags, category, theme, keywords, subtitle text, audio recognition text, or image recognition results.
[0176] For example, added conditions could be displayed as "similar content," "videos with the same theme," "videos with the same tags," or "videos with the same keywords." In response to selecting "similar content" as the added condition, the content features of the second video (which was not initially displayed) are retrieved, and a third video with a content feature similarity greater than a similarity threshold is displayed.
[0177] The added condition can be that the third video belongs to the same publishing entity as the second video that was not canceled from display. Correspondingly, the third video can be a video that shares the same publishing account, publishing entity, publishing organization, publishing media, publishing anchor, publishing lecturer, or publishing merchant as the second video that was not canceled from display.
[0178] The added condition can be that the third video belongs to the same video collection as the remaining second video. Correspondingly, the third video can be a video not currently displayed in the video collection, a video that has been cancelled from display, or a video that is in the same collection chapter, the same collection group, or the same collection directory node as the second video that has not been cancelled. The added condition can also be that the third video is adjacent to the second video that has not been cancelled from its sorting position within the video collection. Correspondingly, the third video can be a video located within a preset number of times before or after the second video that has not been cancelled.
[0179] Adding a condition can be based on whether the video meets a preset popularity criterion. Accordingly, the third-party video can be a video that meets preset criteria based on the number of views, likes, comments, shares, favorites, completions, popularity score, or popularity ranking. Adding a condition can also be based on whether the video meets preset timeliness criteria based on its publication or update time. Accordingly, the third-party video can be a video that was published late, updated late, or is recently published or recently updated.
[0180] Adding a condition can be that the current subject has not watched or has not completed watching the video. Accordingly, the third video can be a video that the current subject has not watched, has a viewing progress of 0, has a viewing progress less than the completion threshold, or has not reached the completion stage. Adding a condition can also be matching the current subject's interests and preferences. Accordingly, the third video can be a video that matches the current subject's historical viewing tags, historical viewing types, followed accounts, favorited videos, liked videos, search terms, or recent viewing behavior.
[0181] Added conditions can include video duration, remaining viewing time, or estimated viewing time meeting preset criteria. Accordingly, the third-party video can be a short video, a long video, a video within a preset duration range, or a video designed for quick user viewing. Added conditions can also include a quality score meeting preset quality criteria. Accordingly, the third-party video can be a video whose clarity, completion rate, interaction rate, content completeness, platform quality score, or manual review score meets preset criteria.
[0182] Adding conditions can be based on the interaction status of the current object or other objects with the video, satisfying preset conditions. Accordingly, the third video can be a video that the current object has favorited, liked, commented on, or shared, or a video with high favorites, comments, or shares from other objects. Adding conditions can also be based on differences between the third video and the second video that is not shown. Accordingly, the third video can be a video that differs from the second video that is not shown in terms of theme, type, target audience, video duration, content format, or viewing scenario, to avoid displaying overly monotonous content.
[0183] Adding conditions or target adding conditions are rule parameters that constrain attributes or filter dimensions of the heterogeneous video data to be added. For example, adding conditions may be manifested as underlying data type isolation characteristics such as "watch only the main feature (with a specific episode number)" or "watch only the extras (without a specific episode number)"; target adding conditions are specific adding conditions for the currently received touch command.
[0184] In some embodiments, in response to a trigger operation on the add control, the terminal packages and sends the feature identifiers of the second video that has not been canceled to the server. The server uses a large language model to perform deep semantic extraction on the multimodal features of the second video that has not been canceled. Based on the deep semantic extraction results, the server generates at least one add condition. The server determines whether the number of at least one add condition is greater than zero. If yes, the server sends the data stream of at least one add condition to the terminal, and the terminal renders and displays at least one add condition in the display layer; if no, the server returns an empty condition status code to the terminal, and the terminal aborts the pop-up rendering logic of the condition selection panel.
[0185] In response to the trigger operation for adding conditions to the target, the terminal sends the type isolation bit parameter corresponding to the target addition condition to the server. The server extracts candidate multimedia data that meets the target addition condition, based on the fact that the second video has not been canceled. The server uses a multidimensional tuple mapping mechanism to calculate a three-dimensional tuple for the candidate multimedia data, including the type isolation bit, the forced integer set number, and the publication timestamp, and performs lexicographical ascending sorting calculation based on the three-dimensional tuple to generate the third video. The server determines whether the total number of data nodes in the third video exceeds a preset video memory threshold. If so, the server performs pagination truncation processing on the third video and sends the truncated third video data packet to the terminal, which decodes and displays the third video in the media stream interface; if not, the server sends the full third video data packet to the terminal, which directly decodes and displays the third video in the media stream interface.
[0186] The video recommendation method provided in this application provides at least one set of conditions after triggering the add control. Only in response to a selection operation for the target add condition is a third video, determined based on the second video (which was not canceled) and meeting the target add condition, displayed. This is equivalent to introducing a pre-level state constraint mechanism, ensuring that network communication requests and decoding instructions are initiated only for specific video data streams that highly match the triggering subject's intent. This precise on-demand data scheduling and isolation filtering blocks network transmission and cache writing of invalid, massive amounts of redundant data at the physical hardware level, significantly reducing the number of invalid parsing operations by the electronic device's central processing unit and significantly improving the memory flow efficiency and underlying operational stability of the electronic device.
[0187] In related technologies, when presenting video list data in a media stream interface, the display sequence number of each video node is pre-created by the object that published the video. The current object cannot cancel the display of video information according to actual usage needs. Furthermore, the video sequence number is pre-created by the video publishing object and will not change.
[0188] In some embodiments, the video information of each second video has a one-to-one corresponding sequence number. When displaying the video information of the second video, the video information of each second video can be displayed in ascending order of sequence number. In response to the cancellation display instruction for the video information of the target second video among the multiple second videos, the sequence number of the second video arranged after the target second video is updated, and the video information of the remaining second videos is re-displayed according to the updated sequence number.
[0189] In some embodiments, the terminal parses the video information of each second video, confirms that the video information of each second video has a one-to-one corresponding sequence number, and renders and displays the video information of each second video in the video display area of the media stream interface in ascending order of sequence number, in a visual style of vertical arrangement or horizontal sliding.
[0190] In response to a command to cancel the display of video information for a target second video among multiple second videos, a hide or collapse animation effect for the target second video's video information is triggered in the display layer, and the video information of the target second video is completely hidden from display on the interface. Simultaneously, the electronic device retrieves the sequence number of the second video following the target second video in the background, performs a numerical subtraction operation, and updates the sequence number of the second video following the target second video.
[0191] The video information located below the target second video display position is shifted upwards to fill the visual gap, and the remaining second videos are redisplayed according to the updated sequence numbers. For example, if the original sequence number of the target second video is 3, the electronic device updates the sequence numbers of the second videos originally numbered 4 and 5 to 3 and 4 and redisplays them.
[0192] See Figure 16 , Figure 16 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 Third, it can display video information from multiple second videos. Figure 16 The video is marked with video name 1601, video name 1602 and video name 1603. In response to the cancellation display instruction for the video information of the target second video (the video corresponding to video name 1602) among multiple second videos, the sequence number of the second video arranged after the target second video can be updated (that is, the sequence number of the video corresponding to video name 1603 and the sequence number of the video corresponding to video name 1604 are updated), and the video information of the remaining second videos is re-displayed according to the updated sequence number.
[0193] The video recommendation method provided in this application displays the video information of each second video in ascending order of the sequence number, based on the existence of a one-to-one correspondence between the video information of each second video and the sequence number. Upon receiving a cancellation command for a target second video, the sequence number of the second videos following the target second video is updated, and the remaining second videos are re-displayed according to the updated sequence number. A dynamic offset and automatic completion mechanism for the sequence number is constructed in the terminal's local memory or underlying linked list, ensuring the continuity and logical rigor of the sequence number at the interface view level. Furthermore, by replacing global interface network data requests and full page redraws with local sequence number offset calculations for the remaining nodes, unnecessary requests for redundant data are effectively avoided, significantly reducing the invalid computational overhead of the electronic device's central processing unit, lowering the peak memory usage, and improving the execution efficiency of the electronic device's underlying data processing and hardware resource scheduling.
[0194] In related technologies, when presenting multiple heterogeneous videos about a target event, a visual display format of unordered arrangement or a single vertical list can be used, lacking a visual structure to support the event's temporal evolution and internal logical evolution. This logical flaw in the interaction flow prevents the terminal from intuitively displaying the logical relationship between the content of the first video and the content of the second video in terms of time dimension (such as the first time point and the second time point) and event dimension (such as the first sub-event, the second sub-event, and the third sub-event). Furthermore, this blind interaction without structural guidance easily leads to the terminal frequently sending invalid video data requests to the server and repeatedly redrawing the entire interface, needlessly consuming the central processing unit's computing resources and running memory of electronic devices.
[0195] In some embodiments, the content of the first video is a first sub-event of the target event occurring at a first time point. A progress axis with at least one node can be displayed, where each node corresponds one-to-one with a sub-event belonging to the target event. In the associated area of each node in the progress axis, video information of a second video corresponding one-to-one with the node is displayed. The content of the second video is at least one of the following: a second sub-event of the target event occurring at the first time point, and a third sub-event of the target event occurring at the second time point. The first and second sub-events may be the same or different. The first and second time points are specific timestamps or time interval identifiers in the development process of the target event, and the second time point is located before or after the first time point in chronological order.
[0196] For example, if the target event is a marathon race, the content of the first video is the first sub-event of the target event occurring at the first point in time. For instance, the first video might show the announcement of the race route for a city marathon on March 15th. In response to a video viewing command, a progress axis with at least one node is displayed. The nodes in the progress axis are arranged chronologically according to the progress of the city marathon event. Each node corresponds one-to-one with a sub-event of the target event, or each node corresponds to one or more sub-events occurring at the same point in time.
[0197] For example, the multiple nodes in the progress axis can correspond to the opening of registration on March 1, the announcement of the race route on March 15, the announcement of the arrangements for picking up race materials on April 1, the announcement of the traffic control plan on April 10, the on-site situation on the day of the race on April 15, and the announcement of the race results on April 16.
[0198] In the progress axis, the associated area of each node displays video information for the second video corresponding to that node. The content of the second video can be the second sub-event of the target event occurring at the first point in time, such as a route highlight analysis, introduction to aid stations along the route, or participant feedback on the route on March 15th. Alternatively, the content can be the third sub-event of the target event occurring at the second point in time, such as registration opening on March 1st, announcement of race material collection arrangements on April 1st, traffic tips on April 10th, the starting line situation on April 15th, interviews with finishers, or instructions on how to check results on April 16th. Thus, the progress axis can display the advancement of the target event at different points in time, and show videos of the target event from different angles at the same point in time, allowing users (i.e., the current participant) to understand the complete development process of the target event according to the event timeline.
[0199] In some embodiments, the terminal sends a request for associated data for a first video to the server. The server extracts the multimodal features of the first video, parses its content, and determines the target event and the first sub-event occurring at a first time point. Based on the target event, a multi-path recall strategy is initiated to obtain a set of candidate videos. A large language model is used to perform deep semantic extraction and timestamp entity recognition on the candidate video set. The server determines whether the content of each candidate video in the set belongs to a sub-event of the target event. If so, the server extracts the time point corresponding to the candidate video, uses a multidimensional tuple mapping algorithm to classify and lexicographically arrange the candidate videos according to time points (first time point, second time point) and sub-event attributes (second sub-event, third sub-event), generating structured data with a progress axis containing nodes. The server marks the candidate videos that meet the criteria as the second video. The server then sends a data packet containing the progress axis data and the video information of the second video to the terminal. The terminal parses the data packet, draws the progress axis and nodes on the front-end interface, and displays the video information of the second video in the associated area. If not, the server removes the candidate videos that do not meet the criteria from the candidate set, blocking the sending process of that part of the video data to avoid invalid data consuming network bandwidth.
[0200] The video recommendation method provided in this application embodiment features a first video whose content is the first sub-event of the target event occurring at a first time point. When displaying the video information of each second video, a progress axis with at least one node is introduced. The associated area of each node on the progress axis displays the video information of the second video corresponding to that node. Simultaneously, the content of the second video is limited to either the second sub-event of the target event occurring at the first time point or the third sub-event occurring at the second time point. This constructs a text-image association structure based on timelines and sub-events, directly achieving accurate serialization, recombination, and presentation of heterogeneous data (horizontal multi-view and vertical timeline) on the front-end interface. This significantly reduces the interaction layers required to obtain the full picture of the target event, effectively curbs duplicate and invalid data requests caused by a lack of guidance, and significantly improves the accuracy and response efficiency of the underlying data flow of electronic devices.
[0201] In some embodiments, in response to a video viewing instruction triggered by a video collection control, a third display style is used to display the video information of the first video, and a fourth display style is used to display the video information of the second video whose sequence number is less than a difference threshold with respect to the sequence number of the first video. The third display style and the fourth display style are different. The third display style is used to indicate the video currently in a playing state, and the fourth display style is used to indicate the video currently in a non-playing state.
[0202] In related technologies, after presenting video information containing multiple second videos, if the current operating entity wishes to synchronously watch and interact with other entities, it can typically only initiate sharing actions for each individual video one by one. This logical flaw in the interaction process forces the terminal to repeatedly respond to separate interactive operations for different individual videos when handling multi-video sharing needs across objects, thus compelling electronic devices to frequently trigger redundant independent network communication requests. This not only leads to low efficiency in human-computer interaction but also needlessly consumes the network communication bandwidth and central processing unit's logical computing resources of electronic devices, increasing the undue load on the electronic devices.
[0203] In some embodiments, after displaying the video information of each second video, a shared viewing control for the video collection is displayed, wherein the shared viewing control is used to invite other objects to watch the videos in the video collection together, and based on the shared viewing control, in response to a shared viewing instruction for other objects, an invitation message for inviting others to watch the video collection together is sent to the other objects.
[0204] In some embodiments, in response to a triggering operation on a co-watching control, at least one candidate object may be displayed; in response to a selection operation on another object among the at least one candidate object, a co-watching instruction for the other object may be triggered; and in response to the co-watching instruction for the other object, an invitation message for inviting the co-watching of the video collection may be sent to the other object.
[0205] In some embodiments, in response to a trigger operation for a shared viewing control, the terminal sends a shared viewing session creation request packet to the server. The server receives the shared viewing session creation request packet and extracts the multidimensional sorting tuple used when pre-generating the video collection (the multidimensional sorting tuple includes a type isolation bit, an integer value for the number of episodes, and a publication timestamp). Based on the multidimensional sorting tuple sequence and a unique session identifier, an invitation message is generated and sent to the terminal. In response to a shared viewing instruction for other objects, the terminal determines whether the network node where the target other object is located is online and active. If so, the terminal calls the underlying persistent transport control protocol socket channel to directly route and push the invitation message to the network node where the other object is located, and the terminal simultaneously opens a synchronous playback status determination thread in its local running memory; if not, the terminal dumps the invitation message into offline message queue data and uses the server's asynchronous message distribution channel to send the invitation message offline to the other object.
[0206] The electronic devices used by other objects render and display the invitation pop-up and accept control in the display layer. In response to the electronic devices where other objects are located detecting the trigger operation on the accept control, the electronic devices where other objects are located execute the collection synchronization decoding instruction, and load and synchronously play the videos in the video collection in the interface of the electronic devices where other objects are located.
[0207] The video recommendation method provided in this application displays a shared viewing control for the video collection after showing the video information of each second video. Based on the shared viewing control, it responds to shared viewing instructions for other objects and sends invitation information to invite them to watch the video collection together. This constructs a global shared entry point for structured video collections, aggregating cumbersome single-point sharing into a single signaling package. This significantly reduces the number of interaction layers in cross-node collaborative scenarios, blocks the concurrent generation of duplicate network data packets from the underlying communication mechanism, significantly reduces the ineffective network transmission overhead of electronic devices, and improves the bandwidth utilization and underlying data scheduling efficiency of electronic devices.
[0208] In related technologies, when different terminals trigger different video nodes within the same video collection, electronic devices often lack targeted data isolation in the background. This forces them to concurrently request and cache all video data streams played by all terminals, or to frequently perform invalid redraws of the entire interface on the front end. This logical flaw in the interaction process causes the central processing unit of the electronic device to handle a large number of redundant data decoding tasks, resulting in excessive memory consumption due to invalid caching. Consequently, it leads to response stuttering and significantly reduces the underlying operating performance and resource utilization efficiency of the electronic device.
[0209] In some embodiments, upon receiving consent information, in response to a selection operation triggered by video information for the second video, if the selected video is consistent with the video selected by other objects, a shared viewing interface is displayed, and the selected video is shown in the shared viewing interface, wherein the consent information is used to indicate that other objects agree to share the video collection based on the invitation information.
[0210] See Figure 17 , Figure 17 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 Fourth, upon receiving consent information, in response to the selection operation triggered by the video information of the second video, if the selected video is the same as the video selected by other objects, the co-viewing interface 1701 is displayed, and the selected video is displayed in the co-viewing interface 1701 (i.e., the selected video is played). Information indicating that the current object is watching the video together with other objects can also be displayed.
[0211] In some embodiments, upon receiving consent information, in response to a selection operation triggered by the video information of the second video, if the selected video is different from the video selected by other objects, the object identifiers and viewing progress of the other objects are displayed in the associated area of the video information of the other objects' selected videos. The selected video is determined by the current terminal in response to the selection operation triggered by the video information of the second video, specifying the audio and video content data to be loaded and decoded.
[0212] See Figure 18 , Figure 18 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 Fifth, upon receiving consent information, in response to the selection operation triggered by the video information of the second video, if the selected video is different from the video selected by other objects, the video being played is displayed, and the object identifier of the other object and the viewing progress (corresponding to the content indicated by box 1801) are displayed in the associated area of the video information of the video selected by the other object.
[0213] In some embodiments, the terminal receives consent information through a network interface and, in response to a selection operation triggered by the video information of the second video, extracts the three-dimensional sorting tuple (including type isolation bit, episode number integer value, and publication timestamp) corresponding to the selected video, and sends the three-dimensional sorting tuple as an identifier parameter to the server. The server receives the identifier parameter and extracts the three-dimensional sorting tuples corresponding to videos selected by other objects from the background session state database. The server determines whether the video identifier of the selected video is consistent with the video identifiers of videos selected by other objects; or, determines whether both the video identifier and the three-dimensional sorting tuple of the selected video are consistent.
[0214] If yes, the server determines that the selected video is the same as the video selected by other objects. The server allocates a synchronous decoding channel for the terminal and other objects, and sends an interface switching signal to the terminal. The terminal parses the interface switching signal, allocates independent video memory space to display the shared viewing interface, and calls the underlying decoder to decode and display the selected video in the shared viewing interface. If no, the server determines that the selected video is different from the video selected by other objects. The server extracts the object identifier of other objects and the real-time synchronized viewing progress data packet and sends it to the terminal. The terminal suspends the rendering process of the shared viewing interface to release the corresponding video memory, parses the viewing progress data packet, calls the basic graphical user interface library, and draws and displays the object identifier of other objects and the viewing progress in the associated area of the video information of the video selected by other objects.
[0215] The video recommendation method provided in this application, upon receiving consent information, responds to a selection operation by determining whether the selected video matches the videos selected by other users. If they match, a shared viewing interface and the selected video are displayed; otherwise, only the object identifiers and viewing progress of other users are displayed in the associated area of the video information of the videos selected by other users. A dynamic memory allocation and network resource scheduling mechanism based on consistency determination is constructed, strictly isolating the underlying processing links of synchronized playback and independent browsing. This avoids redundant loading and decoding operations of irrelevant video data streams, significantly reducing the ineffective power consumption of the underlying hardware of electronic devices while ensuring multi-user interactive collaborative state awareness, thus significantly improving operating performance.
[0216] In some embodiments, while displaying the object identifiers of other objects and the viewing progress, in response to a triggering operation on the object identifiers of other objects, a video follow-up switching instruction is triggered. In response to the video follow-up switching instruction, a shared viewing interface is invoked and displayed in the display layer, and the video selected by other objects is decoded and displayed in the shared viewing interface.
[0217] In related technologies, when sharing a data set containing multiple heterogeneous video nodes in a media streaming interface, it is usually only supported to share a single independent video one by one, or to share the entire set of data as a whole. This causes the terminal to either force the sending of a large amount of redundant and unnecessary video data, resulting in a waste of network bandwidth, or force the operator to repeatedly perform multiple separate interactive operations on a single video when processing the sharing needs of local video data. This leads to low efficiency of human-computer interaction and unnecessarily increases the burden on the central processing unit of electronic devices to handle multiple concurrent network requests.
[0218] To address the aforementioned technical issues, after displaying the video information of each second video, in response to a selection operation involving multiple second videos, the selected second videos are controlled to be in a selected state. The number of second videos in the selected state is less than the number of second videos in the video collection. In response to a sharing instruction for the selected second videos, a new video collection of the selected second videos is created, and the new video collection is shared according to the sharing instruction. The new video collection is a structured subset data set generated based on the selected independent videos, by extracting their metadata and reconstructing the underlying logical association features.
[0219] In some embodiments, in response to a selection operation of video information areas of multiple second videos (e.g., continuous touch clicks on checkbox controls reserved on the sides of video information of multiple second videos), the selected second video is controlled to be in a selected state in the media streaming interface. Visually, the electronic device changes the display style of the checkbox control corresponding to the selected second video from a hollow circle to a solid circle filled with a highlighted color, and overlays a semi-transparent background layer on top of the video information of the selected second video to produce a visual highlighting effect. At this time, the number of second videos in a selected state is less than the number of second videos in the video collection (i.e., partial video extraction is achieved).
[0220] Subsequently, in response to a sharing instruction for the selected second video (e.g., clicking the "Share Selected" virtual button that pops up at the bottom of the interface), a new video collection for the selected second video is created, and a loading animation with the text "Generating new collection link" pops up on the front-end interface; after the loading animation ends, the electronic device invokes the sharing panel according to the sharing instruction and shares the data card or link address of the new video collection to the target application environment, which can be shared with other objects or published as social information.
[0221] In some embodiments, in response to a selection operation for multiple second videos, the terminal pushes the feature identifier of the selected second video into a sharing queue in a doubly linked list in local memory, and updates the local rendering tree to make the selected second video selected. In response to a sharing instruction for the selected second video, the terminal encrypts the data packet containing the feature identifier in the sharing queue and uploads it to the server.
[0222] The server parses the data packet, calls the background counter, and determines whether the number of selected second videos is less than the number of second videos in the video collection. If so, the server determines that the subset reorganization logic is triggered. The server extracts the three-dimensional sorted tuple corresponding to the selected second video from the full database (the three-dimensional sorted tuple includes a type isolation bit, an integer value for the number of episodes, and a publication timestamp). The server performs a rearrangement mapping calculation on the extracted three-dimensional sorted tuple separately, generates a globally unique subset identifier based on the mapping result, and thus creates a new video collection of the selected second videos in the cloud. The server then sends a return data packet containing the subset identifier to the terminal, and the terminal calls the external interface protocol stack to share the new video collection according to the sharing instruction. If not, the server determines that the full sharing logic is triggered. The server directly extracts the root identifier of the original video collection and sends it to the terminal. The terminal blocks the creation process of the new video collection and directly shares the original video collection according to the sharing instruction.
[0223] The video recommendation method provided in this application controls the selected second video to be in a selected state, ensuring that the number of selected second videos is less than the number of second videos in the video collection. Then, in response to a sharing command, a new video collection of the selected second videos is created and shared according to the sharing command. This constructs a fine-grained underlying video data isolation and repackaging mechanism, achieving precise on-demand extraction and aggregation transmission of local data sets. This design compresses the cumbersome multiple single-point sharing network requests into a single signaling flow, not only avoiding the network distribution consumption of invalid and redundant underlying data, but also significantly reducing the concurrent communication load of the electronic device's network communication module, and significantly improving the memory flow efficiency and the accuracy of on-demand transmission of underlying data in the electronic device.
[0224] In related technologies, when displaying videos or video collections shared by other objects to the current object, the current object can only view the shared videos or video collections. If the current object needs to obtain complete data, it can exit the current playback interface and manually initiate a global search network request for the target video collection. The global search network request will frequently trigger redundant underlying data retrieval and full pixel redraw of the search result interface, needlessly consuming the central processing unit computing resources and network communication bandwidth of electronic devices, resulting in the running memory being occupied by invalid search cache, reducing the system resource circulation efficiency and underlying data processing performance of electronic devices.
[0225] To address the aforementioned technical issues, in response to a video viewing command triggered by a video collection control, if the second video included in the video collection is a subset of videos selected by other objects from the target video collection, a second viewing control is displayed. This second viewing control is used to view the target video collection. After displaying the video information of each second video, in response to a trigger operation on the second viewing control, the video information of the second video is switched to the video information of the videos in the target video collection.
[0226] See Figure 19 , Figure 19 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 6. In response to a video viewing command triggered by a video collection control, if the second video included in the video collection is a subset of videos selected by other objects from the target video collection, the second viewing control 1901 is displayed. See also Figure 20 , Figure 20 This is a schematic diagram of the interface provided in the embodiments of this application. Figure 10 7. Acceptance Figure 19 In response to a trigger operation on the second viewing control 1901, the video information of the second video can be displayed, and the video information of the videos in the target video collection can be switched (corresponding to the interface indicated by 2001).
[0227] In some embodiments, when the terminal corresponding to another object responds to a viewing instruction for a new video collection, it can display the video information of the videos in the new video collection and a second viewing control. In response to a trigger operation on the second viewing control, the video information of the videos in the new video collection can be switched to the video information of the second video in the video collection.
[0228] In some embodiments, the video collection includes a subset identifier. In response to a video viewing instruction triggered by a video collection control, the subset identifier is packaged and sent to the server. The server uses the subset identifier to query the backend relational database to determine whether the second video included in the video collection is a subset of videos selected by other objects from the target video collection. If so, the server extracts the root identifier of the target video collection and sends the root identifier of the target video collection and the three-dimensional sorting tuple of the second video (the three-dimensional sorting tuple includes a type isolation bit, an integer value for the number of episodes, and a publication timestamp) to the terminal. The terminal parses the returned data and allocates a handle in the instance layer of the interface front-end to display the second viewing control. If not, the server only sends the three-dimensional sorting tuple of the second video, and the terminal blocks the logic for allocating the graphics rendering handle of the second viewing control.
[0229] After displaying the video information of each of the second videos, the terminal captures the trigger operation for the second viewing control. The terminal sends a network packet to the server carrying the root identifier of the target video collection, requesting full data retrieval. The server uses a large language model to perform deep semantic extraction and multi-dimensional tuple mapping sorting on the target video collection, and sends the sorted full multimedia data stream to the terminal. The terminal then switches from displaying the video information of the second videos to displaying the video information of the videos in the target video collection.
[0230] The video recommendation method provided in this application displays a second viewing control when the second video in the video collection is a subset of videos selected by other objects from the target video collection. After displaying the video information of each second video, in response to a trigger operation on the second viewing control, the video information of the second video is switched to the video information of the videos in the target video collection. This constructs a low-level addressing mapping entry that directly points to the complete target video collection, effectively blocking redundant network concurrent communication packets and invalid interface pixel redrawing caused by repeated manual searches by electronic devices. It also significantly reduces redundant search cache garbage in the running memory, improving the system hardware resource utilization and effective scheduling efficiency of network transmission bandwidth of electronic devices.
[0231] In related technologies, video recommendations are made only for the currently playing video, resulting in a relatively simple method of recommending videos. In order to solve the above technical problems, after displaying the video information of each second video, the second video is played in response to the playback command for the second video. When the number of unplayed second videos decreases to a third number threshold, at least one recommended video collection based on the video collection recommendation is displayed.
[0232] In some embodiments, when the number of unplayed second videos decreases to a third quantity threshold, at least one recommended video collection based on video collection recommendations is displayed. The process involves: obtaining collection features of the video collections; determining candidate video collections from a set of candidate collections based on these features; and determining a recommendation score for the candidate video collections based on at least one of the following: content similarity between candidate video collections and other video collections; matching degree between candidate video collections and the current object's interests; popularity of candidate video collections; freshness of candidate video collections; co-occurrence viewing degree of candidate video collections by the object watching the video collections; continuity between candidate video collections; and event correlation between candidate video collections and corresponding events of the video collections.
[0233] Based on the recommendation score, at least one recommended video collection is determined from the candidate video collections. The collection features include at least one of the following: collection name, collection description, collection tags, collection category, collection theme, collection keywords, collection type, number of videos in the collection, video list, target audience of the collection, collection update time, collection popularity, number of viewers of the collection, number of times the collection has been played, number of times the collection has been liked, number of times the collection has been commented, number of times the collection has been shared, number of times the collection has been saved, number of times the collection has been subscribed, and content features of the videos within the collection. Therefore, when the current viewer is nearing the end of watching the videos in the current collection, other video collections related to the current collection or matching the current viewer's interests can be recommended to the current viewer, improving the efficiency of continuous viewing between video collections.
[0234] In some embodiments, in response to a playback command for the second video, the terminal invokes a local decoding component to decode and play the second video. The terminal initializes and maintains a state counter in its local memory, which is used to decrement the number of unplayed second videos in real time. The terminal invokes an arithmetic logic unit to determine whether the number of unplayed second videos has decreased to a third threshold. If the number of unplayed second videos has decreased to the third threshold, the terminal sends a recommendation data network request packet to the server. The recommendation data network request packet carries a unique identifier for the video collection and current playback progress parameters.
[0235] The server uses unique identifiers to extract multimodal features and three-dimensional sorting tuples (the three-dimensional sorting tuple includes type isolation bits, episode number integer values, and release timestamps) from video collections. These multimodal features are then input into a large language model for deep semantic extraction and node downgrading and cleaning, constructing corresponding entity relationship chains for the same topic. Based on these entity relationship chains, the server initiates a multi-channel parallel recall strategy to obtain a candidate collection pool. The server recalculates the feature correlation score for each collection in the candidate collection pool, selecting at least one recommended video collection whose score meets a safety threshold. The server then sends the data structure stream of at least one recommended video collection to the terminal. The terminal receives the data structure stream, instantiates and displays at least one recommended video collection in the display layer of the media stream interface. If the number of unplayed second videos does not decrease to the third threshold, the terminal blocks the generation and sending of recommendation data network request packets, maintaining the current video decoding and front-end rendering state.
[0236] The video recommendation method provided in this application determines the number of unplayed second videos in real time during the playback of the second video. When the number of unplayed second videos decreases to a third threshold, it dynamically triggers the display of at least one recommended video collection based on the video collection recommendation. Utilizing the third threshold, a precise preloading and interface rendering buffer node is constructed at the underlying level, evenly distributing network requests, parsing, and graphics rendering tasks for newly added recommendation data across the normal decoding and playback cycle of the video collection. This mechanism effectively reduces the instantaneous data processing pressure on the underlying electronic device system, avoids premature loading of invalid and redundant data at the physical hardware level, and significantly improves the memory scheduling efficiency of the electronic device and the operational stability of the underlying computing hardware.
[0237] In related technologies, when playing video collections in a media streaming interface, there is a lack of a categorized playback guidance mechanism for main content and supplementary content with a strict order. As a result, after the first video finishes playing, the terminal forces the next video to be loaded in a single list order, which cannot meet the differentiated needs of different users for watching main content or viewing supplementary content continuously.
[0238] To address the aforementioned technical issues, the second video includes a target video that has an order with the first video, as well as other videos that supplement the content of the first video. The target video can be a video belonging to the same video collection as the first video and located before or after the first video in terms of playback order, content order, time order, chapter order, event development order, or business process order within the video collection. The other videos can be videos used to explain, expand upon, supplement, provide background information, interpret viewpoints, showcase behind-the-scenes moments, or make related recommendations to the content of the first video.
[0239] The second video includes a target video and other videos. The target video is a second video that is arranged in a specific order with the first video. The other videos are supplementary videos that supplement the content of the first video. For example, in the film and television industry, the first video can be a clip from the main feature of a film or television work, and the target video can be the previous clip, the next clip, the previous episode, the next episode, or the complete main feature of that film or television work; other videos can be trailers, behind-the-scenes videos, behind-the-scenes production videos, interviews with the main creators, plot interpretation videos, character analysis videos, or videos introducing the shooting scenes of the film or television work.
[0240] For example, in the field of educational courses, the first video can be a video explaining a certain knowledge point, and the target video can be a course video in the same curriculum system that is before or after the first video, such as the previous lesson, the next lesson, the previous chapter video, or the next chapter video; other videos can be example analysis videos, videos explaining common mistakes, after-class exercise videos, knowledge point summary videos, experimental demonstration videos, or exam question explanation videos for that knowledge point.
[0241] In some embodiments, when the first video finishes playing, a first playback control and a second playback control are displayed. The first playback control is used to play the target video, and the second playback control is used to play other videos. In response to a trigger operation on the first playback control, the target video that is adjacent to the first video in the sorting order and follows the first video is played. In response to a trigger operation on the second playback control, other videos are played. When the other videos finish playing, the target video that is adjacent to the first video in the sorting order and follows the first video is played.
[0242] In some embodiments, the terminal sends a unique identifier of the first video to the server. The server uses a multi-path recall strategy to obtain a set of candidate videos, inputs the set of candidate videos into a large language model, and divides the set of candidate videos into target videos with a sorted order and other videos that supplement the content of the first video. The server uses a three-dimensional tuple mapping algorithm to assign sorted tuples to the target videos and other videos.
[0243] The terminal displays the first video in the media stream interface. The terminal's underlying decoder outputs an interrupt signal indicating that the first video has finished playing. Upon receiving the interrupt signal, the terminal determines whether there are any related videos in the underlying data dictionary. If related videos exist, the terminal calls the basic graphical user interface library to draw the first and second playback controls at preset coordinates on the screen. If no related videos exist, the terminal only draws the first playback control. The terminal captures trigger operations for the first playback control and extracts the data stream of the target video with a higher episode number than the first video and closest to it from the sorted tuple for decoding. The terminal captures trigger operations for the second playback control and extracts the data stream of other videos with a type isolation bit as a supplementary feature for decoding. When other videos finish playing, a callback function is triggered, calling the data stream of the target video for decoding, thereby playing the target video that is adjacent to the first video in the sorted order and follows the first video.
[0244] The video recommendation method provided in this application displays both a first playback control and a second playback control simultaneously upon completion of the first video playback. After playing other videos in response to a trigger operation on the second playback control, and once those other videos have finished playing, it automatically plays the target video that is adjacent to and follows the first video in the order of playback, thus constructing a branch playback logic path with an automatic backtracking mechanism. The mechanism of accurately providing differentiated selections at completion nodes and automatically and seamlessly connecting to the main playback sequence after satisfying the interactive needs of exploring supplementary content significantly reduces the large number of invalid touch commands and redundant network retrieval requests generated by manual searching for the current object. This significantly reduces the invalid computational load on electronic devices and improves the resource flow efficiency of electronic devices when processing complex multimedia chains.
[0245] In related technologies, video collections are fixed, and the current object cannot directly obtain and switch to the data list of other video collections. The current object needs to perform cumbersome operations such as exiting the current playback interface, initiating a global search, and reloading, resulting in low efficiency of human-computer interaction.
[0246] To address the aforementioned technical issues, a target second video is included among the multiple second videos. In response to a trigger operation on the video information of the target second video, the first video is switched to the target second video. If the target second video also belongs to other video collections, a jump control is displayed. In response to a trigger operation on the jump control, the video information of each of the displayed second videos is switched to the video information of videos in other video collections.
[0247] In some embodiments, the terminal extracts the multimedia data stream of the target second video and pushes it into the decoder, switching the display of the first video to the target second video. The terminal sends the feature identifier of the target second video to the server. The server receives the feature identifier of the target second video and determines whether the target second video belongs to other video collections. If yes, the server sends a jump configuration instruction containing the identifier of other video collections to the terminal. The terminal parses the jump configuration instruction and allocates a graphics handle in the video memory to display the jump control. If no, the server returns an empty status code to the terminal, and the terminal blocks the rendering process of the jump control.
[0248] In response to a trigger action on the jump control, the terminal sends the identifiers of other video collections to the server. The server uses a large language model to perform semantic chaining on the candidate videos within the other video collections, and uses a multidimensional tuple mapping algorithm to calculate a three-dimensional sorted tuple for the candidate videos in the other video collections, including a type isolation bit, a forced integer set number, and a publication timestamp. The server performs lexicographical sorting on the three-dimensional sorted tuples, generates video information for the videos in the other video collections, and sends it to the terminal. The terminal then switches the displayed video information of each second video to the video information of the videos in the other video collections.
[0249] The video recommendation method provided in this application, in response to a trigger operation on the video information of a target second video, switches the display of the first video to the display of the target second video. If the target second video also belongs to other video collections, a jump control is displayed. Finally, in response to a trigger operation on the jump control, the video information of each second video is switched to the video information of videos in other video collections. This constructs a quick jump bridge based on the intersection of underlying data between different video collections, avoiding invalid network reload and full interface redraw when browsing across collections. This significantly reduces the invalid computing load of electronic devices and improves the flexibility and overall flow efficiency of underlying data scheduling of electronic devices.
[0250] In some embodiments, if it is determined that the target second video also belongs to other video collections, and the number of other video collections is greater than a preset threshold, a multi-collection panel expansion instruction is executed, and multiple jump controls are displayed in the media stream interface in the style of a floating list. The multiple jump controls correspond to different other video collections.
[0251] In some embodiments, the video recommendation method provided in this application can further determine the video collection by: acquiring the multimodal features of a first video, the multimodal features including at least two of the following: title text, video tags, cover image features, on-screen text recognition, speech recognition, publication time, and video publication object; determining the collection theme anchor point corresponding to the first video based on the multimodal features; recalling a candidate video set based on the collection theme anchor point; and determining multiple second videos from the candidate video set based on the degree of matching between each candidate video in the candidate video set and the collection theme anchor point. The collection theme anchor point is used to characterize the core semantic boundary of the video collection to which the first video belongs, and the collection theme anchor point includes at least one of the following: event entity, work entity, course entity, person entity, location entity, topic entity, series name, chapter name, and episode number information.
[0252] In some embodiments, for each candidate video in the candidate video set, the candidate video’s inclusion score is determined based on at least one of the following: semantic similarity between the candidate video and the theme anchor of the collection, content continuity between the candidate video and the first video, freshness of the candidate video, popularity of the candidate video (see the foregoing description), and duplication between the candidate video and videos already included in the collection. When the inclusion score is greater than the inclusion threshold, the candidate video is determined as the second video.
[0253] In some embodiments, the server inputs at least one of the following text encoding models—title text, video tags, video description, subtitle text, speech-recognized text, and on-screen text—to obtain a candidate video semantic vector; and inputs the theme text corresponding to the collection's theme anchor point into a text encoding model to obtain a theme anchor point semantic vector. The server determines the semantic similarity between the candidate video and the collection's theme anchor point based on the vector distance or distance similarity between the candidate video semantic vector and the theme anchor point semantic vector. For example, the semantic similarity can be determined based on the cosine similarity between the candidate video semantic vector and the theme anchor point semantic vector.
[0254] In some embodiments, the server determines the content continuity between the candidate video and the first video based on at least one of the following: entity overlap, event phase relationship, temporal sequence relationship, chapter sequence relationship, episode continuity relationship, causal relationship, and contextual coherence relationship. Specifically, entity overlap is determined by the number of entities appearing in both the candidate video and the first video; event phase relationship is determined by the event phases corresponding to the candidate video and the first video respectively; temporal sequence relationship is determined by the occurrence times of the content in the candidate video and the first video; chapter sequence relationship or episode continuity relationship is determined by the chapter number or episode number corresponding to the candidate video and the first video; and contextual coherence relationship is determined by the coherence probability obtained after inputting the candidate video summary and the first video summary into a sequence relationship recognition model.
[0255] In some embodiments, the server determines the freshness of a candidate video based on the time difference between the candidate video's publication time and the current time. A smaller time difference indicates higher freshness, and a larger time difference indicates lower freshness. For example, the server can input the time difference between the candidate video's publication time and the current time into a time decay function to obtain the candidate video's freshness.
[0256] In some embodiments, the server determines the degree of duplication between a candidate video and videos already included in the collection based on at least one of the following: title text overlap, subtitle text overlap, video summary overlap, cover image similarity, keyframe image similarity, audio fingerprint similarity, and video segment similarity. When the degree of duplication between a candidate video and any video already included in the collection exceeds a duplication threshold, the server reduces the candidate video's inclusion score in the collection or removes the candidate video from the candidate video set.
[0257] In some embodiments, the server determines the inclusion score of a candidate video based on the semantic similarity between the candidate video and the theme anchor of the collection, the content continuity between the candidate video and the first video, the freshness of the candidate video, the popularity of the candidate video, and the duplication of the candidate video with videos already included in the collection. Semantic similarity, content continuity, freshness, and popularity are positively correlated with the inclusion score, while duplication is negatively correlated with the inclusion score.
[0258] In some embodiments, after determining the second video in the video collection, the server generates a collection theme representation of the video collection based on the multimodal features of the first video and multiple second videos; and determines the theme offset value of each second video based on the degree of deviation between the collection theme representation and each second video. When the theme offset value of a target second video is greater than an offset threshold, the target second video is removed from the video collection, or the target second video is moved to an extended video area of the video collection. This avoids thematic divergence in the video collection due to continuous recall, improving the thematic consistency of the video collection.
[0259] In some embodiments, the server determines a target structure template for the video collection based on the collection type corresponding to the video collection. The target structure template includes at least one of a timeline template, episode template, chapter template, event development template, course knowledge point template, and competition round template. The server maps multiple second videos in the video collection to multiple structure nodes in the target structure template and determines the completeness of the video collection based on the number of structure nodes mapped by the videos and the total number of structure nodes in the target structure template. When the completeness is less than a completeness threshold, the server generates completion recall conditions based on the structure nodes that have not been mapped by the videos and continues to recall supplementary videos based on the completion recall conditions.
[0260] Collection type is used to characterize the way a video collection is organized. For example, when a video collection is a news event collection, the collection type can be event development type; when a video collection is a series or short drama collection, the collection type can be episode type; when a video collection is an educational video collection, the collection type can be chapter type or course knowledge point type; when a video collection is a sports event collection, the collection type can be competition round type.
[0261] The target structure template is used to represent the content structure that a video collection should cover under a corresponding collection type. The target structure template includes multiple structure nodes, each representing a content segment that the video collection should cover. For example, the structure nodes in the event development template may include the event cause, event occurrence, event response, event handling result, and event follow-up impact; the structure nodes in the episode template may include episode 1, episode 2, and episode 3; the structure nodes in the course knowledge point template may include basic concepts, principle explanations, example demonstrations, practical exercises, and summary reviews; and the structure nodes in the competition round template may include group stage, knockout stage, semi-finals, and finals.
[0262] The video recommendation method provided in this application does not construct video collections based solely on a single keyword or a single publishing object. Instead, it automatically filters and structures candidate videos through collection theme anchors, multi-channel recall, collection entry scores, and theme offset filtering. This improves the theme consistency, content completeness, and ranking accuracy of video collections while reducing manual editing. Furthermore, since the server only identifies videos that meet the collection entry score criteria as secondary videos and performs targeted completion recall for missing structural nodes when collection completeness is insufficient, it avoids indiscriminate searching of the entire video library, reducing server-side computational load and network data transmission volume, and improving the processing efficiency of video collection recommendations.
[0263] In some embodiments, the server performs multi-channel recall based on the collection theme anchor to obtain a candidate video set. Multi-channel recall includes at least one of the following: title keyword recall, used to recall videos whose titles contain collection theme keywords; tag recall, used to recall videos with the same or similar video tags; vector recall, used to recall videos whose determined semantic vectors are similar to the semantic vectors of the collection theme anchor; publishing object recall, used to recall videos associated with the first video publishing object or related publishing objects; user co-occurrence behavior recall, used to recall videos that were continuously viewed, favorited, or shared by the same group of objects as the first video; and time window recall, used to recall videos whose publication time or content occurrence time falls within a target time window.
[0264] In some embodiments, the server determines the sorting position of the second video within a video collection based on the collection role, content occurrence time, publication time, episode number identifier, chapter identifier, event stage identifier, collection score, and the interest matching degree of the current object. The collection role is used to characterize whether the second video is a main story video, supplementary video, explanatory video, bonus video, commentary video, or background introduction video.
[0265] When the video collection is of episode type, it is sorted first by episode number identifier; when the video collection is of event type, it is sorted first by event occurrence time or event stage; when the video collection is of course type, it is sorted first by knowledge point dependency relationship or chapter order; when the video collection is of competition type, it is sorted first by competition round or competition time.
[0266] In some embodiments, when multiple second video sets in a video collection originate from the same video publishing object, the server lowers the inclusion score of candidate videos from the same video publishing object or raises the inclusion score of candidate videos from different video publishing objects. This ensures that the second videos in the video collection originate from at least two different video publishing objects.
[0267] To facilitate understanding of the video recommendation method provided in this application's embodiments, the following examples illustrate this method. With the explosive growth of internet video, the fragmented nature of video platforms has become increasingly pronounced. When watching film clips, variety show segments, or sports highlights, users often face problems such as disjointed storylines, difficulty in finding episodes, and a mix of main content and extras. Traditional news or video aggregation methods often rely on manual tagging or simple keyword matching, making it difficult to automatically sort out the category and narrative logic of complex videos, and thus failing to meet users' deep needs for an immersive and continuous viewing experience.
[0268] With the advancements in large language models in natural language processing and multimodal understanding, new technological foundations have been provided for automated semantic disambiguation and deep content extraction. Meanwhile, asynchronous concurrent architectures have demonstrated strong processing capabilities in complex task decomposition and massive data processing. The video recommendation method provided in this application is essentially a video serialization construction method that integrates multimodal features and a two-layer verification technique using large language models, thereby improving the organization, sorting, and presentation capabilities of fragmented videos.
[0269] Among related technologies, video series or compilation generation technologies mainly include the following types of solutions: The first type of solution is aggregation based on manual tagging and hard matching of metadata. It uses hard rule matching based on the series name, episode number or preset tags manually filled in by the creator when uploading. However, this method is highly dependent on the accuracy of manual input, has poor flexibility, and is prone to aggregation failure due to creator omissions or errors in filling in the information.
[0270] The second approach is based on keyword retrieval and inverted indexing of the original text. It extracts core entity words from the original video title and performs text similarity retrieval across the entire database, for example, using BM25 retrieval. This approach has limited intelligence, cannot understand deep semantics, and is prone to false positives due to duplicate titles with different content.
[0271] The third type of approach is recommendation aggregation based on user collaborative filtering and behavioral patterns. It infers the correlation between videos based on a large number of users' continuous clicks and playback navigation behaviors. Although this approach can discover some potential correlations, it performs poorly in the cold start phase and cannot guarantee the rigor of the aggregation results in terms of narrative logic, such as not ensuring that the videos are arranged in episode order.
[0272] Based on the above, it can be seen that the relevant technologies have at least the following problems: First, the logic for mixing heterogeneous videos is complex and error-prone. In actual video compilation businesses, there are usually two types of heterogeneous data: main episodes with a definite number of episodes, and peripheral or bonus content without a definite number of episodes. The relevant technologies usually require writing complex multi-traversal, bucketing, or multi-judgment logic to achieve the requirement of main episodes in ascending order of episode number and bonus content in chronological order without interfering with the display requirements of main episodes. The code maintenance cost is high and boundary errors are prone to occur.
[0273] Second, misaligned data types can cause sorting errors. When extracting video episode numbers, the results are usually in string format. If strings containing numbers are directly sorted in normal lexicographical order, logical errors such as "episode 10 is placed before episode 2" can easily occur.
[0274] Third, the computational efficiency is low. Related technologies often require splitting the main feature and bonus content into two independent lists, applying different sorting rules to each, such as sorting one list by episode number and the other by time, and finally concatenating the lists. This multi-pass processing method increases unnecessary memory overhead and time complexity when facing massive concurrent requests.
[0275] In view of the problems existing in the above-mentioned related technologies, the video recommendation method provided in this application embodiment can solve at least the following technical problems. First, it proposes a heterogeneous sorting algorithm based on multidimensional tuple mapping. By introducing type isolation bits in the underlying data structure, it solves the problem of classification isolation and mixed sorting of main videos and surrounding videos in one go without the need for physical bucketing.
[0276] Second, a type cast is introduced during the tuple mapping process to convert string-formatted set information into integer values, thereby eliminating the set disorder caused by string lexicographical order. Third, the complex heterogeneous sorting requirements are transformed into a standard three-dimensional tuple lexicographical comparison. The sorted tuples are constructed through a single traversal, and a unified sort is performed based on the three-dimensional tuples, thereby improving the system's computational efficiency and robustness.
[0277] The video recommendation method provided in this application aims to offer a heterogeneous video serialization and reassembly method based on multi-path recall and multi-dimensional tuple mapping. This method addresses technical issues encountered during video collection construction, such as heterogeneous data mixing, missing episode metadata, and episode number errors caused by string-based sorting. Through a three-step processing architecture of recall cleaning—large model chain splitting—tuple mapping sorting, it achieves automatic recall, semantic segmentation, episode number extraction, and serialization sorting of video content within the same series or theme, thereby improving the accuracy, computational efficiency, and continuous viewing experience for users in video collection construction.
[0278] The video recommendation method provided in this application addresses technical issues such as heterogeneous data mixing, missing episode metadata, and episode number disorder caused by string sorting during the video collection construction process. It proposes a three-step processing architecture of recall cleaning, large model chain splitting, and tuple mapping sorting.
[0279] First, a multi-channel recall and multi-modal feature completion mechanism. Unlike single-retrieval methods, the video recommendation method provided in this application constructs a multi-channel parallel recall strategy based on title, series name, and video identifier. After obtaining the candidate set, it automatically deduplicates and integrates the underlying multi-modal features of the video, such as optical character recognition (OCR) text on screen and automatic speech recognition (ASR) text on screen, to complete the data, building a high-purity, high-information candidate video pool for subsequent accurate ranking.
[0280] Second, the semantic chain segmentation and degradation extraction strategy based on a large model. The video recommendation method provided in this application introduces a large language model for deep semantic understanding. First, it determines the theme consistency of candidate videos, dividing them into strict chains and relevant chains. The strict chains correspond to the main content, while the relevant chains correspond to peripheral content. Simultaneously, for episode extraction, the video recommendation method provided in this application constructs a degradation extraction strategy based on title, on-screen text OCR, and metadata priority. This dynamically normalizes unstructured episode descriptions into pure numbers or null values, effectively overcoming the problem of episode recognition failure caused by missing title information or incomplete metadata.
[0281] Third, a heterogeneous sorting algorithm based on multidimensional tuple mapping. For the two types of heterogeneous data generated after extraction—main videos with clearly defined episode numbers and surrounding videos without clearly defined episode numbers—the video recommendation method provided in this application abandons the traditional multi-list splitting and splicing logic and designs a multidimensional tuple mapping function. By constructing a three-dimensional tuple containing a type isolation bit, a forced integer episode number, and a publication timestamp for each node, the classification, isolation, and absolutely ordered arrangement of heterogeneous videos can be achieved at the underlying data structure level with only a single traversal. This eliminates episode number disorder caused by string lexicographical order and improves the logical coherence of video playback and system computational efficiency.
[0282] The following describes the application scenarios of the video recommendation method provided in this application embodiment, specifically targeting the automated construction of video series collections in news applications. News applications, as comprehensive information platforms, carry a large amount of video content, including feature reports, analyses of popular dramas, sports event recaps, interview series, etc. When a user watches a video in a news application, the system automatically identifies the theme to which the video belongs (i.e., multiple second videos belong to the same theme as the first video) and displays a fully sorted video collection to the user on the playback page (the video collection is used to view video collections recommended based on the first video), helping users efficiently access serialized content.
[0283] In the scenario of automatically constructing news feature video series, when a user browses the third installment of an in-depth analysis video of a hot topic in a news application, such as a press conference or sports event, the system automatically identifies that the current video belongs to the theme of a news feature. An entry point to all series collections (corresponding to a video collection control) is displayed at the bottom of the playback page. After clicking to enter, the user can see the current episodes, strictly ordered by episode number from smallest to largest (e.g., episode 1, episode 2 through N), along with related commentary, behind-the-scenes footage, interviews, and other supplementary videos. Users can play the entire collection continuously to fully understand the ins and outs of the event.
[0284] In scenarios involving aggregated video series of popular TV drama commentary or sports event recaps, when a user watches the commentary for episode 5 of a popular drama series in a news application, all episodes of the series' commentary are automatically aggregated, distinguishing between the main commentary and related content. The main commentary may include episode 1 commentary, episode 2 commentary, etc., while related content may include actor interviews, behind-the-scenes footage, and summaries of popular online comments. In sports event scenarios, when a user watches highlights of round 10 of a league, all rounds of the league's recap videos are automatically aggregated and sorted in ascending order by round. Related videos such as pre-match predictions, post-match reviews, and commentary from expert commentators are also aggregated and sorted by publication date from newest to oldest.
[0285] When a user browses and plays video content normally in the video recommendation stream, if it is detected that the currently playing video belongs to a certain topic, a series video entry label (i.e., a video collection control) will be automatically attached to the bottom of the video playback interface. This entry is used to guide the user to view the video collection to which the current video belongs.
[0286] Step 1: Trigger a series entry point in the recommendation stream. Users browse and play short videos normally in the news application's video recommendation stream (the first video is displayed in the media stream interface). When it is detected that the currently playing video belongs to a pre-built series collection, a series video entry tag bar is automatically attached to the bottom of the video playback interface (and a video collection control is displayed). The series collection name is displayed in a horizontal bar format. This entry tag does not obscure the main content of the video and does not affect the user's normal viewing experience. It only provides a lightweight indication of the series collection information to which the current video belongs at the bottom of the video. After the user clicks on this entry tag, the series skin display process is triggered (responding to the video viewing command triggered by the video collection control). If the current video is not associated with any series collection, the entry tag is not displayed, and the user experience is uninterrupted.
[0287] Step 2: Series Skin Showcase After the user clicks the "Series Videos" entry tab at the bottom, the recommended videos pause playback, and the interface switches to the series skin display mode. The series skin is presented as a full-screen immersive overlay and includes the following core information elements. First, the series collection (i.e., video collection) title. The series collection title is displayed in larger, bolder font at the top of the skin to emphasize the complete title of the series collection, with keywords highlighted in a striking color to enhance visual impact and recognizability.
[0288] Second, a brief introduction to the series. Below the title, the series introduction displays a summary of the series' content in body text, helping users quickly understand the series' theme. Third, the series background image. The series background image uses a frame from the currently playing video as the series skin's thumbnail, overlaid with a semi-transparent overlay, maintaining visual coherence while ensuring the readability of the text information. A "Continue Watching" button is overlaid above the video thumbnail, allowing users to return to the recommended stream and continue playing the current video.
[0289] Fourth, the "View Full Series Videos" button. This button is located at the bottom of the skin and is presented as a white rounded rectangle with the text "View Full Series Videos". Clicking this button will take the user to the bottom page of the series collection (responding to video viewing commands triggered by the video collection control).
[0290] Step 3: Displaying and Continuously Playing the Series Collection's Bottom Page. After clicking the "View Full Series Videos" button, the user is redirected from the series skin page to the series collection's bottom page, i.e., the collection details page. The top of this page retains the playback area for the current video, while the lower half displays the complete list of collection content (the video collection includes multiple second videos) in the form of a bottom pop-up panel.
[0291] The panel contains the following core elements: First, the collection title bar. The top of the collection title bar displays "Collection {Series Name}", and a close button is provided on the right for users to return to the recommended stream. Second, pagination navigation tabs. When there are many videos in a collection, they are displayed in the form of horizontally sliding pagination tabs, such as "1-10", "11-20", "21-30", "31-40", "41-50", etc., allowing users to quickly jump to the target episode range.
[0292] Third, the video list area. The video list area displays all video entries within the current page in a vertical list format (showing video information for each of the second videos). Each entry includes a serial number, video title, source media (the video information includes the media outlet that published the second video), number of comments, publication time, and a thumbnail and duration tag on the right. Videos in the list are strictly arranged in ascending order by episode number. Users can click on any video entry to jump to playback; after playback ends, it automatically proceeds to the next episode in sequence, providing a continuous playback experience.
[0293] The video recommendation method provided in this application has the following advantages: First, it features fully automated collection construction, eliminating the need for content operators to manually tag or edit collections. Based on multimodal features and semantic understanding using a large language model, it automatically discovers and organizes videos on the same topic, thereby reducing the operational costs of news video content. Second, it intelligently separates main programs from related content. Users can clearly distinguish between main programs and related bonus features on the collection page, avoiding the problem of users getting lost in the viewing progress due to the mixed display of main programs and bonus features in traditional collections, thus improving the efficiency of information video content consumption.
[0294] Third, accurate episode number sorting is guaranteed. The main episodes within the collection are strictly arranged in ascending order of numbers, eliminating the user experience flaw of "episode 10 appearing before episode 2," ensuring the logical correctness of users identifying events chronologically or following content according to the plot. Fourth, a progressive series discovery experience within the recommendation stream. The video recommendation method provided in this application embeds a lightweight "series video" entry tag (and displays a video collection control) within the video recommendation stream. Through a three-step progressive interactive design of "recommendation stream entry trigger—immersive display of series skin—complete browsing of the collection's bottom page," users can naturally discover and delve into series collection content within the recommendation stream without actively searching. The series skin layer serves as an intermediate transition, displaying core information such as the series title, introduction, and background image in a full-screen format. This avoids interrupting the user's browsing rhythm by directly jumping to a page, while stimulating the user's willingness to continue watching through enhanced display, ultimately achieving a seamless transition from fragmented recommendation browsing to systematic follow-up within the news app.
[0295] The video recommendation method provided in this application addresses issues such as the mixing of main content and extras during video collection construction and the easy disorder in episode number sorting. It obtains structured video node states through multi-path recall and preliminary cleaning of large models, and proposes a multi-dimensional tuple sorting algorithm to achieve efficient and accurate sorting of heterogeneous videos at the underlying data structure level.
[0296] The video recommendation method provided in this application includes the following steps: Step 1: Multi-path recall and candidate video cleaning. After receiving the original video trigger request, a multi-path recall strategy is initiated using the title of the original video, such as vector recall or search API recall, to obtain a candidate video set. Subsequently, the candidate set is cleaned and feature-completed, including deduplication based on the title and extraction of multimodal features for each candidate video, such as publication time, OCR on-screen text, and ASR speech text, thereby constructing a clean candidate video pool.
[0297] In step one, multi-channel parallel recall can be performed based on title, series name and video identifier. After obtaining the candidate set, deduplication is automatically performed, and the underlying multimodal features of the video, such as OCR on-screen text and ASR speech text, are integrated to complete the data, thereby building a high-purity and high-information candidate video pool for subsequent accurate sorting.
[0298] Step Two: Chain segmentation and feature extraction based on the large language model. The cleaned candidate video features are assembled into prompt words and input into the large language model. The large language model performs two preliminary tasks based on deep semantic understanding. The first task is chain segmentation. It determines the thematic consistency between the candidate video and the original video, dividing the video into strict chains and relevant chains. Among them, strict chains correspond to the main content, and relevant chains correspond to peripheral content.
[0299] The second task is episode number extraction. Episode number information is dynamically extracted from the video. If a specific episode number exists, a pure numeric string, such as "5", is output; if it's a bonus feature or related content without an episode number, an empty value is output. During episode number extraction, a degradation extraction strategy based on title, on-screen text OCR, and metadata priority can be used to dynamically normalize unstructured episode number descriptions into pure numbers or empty values, thus overcoming the problem of episode number recognition failure caused by missing title information or incomplete metadata.
[0300] Step 3: Based on a heterogeneous sorting algorithm using multidimensional tuples, after chain partitioning and feature extraction, a heterogeneous set containing nodes with and without episode numbers is obtained. To achieve the requirement of sorting the main feature by episode number in ascending order and the extras by time in the same sequence without interfering with the main feature, this heterogeneous set is traversed, and a three-dimensional sorting tuple is calculated and assigned to each node. The entire set is then sorted lexicographically once based on the key value of this tuple, ultimately outputting a logically rigorous video playback chain.
[0301] A multi-dimensional tuple sorting algorithm is proposed for heterogeneous video aggregation. A major technical challenge in constructing video collections lies in handling the sorting of heterogeneous data. Traditional methods such as single-field sorting or multi-list concatenation are inefficient and prone to string sorting anomalies, such as "10" appearing before "2".
[0302] The video recommendation method provided in this application proposes a deterministic multi-dimensional tuple mapping logic. Assume that the episode status of a video node after the preliminary steps is a string or an empty value, and its publication timestamp is a numerical value. A three-dimensional sorted tuple is constructed for each node, with the following mapping rules.
[0303] Rule A: When the episode number is not empty, i.e., the node is a slice with a definite episode number, the first dimension is 0, serving as a type isolation bit. Utilizing the lexicographical property of 0, all slices are forced to be placed at the very beginning of the sequence. The second dimension performs a forced integer conversion, converting the string to an integer, fundamentally avoiding sorting errors caused by string lexicographical order, ensuring that slices are strictly arranged in ascending order from 1, 2, 3 to 10. The third dimension uses a timestamp as a fallback sorting factor.
[0304] Rule B: When the episode number is empty, i.e., the node is a peripheral or bonus episode with no episode number, the first dimension is 1, serving as a type isolation bit. Utilizing the property that the number 1 is greater than 0, all peripheral videos are isolated after the main video without physical table partitioning. The second dimension serves as a placeholder to maintain tuple dimension consistency. The third dimension uses a timestamp to ensure that peripheral videos are sorted by time within the video.
[0305] Finally, the tuple set of all nodes is sorted in ascending lexicographical order using the underlying standard language. This algorithm abstracts complex business rules into mathematical mappings. It constructs a three-dimensional sorted tuple for each video node through a single traversal and performs a unified sorting of the video node set based on the three-dimensional sorted tuple, thereby achieving the classification, isolation, and ordered arrangement of heterogeneous video sequences.
[0306] When a user watches a video in the recommendation stream, the system first determines whether the video can trigger the creation of a series collection or whether it belongs to an already constructed series collection. If it can trigger the collection or a corresponding collection already exists, the system performs multi-path retrieval based on information such as the original video's title, series name, and video identifier, and combines OCR on-screen text, ASR voice text, and release time with multimodal features for completion and cleaning.
[0307] Subsequently, a large language model was used to perform semantic discrimination on the candidate videos to determine whether the candidate videos and the original videos belonged to the same topic, and to perform strict chain and related chain division on the candidate videos. For videos whose episode number or episode number could be identified, their episode number information was extracted; for videos without a clear episode number, their episode number was set to a null value.
[0308] During the sorting phase, instead of splitting the main video and surrounding videos into multiple lists for separate sorting, a unified three-dimensional sorting tuple is constructed for each video node. A type isolation bit distinguishes between the main video and surrounding videos, a forced integer episode number ensures the main videos are arranged in ascending numerical order, and a publication timestamp guarantees the chronological order within the same video category. Finally, a unified sorting based on the three-dimensional sorting tuple yields a complete, coherent, and logically correct video collection playback chain, which is displayed to users in the product interface through series entry points, series skins, and the collection's underlying page.
[0309] The video recommendation method provided in this application provides the following beneficial effects from both product experience and technical implementation perspectives: First, it reduces content operation costs. The video recommendation method provided in this application combines multimodal feature recognition with large language model semantic understanding to achieve fully automated construction of video series collections. This eliminates the need for manual tagging or editing of collections, significantly reducing the manpower costs and response time for operating informational video content.
[0310] Second, it improves content consumption efficiency. The video recommendation method provided in this application solves the problem of users getting lost in the viewing progress caused by the mixed display of heterogeneous videos in traditional compilations by intelligently partitioning the main content and extras. This allows users to quickly locate target content and effectively improves the completion rate and user retention rate of information videos. Third, it ensures the correctness of content consumption logic. The video recommendation method provided in this application is based on a heterogeneous video serialization and sorting algorithm using multi-dimensional tuple mapping. It abstracts complex business rules into mathematical mappings, and only requires a single traversal to complete the classification, isolation, and ordered arrangement of heterogeneous videos. This eliminates the issue of episode number sorting errors caused by string lexicographical order, such as "episode 10" being placed before "episode 2", thereby ensuring the logical correctness of users determining events according to the timeline or following content according to the plot.
[0311] Fourth, it enables seamless content discovery within the recommendation stream. The video recommendation method provided in this application adopts a three-step progressive interactive design of "recommendation stream entry trigger - immersive display of series skins - complete browsing of the collection's underlying page". Without interrupting the user's browsing rhythm, it guides users to naturally transition from fragmented single-video recommendation stream consumption to systematic series collection viewing, significantly improving the user's viewing time and content consumption depth in a single session.
[0312] Fifth, improve computational efficiency. The multidimensional tuple sorting algorithm of the video recommendation method provided in this application unifies the classification and sorting of heterogeneous videos into a single lexicographical comparison, eliminating the need for multiple table splitting, multiple sorting and reassembly. When implemented using a general sorting algorithm, the overall sorting process typically has a time complexity of O(n log n) and a space complexity of O(n), demonstrating good performance scalability in large-scale video collection construction scenarios.
[0313] The following description continues to illustrate the exemplary structure of the video recommendation device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software module stored in the video recommendation device 455 in the memory 450 may include: Display module 4551 is used to display the first video in the media stream interface and to display the video collection control; The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same theme as the first video. Response module 4552 is used to respond to a video viewing command triggered based on the video collection control and display video information of each of the second videos; The video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to multiple pieces of video information is different.
[0314] In some embodiments, the response module 4552 is further configured to, before displaying the video information of each of the second videos, in response to a trigger operation on the video collection control, display collection information of the video collection and display a first viewing control for viewing the video information of the second videos; wherein the collection information includes at least one of the following: the name of the video collection, the cover of the video collection, and a description of the video collection; and in response to a trigger operation on the first viewing control, trigger the video viewing instruction.
[0315] In some embodiments, the first video belongs to different candidate video collections, and different candidate video collections correspond to different types. The response module 4552 is further configured to, in response to a trigger operation on the video collection control, display multiple candidate types, wherein the candidate type is the type of the candidate video collection to which the first video belongs, and the multiple candidate types include the target type of the video collection; and in response to a trigger operation on the target type among the multiple candidate types, display collection information of the video collection of the target type.
[0316] In some embodiments, the display module 4551 is further configured to display at least one of the following information for each candidate type: the number of objects watching videos in the video collection of the candidate type, the recommendation level for the video collection of the candidate type, and the update time of the video collection of the candidate type.
[0317] In some embodiments, the collection information of the video collection is displayed on the information display interface. The display module 4551 is further configured to display the video collection control using a first display style, the first display style being used to guide and trigger the video collection control. The response module 4552 is further configured to, in response to a trigger operation on the video collection control for the first display style, display a playback control for playing the first video in the information display interface; in response to a trigger operation on the playback control, return to the media stream interface to play the first video, and display the video collection control using a second display style; wherein the second display style is different from the first display style, and the second display style is used to indicate that the video collection control has been triggered.
[0318] In some embodiments, the response module 4552 is further configured to, in response to a trigger operation on the video collection control, display the viewing progress of the current object watching the video collection; when the viewing progress reaches a preset viewing progress, display the virtual reward resources corresponding to the preset viewing progress, and display a claim control for claiming the virtual reward resources.
[0319] In some embodiments, the response module 4552 is further configured to, in response to a video viewing instruction triggered based on the video collection control, display the first video in a first display area of the media stream interface; if the number of the second videos is less than or equal to a first quantity threshold, display video information of each of the second videos in a second display area of the media stream interface; if the number of the second videos is greater than the first quantity threshold, display a plurality of filter tags in the second display area of the media stream interface, the plurality of filter tags including a target filter tag; and, in response to a triggering operation for the target filter tag, display video information of each of the second videos filtered according to the target filter tag in the second display area.
[0320] In some embodiments, the display module 4551 is further configured to display a setting control before displaying the video information of each of the second videos in response to a video viewing instruction triggered based on the video collection control. The setting control is used to set the filtering criteria for filtering the second videos. The response module 4552 is also configured to respond to a setting instruction triggered based on the setting control, and if the number of the second videos is greater than the first number threshold, display the filtering criteria set by the setting instruction, and display multiple filtering labels after filtering the second videos according to the filtering criteria in the second display area of the media stream interface.
[0321] In some embodiments, the display module 4551 is further configured to display the video collection control when at least one of the following conditions is met: the duration of playing the first video is greater than a duration threshold; or an interactive operation is performed on the first video.
[0322] In some embodiments, the display module 4551 is further configured to, when displaying the first video, if there is a video collection including the first video, display a collection switch; wherein the collection switch is used to enable the video collection mode; and when the video collection mode is enabled based on the collection switch, display a video collection control.
[0323] In some embodiments, the response module 4552 is further configured to, after displaying the video information of each of the second videos, cancel the display of the video information of the target second video in response to a cancellation display instruction for the video information of the target second video among the plurality of second videos; and when the number of video information of the second videos that have not been canceled decreases to a second quantity threshold, display an add control; wherein the add control is configured to add videos to be displayed based on the content of the second videos that have not been canceled.
[0324] In some embodiments, the response module 4552 is further configured to, after the display of the add control, in response to a triggering operation for the add control, display at least one add condition; wherein the add condition is a condition satisfied by the added video; and in response to a selection operation for a target add condition among the at least one add condition, display an added third video; wherein the third video is determined based on the fact that the second video was not canceled from display and meets the target add condition.
[0325] In some embodiments, each of the second video's video information has a one-to-one corresponding sequence number. The display module 4551 is further configured to display the video information of each of the second videos in ascending order of the sequence number. The response module 4552 is further configured to, in response to a cancellation display instruction for the video information of a target second video among the plurality of second videos, update the sequence number of the second videos arranged after the target second video, and re-display the video information of the remaining second videos according to the updated sequence number.
[0326] In some embodiments, the content of the first video is a first sub-event of the target event occurring at a first time point. The display module 4551 is further configured to display a progress axis with at least one node, wherein each node has a corresponding sub-event, and the sub-event belongs to the target event. In the associated area of each node in the progress axis, video information of the second video corresponding to the node is displayed. The content of the second video is at least one of the following: a second sub-event of the target event occurring at the first time point, and a third sub-event of the target event occurring at a second time point.
[0327] In some embodiments, the response module 4552 is further configured to display a shared viewing control for the video collection after displaying the video information of each of the second videos; wherein the shared viewing control is configured to invite other objects to watch the videos in the video collection together; based on the shared viewing control, in response to a shared viewing instruction for the other objects, an invitation message for inviting them to watch the video collection together is sent to the other objects.
[0328] In some embodiments, the response module 4552 is further configured to, upon receiving consent information, respond to a selection operation triggered by video information for the second video, and if the selected video is consistent with the video selected by the other object, display a shared viewing interface and display the selected video in the shared viewing interface, wherein the consent information is used to indicate that the other object agrees to share the video collection based on the invitation information; Upon receiving the consent information, in response to the selection operation triggered by the video information of the second video, if the selected video is inconsistent with the video selected by the other object, the selected video is displayed, and the object identifier and viewing progress of the other object are displayed in the associated area of the video information of the video selected by the other object.
[0329] In some embodiments, the response module 4552 is further configured to, after displaying the video information of each of the second videos, control the selected second video to be in a selected state in response to the selection operation of the plurality of second videos; wherein the number of second videos in the selected state is less than the number of second videos in the video collection; and in response to the sharing instruction for the second video in the selected state, create a new video collection of the second video in the selected state, and share the new video collection according to the sharing instruction.
[0330] In some embodiments, the response module 4552 is further configured to respond to a video viewing instruction triggered based on the video collection control, and if the second video included in the video collection is a portion of the videos selected by other objects from the target video collection, display a second viewing control; wherein the second viewing control is used to view the target video collection; The response module 4552 is further configured to, after displaying the video information of each of the second videos, in response to a trigger operation on the second viewing control, display the video information of the second videos and switch to the video information of the videos in the target video collection.
[0331] In some embodiments, the response module 4552 is further configured to, after displaying video information for each of the second videos, play the second video in response to a playback instruction for the second video; and when the number of unplayed second videos decreases to a third quantity threshold, display at least one recommended video collection based on the video collection.
[0332] In some embodiments, the second video includes a target video that has an arrangement order with the first video, and other videos that supplement the content of the first video; the response module 4552 is further configured to display a first playback control and a second playback control after the first video is displayed in the media stream interface and the first video finishes playing; wherein, the first playback control is used to play the target video, and the second playback control is used to play the other videos; in response to a trigger operation on the first playback control, the target video that is adjacent to the first video in the arrangement order and follows the first video is played; in response to a trigger operation on the second playback control, the other videos are played; when the other videos finish playing, the target video that is adjacent to the first video in the arrangement order and follows the first video is played.
[0333] In some embodiments, the plurality of second videos includes a target second video. The response module 4552 is further configured to, after displaying the video information of each of the second videos, in response to a trigger operation on the video information of the target second video, switch the display of the first video to display the target second video; if the target second video also belongs to other video collections, display a jump control; and in response to a trigger operation on the jump control, switch the displayed video information of each of the second videos to the video information of videos in the other video collections.
[0334] This application provides a computer program product, which includes a computer program or computer-executable instructions. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the video recommendation method provided in this application. For example, ... Figure 3 The illustrated video recommendation method involves an electronic device's processor reading a computer program or computer-executable instructions from a computer-readable storage medium, executing the computer program or computer-executable instructions, and causing the electronic device to perform the video recommendation method described in the embodiments of this application.
[0335] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the video recommendation method provided in this application. For example, ... Figure 3 The video recommendation method shown.
[0336] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0337] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0338] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be part of a file of other programs or data, such as in one or more scripts stored in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files storing one or more modules, subroutines, or code sections).
[0339] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0340] In the video recommendation method provided in this application embodiment, a first video is displayed in the media stream interface, and a video collection control is also displayed. The video collection control is used to view a video collection recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same topic as the first video. In response to a video viewing instruction triggered based on the video collection control, video information of each second video is displayed. The video information includes the video publishing object that published the second video, and at least one of the video publishing objects corresponding to the multiple video information is different.
[0341] The video recommendation method provided in this application displays a video collection control instead of preloading the media stream data or video information of all second videos in the video collection when displaying the first video. This effectively saves network bandwidth and memory consumption during the operation of the electronic device, ensuring the smooth playback of the first video. Responding to video viewing instructions triggered by the video collection control, the video information of each second video is displayed, which is equivalent to on-demand loading. This reduces graphics rendering calculations, thereby reducing processor power consumption and computational load. Furthermore, responding to video viewing instructions triggered by the video collection control, video information of second videos belonging to the same topic but from different video publishers can be retrieved in batches. This avoids the electronic device initiating multiple independent search requests or multiple network handshakes to obtain videos published by different video publishers, significantly reducing the frequency of network input and output interactions of the electronic device and improving data acquisition efficiency. Moreover, since the first video and the second videos in the recommended video collection belong to the same topic, the accuracy of video recommendations is improved.
[0342] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A video recommendation method, characterized in that, The method includes: Display the first video in the media stream interface, and also display the video collection control; The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same theme as the first video. In response to a video viewing command triggered by the video collection control, video information of each of the second videos is displayed; The video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to multiple pieces of video information is different.
2. The method according to claim 1, characterized in that, Before displaying the video information of each of the second videos, the method further includes: In response to a trigger operation on the video collection control, the collection information of the video collection is displayed, and a first viewing control for viewing the video information of the second video is displayed; The collection information includes at least one of the following: the name of the video collection, the cover of the video collection, and a brief description of the video collection; In response to a triggering operation on the first viewing control, the video viewing instruction is triggered.
3. The method according to claim 2, characterized in that, The first video belongs to different candidate video collections, and different candidate video collections correspond to different types. The step of displaying the collection information of the video collection in response to a trigger operation on the video collection control includes: In response to a trigger operation on the video collection control, multiple candidate types are displayed. The candidate types are the types of the candidate video collection to which the first video belongs, and the multiple candidate types include the target type of the video collection. In response to a trigger operation for the target type among the plurality of candidate types, the collection information of the video collection of the target type is displayed.
4. The method according to claim 3, characterized in that, The method further includes: For each candidate type, display at least one of the following information: The number of people watching videos in the candidate video collection, the recommendation level for the candidate video collection, and the update time of the candidate video collection.
5. The method according to claim 2, characterized in that, The video collection information is displayed on the information display interface, and the video collection display control includes: The video collection control is displayed using a first display style, which is used to guide and trigger the video collection control. The method further includes: In response to a trigger operation on the video collection control for the first display style, a playback control for playing the first video is displayed in the information display interface; In response to a trigger operation on the playback control, the system returns to the media stream interface to play the first video and displays the video collection control using a second display style. The second display style is different from the first display style, and the second display style is used to indicate that the video collection control has been triggered.
6. The method according to claim 2, characterized in that, The method further includes: In response to a trigger operation on the video collection control, display the viewing progress of the current object watching the video collection; When the viewing progress reaches the preset viewing progress, the virtual reward resources corresponding to the preset viewing progress are displayed, and a claiming control for claiming the virtual reward resources is displayed.
7. The method according to claim 1, characterized in that, The step of displaying video information for each of the second videos in response to a video viewing command triggered by the video collection control includes: In response to a video viewing command triggered by the video collection control, the first video is displayed in the first display area of the media stream interface; If the number of the second videos is less than or equal to the first number threshold, then the video information of each of the second videos is displayed in the second display area of the media stream interface; If the number of the second video is greater than the first number threshold, then multiple filter tags are displayed in the second display area of the media stream interface, including the target filter tag; In response to a trigger operation targeting the target filter tag, video information of each of the second videos filtered according to the target filter tag is displayed in the second display area.
8. The method according to claim 7, characterized in that, Before displaying the video information of each of the second videos in response to a video viewing instruction triggered based on the video collection control, the method further includes: Display settings controls, which are used to set the filtering criteria for filtering the second video; If the number of the second videos exceeds the first quantity threshold, then multiple filter tags are displayed in the second display area of the media stream interface, including: In response to a setting instruction triggered by the setting control, if the number of the second videos is greater than the first quantity threshold, the filtering criteria set by the setting instruction are displayed, and multiple filtering labels after filtering the second videos according to the filtering criteria are displayed in the second display area of the media stream interface.
9. The method according to claim 1, characterized in that, The video collection display control includes: The video collection control is displayed when at least one of the following conditions is met: The duration of playing the first video is greater than the duration threshold; Perform interactive operations on the first video.
10. The method according to claim 1, characterized in that, The method further includes: When the first video is displayed, if there is a video collection that includes the first video, the collection switch is displayed; The collection switch is used to enable the video collection mode; The video collection display control includes: When the video collection mode is enabled based on the collection switch, the video collection control is displayed.
11. The method according to claim 1, characterized in that, After displaying the video information of each of the second videos, the method further includes: In response to a cancellation display instruction for video information of a target second video among the plurality of second videos, the display of video information of the target second video is cancelled; When the amount of video information of the second video that has not been canceled decreases to a second quantity threshold, the increase control is displayed; The add control is used to add a video to be displayed based on the content of the second video that has not been canceled.
12. The method according to claim 11, characterized in that, After adding the control to the display, the method further includes: In response to a triggering operation on the added control, at least one added condition is displayed; The added condition refers to the condition that the added video must satisfy; In response to a selection operation for a target addition condition among the at least one addition condition, an additional third video is displayed; The third video is determined based on the fact that the second video was not canceled from display, and it meets the target addition condition.
13. The method according to claim 11, characterized in that, Each of the second videos has a unique serial number, and the display of the video information for each second video includes: The video information of each of the second videos is displayed in ascending order of the serial numbers; The method further includes: In response to a cancellation instruction for the video information of a target second video among the plurality of second videos, the sequence number of the second videos arranged after the target second video is updated, and the video information of the remaining second videos is re-displayed according to the updated sequence number.
14. The method according to claim 1, characterized in that, The content of the first video is the first sub-event of the target event occurring at a first time point, and the video information for displaying each of the second videos includes: Display a progress axis with at least one node, wherein each node has a corresponding sub-event, and the sub-event belongs to the target event; In the associated area of each node in the progress axis, the video information of the second video corresponding to each node is displayed; The content of the second video is at least one of the following: a second sub-event of the target event occurring at the first time point, and a third sub-event of the target event occurring at the second time point.
15. The method according to any one of claims 1-14, characterized in that, After displaying the video information of each of the second videos, the method further includes: Display common viewing controls for the aforementioned video collection; The shared viewing control is used to invite other objects to watch the videos in the video collection together. Based on the shared viewing control, in response to a shared viewing instruction for the other objects, an invitation message is sent to the other objects to invite them to watch the video collection together.
16. The method according to claim 15, characterized in that, The method further includes: Upon receiving consent information, in response to a selection operation triggered by the video information of the second video, if the selected video is consistent with the video selected by the other object, a shared viewing interface is displayed, and the selected video is displayed in the shared viewing interface, wherein the consent information is used to indicate that the other object agrees to share the video collection based on the invitation information; Upon receiving the consent information, in response to the selection operation triggered by the video information of the second video, if the selected video is inconsistent with the video selected by the other object, the selected video is displayed, and the object identifier and viewing progress of the other object are displayed in the associated area of the video information of the video selected by the other object.
17. The method according to any one of claims 1-14, characterized in that, After displaying the video information of each of the second videos, the method further includes: In response to a selection operation on the plurality of second videos, the selected second video is controlled to be in a selected state; The number of second videos that are selected is less than the number of second videos in the video collection. In response to a sharing instruction for a second video that is selected, a new video collection of the second video that is selected is created, and the new video collection is shared according to the sharing instruction.
18. The method according to any one of claims 1-14, characterized in that, The method further includes: In response to a video viewing instruction triggered by the video collection control, if the second video included in the video collection is a subset of videos selected by other objects from the target video collection, the second viewing control is displayed; The second viewing control is used to view the target video collection; After displaying the video information of each of the second videos, the method further includes: In response to a trigger operation on the second viewing control, the video information of the second video will be displayed, and then switched to the video information of the videos in the target video collection.
19. The method according to any one of claims 1-14, characterized in that, After displaying the video information of each of the second videos, the method further includes: In response to a playback command for the second video, the second video is played; When the number of unplayed second videos decreases to a third threshold, at least one recommended video collection based on the video collection is displayed.
20. The method according to any one of claims 1-14, characterized in that, The second video includes a target video that has an order with the first video, as well as other videos that supplement the content of the first video; After displaying the first video in the media streaming interface, the method further includes: When the first video finishes playing, the first playback control and the second playback control are displayed; The first playback control is used to play the target video, and the second playback control is used to play the other videos; In response to a trigger operation on the first playback control, play the target video that is adjacent to the first video and follows the first video in the arrangement order; In response to a trigger operation on the second playback control, the other video is played; After the other videos have finished playing, the target video that is adjacent to the first video and follows the first video in the order of arrangement will be played.
21. The method according to any one of claims 1-14, characterized in that, The plurality of second videos includes a target second video, and after displaying the video information of each of the second videos, the method further includes: In response to a trigger operation on video information of the target second video, the display of the first video will be switched to display of the target second video; If the target second video belongs to other video collections, a jump control is displayed; In response to a trigger operation on the jump control, the video information of each of the second videos displayed will be switched to the video information of the other videos in the video collection.
22. A video recommendation device, characterized in that, The device includes: The display module is used to display the first video in the media stream interface and to display the video collection control; The video collection control is used to view video collections recommended based on the first video. The video collection includes multiple second videos, and the multiple second videos belong to the same theme as the first video. The response module is used to respond to a video viewing command triggered based on the video collection control and display the video information of each of the second videos; The video information includes the video publishing object that publishes the second video, and at least one of the video publishing objects corresponding to multiple pieces of video information is different.
23. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the video recommendation method according to any one of claims 1 to 21.
24. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the video recommendation method according to any one of claims 1 to 21.
25. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the video recommendation method according to any one of claims 1 to 21.