A server and media recommendation method

By training a media asset recommendation model and combining media asset attributes and user behavior, text sequences are generated for media asset recommendation, which solves the problem of low accuracy in existing media asset recommendation technologies and achieves the effect of accurate recommendation.

CN116150472BActive Publication Date: 2026-05-12JUHAOKAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JUHAOKAN TECH CO LTD
Filing Date
2022-09-21
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, media asset recommendation methods for smart TVs only recommend media assets based on user profiles, resulting in low accuracy. There is an urgent need for more precise recommendations based on media asset content and user behavior.

Method used

By training a media asset recommendation model, the system obtains the attribute information of media assets, including tags, names, and actors, generates text sequences, inputs them into the pre-trained media asset recommendation model, and outputs recommended media assets that match the attribute information. The system then combines this with the user's historical viewing records to make accurate recommendations.

Benefits of technology

It improves the accuracy and effectiveness of media asset recommendations, ensuring that the recommended media assets are relevant to user interests and content, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150472B_ABST
    Figure CN116150472B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a kind of server and media recommendation method, relate to media recommendation technical field.Therein, server includes: controller, is configured as: receiving the first media identification sent by client;Obtain the first attribute information of the first media indicated by first media identification, first attribute information includes first media label, first media name, first media description, first media actor;According to first attribute information, determine the first text sequence of first media;First text sequence is input into pre-trained media recommendation model, obtains the recommended media matched with first attribute information that media recommendation model outputs;Obtain the first media indicated by first media identification;Recommended media and first media are sent to client.The embodiment of the present disclosure is used to improve the accuracy of media recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of media asset recommendation technology, and more particularly to a server and a media asset recommendation method. Background Technology

[0002] Smart TVs are essential devices for watching movies, TV shows, variety shows, news, and other media. To facilitate user viewing, some smart TVs offer media recommendations. In this technology, recommended media is selected from a media database based on user profiles. These user profiles include data predicting the user's interests based on their historical viewing history. However, recommending media solely based on user profiles lacks accuracy. Currently, there is a pressing need for media recommendation models that provide precise recommendations based on both content and user behavior. Summary of the Invention

[0003] To address, or at least partially address, the aforementioned technical problems, this disclosure provides a server and a media asset recommendation method. This method can train a media asset recommendation model that accurately recommends media assets to users based on media asset content and user behavior, thereby improving the recommendation effect.

[0004] To achieve the above objectives, the technical solutions provided by the embodiments of this disclosure are as follows:

[0005] In a first aspect, this disclosure provides a display device, comprising:

[0006] The controller is configured to receive the first media asset identifier sent by the client;

[0007] Obtain the first attribute information of the first media asset indicated by the first media asset identifier. The first attribute information includes the first media asset tag, the first media asset name, the first media asset description, and the first media asset actor.

[0008] Determine the first text sequence of the first media asset based on the first attribute information;

[0009] Input the first text sequence into the pre-trained media asset recommendation model, and obtain the recommended media assets output by the media asset recommendation model that match the first attribute information;

[0010] Obtain the first media asset indicated by the first media asset identifier;

[0011] Recommended media assets and first media assets will be sent to the client.

[0012] Secondly, this disclosure provides a media asset recommendation method, including:

[0013] Receive the first media asset identifier sent by the client;

[0014] Obtain the first attribute information of the first media asset indicated by the first media asset identifier. The first attribute information includes the first media asset tag, the first media asset name, the first media asset description, and the first media asset actor.

[0015] Determine the first text sequence of the first media asset based on the first attribute information;

[0016] Input the first text sequence into the pre-trained media asset recommendation model, and obtain the recommended media assets output by the media asset recommendation model that match the first attribute information;

[0017] Obtain the first media asset indicated by the first media asset identifier;

[0018] Recommended media assets and first media assets will be sent to the client.

[0019] Thirdly, this disclosure provides a computer-readable storage medium, including: storing a computer program on the computer-readable storage medium, wherein when the computer program is executed by a processor, it implements the media asset recommendation method as shown in the second aspect.

[0020] Fourthly, this disclosure provides a computer program product comprising a computer program that, when run on a computer, causes the computer to implement the media asset recommendation method as described in the second aspect.

[0021] The technical solution provided in this disclosure has the following advantages compared with the prior art:

[0022] This disclosure provides a server and a media asset recommendation method. The server first receives a first media asset identifier sent by a client. Then, based on the first media asset identifier, it obtains first attribute information of the first media asset corresponding to the first media asset identifier. This attribute information includes a first media asset tag, a first media asset name, a first media asset description, and first media asset actors. Next, it determines a first text sequence of the first media asset based on the first attribute information. This first text sequence is then input into a pre-trained media asset recommendation model, which is pre-trained based on the media asset attribute information and associated preferred media assets. The media asset attribute information reflects the media asset content, and the associated preferred media assets are determined based on the user's historical viewing records. The server obtains recommended media assets that match the first attribute information output by the media asset recommendation model, and then obtains the first media asset corresponding to the first media asset identifier. This first media asset and the recommended media assets are sent together to the client. This achieves targeted acquisition and recommendation of content-related media assets to the client based on the content of the first media asset corresponding to the first media asset identifier. By combining user behavior and media asset content for accurate recommendations, the recommendation effect is improved. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0024] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1A These are schematic diagrams of scenarios provided in some embodiments of the present disclosure;

[0026] Figure 1B A schematic diagram of the interface of display device 200;

[0027] Figure 2 A configuration block diagram of the control device 100 provided in an embodiment of this disclosure;

[0028] Figure 3 This is a hardware configuration block diagram of a display device 200 provided in an embodiment of the present disclosure;

[0029] Figure 4 This is a schematic diagram of the software configuration in a display device 200 according to one or more embodiments of the present disclosure;

[0030] Figure 5 A schematic diagram showing the icon control interface of an application in a display device 200 provided in this embodiment of the present disclosure;

[0031] Figure 6 A schematic flowchart illustrating a media asset recommendation method provided in this embodiment of the disclosure;

[0032] Figure 7 A flowchart illustrating the training process of the media asset recommendation model provided in this embodiment of the disclosure;

[0033] Figure 8 An exemplary knowledge graph is shown in an embodiment of this disclosure;

[0034] Figure 9 This is a schematic diagram of the server structure provided in an embodiment of this disclosure. Detailed Implementation

[0035] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0036] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0037] Currently, there are two main approaches to media asset recommendation technologies. One is based on user viewing behavior modeling. This method analyzes user profiles by acquiring historical viewing records, identifies media asset characteristics that users are interested in, and then matches and recommends media assets from a database based on these characteristics. The other is based on film content. This method acquires characteristic parameters reflecting film content, such as text descriptions and keywords, models each film to output a vector, and uses this vector to search for and recommend related films during media asset recommendation. However, the process of modeling and vectorizing each film often uses supervised learning methods, which is time-consuming, labor-intensive, and has high manual costs, resulting in poor media asset recommendation performance.

[0038] To address the aforementioned issues, this disclosure provides a server and a media asset recommendation method. The server first receives a first media asset identifier sent by a client. Then, based on the first media asset identifier, it obtains first attribute information of the first media asset corresponding to the first media asset identifier. This attribute information includes a first media asset tag, a first media asset name, a first media asset description, and first media asset actors. Next, based on the first attribute information, it determines a first text sequence of the first media asset. This first text sequence is then input into a pre-trained media asset recommendation model, which is pre-trained based on user behavior and media asset content. The model then obtains recommended media assets that match the first attribute information, and subsequently obtains the first media asset corresponding to the first media asset identifier. Both the first media asset and the recommended media assets are sent to the client. This achieves targeted acquisition and recommendation of media assets related to the content of the first media asset corresponding to the first media asset identifier, combining user behavior and media asset content for accurate recommendations and improving recommendation effectiveness.

[0039] Figure 1A These are schematic diagrams illustrating scenarios from some embodiments provided in this disclosure. For example... Figure 1A As shown in the figure, the device includes a control device 100, a display device 200, a smart device 300, and a server 400. Users can operate the display device 200 through the smart device 300 or the control device 100 to play media resources on the display device 200.

[0040] Taking the example of a user operating the display device 200 through the control device 100, in a scenario where the user uses the display device 200 to play a video, such as... Figure 1B As shown, Figure 1BThe diagram shows the interface of the display device 200. The user inputs information on the interface through the control device 100 and sends a query command to the display device 200. The query command includes a first media asset identifier and is used to instruct the user to obtain the first media asset indicated by the first media asset identifier and the recommended media asset.

[0041] When the display device 200 receives the query instruction, it sends a request instruction to the server 400 in response to the query instruction. The request instruction includes a first media asset identifier and is used to request the first media asset indicated by the first media asset identifier and related recommended media assets.

[0042] Server 400 receives a request instruction from display device 200 acting as a client. In response, it obtains a first media asset identifier. Based on this identifier, server 400 queries a database for the first attribute information of the first media asset indicated by the identifier. This first attribute information includes a first media asset tag, name, description, and actors. Based on this first attribute information, server 400 determines a first text sequence for the first media asset. This text sequence is then input into a pre-trained media asset recommendation model to obtain recommended media assets that match the first attribute information. Finally, server 400 retrieves the first media asset indicated by the identifier and returns it along with the recommended media assets to display device 200. This satisfies the user's request for the first media asset and accurately recommends related media assets to the user. The media asset recommendation model is pre-trained based on the user's historical viewing behavior and media content. Therefore, the recommended media assets sent to the display device accurately grasp the user's interests and ensure that the content of the recommended media assets is relevant to the first media asset, resulting in high recommendation accuracy and good recommendation effect.

[0043] In some embodiments, the control device 100 may be a remote control, and communication between the remote control and the display device may include infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. Users can input user commands through buttons on the remote control, voice input, control panel input, etc., to control the display device 200. In some embodiments, mobile terminals, tablet computers, computers, laptops, and other smart devices may also be used to control the display device 200.

[0044] In some embodiments, the smart device 300 can install software applications with the display device 200 to achieve connection and communication via network communication protocols, enabling one-to-one control operations and data communication. Audio and video content displayed on the smart device 300 can also be transmitted to the display device 200 for synchronized display. The display device 200 also communicates with the server 400 via various communication methods. The display device 200 can communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 can provide various content and interactive features to the display device 200. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing broadcast television reception functions, the display device 200 can also be equipped with a smart network television function that provides computer support.

[0045] Figure 2 This is a configuration block diagram of the control device 100 provided in an embodiment of this disclosure. (See diagram below.) Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input commands and convert them into commands that the display device 200 can recognize and respond to, acting as an intermediary for interaction between the user and the display device 200. The communication interface 130 is used for external communication and includes at least one of a Wi-Fi chip, a Bluetooth module, NFC, or a replacement module. The user input / output interface 140 includes at least one of a microphone, a touchpad, a sensor, buttons, or a replacement module.

[0046] Figure 3 This is a hardware configuration block diagram of a display device 200 provided in an embodiment of this disclosure. For example... Figure 3 The display device 200 includes: a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, etc. The controller 250 includes a central processing unit, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to nth interface for input / output. The display 260 can be at least one of a liquid crystal display, an OLED display, a touch display, and a projection display, and can also be a projection device and a projection screen. The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or signals interacting with the external environment. The controller 250 and the tuner / demodulator 210 can be located in different separate devices; that is, the tuner / demodulator 210 can also be an external device of the main device where the controller 250 is located, such as an external set-top box.

[0047] In some embodiments, the display device described above is a terminal device with display function, such as a television, mobile phone, computer, or learning machine.

[0048] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200. The user can input user commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input user commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0049] The output interface (monitor 260, and / or audio output interface 270) is configured to output user interaction information;

[0050] Communicator 220 is used to communicate with server 400 or other devices.

[0051] This disclosure provides a server, the server comprising:

[0052] Controller 250 is configured to receive a first media asset identifier sent by the client;

[0053] Obtain the first attribute information of the first media asset indicated by the first media asset identifier. The first attribute information includes the first media asset tag, the first media asset name, the first media asset description, and the first media asset actor.

[0054] Determine the first text sequence of the first media asset based on the first attribute information;

[0055] Input the first text sequence into the pre-trained media asset recommendation model, and obtain the recommended media assets output by the media asset recommendation model that match the first attribute information;

[0056] Obtain the first media asset indicated by the first media asset identifier;

[0057] Recommended media assets and first media assets will be sent to the client.

[0058] The aforementioned server utilizes a media asset recommendation model trained based on media asset attribute information and associated preferred media assets to achieve accurate recommendations. The media asset attribute information reflects the content of the media asset, while the associated preferred media assets are determined based on the user's historical viewing history. During the recommendation process, the server first obtains the first media asset identifier sent by the client, then obtains the first attribute information of the first media asset indicated by the first media asset identifier, determines the first text sequence based on the first attribute information, and inputs the first text sequence into the media asset recommendation model. This yields recommended media assets output by the model that match the first attribute information, thus improving the accuracy and effectiveness of the recommendations.

[0059] In some embodiments, the controller 250 is configured to determine the first text sequence of the first media asset based on the first attribute information, and to: obtain a target template set; determine the first text sequence of the first media asset based on the first media asset name, the first media asset actor, the first media asset description and the target template set.

[0060] In some embodiments, the controller 250 is configured to: obtain second attribute information of multiple second media assets from a database; wherein the second attribute information of the second media assets includes second media asset tags, second media asset names, second media asset descriptions, and second media asset actors; perform sentence segmentation on the second media asset description of each of the multiple second media assets to obtain a sentence sequence; wherein the second media asset description of each second media asset corresponds to a sentence in the sentence sequence; remove the second media asset name and second media asset actor included in each sentence in the sentence sequence to obtain a candidate template set, wherein each second media asset corresponds to a candidate template in the candidate template set; and vectorize and cluster the candidate templates included in the candidate template set to determine the target template set.

[0061] In some embodiments, the controller 250, training a media asset recommendation model, is configured to: acquire a first dataset; the first dataset includes second text sequences of multiple second media assets and second media asset labels of multiple second media assets; acquire a second dataset; the second dataset includes positive samples and negative samples, the positive samples include preference media assets with first-order association, and the negative samples include preference media assets with multi-order association; determine a first loss function corresponding to the first dataset and a second loss function corresponding to the second dataset;

[0062] The first dataset is input into the initial media asset recommendation model to obtain the predicted media asset labels output by the initial media asset recommendation model; and the second dataset is input into the initial media asset recommendation model to obtain the predicted associated media assets output by the initial media asset recommendation model; a first loss value is calculated based on the second media asset labels, the predicted media asset labels, and the first loss function; and a second loss value is calculated based on positive samples, negative samples, predicted associated media assets, and the second loss function; a target loss value is determined based on the first loss value and the second loss value; and the model parameters in the initial media asset recommendation model are adjusted based on the target loss value until a converged media asset recommendation model is obtained, wherein the input of the media asset recommendation model is the text sequence corresponding to the attribute information of the media asset, and the output is the recommended media asset that matches the attribute information of the media asset.

[0063] In some embodiments, the controller 250 is configured to: acquire a target template set and second attribute information of multiple second media assets; determine a second text sequence of multiple second media assets using the target template set for the second media asset name, second media asset actor, and second media asset description of each of the multiple second media assets; and obtain the first dataset based on the second text sequence of the multiple second media assets and the second media asset tags of the multiple second media assets.

[0064] In some embodiments, the controller 250 is configured to acquire a second dataset by: acquiring second media asset identifiers of multiple preferred media assets viewed in the past; determining the correlation coefficients between the multiple preferred media assets; constructing a knowledge graph based on the correlation coefficients, wherein the knowledge graph has preferred media assets as nodes and correlation coefficients between preferred media assets as edges; determining positive samples and negative samples based on the knowledge graph, wherein positive samples include preferred media assets with first-order correlations and negative samples include preferred media assets with multi-order correlations; and obtaining a second dataset based on the positive samples and negative samples.

[0065] In some embodiments, before acquiring the second media asset identifier of multiple preferred media assets viewed in the past, the controller 250 is further configured to: acquire user information sent by the client;

[0066] The controller 250, which acquires the second media asset identifier of multiple preferred media assets viewed in the past, is configured to: acquire the second media asset identifier of multiple preferred media assets viewed in the past corresponding to user information.

[0067] In some embodiments, the number of media asset recommendation models is N, and each media asset recommendation model has a corresponding relationship with user information;

[0068] Before the controller 250 inputs the first text sequence into the pre-trained media asset recommendation model and obtains the recommended media assets output by the media asset recommendation model that match the first attribute information, it is also configured to: obtain user information sent by the client; and determine the media asset recommendation model corresponding to the user information.

[0069] like Figure 4 As shown, Figure 4 This is a schematic diagram of the software configuration in a display device 200 according to one or more embodiments of the present disclosure, such as... Figure 4 As shown, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android Runtime and System Library layer (referred to as the "System Runtime Layer"), and the Kernel layer. The Kernel layer contains at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, Wi-Fi driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.

[0070] like Figure 5 As shown, Figure 5 This is a schematic diagram showing the icon control interface of an application in a display device 200 provided in this embodiment of the disclosure. The application layer includes at least one application that can display corresponding icon controls on the display, such as: live TV application icon control, video-on-demand application icon control, media center application icon control, application center icon control, game application icon control, etc.

[0071] In some embodiments, a live TV application can provide live TV from different signal sources. For example, the live TV application can provide a TV signal using input from cable television, terrestrial broadcasting, satellite services, or other types of live TV services. Furthermore, the live TV application can display the video of the live TV signal on display device 200.

[0072] In some embodiments, a video-on-demand application may provide video from different storage sources. Unlike live TV applications, video-on-demand provides video display from certain storage sources. For example, video-on-demand may come from a cloud storage server or from local hard drive storage containing existing video programs.

[0073] In some embodiments, a media center application may be an application that provides playback of various multimedia content. For example, a media center may provide services that, unlike live TV or video-on-demand, allow users to access various images or audio through the media center application.

[0074] In some embodiments, the application center may provide a storage for various applications. An application may be a game, an application, or other applications related to a computer system or other device but capable of running on a smart TV. The application center may obtain these applications from various sources, store them in local storage, and then make them runnable on the display device 200.

[0075] To enhance the user experience of smart TVs, they typically include a recommendation feature. For example, when displaying the homepage, the monitor in display device 200 recommends media assets through multiple recommendation slots set in the various channels of the channel bar.

[0076] To illustrate this solution in more detail, the following will use examples to illustrate it. Figure 6 To explain, it is understandable that Figure 6 The steps involved may include more or fewer steps in actual implementation, and the order of these steps may also be different, depending on whether the media asset recommendation method provided in the embodiments of this disclosure can be implemented.

[0077] like Figure 6 As shown, Figure 6 This is a flowchart illustrating a media asset recommendation method provided in an embodiment of the present disclosure. The method includes the following steps S601~S606:

[0078] S601, Receive the first media asset identifier sent by the client.

[0079] The first media asset identifier can be the media asset name, such as "Movie 1", or it can be the first media asset's identity document (ID). This disclosure does not limit this.

[0080] S602. Obtain the first attribute information of the first media asset indicated by the first media asset identifier.

[0081] The first attribute information describes the basic attributes of the first media asset. The first attribute information may include the first media asset tag, the first media asset name, the first media asset description, the first media asset actors, and may also include other information that can describe the attributes of the first media asset; this disclosure does not impose any limitations on this.

[0082] The first media asset tag is attribute information that is directly related to the content of the first media asset. For example, the first media asset tag for "Movie 1" is "Action / War". It can be understood that one media asset can correspond to multiple media asset tags.

[0083] The first media asset name can contain characters such as series, spin-off, etc.

[0084] The first media asset description can be a brief introduction or summary of the first media asset. For example, the first media asset description for "Movie 1" is: "It tells the story of the protagonist who is experiencing the lowest point in his life and originally wanted to drift at sea...".

[0085] The first media asset actor is the name of the person participating in the first media asset, or it can be the name of the character in the first media asset. For example, the first media asset actor of the first media asset "Movie 1" includes "Actor 1, Actor 2".

[0086] In some embodiments, the server retrieves the first attribute information of the first media asset indicated by the first media asset identifier from the database based on the first media asset identifier.

[0087] S603. Determine the first text sequence of the first media asset based on the first attribute information.

[0088] The first text sequence is a sequence of texts with a fixed format, which is also the input to the media asset recommendation model. The fixed format refers to the target template set obtained by learning from a large number of media assets. This disclosure obtains the first text sequence of the first media asset based on the target template set and the first attribute information, and uses it as the input to the media asset recommendation model to reflect the media asset content of the first media asset, thereby achieving accurate recommendation of the media asset content.

[0089] In some embodiments, during the process of determining the first text sequence of the first media asset based on the first attribute information, the server first obtains a target template set, and then, based on the target template set and according to the first media asset name, the first media asset actor, and the first media asset description, determines the first text sequence of the first media asset. The target template set is constructed based on a large amount of attribute information from second media assets, and includes one or more templates.

[0090] For example, any template A in the target template set is: "a1 starring in 'a2': a3...". The first media asset "Movie 1" uses template A to generate a first text in the first text sequence based on the first media asset name a2: "Movie 1", the first media asset actors a1: "Actor 1, Actor 2" and the first media asset description a3: "It tells the story of the protagonist who is experiencing the lowest point in his life and originally wanted to drift at sea...".

[0091] In some embodiments, the process of constructing the target template set is as follows: First, second attribute information of multiple second media assets is obtained from the database. The second attribute information includes second media asset tags, second media asset names, second media asset descriptions, and second media asset actors. Second, the second media asset description of each media asset in the multiple second media assets is segmented to obtain a segmented sequence. The specific segmentation method can be referred to the prior art, and will not be elaborated here. The second media asset description of each second media asset corresponds to a segment in the segmented sequence. Then, the second media asset name and second media asset actor included in each segment in the segmented sequence are removed to obtain a candidate template set. Each second media asset corresponds to a candidate template in the candidate template set. Further, the candidate templates included in the candidate template set are vectorized and clustered to determine the target template set.

[0092] For example, firstly, the second attribute information of multiple second media assets is retrieved from the database. Taking one of the second media assets j as an example, the second media asset description of second media asset j is retrieved. Segment by clause to obtain a sentence sequence =[ ], For each clause in the clause sequence, remove the included second media asset name. Second Media Actor Obtain candidate templates corresponding to the clauses. Iterate through multiple secondary media assets to obtain a set of candidate templates. Further use of word frequency vectorization to transform the candidate template set The templates are vectorized and then clustered using the k-means clustering algorithm (kmeans) to remove duplicate candidate templates, resulting in a target template set where each template is independent. The clustered candidate template set can also be manually fine-tuned to obtain the final target template set. The target template set is stored in the server's database.

[0093] The above embodiments construct a target template set and then use this target template set to determine the first text sequence corresponding to the first media asset. Using the target template set unifies the input format of the media asset recommendation model, facilitating processing and enabling the media asset recommendation model to perform recommendations efficiently.

[0094] S604. Input the first text sequence into the pre-trained media asset recommendation model, and obtain the recommended media assets output by the media asset recommendation model that match the first attribute information.

[0095] The media asset recommendation model is trained based on media asset attribute information and associated preferred media assets. The media asset attribute information reflects the content of the media asset, and the associated preferred media assets are determined based on the user's historical viewing records.

[0096] In some embodiments, the number of pre-trained media asset recommendation models is N, and each media asset recommendation model corresponds to user information. This means that different users have different media asset recommendation models. Before inputting the first text sequence into the pre-trained media asset recommendation model and obtaining the recommended media assets output by the model that match the first attribute information, the user information sent by the client is obtained to determine the media asset recommendation model corresponding to the user information. This allows for targeted media asset recommendations, resulting in higher accuracy and better recommendation performance.

[0097] The following will introduce the training process of the media asset recommendation model:

[0098] like Figure 7 As shown, Figure 7 This is a flowchart illustrating the training process of the media asset recommendation model provided in this embodiment of the disclosure. The training process of the media asset recommendation model includes the following steps S701 to S707:

[0099] S701. Obtain the first dataset.

[0100] The first dataset includes multiple second-text sequences and tags for second-media assets. The second-text sequences reflect the content of the second-media assets, while the tags reflect their types. This first dataset represents the correspondence between the content and type of second-media assets and is used to train a multi-tag recognition task for media asset recommendation models.

[0101] In some embodiments, a target template set and second attribute information of multiple second media assets are obtained, and then the second text sequence of multiple second media assets is determined using the target template set for the second media asset name, second media asset actor, and second media asset description of each of the multiple second media assets.

[0102] In determining the second text sequence, this disclosure provides an implementation method that uses each template in the target template set to determine the second text sequence of multiple second media assets. For example, if the target template set includes i templates, then the second text sequence of the second media asset j includes i second texts.

[0103] This disclosure also provides an implementation method in which a template is randomly selected from a target template set with a preset probability using a random sampling method, and the template is used to determine the second text of multiple second media assets to obtain a second text sequence.

[0104] S702, Obtain the second dataset.

[0105] The second dataset includes positive and negative samples. Positive samples contain preference media assets with first-order associations, while negative samples contain preference media assets with multi-order associations. It should be noted that multi-order media assets refer to N-order media assets other than first-order media assets, such as second-order and third-order media assets. The second dataset is used to reflect the associations between preference media assets viewed historically by users.

[0106] In some embodiments, during the acquisition of the second dataset, the second media asset identifiers of multiple historically viewed preferred media assets are first obtained. These identifiers are typically stored in a database in tabular form. Then, the correlation coefficients between the multiple preferred media assets are determined. A knowledge graph is then constructed based on these correlation coefficients, with preferred media assets as nodes and the correlation coefficients between them as edges. Further, based on the knowledge graph, preferred media assets with first-order correlations are classified as positive samples, and preferred media assets with multi-order correlations (excluding first-order correlations) are classified as negative samples. The positive and negative samples are then used as the second dataset. This can be understood as the positive samples containing directly correlated preferred media assets, while the negative samples contain preferred media assets that are correlated but not directly related.

[0107] In some embodiments, hyperparameters are set during the process of determining the correlation coefficients between multiple preferred media assets. This represents the minimum number of times two preferred media assets have co-occurred, which is the minimum number of times they have appeared together across different users' viewing histories. Preferred media asset p and preferred media asset q appear simultaneously... In the viewing history of individual users, As the co-occurrence frequency of preferred media asset p and preferred media asset q, if Greater than hyperparameters If we determine that the media preference p and media preference q are related, it can be understood that the relationship between media preference p and media preference q is bidirectional.

[0108] The correlation coefficient between preferred media asset p and preferred media asset q is calculated according to the following formula (1):

[0109] (1)

[0110] The correlation coefficient between preferred media asset q and preferred media asset p is calculated according to the following formula (2):

[0111] (2)

[0112] In some embodiments, after determining the correlation coefficients between multiple preference media assets, a knowledge graph is constructed, and then, for preference media asset p, a list of preference media assets that are first-order correlated with preference media asset p is determined. ,by For probability from media asset list Positive samples were obtained through sampling; furthermore, a list of preference media assets that are multi-level associated with preference media asset p was determined. From the media asset list Negative samples are obtained through sampling.

[0113] For example, the second media asset identifier for multiple preferred media assets viewed historically includes: Movie 1, Movie 2, Movie 3, Movie 4, and Movie 5. The correlation coefficient between Movie 1 and Movie 4 is determined to be 0.8, the correlation coefficient between Movie 1 and Movie 2 is 0.2, the correlation coefficient between Movie 2 and Movie 4 is 0.3, the correlation coefficient between Movie 2 and Movie 1 is 0.3, the correlation coefficient between Movie 2 and Movie 3 is 0.4, and the correlation coefficient between Movie 4 and Movie 5 is 0.6. The constructed knowledge graph is as follows: Figure 8 As shown, the knowledge graph uses preferred media assets as nodes and the association coefficients between preferred media assets as edges. Further, positive and negative samples are determined based on the knowledge graph. For the preferred media asset "Movie 1," positive samples are obtained as (Movie 1, Wolf Warrior) or (Movie 1, Movie 4); negative samples are as (Movie 1, Movie 3) or (Movie 1, Movie 5). Typically, there may be more than one user on the client side. With user authorization, the client can obtain and send user information to the server. The server obtains the user information sent by the client and then obtains the second media asset identifiers corresponding to the user's historical viewing preferences. This allows the second dataset to accurately reflect the user's interests and preferences. A media asset recommendation model is trained based on this second dataset, making the model more targeted and accurate in recommending media assets to users.

[0114] After obtaining the first and second datasets, an initial media asset recommendation model is determined, and various parameters of the initial media asset recommendation model are set. For details on the selection of the initial media asset recommendation model and the setting of model parameters, please refer to the existing technology, which will not be elaborated here.

[0115] In some embodiments, micro-F1 is used to determine the score for each tag. and the total score of all tags ,according to Set the weight parameters for the initial media asset recommendation model.

[0116] S703. Determine the first loss function corresponding to the first dataset and the second loss function corresponding to the second dataset.

[0117] In some embodiments, the first loss function corresponds to the first dataset. It can be the focal loss function.

[0118] The second loss function corresponding to the second dataset The following formulas (3), (4), and (5) are used to calculate the results:

[0119] (3)

[0120] (4)

[0121] (5)

[0122] in, These are preference media resources with first-order association in positive samples. It is the preference information of second-order association in negative samples. This refers to the preference media resources for third-order and higher-order associations in negative samples. It is the loss function corresponding to the second-order correlation of preference media resources. It is the loss function corresponding to preference mediators with third-order and higher associations. Batchsize represents the number of preference mediators in the negative samples.

[0123] S704. Input the first dataset into the initial media asset recommendation model and obtain the predicted media asset labels output by the initial media asset recommendation model; and input the second dataset into the initial media asset recommendation model and obtain the predicted associated media assets output by the initial media asset recommendation model.

[0124] S705, calculate a first loss value based on the second media asset label, the predicted media asset label, and the first loss function; and calculate a second loss value based on positive samples, negative samples, the predicted associated media assets, and the second loss function.

[0125] S706. Determine the target loss value based on the first loss value and the second loss value.

[0126] The target loss value is the sum of the first loss value and the second loss value.

[0127] S707. Adjust the model parameters in the initial media asset recommendation model according to the target loss value until a converged media asset recommendation model is obtained.

[0128] The media asset recommendation model takes as input a text sequence corresponding to the attribute information of the media asset and outputs recommended media assets that match the attribute information of the media asset.

[0129] In some embodiments, a loss threshold is preset. If the calculated target loss value is greater than the loss threshold, the model parameters of the current media asset recommendation model are adjusted. If the target loss value is less than or equal to the loss threshold, it indicates that the current media asset recommendation model has converged, and the training of the media asset recommendation model is complete.

[0130] The media asset recommendation model is trained through the above steps S701~S707. The media asset recommendation model takes media asset attribute information as input and outputs recommended media assets that match the media asset attributes. The recommended media assets are predicted based on the user's historical viewing behavior and are related to the first media asset that the user expects to obtain in terms of content. The media asset recommendation model can simultaneously meet the requirements of accurate recommendation based on user behavior and media asset content.

[0131] S605. Obtain the first media asset indicated by the first media asset identifier.

[0132] S606. Send the recommended media assets and the first media asset to the client.

[0133] Sending recommended media assets and primary media assets to the client satisfies users' need to access primary media assets and accurately recommends recommended media assets that are highly relevant to primary media assets in terms of content and user behavior, thereby achieving the goal of precise recommendation and improving the effectiveness of media asset recommendation.

[0134] In summary, this disclosure provides a media asset recommendation method. This method utilizes a pre-trained media asset recommendation model, which is trained based on media asset attribute information and associated preferred media assets. The media asset attribute information reflects the media asset content, and the associated preferred media assets are determined based on the user's historical viewing records. The method involves receiving a first media asset identifier from the client, then obtaining the first attribute information of the first media asset corresponding to the first media asset identifier. This attribute information includes a first media asset tag, a first media asset name, a first media asset description, and a first media asset actor. Next, a first text sequence of the first media asset is determined based on the first attribute information. This first text sequence is then input into the pre-trained media asset recommendation model to obtain recommended media assets that match the first attribute information, as output by the model. This process retrieves the first media asset corresponding to the first media asset identifier and sends both the first media asset and the recommended media assets to the client. This achieves targeted acquisition and recommendation of media assets related to the content of the first media asset corresponding to the first media asset identifier, combining user behavior and media asset content for accurate recommendations and improving recommendation effectiveness.

[0135] like Figure 9 As shown, Figure 9The present invention provides a schematic diagram of the structure of a server, which includes a processor 901, a memory 902, and a computer program stored on the memory 902 and executable on the processor 901. The computer program can be implemented by the processor 901 in various processes of the above-described media asset recommendation method and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0136] This disclosure provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the media asset recommendation method described above and achieves the same technical effect. To avoid repetition, further details are omitted here.

[0137] The computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0138] This disclosure provides a computer program product that includes a computer program that, when run on a computer, causes the computer to implement the aforementioned media asset recommendation method.

[0139] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments above is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of the embodiments suitable for specific application considerations.

Claims

1. A server, characterized in that, include: The controller is configured to receive the first media asset identifier sent by the client; Obtain the first attribute information of the first media asset indicated by the first media asset identifier, wherein the first attribute information includes the first media asset tag, the first media asset name, the first media asset description, and the first media asset actor; Determine the first text sequence of the first media asset based on the first attribute information; Input the first text sequence into a pre-trained media asset recommendation model, and obtain recommended media assets that match the first attribute information output by the media asset recommendation model; Obtain the first media asset indicated by the first media asset identifier; Send the recommended media asset and the first media asset to the client; The training process of the media asset recommendation model includes: acquiring a first dataset; the first dataset includes second text sequences of multiple second media assets and second media asset tags of multiple second media assets; the first dataset represents the correspondence between the content and type of the second media assets, and is used to train the multi-tag recognition task of the media asset recommendation model; The process involves: obtaining second media asset identifiers for multiple historically viewed preferred media assets; determining the correlation coefficients between these preferred media assets based on their co-occurrence frequency; constructing a knowledge graph based on these correlation coefficients, where preferred media assets are nodes and the correlation coefficients between them are edges; determining positive and negative samples based on the knowledge graph, where positive samples include preferred media assets with first-order correlations and negative samples include preferred media assets with multi-order correlations; obtaining a second dataset based on the positive and negative samples; and using this second dataset to reflect the correlation relationships between preferred media assets viewed historically by the user. Determine the first loss function corresponding to the first dataset and the second loss function corresponding to the second dataset; Input the first dataset into the initial media asset recommendation model to obtain the predicted media asset labels output by the initial media asset recommendation model; and input the second dataset into the initial media asset recommendation model to obtain the predicted associated media assets output by the initial media asset recommendation model. A first loss value is calculated based on the second media asset label, the predicted media asset label, and the first loss function; and a second loss value is calculated based on the positive sample, the negative sample, the predicted associated media asset, and the second loss function. Determine the target loss value based on the first loss value and the second loss value; The model parameters in the initial media asset recommendation model are adjusted according to the target loss value until a converged media asset recommendation model is obtained. The input of the media asset recommendation model is the text sequence corresponding to the attribute information of the media asset, and the output is the recommended media asset that matches the attribute information of the media asset.

2. The server according to claim 1, characterized in that, The controller, which determines the first text sequence of the first media asset based on the first attribute information, is configured as follows: Obtain the target template set; Based on the first media asset's first media asset name, first media asset actor, first media asset description, and the target template set, determine the first text sequence of the first media asset.

3. The server according to claim 2, characterized in that, The controller, having acquired the target template set, is configured as follows: Retrieve secondary attribute information from multiple secondary media assets from the database; the secondary attribute information of the secondary media assets includes secondary media asset tags, secondary media asset names, secondary media asset descriptions, and secondary media asset actors; The second media asset description of each of the plurality of second media assets is segmented into sentences to obtain a sentence sequence; wherein, the second media asset description of each second media asset corresponds to a sentence in the sentence sequence; Remove the second media asset name and second media asset actor from each sentence in the sentence sequence to obtain a candidate template set, wherein each second media asset corresponds to a candidate template in the candidate template set; The candidate templates included in the candidate template set are vectorized and clustered to determine the target template set.

4. The server according to claim 1, characterized in that, The controller, acquiring the first dataset, is configured as follows: Obtain the target template set and the second attribute information of multiple second media assets; For each of the plurality of secondary media assets, including its name, actor, and description, the target template set is used to determine the second text sequence of the plurality of secondary media assets. The first dataset is obtained based on the second text sequence of the plurality of second media assets and the second media asset labels of the plurality of second media assets.

5. The server according to claim 1, characterized in that, Before acquiring the second media asset identifier of multiple preferred media assets viewed in the past, the controller is also configured to: Obtain the user information sent by the client; The controller, which acquires the second media asset identifier of multiple preferred media assets viewed in the past, is configured as follows: Obtain the second media asset identifier of multiple preferred media assets in the user's historical viewing history corresponding to the user information.

6. The server according to claim 1, characterized in that, There are N media asset recommendation models, and each media asset recommendation model has a corresponding relationship with user information; The controller, before inputting the first text sequence into the pre-trained media asset recommendation model and obtaining the recommended media assets output by the media asset recommendation model that match the first attribute information, is further configured as follows: Obtain the user information sent by the client; Determine the media asset recommendation model corresponding to the user information.

7. A media asset recommendation method, characterized in that, include: Receive the first media asset identifier sent by the client; Obtain the first attribute information of the first media asset indicated by the first media asset identifier, wherein the first attribute information includes the first media asset tag, the first media asset name, the first media asset description, and the first media asset actor; Determine the first text sequence of the first media asset based on the first attribute information; Input the first text sequence into a pre-trained media asset recommendation model, and obtain recommended media assets that match the first attribute information output by the media asset recommendation model; Obtain the first media asset indicated by the first media asset identifier; Send the recommended media asset and the first media asset to the client; The training process of the media asset recommendation model includes: acquiring a first dataset; the first dataset includes second text sequences of multiple second media assets and second media asset tags of multiple second media assets; the first dataset represents the correspondence between the content and type of the second media assets, and is used to train the multi-tag recognition task of the media asset recommendation model; The process involves: obtaining second media asset identifiers for multiple historically viewed preferred media assets; determining the correlation coefficients between these preferred media assets based on their co-occurrence frequency; constructing a knowledge graph based on these correlation coefficients, where preferred media assets are nodes and the correlation coefficients between them are edges; determining positive and negative samples based on the knowledge graph, where positive samples include preferred media assets with first-order correlations and negative samples include preferred media assets with multi-order correlations; obtaining a second dataset based on the positive and negative samples; and using this second dataset to reflect the correlation relationships between preferred media assets viewed historically by the user. Determine the first loss function corresponding to the first dataset and the second loss function corresponding to the second dataset; Input the first dataset into the initial media asset recommendation model to obtain the predicted media asset labels output by the initial media asset recommendation model; and input the second dataset into the initial media asset recommendation model to obtain the predicted associated media assets output by the initial media asset recommendation model. A first loss value is calculated based on the second media asset label, the predicted media asset label, and the first loss function; and a second loss value is calculated based on the positive sample, the negative sample, the predicted associated media asset, and the second loss function. Determine the target loss value based on the first loss value and the second loss value; The model parameters in the initial media asset recommendation model are adjusted according to the target loss value until a converged media asset recommendation model is obtained. The input of the media asset recommendation model is the text sequence corresponding to the attribute information of the media asset, and the output is the recommended media asset that matches the attribute information of the media asset.

8. The method according to claim 7, characterized in that, Determining the first text sequence of the first media asset based on the first attribute information includes: Obtain the target template set; Based on the first media asset's first media asset name, first media asset actor, first media asset description, and the target template set, determine the first text sequence of the first media asset.