Multimedia Recommendation Method, Device, Equipment and Storage Medium

The multimedia recommendation method enhances accuracy by fusing content and position features from historical playback data to derive user interest features, addressing the mismatch between popular and user-interested content in existing systems.

CN112148899BActive Publication Date: 2025-07-15SHENZHEN YAYUE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011103343.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-15
Publication Date
2025-07-15
Estimated Expiration
2040-10-15

AI Technical Summary

Technical Problem

Among the existing multimedia recommendation technologies, popular multimedia data does not match user interests, resulting in low recommendation accuracy and poor recommendation effect.

Method used

By obtaining the historical playback sequence of the target user, integrating the content and position characteristics of the multimedia data, generating interest characteristics, and using the multimedia recommendation model for recommendation.

Benefits of technology

It improves the accuracy of multimedia recommendations and makes the recommended multimedia data more in line with user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112148899B_ABST
    Figure CN112148899B_ABST
Patent Text Reader

Abstract

The present application provides a multimedia recommendation method, device, equipment and storage medium, belonging to the field of Internet technology. The method includes: obtaining a historical playback sequence corresponding to a target user identifier, where the historical playback sequence includes a plurality of multimedia data played by a terminal logged in with the target user identifier, and the plurality of multimedia data are arranged in the playback order; fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, where the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence; fusing the plurality of multimedia features in the arrangement order of the plurality of multimedia data to obtain the interest feature of the target user identifier; and recommending multimedia data for the target user identifier according to the interest feature. The above method can improve the accuracy of recommended multimedia.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to a multimedia recommendation method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of Internet technologies, the quantity of multimedia data such as videos and audios has increased rapidly, resulting in a data overload problem. In the face of this problem, multimedia recommendation technologies have emerged, so as to be able to actively recommend multimedia data that a user may be interested in to the user. In related technologies, generally popular multimedia data is recommended to the user. However, the popular multimedia data may not conform to the user's interests, resulting in low accuracy of multimedia recommendation and poor recommendation effects. Summary of the Invention

[0003] Embodiments of this application provide a multimedia recommendation method, apparatus, device, and storage medium, which can improve the accuracy of recommended multimedia. The technical solution is as follows:

[0004] On the one hand, a multimedia recommendation method is provided. The method includes:

[0005] Obtain a historical playback sequence corresponding to a target user identifier. The historical playback sequence includes multiple multimedia data played by a terminal that logs in to the target user identifier, and the multiple multimedia data are arranged in the playback order;

[0006] Fuse the content feature and the position feature of each multimedia data to obtain a multimedia feature of each multimedia data. The content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence;

[0007] Fuse multiple multimedia features in the arrangement order of the multiple multimedia data to obtain an interest feature of the target user identifier;

[0008] Recommend multimedia data for the target user identifier according to the interest feature.

[0009] In a possible implementation manner, the fusing the content feature and the position feature of each multimedia data to obtain a multimedia feature of each multimedia data includes:

[0010] Call a feature fusion layer in a multimedia recommendation model to fuse the content feature and the position feature of each multimedia data to obtain a multimedia feature of each multimedia data;

[0011] The fusing the multiple multimedia features in the arrangement order of the multiple multimedia data to obtain an interest feature of the target user identifier includes:

[0012] Invoke the interest extraction layer in the multimedia recommendation model, and fuse the multiple multimedia features according to the arrangement order of the multiple multimedia data to obtain the interest features of the target user identifier.

[0013] In another possible implementation, the invoking the interest extraction layer in the multimedia recommendation model, and fusing the multiple multimedia features according to the arrangement order of the multiple multimedia data to obtain the interest features of the target user identifier includes:

[0014] Invoke the interest extraction layer, and determine the relevance between each multimedia feature and the target multimedia feature according to the arrangement order of the multiple multimedia data, and extract the hidden layer multimedia features from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the multiple multimedia features;

[0015] Fuse the hidden layer multimedia features of the multiple multimedia data to obtain the interest features of the target user identifier.

[0016] In another possible implementation, the training process of the multimedia recommendation model includes:

[0017] Obtain the content features of multiple sample multimedia data, the first sample content features of the positive sample data, and the second sample content features of the negative sample data in the first sample play sequence, where the positive sample data is the multimedia data that has been recommended and played after the first sample play sequence, and the negative sample is the multimedia data that has been recommended but not played;

[0018] Invoke the multimedia recommendation model, and obtain the first interest features according to the content features of the multiple sample multimedia data in the first sample play sequence;

[0019] Train the multimedia recommendation model according to the first interest features, the first sample content features, and the second sample content features.

[0020] In another possible implementation, the training the multimedia recommendation model according to the first interest features, the first sample content features, and the second sample content features includes:

[0021] Obtain the first similarity between the first interest features and the first sample content features and the second similarity between the first interest features and the second sample content features;

[0022] Train the multimedia recommendation model according to the first similarity and the second similarity.

[0023] In another possible implementation, before training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature, the training process of the multimedia recommendation model further includes:

[0024] Obtain a second sample playback sequence;

[0025] Perform a masking process on the first sample multimedia data in the second sample playback sequence, where the masking process refers to masking the content of the first sample multimedia data;

[0026] Invoke the multimedia recommendation model, and predict the second sample multimedia data of the masked content according to the masked second sample playback sequence;

[0027] Train the multimedia recommendation model according to the first sample multimedia data and the second sample multimedia data.

[0028] In another possible implementation, before fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, the method further includes:

[0029] Invoke the content extraction layer in the multimedia recommendation model to query the content features of the multiple multimedia data from the multimedia database;

[0030] Invoke the position extraction layer in the multimedia recommendation model to obtain the position features of the multiple multimedia data.

[0031] In another possible implementation, before fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, the method further includes:

[0032] Invoke a content recognition model to perform feature extraction on each multimedia data to obtain the content feature of each multimedia data.

[0033] In another possible implementation, invoking the content recognition model to perform feature extraction on each multimedia data to obtain the content feature of each multimedia data includes:

[0034] Invoke the content recognition model to perform multi-modal feature extraction on each multimedia data to obtain the multi-modal content feature of each multimedia data;

[0035] Wherein, the multi-modal content feature includes at least two of the content feature of the multimedia title, the content feature of the multimedia picture, or the content feature of the multimedia audio.

[0036] In another possible implementation, the training process of the content recognition model includes:

[0037] Obtain sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data;

[0038] Invoke the content recognition model to extract features from the sample multimedia data to obtain the content features of the sample multimedia data;

[0039] Invoke a content classification model to classify the content features to obtain a predicted label;

[0040] Train the content recognition model and the content classification model according to the sample label and the predicted label.

[0041] In another possible implementation, the invoking the content recognition model to extract features from the sample multimedia data to obtain the content features of the sample multimedia data includes:

[0042] Invoke the content recognition model to perform multi-modal feature extraction on the sample multimedia data to obtain the multi-modal content features of the sample multimedia data;

[0043] Wherein, the multi-modal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

[0044] In another possible implementation, the recommending multimedia data for the target user identifier according to the interest features includes:

[0045] Obtain the content features of at least one multimedia data to be recommended;

[0046] According to the interest features and the content features of the at least one multimedia data to be recommended, obtain target multimedia data that matches the interest features;

[0047] Send the target multimedia data to the terminal.

[0048] In another possible implementation, the obtaining target multimedia data that matches the interest features according to the interest features and the content features of the at least one multimedia data to be recommended includes:

[0049] Invoke the feature matching layer in the multimedia recommendation model, and according to the interest features and the content features of the at least one multimedia data to be recommended, obtain target multimedia data that matches the interest features.

[0050] On the other hand, a multimedia recommendation device is provided, and the device includes:

[0051] A sequence acquisition module, configured to acquire a historical playback sequence corresponding to a target user identifier, where the historical playback sequence includes a plurality of multimedia data played by a terminal logged in with the target user identifier, and the plurality of multimedia data are arranged in the playback order;

[0052] A first fusion module, configured to fuse the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, where the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence;

[0053] A second fusion module, configured to fuse a plurality of multimedia features in the arrangement order of the plurality of multimedia data to obtain the interest feature of the target user identifier;

[0054] A multimedia recommendation module, configured to recommend multimedia data for the target user identifier according to the interest feature.

[0055] In a possible implementation manner, the first fusion module is configured to call a feature fusion layer in a multimedia recommendation model to fuse the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data;

[0056] The second fusion module is configured to call an interest extraction layer in the multimedia recommendation model to fuse the plurality of multimedia features in the arrangement order of the plurality of multimedia data to obtain the interest feature of the target user identifier.

[0057] In another possible implementation manner, the second fusion module is configured to call the interest extraction layer to determine the relevance between each multimedia feature and a target multimedia feature in the arrangement order of the plurality of multimedia data, and extract a hidden layer multimedia feature from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the plurality of multimedia features; fuse the hidden layer multimedia features of the plurality of multimedia data to obtain the interest feature of the target user identifier.

[0058] In another possible implementation manner, the training process of the multimedia recommendation model includes:

[0059] Acquire the content features of a plurality of sample multimedia data in a first sample playback sequence, the first sample content feature of positive sample data, and the second sample content feature of negative sample data, where the positive sample data is multimedia data that has been recommended and played after the first sample playback sequence, and the negative sample is multimedia data that has been recommended but not played;

[0060] Call the multimedia recommendation model, and obtain a first interest feature according to the content features of multiple sample multimedia data in the first sample play sequence;

[0061] Train the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature.

[0062] In another possible implementation manner, the training of the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature includes:

[0063] Obtain a first similarity between the first interest feature and the first sample content feature, and a second similarity between the first interest feature and the second sample content feature;

[0064] Train the multimedia recommendation model according to the first similarity and the second similarity.

[0065] In another possible implementation manner, before training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature, the training process of the multimedia recommendation model further includes:

[0066] Obtain a second sample play sequence;

[0067] Perform a masking process on the first sample multimedia data in the second sample play sequence, where the masking process refers to masking the content of the first sample multimedia data;

[0068] Call the multimedia recommendation model, and predict the second sample multimedia data with the masked content according to the masked second sample play sequence;

[0069] Train the multimedia recommendation model according to the first sample multimedia data and the second sample multimedia data.

[0070] In another possible implementation manner, the device further includes:

[0071] A first feature acquisition module, configured to call a content extraction layer in the multimedia recommendation model to query the content features of the multiple multimedia data from a multimedia database;

[0072] A second feature acquisition module, configured to call a position extraction layer in the multimedia recommendation model to obtain the position features of the multiple multimedia data.

[0073] In another possible implementation manner, the device further includes:

[0074] A content feature recognition module, which is used to call a content recognition model to extract features from each of the multimedia data, so as to obtain the content features of each of the multimedia data.

[0075] In another possible implementation, the content feature recognition module is used to call the content recognition model to perform multimodal feature extraction on each of the multimedia data, so as to obtain the multimodal content features of each of the multimedia data; wherein, the multimodal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

[0076] In another possible implementation, the training process of the content recognition model includes:

[0077] Obtain sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data;

[0078] Call the content recognition model to extract features from the sample multimedia data, so as to obtain the content features of the sample multimedia data;

[0079] Call a content classification model to classify the content features to obtain a predicted label;

[0080] Train the content recognition model and the content classification model according to the sample label and the predicted label.

[0081] In another possible implementation, the calling the content recognition model to extract features from the sample multimedia data, so as to obtain the content features of the sample multimedia data, includes:

[0082] Call the content recognition model to perform multimodal feature extraction on the sample multimedia data, so as to obtain the multimodal content features of the sample multimedia data;

[0083] Wherein, the multimodal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

[0084] In another possible implementation, the multimedia recommendation module includes:

[0085] A third feature acquisition unit, which is used to acquire the content features of at least one multimedia data to be recommended;

[0086] A feature matching unit, which is used to obtain target multimedia data that matches the interest feature according to the interest feature and the content features of the at least one multimedia data to be recommended;

[0087] A multimedia sending unit, configured to send the target multimedia data to the terminal.

[0088] In another possible implementation, the feature matching unit is configured to call a feature matching layer in a multimedia recommendation model, and obtain target multimedia data that matches the interest feature according to the interest feature and the content features of at least one multimedia data to be recommended.

[0089] On the other hand, a computer device is provided, which includes a processor and a memory. At least one program code is stored in the memory, and the program code is loaded and executed by the processor to implement the operations performed in the multimedia recommendation method in any of the above possible implementations.

[0090] On the other hand, a computer-readable storage medium is provided. At least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by a processor to implement the operations performed in the multimedia recommendation method in any of the above possible implementations.

[0091] In still another aspect, a computer program product or a computer program is provided. The computer program product or the computer program includes computer program code, and the computer program code is stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the operations performed in the multimedia recommendation method in the above various alternative implementations.

[0092] The beneficial effects brought by the technical solution provided in the embodiments of this application at least include:

[0093] The technical solution provided in the embodiments of this application recommends multimedia data to a user by mining the user's interest features. Among them, when obtaining the interest features, the content features and position features of multiple multimedia data in the user's historical play sequence are fused to obtain multiple multimedia features, and the multiple multimedia features are fused to obtain the user's interest features. Since this interest feature considers the content of multiple multimedia data played by the user and the influence of the play order of the multiple media data on the multimedia data that the user wants to watch next, this interest feature is more matched with the user, so that the multimedia data recommended to the user according to this interest feature can better meet the user's interests and improve the recommendation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0095] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0096] Figure 2 is a flowchart of a multimedia recommendation method provided by an embodiment of the present application;

[0097] Figure 3 is a flowchart of a training method for a multimedia recommendation model provided by an embodiment of the present application;

[0098] Figure 4 is a schematic diagram of a multimedia recommendation model provided by an embodiment of the present application;

[0099] Figure 5 is a flowchart of a multimedia recommendation method provided by an embodiment of the present application;

[0100] Figure 6 is a schematic diagram of a multimedia recommendation interface provided by an embodiment of the present application;

[0101] Figure 7 is a block diagram of a multimedia recommendation device provided by an embodiment of the present application;

[0102] Figure 8 is a block diagram of a multimedia recommendation device provided by an embodiment of the present application;

[0103] Figure 9 is a schematic diagram of the structure of a terminal provided by an embodiment of the present application;

[0104] Figure 10 is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners

[0105] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0106] As used in this application, the terms "first", "second", "third", "fourth", etc. may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, a first sample content feature may be referred to as a sample content feature, and similarly, a second sample content feature may be referred to as a first sample content feature.

[0107] As used in this application, the terms "at least one", "multiple", "each", "any one", where at least one includes one, two or more than two, multiple includes two or more than two, and each refers to each one of the corresponding multiple, and any one refers to any one of the multiple. For example, multiple multimedia data include 3 multimedia data, and each refers to each one of these 3 multimedia data, and any one refers to any one of these 3 multimedia data, which can be the first, the second, or the third.

[0108] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of this application. Refer to Figure 1 , this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network. Optionally, the terminal 101 is a computer, a mobile phone, a tablet computer or other terminal. Optionally, the server 102 is the background server of the target application or a cloud server providing services such as cloud computing and cloud storage.

[0109] Optionally, a target application provided by the server 102 is installed on the terminal 101, and the terminal 101 can implement functions such as data transmission and message interaction through this target application. Optionally, the target application is a target application in the operating system of the terminal 101, or a target application provided by a third party. The target application has the function of recommending multimedia data. Optionally, the target application has the function of recommending videos, the function of recommending music, etc. Of course, the target application can also have the function of recommending other information, and this application does not limit this. Optionally, the target application is a short video application, a music application, a search application, a chat application or other applications, and this application does not limit this.

[0110] In an embodiment of this application, the server 102 is used to recommend multimedia data to the terminal 101, and the terminal 101 is used to play the recommended multimedia data.

[0111] The multimedia recommendation method provided by this application can be applied to scenarios where any multimedia data is recommended.

[0112] For example, in a scenario where a user plays short videos using a terminal, the user plays their favorite short videos through the terminal. When the user has played multiple short videos, a play sequence of the short videos will be obtained. Then, the method provided in the embodiments of the present application can be adopted to recommend short videos that the user is interested in according to the play sequence of the short videos.

[0113] For example, in a scenario where a user plays music using a terminal, the user plays their favorite music through the terminal. When the user has played multiple pieces of music, a play sequence of the music will be obtained. Then, the method provided in the embodiments of the present application can be adopted to recommend music that the user likes according to the play sequence of the music.

[0114] For example, in a scenario where a user plays a live broadcast using a terminal, the user plays their favorite live broadcast through the terminal. After the user has played the live broadcasts of multiple live broadcast rooms, a play sequence of the live broadcast rooms will be obtained. Then, the method provided in the embodiments of the present application can be adopted to recommend live broadcast rooms that the user is interested in according to the play sequence of the live broadcast rooms.

[0115] Figure 2 is a flowchart of a multimedia recommendation method provided in the embodiments of the present application. The execution entity is a computer device. Refer to Figure 2 This embodiment includes:

[0116] 201. Obtain a historical play sequence corresponding to a target user identifier. The historical play sequence includes multiple multimedia data played by a terminal logged in with the target user identifier, and the multiple multimedia data are arranged in the play order.

[0117] 202. Integrate the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data. The content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical play sequence.

[0118] 203. Integrate the multiple multimedia features in the arrangement order of the multiple multimedia data to obtain the interest feature of the target user identifier.

[0119] 204. Recommend multimedia data for the target user identifier according to the interest feature.

[0120] The technical solution provided by the embodiments of the present application recommends multimedia data to users by mining the interest characteristics of users and based on the interest characteristics of users. Among them, when obtaining the interest characteristics, the content characteristics and location characteristics of multiple multimedia data in the user's historical play sequence are fused to obtain multiple multimedia characteristics, and the multiple multimedia characteristics are fused to obtain the user's interest characteristics. Since this interest characteristic takes into account the content of multiple multimedia data played by the user and the influence of the play order of the multiple media data on the multimedia data that the user wants to watch next, therefore, this interest characteristic is more matched with the user, making the multimedia data recommended for the user according to this interest characteristic more in line with the user's interests and improving the recommendation accuracy.

[0121] The above multimedia recommendation method can be implemented through a multimedia recommendation model. Before using the multimedia recommendation model for multimedia recommendation, the multimedia recommendation model needs to be trained first. The following embodiments are the training methods of the multimedia recommendation model provided by the present application.

[0122] Figure 3 It is a flowchart of a training method of a multimedia recommendation model provided by an embodiment of the present application. In this embodiment, the execution entity is taken as an example of a server for description. See Figure 3 This embodiment includes:

[0123] 301. The server obtains a second sample play sequence.

[0124] It should be noted that this embodiment describes the training method of the multimedia recommendation model, and this multimedia recommendation model is used to recommend multimedia data to users based on the multimedia data played by the users. Optionally, the multimedia data is video data, audio data, etc., and the embodiments of the present application do not limit this.

[0125] The samples for training the multimedia recommendation model include the second sample play sequence of the sample user identifier. The second sample play sequence includes multiple multimedia data played by the terminal logging in to the sample user identifier, and the multiple multimedia data are arranged in the play order. Since the play order of the multimedia data by the user has an impact on the multimedia data that the user wants to watch next, therefore, the multiple multimedia data in the second sample play sequence are arranged in the play order, so that the multimedia recommendation model can learn the characteristics related to the play order of the multimedia data, and thus the obtained interest characteristics are more accurate.

[0126] Optionally, the second sample play sequence includes a reference number of sample multimedia data. For example, it includes 10 sample multimedia data. The sample user identifier is used to represent the identity of the sample user. Optionally, the sample user identifier is the mobile phone number, ID number, etc. of the sample user.

[0127] Optionally, the server obtains a second sample play sequence according to the behavior log of the sample user identifier. For example, the server obtains the identifiers of multiple multimedia data and the play time corresponding to each multimedia data from the behavior log of the sample user identifier, and generates a second sample play sequence according to the order of the identifiers of the multiple multimedia data and the play time corresponding to each multimedia data. Optionally, the play time is the start play time of the multimedia data, or the end play time of the multimedia data, or other time, and the embodiments of the present application do not limit this.

[0128] Optionally, when the terminal logged in with the sample user identifier is running, it will generate a behavior log of the user according to the detected user operations. The behavior log includes the multimedia play data of the user. For example, the identifier of the multimedia data played by the user, the time of playing the multimedia data, etc. Correspondingly, the server can obtain a second sample play sequence according to the behavior log of the user.

[0129] 302. The server performs a masking process on the first sample multimedia data in the second sample play sequence. The masking process means masking the content of the first sample multimedia data.

[0130] Optionally, the server performs a random masking (mask) process on the first sample multimedia data in the second sample play sequence, where the first sample multimedia data is any multimedia data in the second sample play sequence.

[0131] It should be noted that after the masking process is performed on the first sample multimedia data in the second sample play sequence, the number and arrangement order of the multiple multimedia data in the second sample play sequence will not change. Only the content of the first sample multimedia data is masked, and the arrangement position of the first sample multimedia data remains.

[0132] For example, if the second sample play sequence is "Multimedia Data A, Multimedia Data B, Multimedia Data C", then after the masking process is performed on Multimedia Data B, the obtained second sample play sequence is "Multimedia Data A, Multimedia Data, Multimedia Data C".

[0133] 303. The server invokes the multimedia recommendation model and predicts the second sample multimedia data with the masked content according to the masked second sample play sequence.

[0134] After the first sample multimedia data in the second sample playback sequence is masked, the server obtains the content features of other multimedia data in the second sample playback sequence and the position features of each multimedia data in the second sample playback sequence, and predicts the second sample multimedia data with the masked content according to the content features of the other multimedia data and the position features of each multimedia in the second sample playback sequence. Among them, the content feature of the multimedia data represents the content of the multimedia data, and the position feature of the multimedia data represents the position of the multimedia data in the second sample playback sequence.

[0135] 304. The server trains a multimedia recommendation model based on the first sample multimedia data and the second sample multimedia data.

[0136] Optionally, if the first sample multimedia data is different from the second sample multimedia data, the parameters of the multimedia recommendation model are adjusted to make the first sample multimedia data the same as the second sample multimedia data. Optionally, if the first sample multimedia data is the same as the second sample multimedia data, the multimedia recommendation model is continuously trained through a new second sample playback sequence, that is, steps 301-304 are repeated until the number of trained second sample playback sequences reaches a reference number, and a trained multimedia recommendation model is obtained.

[0137] Optionally, step 303 includes: calling the multimedia recommendation model to predict the sample multimedia data with the masked content in a reference number of second sample playback sequences, obtaining a reference number of predicted multimedia data. Correspondingly, step 304 includes: determining the prediction accuracy rate of the multimedia recommendation model according to the reference number of sample multimedia data with the masked content and the reference number of predicted multimedia data. If the prediction accuracy rate does not reach the reference threshold, the multimedia recommendation model is continuously trained until the prediction accuracy rate of the multimedia recommendation model reaches the reference threshold. Among them, the reference number is greater than 1.

[0138] In the embodiment of the present application, steps 301-304 are the pre-training process of the multimedia recommendation model. During the pre-training process, the multimedia data in the sample playback sequence is masked, and the multimedia recommendation model is called to predict the multimedia data with the masked content in the sample playback sequence, so that the model learns the dependency relationship between multiple multimedia data in the playback sequence, which is more conducive to the multimedia recommendation model to learn the context features of the multimedia data. The following steps 305-307 are the fine-tuning process of the multimedia recommendation model.

[0139] It should be noted that steps 301-304 are optional steps. It is also possible not to execute steps 301-304, but directly train the multimedia recommendation model through the following steps 305-307.

[0140] 305. The server obtains the content features of multiple sample multimedia data in the first sample playback sequence, the first sample content features of the positive sample data, and the second sample content features of the negative sample data.

[0141] The first sample playback sequence includes multiple multimedia data played by a terminal logged in with the first sample user identifier, and the multiple multimedia data are arranged in the playback order. The positive sample data are the multimedia data that have been recommended to the first sample user and that the first sample user has played after the first sample playback sequence. The negative samples are the multimedia data that have been recommended to the first sample user but that the first sample user has not played.

[0142] For example, if multimedia data A and multimedia data B are recommended to the first sample user, and the first sample user plays multimedia data A after the first sample playback sequence and does not play multimedia data B, it indicates that the first sample user is interested in multimedia data A and not interested in multimedia data B. Then, multimedia data A can be used as positive sample data, and multimedia data B can be used as negative sample data.

[0143] Optionally, the server obtains the playback sequence of multimedia data from the user's behavior log. Exemplarily, the server obtains the historical playback sequence within a reference time range from the current time. For example, it obtains the playback sequence within the last 1 hour, takes the last multimedia data in the playback sequence as the positive sample data, and forms the first sample playback sequence with the other multiple multimedia data in the playback sequence in the playback order. Optionally, after the server obtains the playback sequence of multimedia data from the user's behavior log, it first filters out the dirty data therein, and then takes the last multimedia data in the playback sequence as the positive sample data, and forms the first sample playback sequence with the other multiple multimedia data in the playback sequence in the playback order. Among them, the dirty data include multimedia data with a playback duration less than the reference duration, multimedia data with a ratio of the playback duration to the total duration less than the reference ratio, etc. Optionally, the reference duration is any duration. For example, the reference duration is 10 seconds. Optionally, the reference ratio is any ratio. For example, the reference ratio is 0.3.

[0144] Optionally, the server randomly selects one from the multimedia data that have been recommended to the first sample user but that the first sample user has not played as the negative sample data. Or, the server randomly selects one from the dirty data corresponding to the first sample user as the negative sample data.

[0145] Reference Figure 4 , Figure 4Schematic diagram of a multimedia recommendation model. Among them, the multimedia recommendation model includes a content extraction layer (Embedding layer) 401. Optionally, the implementation method for the server to obtain the content features of multiple sample multimedia data, the first sample content feature of the positive sample data, and the second sample content feature of the negative sample data in the first sample playback sequence is as follows: The server calls the content extraction layer 401 in the multimedia recommendation model to query the content features of multiple multimedia data, the first sample content feature of the positive sample data, and the second sample content feature of the negative sample data in the first sample playback sequence from the multimedia database. Continue to refer to Figure 4 , where f1, f2, ···, fn represent the content features of each multimedia data in the sample playback sequence, and fn+1 represents the content feature of the positive sample data or the negative sample data.

[0146] Among them, the multimedia database is used to store the content features of multimedia data. Optionally, the content features of the multimedia data in the multimedia database are obtained through a content recognition model and then stored in the multimedia database. In this way, when the server needs to obtain the content features of multimedia data, it can directly query from the multimedia database, with high efficiency.

[0147] 306. The server calls the multimedia recommendation model and obtains a first interest feature according to the content features of multiple sample multimedia data in the first sample playback sequence.

[0148] The first interest feature represents the features of the multimedia data that the first sample user is interested in. Optionally, the implementation method for obtaining the first interest feature includes the following steps (1)-(3):

[0149] (1) Continue to refer to Figure 4 , the multimedia recommendation model further includes a position extraction layer 402. The server calls the position extraction layer 402 to obtain the position features of multiple sample multimedia data.

[0150] Among them, the position feature represents the position of the sample multimedia data in the first sample playback sequence. The implementation method for the server to call the position extraction layer 402 to obtain the position features of multiple sample multimedia data is as follows: The server inputs the first sample playback sequence into the multimedia recommendation model, calls the position extraction layer 402 in the multimedia recommendation model, and obtains the position features of the multiple sample multimedia data according to the arrangement order of the multiple sample multimedia data in the first sample playback sequence. Figure 4 In, p1, p2, ···, pn represent the position features of each multimedia data in the sample playback sequence.

[0151] For example, if there are three sample multimedia data in the first sample playback sequence, the position feature of the first sample multimedia data is {1, 0, 0}, the position feature of the second sample multimedia data is {0, 1, 0}, and the position feature of the third sample multimedia data is {0, 0, 1}.

[0152] (2) Continuing to refer to Figure 4 , the multimedia recommendation model further includes a feature fusion layer 403. The server calls the feature fusion layer 403 to fuse the content feature and the position feature of each sample multimedia data to obtain the multimedia feature of each sample multimedia data.

[0153] Optionally, the server calls the feature fusion layer 403 in the multimedia recommendation model to sum the content feature and the position feature of each sample multimedia data to obtain the multimedia feature of each sample multimedia data. Alternatively, the content feature and the position feature of each sample multimedia data are fused in other ways to obtain the multimedia feature of each sample multimedia data, which is not limited in the embodiments of the present application.

[0154] (3) Continuing to refer to Figure 4 , the multimedia recommendation model further includes an interest extraction layer 404. The server calls the interest extraction layer 404 to fuse multiple multimedia features in the arrangement order of multiple sample multimedia data to obtain the first interest feature of the first user identifier.

[0155] Optionally, the implementation manner of this step is: the server calls the interest extraction layer 404 to determine the relevance between each sample multimedia feature and the target multimedia feature in the arrangement order of multiple sample multimedia data, and extracts the hidden layer multimedia feature from the target multimedia feature according to the determined relevance. The target multimedia feature is any one of multiple sample multimedia features; the server fuses the hidden layer multimedia features of multiple sample multimedia data to obtain the first interest feature of the first user identifier.

[0156] Among them, the hidden layer multimedia feature is the representation of the multimedia feature in the interest extraction layer. In the interest extraction layer, for each target multimedia feature, the relevance between each sample multimedia feature before the target multimedia feature and the target multimedia feature, and the relevance between each sample multimedia feature after the target multimedia feature and the target multimedia feature will be determined. The greater the relevance between the sample multimedia feature and the target multimedia feature, the greater the attention weight of the sample multimedia feature when extracting the hidden layer multimedia feature from the target multimedia feature, so that the extracted hidden layer multimedia feature can include the context feature of the target multimedia feature. In this way, after fusing the hidden layer multimedia features of multiple sample multimedia data, the obtained interest feature is more accurate.

[0157] Optionally, the interest extraction layer 404 includes multiple layers of multi-head self-attention sub-layers, and the number of layers of the multi-head self-attention sub-layers is set as required.

[0158] Optionally, the server invokes each multi-head self-attention sub-layer, determines the relevance between each sample multimedia feature in each multi-head self-attention sub-layer and the target multimedia feature according to the arrangement order of multiple sample multimedia data, and extracts the hidden-layer multimedia feature from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the multiple sample multimedia features; the server fuses the hidden-layer multimedia features of the multiple sample multimedia data to obtain the first interest feature of the first user identifier. The interest feature fused by multiple multi-head self-attention sub-layers is more accurate.

[0159] 307. The server trains the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature.

[0160] The implementation manner of this step is: the server obtains the first similarity between the first interest feature and the first sample content feature and the second similarity between the first interest feature and the second sample content feature; and trains the multimedia recommendation model according to the first similarity and the second similarity.

[0161] Optionally, continue to refer to Figure 4 , the multimedia recommendation model further includes a feature matching layer 405, and the server invokes the feature matching layer 405 to determine the first similarity between the first interest feature and the first sample content feature and the second similarity between the first interest feature and the second sample content feature.

[0162] Continue to refer to Figure 4 , the multimedia recommendation model further includes a perceptron layer 406. Before the server determines the first similarity between the first interest feature and the first sample content feature and the second similarity between the first interest feature and the second sample content feature, it needs to first invoke the perceptron layer 406 to convert the feature dimensions of the first sample content feature and the second sample content feature to be the same as the feature dimension of the interest feature. In this way, the first similarity and the second similarity can be determined in the same dimension space, and the determined similarity is more accurate. Optionally, the server invokes the perceptron layer 406 to perform a linear mapping on the first sample content feature and the second sample content feature, and maps the feature dimensions of the first sample content feature and the second sample content feature to be the same as the feature dimension of the interest feature.

[0163] Optionally, the implementation manner in which the server trains the multimedia recommendation model according to the first similarity and the second similarity is as follows: The server adjusts the parameters of the multimedia recommendation model so that the first similarity obtained after the adjustment increases, the second similarity decreases, and the difference between the first similarity and the second similarity reaches a reference threshold. Since the first sample content feature is the content feature of the sample multimedia data that the user is interested in, and the second sample content feature is the content feature of the sample multimedia data that the user is not interested in, an increase in the first similarity and a decrease in the second similarity indicate that the interest feature obtained by the multimedia recommendation model is more matched with the multimedia data that the user is interested in and less matched with the multimedia data that the user is not interested in. That is, the obtained interest feature is more accurate. Therefore, when the multimedia recommendation model is subsequently used, the accuracy of multimedia recommendation based on the user's interest feature is higher.

[0164] Optionally, the first similarity and the second similarity are cosine similarities. Of course, they can also be other similarities, and the embodiments of the present application do not limit this. Optionally, the server obtains the loss value of the multimedia recommendation model according to the first similarity and the second similarity, and trains the multimedia recommendation model according to this loss value. For example, the server adjusts the parameters of the multimedia recommendation model so that the loss value obtained after the adjustment decreases and is less than the reference loss value. Optionally, the loss value of the multimedia recommendation model is obtained through the following formula (1).

[0165] L = argmin max(0, d(f(u), g(p)) - d(f(u), g(n)) - 1) (1)

[0166] Wherein, L represents the loss value of the multimedia recommendation model, f(u) represents the first interest feature, g(p) represents the first sample content feature of the positive sample data, and g(n) represents the second sample content feature of the negative sample data. d(f(u), g(p)) represents the first similarity, that is, the similarity between the first interest feature and the first sample content feature. d(f(y), g(n)) represents the second similarity, that is, the similarity between the first interest feature and the second sample content feature.

[0167] Optionally, the above multimedia recommendation model is BERT (Bidirectional Encoder Representation from Transformers, the bidirectional encoder representation of Transformer (a model)). Of course, the above multimedia recommendation model can also be other models, and the embodiments of the present application do not limit this.

[0168] In the embodiments of the present application, the interest characteristics of a user are obtained according to the content characteristics of multiple sample multimedia data, and a multimedia recommendation model is trained according to the interest characteristics of the user, the content characteristics of positive sample data, and the content characteristics of negative sample data, so that the multimedia recommendation model can learn the relationship between the content characteristics of positive sample data and the interest characteristics of the user, and the relationship between the content characteristics of negative sample data and the interest characteristics of the user. That is, the multimedia recommendation model can learn the relationship between the multimedia data that the user is interested in and the interest characteristics of the user, and the relationship between the multimedia data that the user is not interested in and the interest characteristics of the user, so as to have the ability to recommend multimedia data that the user may be interested in according to the interest characteristics of the user and the content characteristics of the multimedia data to be recommended.

[0169] Moreover, before training the multimedia recommendation model according to the interest characteristics of the user, the content characteristics of positive sample data, and the content characteristics of negative sample data, the multimedia data in the sample play sequence is masked, and the multimedia recommendation model is called to predict the multimedia data with the masked content in the sample play sequence, so that the multimedia recommendation model can quickly learn the dependency relationship between multiple multimedia data in the play sequence and the context characteristics of the multimedia data, thereby optimizing the model parameters. Then, the multimedia recommendation model is further trained according to the interest characteristics of the user and the content characteristics of positive and negative sample data, which can further optimize the model parameters, and the recommendation accuracy of the obtained multimedia recommendation model is higher. Moreover, through these two-step training methods, the learning ability of the model is rapidly improved, so that a good training effect can be achieved with less sample data, reducing the training cost of the model.

[0170] Figure 5 It is a flowchart of a multimedia recommendation method provided by the embodiments of the present application. In this embodiment, the execution entity is taken as an example of a server for description. Refer to Figure 5 , this embodiment includes:

[0171] 501. The server obtains the historical play sequence corresponding to the target user identifier.

[0172] The historical play sequence includes multiple multimedia data played by the terminal logged in with the target user identifier, and the multiple multimedia data are arranged in the play order. The user identifier is used to represent the user identity. Optionally, the user identifier is the user's mobile phone number, ID number, etc. Optionally, the target user identifier is any user identifier.

[0173] Optionally, the server obtains the historical playback sequence according to the behavior log of the target user. Exemplarily, the server obtains the historical playback sequence within a reference time range from the current time. For example, the server obtains the historical playback sequence within the most recent 1 hour. Optionally, the historical playback sequence includes a reference number of multimedia data. For example, it includes 10 multimedia data.

[0174] Among them, the implementation manner of the server obtaining the historical playback sequence according to the behavior log of the target user is similar to the implementation manner of the server obtaining the second sample playback sequence according to the behavior log of the sample user identifier in step 301 above, and will not be elaborated here.

[0175] Optionally, after the server obtains the original playback sequence of the multimedia data from the user's behavior log, it first filters out the dirty data therein to obtain the historical playback sequence. Among them, the dirty data includes multimedia data with a playback duration less than the reference duration, multimedia data with a ratio of the playback duration to the total duration less than the reference ratio, etc. Optionally, the reference duration is any duration. For example, the reference duration is 10 seconds. Optionally, the reference ratio is any ratio. For example, the reference ratio is 0.3.

[0176] Optionally, the timing for the server to obtain the historical playback sequence corresponding to the target user identifier is: the server obtains the historical playback sequence corresponding to the target user identifier in response to receiving a multimedia recommendation request sent by the terminal. Among them, the multimedia recommendation request is sent by the terminal in response to an operation of obtaining multimedia data. Optionally, the operation of obtaining multimedia data is a trigger operation on a recommendation option in the multimedia recommendation interface, or a sliding operation in the multimedia recommendation interface, or an operation of opening the target application, etc. Of course, it can also be other operations, and the embodiments of the present application do not limit this.

[0177] Reference Figure 6 , is a schematic diagram of the multimedia recommendation interface 602. Among them, taking the multimedia data as a video as an example for illustration. The multimedia recommendation interface includes a recommendation option 601. The terminal sends a multimedia recommendation request to the server in response to a trigger operation on the recommendation option 601. Or, the terminal sends a multimedia recommendation request to the server in response to a sliding operation in the multimedia recommendation interface 602, such as a swiping up operation.

[0178] 502. The server obtains the content feature and the position feature of each multimedia data.

[0179] Among them, the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence.

[0180] Optionally, the server obtains the position feature of each multimedia data according to the arrangement order of each multimedia data in the historical playback sequence. For example, if there are 3 multimedia data in the historical playback sequence, the position feature of the first multimedia data is {1, 0, 0}, the position feature of the second multimedia data is {0, 1, 0}, and the position feature of the third multimedia data is {0, 0, 1}.

[0181] Optionally, the server queries the content feature of each multimedia data from the multimedia database. The multimedia database is used to store the content features of multimedia data. Optionally, the content features of the multimedia data in the multimedia database are stored in the multimedia database after being obtained by the content recognition model. In this way, when the server needs to obtain the content features of the multimedia data, it can directly query from the multimedia database, with high efficiency. Or, the server directly calls the content recognition model to extract the features of each multimedia data to obtain the content features of each multimedia data.

[0182] In a possible implementation, the server calls the content recognition model to extract the features of each multimedia data to obtain the content features of each multimedia data, including: the server calls the content recognition model to perform multi-modal feature extraction on each multimedia data to obtain the multi-modal content features of each multimedia data.

[0183] The multi-modal feature refers to describing the basic attributes of an object through multiple expression methods. For example, by using multiple expression methods such as text, image, and audio to describe an object, the text feature, image feature, and audio feature of the object can be obtained. In the embodiments of the present application, the multi-modal content feature is to describe the content feature of the multimedia data through multiple expression methods. Optionally, the multi-modal content feature includes at least two of the content features of the multimedia title, the content features of the multimedia screen, or the content features of the multimedia audio. Optionally, the multimedia screen includes a background screen and subtitles. Correspondingly, the content features of the multimedia screen include the content features of the multimedia background screen and the content features of the subtitles.

[0184] For example, if the multimedia data is a short video, the multi-modal content features of the short video can include the content features in the title of the short video, the content features in the background screen of the short video, the content features of the subtitles of the short video, the content features in the background audio of the short video, etc.

[0185] In the embodiments of the present application, by performing multi-modal feature extraction on each multimedia data to obtain the multi-modal content features of each multimedia data, since the multi-modal content features can describe the content of the multimedia data from multiple angles, the representation of the content of the multimedia data is more accurate.

[0186] In a possible implementation, the training process of the content recognition model includes the following (A)-(D):

[0187] (A) The server obtains sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data.

[0188] (B) The server calls the content recognition model to extract features from the sample multimedia data to obtain the content features of the sample multimedia data.

[0189] Optionally, this step includes: The server calls the content recognition model to perform multimodal feature extraction on the sample multimedia data to obtain the multimodal content features of the sample multimedia data.

[0190] (C) The server calls the content classification model to classify the content features to obtain a predicted label.

[0191] (D) The server trains the content recognition model and the content classification model according to the sample label and the predicted label.

[0192] The implementation of this step is as follows: If the predicted label is different from the sample label, the server adjusts the parameters of the content recognition model and the parameters of the content classification model so that the obtained predicted label is the same as the sample label. Optionally, if the predicted label is the same as the sample label, the content recognition model and the content classification model are continuously trained with new sample information until the number of training sample information reaches a reference quantity, and the trained content recognition model and content classification model are obtained.

[0193] Alternatively, the content recognition model and the content classification model are called, a reference quantity of predicted labels are obtained according to a reference quantity of sample information, the classification accuracy of the content classification model is determined by the reference quantity of predicted labels and the corresponding sample labels, and if the classification accuracy does not reach a reference threshold, the content recognition model and the content classification model are continuously trained until the classification accuracy of the content classification model reaches the reference threshold. Among them, the reference quantity is greater than 1.

[0194] In the embodiments of the present application, the content classification model is used to assist in training the content recognition model. Since the content classification model performs content classification based on the content features extracted by the content recognition model, if the classification accuracy of the content classification model is higher, it means that the content features extracted by the content recognition model are more accurate. In this way, the feature extraction effect of the content recognition model can be intuitively shown through the classification accuracy, and the content features extracted by the obtained content recognition model are more accurate.

[0195] In a possible implementation, the server obtains the content features and location features of multimedia data through a multimedia recommendation model. The implementation is as follows: The server calls the content extraction layer in the multimedia recommendation model to query the content features of multiple multimedia data from the multimedia database; and calls the location extraction layer in the multimedia recommendation model to obtain the location features of multiple multimedia data.

[0196] Among them, the implementation of the server calling the location extraction layer in the multimedia recommendation model to obtain the location features of multiple multimedia data is as follows: The server inputs the historical playback sequence into the multimedia recommendation model, calls the location extraction layer in the multimedia recommendation model, and obtains the location features of the multiple sample multimedia data according to the arrangement order of the multiple multimedia data in the historical playback sequence.

[0197] It should be noted that, optionally, the content features of the multimedia data stored in the above multimedia database are the multi-modal content features of the multimedia data.

[0198] 503. The server fuses the content features and location features of each multimedia data to obtain the multimedia features of each multimedia data.

[0199] Optionally, the implementation of this step is as follows: The server sums the content features and location features of each multimedia data to obtain the multimedia features of each multimedia data. Or, the content features and location features of each multimedia data are fused by other means to obtain the multimedia features of each multimedia data, and the embodiments of the present application do not limit this.

[0200] Optionally, the implementation of this step is as follows: The server calls the feature fusion layer in the multimedia recommendation model to fuse the content features and location features of each multimedia data to obtain the multimedia features of each multimedia data. For example, the server calls the feature fusion layer in the multimedia recommendation model to perform vector summation on the content features and location features of each multimedia data to obtain the multimedia features of each multimedia data.

[0201] 504. The server fuses the multiple multimedia features according to the arrangement order of the multiple multimedia data to obtain the interest features of the target user identifier.

[0202] Optionally, the implementation of this step is as follows: The server calls the interest extraction layer in the multimedia recommendation model to fuse the multiple multimedia features according to the arrangement order of the multiple multimedia data to obtain the interest features of the target user identifier.

[0203] Optionally, the server invokes the interest extraction layer in the multimedia recommendation model, and fuses multiple multimedia features in the arranged order of the multiple multimedia data to obtain the interest feature of the target user identifier, including: the server invokes the interest extraction layer, determines the relevance between each multimedia feature and the target multimedia feature in the arranged order of the multiple multimedia data, extracts the hidden layer multimedia feature from the target multimedia feature according to the determined relevance, and the target multimedia feature is any one of the multiple multimedia features; fuses the hidden layer multimedia features of the multiple multimedia data to obtain the interest feature of the target user identifier.

[0204] Among them, the hidden layer multimedia feature is the representation of the multimedia feature in the interest extraction layer. In the interest extraction layer, for each target multimedia feature, the relevance between each multimedia feature before and after the target multimedia feature and the multimedia feature is determined. The greater the relevance between the multimedia feature and the target multimedia feature, the greater the attention weight of the multimedia feature when extracting the hidden layer multimedia feature from the target multimedia feature, so that the extracted hidden layer multimedia feature can include the context feature of the target multimedia feature. In this way, after fusing the hidden layer multimedia features of the multiple multimedia data, the obtained interest feature is more accurate.

[0205] 505. The server obtains the content features of at least one multimedia data to be recommended.

[0206] Optionally, the server queries the content features of the at least one multimedia data to be recommended from the multimedia database. Optionally, the server invokes the content extraction layer in the multimedia recommendation model to query the content features of the at least one multimedia data to be recommended from the multimedia database. In this way, the efficiency of obtaining content features is high.

[0207] Optionally, the server invokes the content recognition model to perform feature extraction on each multimedia data to be recommended to obtain the content features of each multimedia data to be recommended. Optionally, the server invokes the content recognition model to perform multi-modal feature extraction on each multimedia data to be recommended to obtain the multi-modal content features of each multimedia data to be recommended.

[0208] In the embodiment of the present application, multi-modal feature extraction is performed on each multimedia data to be recommended to obtain the multi-modal content features of each multimedia data to be recommended. Since the multi-modal content features can describe the content of the multimedia data to be recommended from multiple perspectives, the representation of the content of the multimedia data to be recommended is more accurate.

[0209] 506. The server obtains the target multimedia data that matches the interest feature according to the interest feature and the content features of at least one multimedia data to be recommended.

[0210] Optionally, this step is implemented as follows: The server invokes the feature matching layer in the multimedia recommendation model, and based on the interest features and the content features of at least one multimedia data to be recommended, obtains the target multimedia data that matches the interest features.

[0211] Optionally, the implementation manner in which the server invokes the feature matching layer in the multimedia recommendation model and obtains the target multimedia data that matches the interest features based on the interest features and the content features of at least one multimedia data to be recommended is as follows: The server invokes the feature matching layer in the multimedia recommendation model, and based on the interest features and the content features of at least one multimedia data to be recommended, obtains the similarity between the interest features and the content features of each multimedia data to be recommended, and selects the target multimedia data from at least one multimedia data to be recommended according to the similarity. The similarity corresponding to the target multimedia data is greater than the similarity corresponding to other multimedia data to be recommended.

[0212] Continue to refer to Figure 4 , before the server determines the similarity between the interest features and the content features of at least one multimedia data to be recommended, it needs to first invoke the perceptron layer 406 to convert the feature dimension of the content features of at least one multimedia data to be recommended into the same feature dimension as the interest features. In this way, the similarity determined in the same dimension space is more accurate.

[0213] Among them, the implementation manner in which the server selects the target multimedia data from at least one multimedia data to be recommended according to the similarity is as follows: The server selects at least one target multimedia data whose similarity is greater than the reference similarity from at least one multimedia data to be recommended. Alternatively, the server selects a reference number of target multimedia data from at least one multimedia data to be recommended. Among them, the reference similarity and the reference number are set to any values according to needs.

[0214] Continue to refer to Figure 4 , the multimedia recommendation model further includes an output layer 407. Correspondingly, the server invokes the output layer 407 and selects the target multimedia data from at least one multimedia data to be recommended according to the similarity. In the embodiments of the present application, after obtaining the user's interest features, based on the similarity between the content features of the multimedia data to be recommended and the user's interest features, videos are recommended to the user, so that new multimedia data to be recommended can also be recommended to users who may be interested in it by virtue of its own content features, solving the cold start problem of multimedia data and expanding the applicable scope of the solution. In addition, in the embodiments of the present application, based on the similarity between the content features of the multimedia data to be recommended and the user's interest features, videos are recommended to the user, which is beneficial to mining long-tail content, that is, mining and utilizing some currently relatively unpopular but content-rich multimedia data, thereby promoting the wide spread of multimedia data.

[0215] 507. The server sends target multimedia data to the terminal logged in with the target user identifier.

[0216] After receiving the target multimedia data, the terminal can play the target multimedia data according to the detected user operation.

[0217] Referring to Table 1, it is a performance comparison between the multimedia recommendation model in the embodiments of the present application and other multimedia recommendation models. The evaluation method is to call each multimedia recommendation model, predict the second half of the playback sequence based on the first half of the playback sequence in the multimedia data playback sequence, and then count the performance indicators of each multimedia recommendation model. Among them, BERT4CVRP is the multimedia recommendation model in the present application, and the other two models are the CDML (Collaborative deepmetric learning for video understanding) model and the DMLBBS (DeepMetric Learning Beyond Binary Supervision) model. HR (Hit Ratio) refers to the ratio of the number of correctly predicted multimedia data by the multimedia recommendation model to the total number of multimedia data in the second half of the playback sequence. NDCG (Normalized Discounted cumulative gain) is an index for evaluating the sorting accuracy of multiple predicted multimedia data. @5 means taking the first 5 multimedia data in the second half of the playback sequence to form an evaluation sequence, and @10 means taking the first 10 multimedia data in the second half of the playback sequence to form an evaluation sequence. It can be seen from Table 1 that the performance of the multimedia recommendation model in the present application is better than that of the other two multimedia recommendation models.

[0218] Table 1

[0219] Model NDCG@5 NDCG@10 HR@5 HR@10 CDML 0.6247 0.6278 0.4644 0.4563 DMLBBS 0.5658 0.5667 0.3558 0.3479 BERT4CVRP 0.7565 0.7664 0.4824 0.4955

[0220] The technical solution provided by the embodiments of the present application recommends multimedia data to users by mining the interest characteristics of users. Among them, when obtaining the interest characteristics, the content characteristics and position characteristics of multiple multimedia data in the user's historical playback sequence are fused to obtain multiple multimedia characteristics, and these multiple multimedia characteristics are fused to obtain the user's interest characteristics. Since this interest characteristic considers the content of multiple multimedia data played by the user and the influence of the playback order of these multiple media data on the multimedia data that the user wants to watch next, this interest characteristic is more matched with the user, so that the multimedia data recommended to the user according to this interest characteristic can better meet the user's interests and improve the recommendation accuracy.

[0221] It should be noted that, in the above embodiments, the server is taken as an example of the execution entity. Optionally, the execution entity may also be other computer devices, such as terminals.

[0222] Figure 7 It is a block diagram of a multimedia recommendation device provided by an embodiment of the present application. Refer to Figure 7 The device includes:

[0223] A sequence acquisition module 701, configured to acquire a historical playback sequence corresponding to a target user identifier, where the historical playback sequence includes multiple multimedia data played by a terminal logged in with the target user identifier, and the multiple multimedia data are arranged in the playback order;

[0224] A first fusion module 702, configured to fuse the content feature and the position feature of each multimedia data to obtain a multimedia feature of each multimedia data, where the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence;

[0225] A second fusion module 703, configured to fuse multiple multimedia features in the arrangement order of the multiple multimedia data to obtain an interest feature of the target user identifier;

[0226] A multimedia recommendation module 704, configured to recommend multimedia data for the target user identifier according to the interest feature.

[0227] In a possible implementation manner, the first fusion module 702 is configured to call a feature fusion layer in the multimedia recommendation model to fuse the content feature and the position feature of each multimedia data to obtain a multimedia feature of each multimedia data;

[0228] The second fusion module 703 is configured to call an interest extraction layer in the multimedia recommendation model to fuse multiple multimedia features in the arrangement order of the multiple multimedia data to obtain an interest feature of the target user identifier.

[0229] In another possible implementation manner, the second fusion module 703 is configured to call an interest extraction layer to determine the relevance between each multimedia feature and a target multimedia feature in the arrangement order of the multiple multimedia data, extract a hidden layer multimedia feature from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the multiple multimedia features; fuse the hidden layer multimedia features of the multiple multimedia data to obtain an interest feature of the target user identifier.

[0230] In another possible implementation manner, the training process of the multimedia recommendation model includes:

[0231] Obtain the content features of multiple sample multimedia data in the first sample playback sequence, the first sample content features of the positive sample data, and the second sample content features of the negative sample data. The positive sample data is the multimedia data that has been recommended and played after the first sample playback sequence, and the negative sample is the multimedia data that has been recommended but not played.

[0232] Invoke the multimedia recommendation model, and obtain the first interest feature according to the content features of multiple sample multimedia data in the first sample playback sequence.

[0233] Train the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature.

[0234] In another possible implementation, training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature includes:

[0235] Obtain the first similarity between the first interest feature and the first sample content feature, and the second similarity between the first interest feature and the second sample content feature.

[0236] Train the multimedia recommendation model according to the first similarity and the second similarity.

[0237] In another possible implementation, before training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature, the training process of the multimedia recommendation model further includes:

[0238] Obtain the second sample playback sequence.

[0239] Perform a masking process on the first sample multimedia data in the second sample playback sequence. The masking process refers to masking the content of the first sample multimedia data.

[0240] Invoke the multimedia recommendation model, and predict the second sample multimedia data with the masked content according to the masked second sample playback sequence.

[0241] Train the multimedia recommendation model according to the first sample multimedia data and the second sample multimedia data.

[0242] In another possible implementation, referring to Figure 8 , the apparatus further includes:

[0243] The first feature acquisition module 705 is used to call the content extraction layer in the multimedia recommendation model to query the content features of multiple multimedia data from the multimedia database.

[0244] The second feature acquisition module 706 is used to call the location extraction layer in the multimedia recommendation model to obtain the location features of multiple multimedia data.

[0245] In another possible implementation, continuing to refer to Figure 8 , the device further includes:

[0246] A content feature recognition module 707, configured to call a content recognition model to extract features from each multimedia data, so as to obtain the content features of each multimedia data.

[0247] In another possible implementation, the content feature recognition module 707 is configured to call a content recognition model to perform multimodal feature extraction on each multimedia data, so as to obtain the multimodal content features of each multimedia data; wherein, the multimodal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

[0248] In another possible implementation, the training process of the content recognition model includes:

[0249] Obtain sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data;

[0250] Call a content recognition model to extract features from the sample multimedia data, so as to obtain the content features of the sample multimedia data;

[0251] Call a content classification model to classify the content features to obtain a predicted label;

[0252] Train the content recognition model and the content classification model according to the sample label and the predicted label.

[0253] In another possible implementation, calling a content recognition model to extract features from the sample multimedia data, so as to obtain the content features of the sample multimedia data, includes:

[0254] Call a content recognition model to perform multimodal feature extraction on the sample multimedia data, so as to obtain the multimodal content features of the sample multimedia data;

[0255] Wherein, the multimodal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

[0256] In another possible implementation, continuing to refer to Figure 8 , the multimedia recommendation module 704 includes:

[0257] A third feature acquisition unit 7041, configured to acquire the content features of at least one multimedia data to be recommended;

[0258] A feature matching unit 7042, configured to obtain target multimedia data that matches the interest feature according to the interest feature and the content features of at least one multimedia data to be recommended.

[0259] A multimedia sending unit 7043, configured to send the target multimedia data to a terminal.

[0260] In another possible implementation, a feature matching unit 7041 is configured to call a feature matching layer in a multimedia recommendation model, and obtain target multimedia data that matches the interest feature according to the interest feature and the content features of at least one multimedia data to be recommended.

[0261] The technical solution provided by the embodiments of the present application recommends multimedia data to a user by mining the user's interest features. Among them, when obtaining the interest feature, the content features and position features of multiple multimedia data in the user's historical play sequence are fused to obtain multiple multimedia features, and the multiple multimedia features are fused to obtain the user's interest feature. Since this interest feature takes into account the content of multiple multimedia data played by the user and the influence of the play order of the multiple media data on the multimedia data that the user wants to watch next, this interest feature is more matched with the user, so that the multimedia data recommended to the user according to this interest feature can better meet the user's interests and improve the recommendation accuracy.

[0262] It should be noted that when the multimedia recommendation device provided in the above embodiment performs multimedia recommendation, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the multimedia recommendation device provided in the above embodiment and the embodiment of the multimedia recommendation method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0263] Figure 9 The block diagram of a terminal 900 provided by an exemplary embodiment of the present application is shown. The terminal 900 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 900 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0264] Generally, the terminal 900 includes a processor 901 and a memory 902.

[0265] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0266] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one program code, and the at least one program code is used to be executed by the processor 901 to implement the multimedia recommendation method provided in the method embodiments of the present application.

[0267] In some embodiments, the terminal 900 may further optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning component 908, and a power supply 909.

[0268] The peripheral device interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0269] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0270] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905. The touch signals can be input to the processor 901 as control signals for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, which is provided on the front panel of the terminal 900; in other embodiments, there may be at least two display screens 905, which are respectively provided on different surfaces of the terminal 900 or are in a foldable design; in other embodiments, the display screen 905 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 900. Even, the display screen 905 can also be set as an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 905 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0271] The camera module 906 is used to collect images or videos. Optionally, the camera module 906 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0272] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 901 for processing, or input to the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.

[0273] The positioning component 908 is used to locate the current geographical location of the terminal 900 to achieve navigation or LBS (Location Based Service). The positioning component 908 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.

[0274] The power supply 909 is used to supply power to each component in the terminal 900. The power supply 909 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 909 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0275] In some embodiments, the terminal 900 further includes one or more sensors 910. The one or more sensors 910 include but are not limited to: an acceleration sensor 911, a gyroscope sensor 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.

[0276] The acceleration sensor 911 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 900. For example, the acceleration sensor 911 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 901 can control the display screen 905 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 911. The acceleration sensor 911 can also be used for collecting game or user's motion data.

[0277] The gyroscope sensor 912 can detect the body direction and rotation angle of the terminal 900. The gyroscope sensor 912 can cooperate with the acceleration sensor 911 to collect the 3D actions of the user on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0278] The pressure sensor 913 can be disposed on the side frame of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side frame of the terminal 900, it can detect the holding signal of the user on the terminal 900, and the processor 901 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 905. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0279] The fingerprint sensor 914 is used to collect the fingerprint of the user. The processor 901 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 can identify the user's identity according to the collected fingerprint. When the identified user identity is a trusted identity, the processor 901 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 914 can be disposed on the front, back, or side of the terminal 900. When there is a physical button or a manufacturer logo on the terminal 900, the fingerprint sensor 914 can be integrated with the physical button or the manufacturer logo.

[0280] The optical sensor 915 is used to collect the ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera module 906 according to the ambient light intensity collected by the optical sensor 915.

[0281] The proximity sensor 916, also known as the distance sensor, is usually disposed on the front panel of the terminal 900. The proximity sensor 916 is used to collect the distance between the user and the front of the terminal 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from the lit state to the off state; when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from the off state to the lit state.

[0282] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the terminal 900, and it may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.

[0283] Figure 10 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. Among them, at least one program code is stored in the memory 1002, and the at least one program code is loaded and executed by the processor 1001 to implement the multimedia recommendation methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0284] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the operations performed in the multimedia recommendation method of the above embodiment.

[0285] An embodiment of the present application also provides a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by the processor to implement the operations performed in the multimedia recommendation method of the above embodiment.

[0286] The embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer program code, and the computer program code is stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the operations performed in the multimedia recommendation method in the above various optional implementation manners.

[0287] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The described program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0288] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multimedia recommendation method, characterized in that, The method includes: Obtaining a historical playback sequence corresponding to a target user identifier, where the historical playback sequence includes a plurality of multimedia data played by a terminal logged in with the target user identifier, and the plurality of multimedia data are arranged in the playback order; Fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, where the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence; Invoking an interest extraction layer in the multimedia recommendation model, determining the relevance between the multimedia feature of each multimedia data and the target multimedia feature in the arrangement order of the plurality of multimedia data, extracting a hidden-layer multimedia feature from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the plurality of multimedia features; fusing the hidden-layer multimedia features of the plurality of multimedia data to obtain the interest feature of the target user identifier; in the interest extraction layer, for each target multimedia feature, determining the relevance between each multimedia feature before and after the target multimedia feature and the target multimedia feature, the greater the relevance between the multimedia feature and the target multimedia feature, the greater the attention weight on the target multimedia feature when extracting the hidden-layer multimedia feature from the target multimedia feature, so that the extracted hidden-layer multimedia feature includes the context feature of the target multimedia feature; Recommending multimedia data for the target user identifier according to the interest feature.

2. The method according to claim 1, characterized in that The fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data includes: Invoking a feature fusion layer in the multimedia recommendation model to fuse the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data.

3. The method according to claim 2, characterized in that, The training process of the multimedia recommendation model includes: Obtaining the content features of a plurality of sample multimedia data in a first sample playback sequence, the first sample content feature of positive sample data, and the second sample content feature of negative sample data, where the positive sample data is multimedia data that has been recommended and played after the first sample playback sequence, and the negative sample is multimedia data that has been recommended but not played; Invoking the multimedia recommendation model to obtain a first interest feature according to the content features of the plurality of sample multimedia data in the first sample playback sequence; Training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature.

4. The method according to claim 3, wherein The training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature includes: Obtaining a first similarity between the first interest feature and the first sample content feature and a second similarity between the first interest feature and the second sample content feature; Training the multimedia recommendation model according to the first similarity and the second similarity.

5. The method according to claim 3, characterized in that, Before training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature, the training process of the multimedia recommendation model further includes: Obtain a second sample play sequence; Perform a masking process on the first sample multimedia data in the second sample play sequence, where the masking process refers to masking the content of the first sample multimedia data; Invoke the multimedia recommendation model, and predict the second sample multimedia data of the masked content according to the masked second sample play sequence; Train the multimedia recommendation model according to the first sample multimedia data and the second sample multimedia data.

6. The method according to claim 1, characterized in that, Before fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, the method further includes: Invoke the content extraction layer in the multimedia recommendation model to query the content features of the multiple multimedia data from the multimedia database; Invoke the position extraction layer in the multimedia recommendation model to obtain the position features of the multiple multimedia data.

7. The method according to claim 1, characterized in that, Before fusing the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, the method further includes: Invoke a content recognition model to perform feature extraction on each multimedia data to obtain the content feature of each multimedia data.

8. The method according to claim 7, wherein The invoking the content recognition model to perform feature extraction on each multimedia data to obtain the content feature of each multimedia data includes: Invoke the content recognition model to perform multi-modal feature extraction on each multimedia data to obtain the multi-modal content feature of each multimedia data; Wherein, the multi-modal content feature includes at least two of the content feature of the multimedia title, the content feature of the multimedia picture, or the content feature of the multimedia audio.

9. The method according to claim 7, wherein The training process of the content recognition model includes: Obtain sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data; Invoke the content recognition model to perform feature extraction on the sample multimedia data to obtain the content feature of the sample multimedia data; Invoke a content classification model to classify the content feature to obtain a predicted label; Train the content recognition model and the content classification model according to the sample label and the predicted label.

10. The method according to claim 9, characterized in that, The invoking the content recognition model to perform feature extraction on the sample multimedia data to obtain the content feature of the sample multimedia data includes: Invoke the content recognition model to perform multi-modal feature extraction on the sample multimedia data to obtain the multi-modal content feature of the sample multimedia data; Wherein, the multi-modal content feature includes at least two of the content feature of the multimedia title, the content feature of the multimedia picture, or the content feature of the multimedia audio.

11. The method according to claim 1, wherein The recommending multimedia data for the target user identifier according to the interest feature includes: Obtain the content features of at least one multimedia data to be recommended; Obtain target multimedia data that matches the interest feature according to the interest feature and the content feature of the at least one multimedia data to be recommended; Send the target multimedia data to the terminal.

12. A multimedia recommendation device, characterized in that, The device includes: A sequence acquisition module, configured to acquire a historical playback sequence corresponding to a target user identifier, where the historical playback sequence includes a plurality of multimedia data played by a terminal logged in with the target user identifier, and the plurality of multimedia data are arranged in the playback order; A first fusion module, configured to fuse the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data, where the content feature represents the content of the multimedia data, and the position feature represents the position of the multimedia data in the historical playback sequence; A second fusion module, configured to call an interest extraction layer in the multimedia recommendation model, determine the relevance between the multimedia feature of each multimedia data and the target multimedia feature according to the arrangement order of the plurality of multimedia data, extract hidden layer multimedia features from the target multimedia feature according to the determined relevance, where the target multimedia feature is any one of the plurality of multimedia features; fuse the hidden layer multimedia features of the plurality of multimedia data to obtain the interest feature of the target user identifier; in the interest extraction layer, for each target multimedia feature, determine the relevance between each multimedia feature before and after the target multimedia feature and the target multimedia feature, the greater the relevance between the multimedia feature and the target multimedia feature, the greater the attention weight for the target multimedia feature when extracting the hidden layer multimedia features from the target multimedia feature, so that the extracted hidden layer multimedia features include the context features of the target multimedia feature; A multimedia recommendation module, configured to recommend multimedia data for the target user identifier according to the interest feature.

13. The device according to claim 12, characterized in that, The first fusion module is configured to call a feature fusion layer in the multimedia recommendation model to fuse the content feature and the position feature of each multimedia data to obtain the multimedia feature of each multimedia data.

14. The device according to claim 13, characterized in that, The training process of the multimedia recommendation model includes: Obtain the content features of a plurality of sample multimedia data in a first sample playback sequence, the first sample content features of positive sample data, and the second sample content features of negative sample data, where the positive sample data is multimedia data that has been recommended and played after the first sample playback sequence, and the negative sample is multimedia data that has been recommended but not played; Call the multimedia recommendation model to obtain a first interest feature according to the content features of the plurality of sample multimedia data in the first sample playback sequence; Train the multimedia recommendation model according to the first interest feature, the first sample content features, and the second sample content features.

15. The device according to claim 14, wherein, The training of the multimedia recommendation model according to the first interest feature, the first sample content features, and the second sample content features includes: Obtain a first similarity between the first interest feature and the first sample content feature and a second similarity between the first interest feature and the second sample content feature; Train the multimedia recommendation model according to the first similarity and the second similarity.

16. The device according to claim 15, characterized in that, Before training the multimedia recommendation model according to the first interest feature, the first sample content feature, and the second sample content feature, the training process of the multimedia recommendation model further includes: Obtain a second sample play sequence; Perform a masking process on the first sample multimedia data in the second sample play sequence, where the masking process refers to masking the content of the first sample multimedia data; Invoke the multimedia recommendation model, and predict the second sample multimedia data of the masked content according to the masked second sample play sequence; Train the multimedia recommendation model according to the first sample multimedia data and the second sample multimedia data.

17. The device according to claim 12, characterized in that, The apparatus further includes: A first feature acquisition module, configured to invoke a content extraction layer in the multimedia recommendation model to query the content features of the multiple multimedia data from a multimedia database; A second feature acquisition module, configured to invoke a position extraction layer in the multimedia recommendation model to obtain the position features of the multiple multimedia data.

18. The device according to claim 12, characterized in that, The apparatus further includes: A content feature recognition module, configured to invoke a content recognition model to perform feature extraction on each multimedia data to obtain the content features of each multimedia data.

19. The device according to claim 18, characterized in that, The content feature recognition module is configured to invoke the content recognition model to perform multi-modal feature extraction on each multimedia data to obtain the multi-modal content features of each multimedia data; wherein, the multi-modal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

20. The device according to claim 18, characterized in that The training process of the content recognition model includes: Obtain sample information, where the sample information includes sample multimedia data and a sample label, and the sample label is used to describe the content of the sample multimedia data; Invoke the content recognition model to perform feature extraction on the sample multimedia data to obtain the content features of the sample multimedia data; Invoke a content classification model to classify the content features to obtain a predicted label; Train the content recognition model and the content classification model according to the sample label and the predicted label.

21. The device according to claim 20, characterized in that, The invoking the content recognition model to perform feature extraction on the sample multimedia data to obtain the content features of the sample multimedia data includes: Invoke the content recognition model to perform multi-modal feature extraction on the sample multimedia data to obtain the multi-modal content features of the sample multimedia data; Wherein, the multi-modal content features include at least two of the content features of the multimedia title, the content features of the multimedia picture, or the content features of the multimedia audio.

22. The device according to claim 12, characterized in that, The multimedia recommendation module includes: A third feature acquisition unit, configured to obtain the content features of at least one multimedia data to be recommended; A feature matching unit, configured to obtain target multimedia data that matches the interest feature according to the interest feature and the content features of the at least one multimedia data to be recommended; A multimedia sending unit, configured to send the target multimedia data to the terminal.

23. The device according to claim 22, characterized in that, The feature matching unit is configured to call a feature matching layer in a multimedia recommendation model, and obtain target multimedia data that matches the interest feature according to the interest feature and the content features of the at least one multimedia data to be recommended.

24. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program code is stored in the memory, and the program code is loaded and executed by the processor to implement the operations performed by the multimedia recommendation method according to any one of claims 1 to 11.

25. A computer-readable storage medium, characterized in that, At least one program code is stored in the storage medium, and the program code is loaded and executed by a processor to implement the operations performed by the multimedia recommendation method according to any one of claims 1 to 11.

26. A computer program product, characterized in that, The computer program product includes computer program code, and the computer program code is stored in a computer-readable storage medium; a processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the operations performed by the multimedia recommendation method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Content recommendation method, device and equipment and storage medium

    CN111680217A