Multimedia data processing method, device, equipment and storage medium

By calculating the interaction weight of multimedia data and the media behavior data of the target object, and using the collaborative filtering algorithm to generate the target interest level, the problem of low accuracy in multimedia data push is solved and accurate personalized recommendations are achieved.

CN114860968BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210514671.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-09-16
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

In the existing technology of multimedia data push, the uniform distribution of users' media behavior data leads to low push accuracy and inability to achieve accurate recommendations.

Method used

By obtaining the media attribute information of multimedia data and the media behavior data of the target object, the interaction weight is calculated, and the target interest degree is generated using the collaborative filtering algorithm to perform personalized recommendations.

Benefits of technology

The accuracy of multimedia data push is improved, accurate personalized recommendations are achieved, and the problem of low push accuracy when media behavior data is evenly distributed is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860968B_ABST
    Figure CN114860968B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a multimedia data processing method, apparatus, device and storage medium, which are applied to the fields of artificial intelligence and Internet of Vehicles, wherein the method includes: determining a target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij Based on P relevant information, M media attribute information, and M media behavior data corresponding to N target objects, interaction weights for the N target objects with respect to the M media data are generated. Based on the interaction weights for the N target objects with respect to the M media data, the M media attribute information and the M media behavior data are collaboratively processed to obtain target interest levels for the N target objects with respect to the M media data. Based on the target interest levels, multimedia data are recommended to the N target objects. This application can improve the accuracy of multimedia data push.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to fields such as artificial intelligence and vehicle networking technology, and in particular to a multimedia data processing method, device, equipment and storage medium. Background Art

[0002] With the development of big data, various multimedia platforms (such as audio and video platforms, gaming platforms, etc.) have gradually introduced personalized push algorithms based on big data to push multimedia data to users, improving the convenience of users accessing multimedia data. In the process of pushing multimedia data, it is usually necessary to push multimedia data to users based on the user's media behavior data related to the multimedia data. In practice, it has been found that if a user's media behavior data related to multimedia data is relatively evenly distributed, this recommendation method will recommend all multimedia data to the user, resulting in relatively low multimedia data push accuracy. Summary of the Invention

[0003] The embodiments of the present application provide a multimedia data processing method, apparatus, device, and storage medium to improve the push accuracy of multimedia data.

[0004] An embodiment of the present application provides a method for processing multimedia data, including:

[0005] Acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; one target object and one multimedia data together correspond to one piece of media behavior data;

[0006] Determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M;

[0007] If the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated according to the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the relevant information P is the interaction weight of the N target objects with respect to the M multimedia data. ij Belonging to the P relevant information;

[0008] According to the interaction weights of the N target objects with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest levels of the N target objects with respect to the M multimedia data. According to the target interest levels of the N target objects with respect to the M multimedia data, multimedia data are recommended to the N target objects respectively.

[0009] An embodiment of the present application provides a multimedia data processing device, including:

[0010] an acquisition module, configured to acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; one target object and one multimedia data together correspond to one piece of media behavior data;

[0011] Determination module, used to determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M;

[0012] The generating module is configured to generate the interaction weights of the N target objects with respect to the M multimedia data respectively based on the P relevant information, the M media attribute information and the M media behavior data respectively corresponding to the N target objects after the relevant information corresponding to the M multimedia data is determined; P is the product of M and N, and the relevant information P ij Belonging to the P relevant information;

[0013] The recommendation module is used to collaboratively process the M media attribute information and the M media behavior data corresponding to the N target objects based on the interaction weights of the N target objects with respect to the M multimedia data, obtain the target interest levels of the N target objects with respect to the M multimedia data, and recommend multimedia data to the N target objects based on the target interest levels of the N target objects with respect to the M multimedia data.

[0014] On one hand, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0015] On the one hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described are implemented.

[0016] On the one hand, an embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the method when executed by a processor.

[0017] In this application, the computer device can obtain the media attribute information corresponding to M multimedia data, and the media behavior data of N target objects respectively about the M multimedia data in the first time period, and determine the target object N. i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , the relevant information P ij Used to reflect the target user N i Media behavior data and multimedia data M j Further, if the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated based on the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects. i About multimedia data M j The interaction weight is mainly used to reflect the target object N i About multimedia data M j The interaction parameters (such as clicks, playback frequency), and the preference (i.e., popularity) of multimedia data are calculated. Then, according to the interaction weights of the N target objects respectively with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest of the N target objects respectively with respect to the M multimedia data. It can be seen that in the process of obtaining the target interest of the target object with respect to multimedia data, the introduction of interaction weights is conducive to expanding the distinction between the media behavior data of different target objects with respect to multimedia data, and can improve the accuracy of obtaining the target interest. By recommending multimedia data to the N target objects respectively based on their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 This is a schematic diagram of the architecture of a multimedia data processing system provided by this application;

[0020] Figure 2 This is a flowchart of a multimedia data processing method provided by this application;

[0021] Figure 3 This is a flowchart of a multimedia data processing method provided by this application;

[0022] Figure 4 This is a schematic diagram of a scenario for obtaining the initial interest of a target object in multimedia data provided by the present application;

[0023] Figure 5 This is a flowchart of a multimedia data processing method provided by this application;

[0024] Figure 6 This is a schematic diagram of a training scenario for a first initial media recognition model provided by this application;

[0025] Figure 7 This is a schematic diagram of a training scenario for a second initial media recognition model provided by this application;

[0026] Figure 8 This is a schematic structural diagram of a multimedia data processing device provided in an embodiment of the present application;

[0027] Figure 9 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] The present application relates to artificial intelligence. For example, the present application mainly relates to machine learning technology in artificial intelligence. Machine learning technology is used to analyze the media attribute information of M multimedia data and the media behavior data of N target objects regarding the M multimedia data, and the target interest of the N target objects regarding the M multimedia data is obtained. According to the target interest of the N target objects regarding the M multimedia data, multimedia data are recommended to the N target objects respectively. This can improve the push accuracy of multimedia data and realize the precise push of multimedia data.

[0030] Understandably, AI technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.

[0031] As you can see, machine learning (ML) is a multidisciplinary field, encompassing probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0032] In order to facilitate a clearer understanding of the present application, the media data processing system for implementing the media data processing method of the present application is first introduced. Figure 1 As shown, the media data processing system includes a server 10 and a terminal cluster. The terminal cluster may include one or more terminals. The number of terminals is not limited here. Figure 1 As shown, the terminal cluster can specifically include terminal 1, terminal 2, ..., terminal n; it can be understood that terminal 1, terminal 2, terminal 3, ..., terminal n can all be connected to the server 10 through the network, so that each terminal can exchange data with the server 10 through the network connection.

[0033] The terminal is installed with a multimedia platform that provides multimedia data to users. The multimedia platform may include but is not limited to: game application download platform, short video platform, content publishing platform, audio and video playback platform (such as audio playback application, audio radio station) and shopping platform, etc. The terminal can be used to recommend multimedia data to users based on the user's media behavior data about the multimedia data and the media attribute information of the multimedia data.

[0034] It is understandable that the specific content of multimedia data varies across different multimedia platforms. For example, on a game application download platform, multimedia data may refer to a game application, such as a single-player game, online game, mobile game, or mini-game; on a short video platform, multimedia data may refer to a video clip; on an audio and video playback platform, multimedia data may refer to a film or television work, audio data, etc.; on a shopping platform, multimedia data may refer to products or services sold on the shopping platform; and on a content publishing platform, multimedia data may refer to a literary work, a piece of news information, a travelogue, etc.

[0035] It is understood that user media behavior data related to multimedia data includes one or more of historical operation information and payment information. For example, if the multimedia data is audio data, the historical operation information includes at least one of click, comment, favorite, play, cancel, pause, stop, and exit. Payment information includes at least one of payment amount, number of payments, recharge amount, number of recharges, the time between the last payment and the current moment, and the time between the last recharge and the payment. Media attribute information of multimedia data is basic attribute information of the multimedia data. For example, if the multimedia data is audio data, the media attribute information of the audio data includes at least one of average play duration, average number of play times, total number of play times, audio data length, audio category, language type, and generation time.

[0036] It is understandable that the server 10 may refer to a device used to provide back-end services for the multimedia platform. For example, the server 10 may be used to review, arrange, and process multimedia data published by users on the multimedia platform. The server 10 may also be used to store users' media behavior data and media attribute information of multimedia data.

[0037] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of at least two physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can specifically refer to a vehicle-mounted terminal, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a screen speaker, a smart watch, etc., but is not limited to these. Each terminal and server can be directly or indirectly connected via wired or wireless communication. At the same time, the number of terminals and servers can be one or at least two, and this application does not limit this.

[0038] based on Figure 1 The multimedia data processing system in the present application can realize the multimedia data processing method in the present application. The method is suitable for multimedia data recommendation scenarios where the scale of multimedia data is relatively large and the target object's media behavior data such as click rate and play rate of multimedia data is relatively evenly distributed, such as Internet of Vehicles song recommendation scenario, video data recommendation scenario, product recommendation scenario and game application recommendation scenario, etc. Figure 2 As shown, Figure 2 In the description, the multimedia data is audio data as an example, and the multimedia data processing method includes the following steps S1 to S9:

[0039] S1. Data input and label construction phase. The server can obtain log data from the multimedia data platform, which includes media behavior data of L users regarding H multimedia data in the T-2 time period, media behavior data of L users regarding H multimedia data in the T-1 time period, media attribute information corresponding to the H multimedia data, and media behavior data of N users regarding M multimedia data in the T time period. Media behavior data includes clicks, evaluations, favorites, plays, cancel plays, pauses, stops, exits, play duration, number of plays, payment amount, number of payments, recharge amount, number of recharges, the time from the last payment to the current moment, the time between the last recharge and payment, etc.; media attribute information includes: average play duration, average number of plays, average UV, unit price, language type, style, etc. of media data. Among them, the media behavior data of L users regarding H multimedia data in the T-1 time period and the T-2 time period, and the media attribute information corresponding to the H multimedia data are used in the training stage of the media recognition model, and the media behavior data of N users regarding M multimedia data in the T time period are used in the recommendation stage of multimedia data, with L users as sample objects and N users as target objects.

[0040] Label construction phase: Based on the media behavior data of L users on H multimedia data in the T-1 time period, determine the labeling interest of L users on H multimedia data. For example, if user L e About multimedia data H in time period T-1 r Media behavior data reflects: User L e For multimedia data H r Execute the click operation and play operation, and use the first marked interest value as the user L e About multimedia data H r The annotation interest of user L e About multimedia data H in time period T-1 r Media behavior data reflects: User L e For multimedia data H r If a click operation is performed but a play operation is not performed, the second marked interest value is used as the user L e About multimedia data H r The first annotation interest degree is greater than the second annotation interest degree. For example, the first annotation interest degree is 1 and the second annotation interest degree is 0. When user L e When the annotation interest degree is 1, it means that user L e For multimedia data H r positive sample, when user L e When the annotation interest degree is 0, it means that user L e For multimedia data H r By determining the user's annotated interest in multimedia data based on the user's media behavior data in the T-1 period, manual annotation is not required, which improves the accuracy and efficiency of obtaining annotated interest.

[0041] S2, sample construction phase: Use the media behavior feature data of L users in the T-2 time period, the media attribute information of H multimedia data, and the annotated interest of L users in the H multimedia data in the T-1 time period to construct user sample data. For example, randomly divide the L users into K training objects and G test objects according to the division ratio. For example, according to general experience, randomly divide the sample into: training object: test object = 8:2 (that is, randomly divide the training object and test object in the ratio of 8:2). Further, the media behavior data of the training object is divided into third media behavior data with dense features and fourth media behavior data with sparse features, and the media attribute information of the H media data is divided into third media attribute information with dense features and fourth media attribute information with sparse features. The third media behavior data with sparse features and the third media attribute information with sparse features are used to train the deep learning layer of the first initial media recognition model. For example, the deep learning layer can be a deep learning network Deep. The fourth media behavior data with dense features, the fourth media attribute information with dense features, and the user's annotated interest in multimedia data are used to train the first interest recognition layer of the first initial media recognition model. For example, the first interest recognition layer can be a linear network Wide. Similarly, the media behavior data corresponding to the test object can be divided into fifth media behavior data with sparse features and sixth media behavior data with dense features. The media behavior data with sparse features are used to test the deep learning layer, and the media behavior data with dense features and the annotated interest are used to test the first interest recognition side. Similarly, the media behavior data and multimedia information of N users in a time period of T are divided into sparse features and dense features. The sparse features are used for prediction of the deep learning layer, and the dense features are used for prediction of the first interest recognition layer.

[0042] S3, score calculation stage. The fourth media behavior data and fourth media attribute information with sparse features are used to train the deep learning layer of the first initial media recognition model to obtain key media behavior data (i.e., embadding) and key media attribute information. The key media behavior data, key media attribute information, the third media behavior data of the training object, the third media attribute information, and the annotated interest of the test object in multimedia data obtained by the deep learning layer are used to train the first interest recognition layer of the first initial media recognition model. The model parameters of the adjusted (i.e., trained) first initial media recognition model are obtained by the gradient descent method. The trained first initial media recognition model is tested using the test object. If the interest recognition performance parameters (recall rate, precision rate, AUC, etc.) meet the evaluation effect, the trained first initial media recognition model, the probability score of the training object, and the probability score of the test object are saved. If the evaluation effect is not achieved, this step is repeated until the trained first initial media recognition model meets the evaluation effect.

[0043] Furthermore, the probability score obtained by the first initial media recognition model trained in step S2 is multiplied by a constant C, and the score intervals {S1, S2, ..., S H}, where Si = (Cpi-1, Cpi] is the user's rating of the i-th level of multimedia data [1,…,H represents the user's rating of 1,…,H levels of multimedia data], and pi represents the i-th probability interval. The rating interval is converted into a Rating Data matrix (where the behavior is the user, the column is the multimedia data, and the Rating Data with rows and columns intersecting is the user's rating of the multimedia data). The values ​​in the Rating Data matrix (abbreviated as RD matrix) are determined as the initial interest of the L users in the H multimedia data.

[0044] S4, the collaborative filtering algorithm construction stage based on the interaction between users and multimedia data. Weight constraints are imposed on users and multimedia data respectively, and a collaborative filtering algorithm model based on the interaction between users and multimedia data is constructed. The algorithm model is as follows: Formula (1), Formula (2-1), Formula (2-2), Formula (2-3), Formula (3), Formula (4-1), and Formula (4-2).

[0045] S5, average rating calculation stage. Based on the RD matrix in S3, the mean of all data in the RD matrix is ​​calculated as the overall average interest of the H multimedia data. The average rating of each user for all multimedia data in the RD matrix is ​​calculated to obtain the object average interest of each multimedia data. The average rating of each multimedia data in the matrix is ​​calculated to obtain the media average interest of the multimedia data.

[0046] S6: Calculation of interaction weights between users and multimedia data. Input the sample object from Step 2 and substitute the weight formulas and their constraints in Formulas (1), (2-1), (2-2), and (2-3). Using eigendecomposition, an interaction weight matrix is ​​obtained. The interaction weight matrix includes the interaction weights of each sample object with respect to each multimedia data item. Simultaneously, the eigenvalue vectors of the interaction weight matrix are obtained using eigendecomposition, and the mean of the eigenvalue vectors is calculated to obtain the average eigenvalue.

[0047] S7. Training of the second initial recognition model based on user interaction with multimedia data. Substitute the RD matrix and the interaction weight matrix of the sample objects with respect to the media data into formulas (3), (4-1), and (4-2) to obtain the object standard feature matrix (i.e., the user-side scoring matrix) and the media standard feature matrix (the multimedia data-side scoring matrix).

[0048] S8, the second initial recognition model testing phase based on the interaction between the user and the multimedia data. Input the interaction weight matrix obtained in S6, the object standard matrix U, the media standard matrix V, the average characteristic root, the overall average interest obtained in S5, the object average interest, and the media average interest to the second initial prediction model to obtain the target interest of the sample object with respect to the multimedia data. According to the RD matrix in S3 and the target interest of the sample object with respect to the multimedia data, the interest recognition error of the second initial media recognition model is calculated. If the interest recognition error is less than the error threshold, it is determined that the second media recognition model is in a convergence state, and the second media recognition model is obtained. Otherwise, repeat S2 to S8 above until the interest recognition error of the second initial media recognition model is less than the error threshold.

[0049] S9, multimedia recommendation phase based on user interaction with multimedia data: Based on the media behavior data of N target objects with respect to M multimedia data within a time period T and the media attribute information of the M multimedia data, the initial interest level of each target object with respect to each multimedia data (see steps S2-S3) and the interaction weight of each target object with respect to each multimedia data are calculated. Furthermore, based on the interaction weight, the initial interest level, the media behavior data of the N target objects with respect to the M multimedia data within a time period T and the media attribute information of the M multimedia data, the target interest level of each target object with respect to the multimedia data is calculated, and the multimedia data are sorted according to the target interest level, and the sorted multimedia data are recommended to the user.

[0050] It is understandable that the first media identification model and the second media identification model in the present application may refer to media identification models based on collaborative filtering, where collaborative filtering (CF) includes at least one of multimedia data-based collaborative filtering and user-based collaborative filtering. User-based collaborative filtering is a method of analyzing user interests, finding similar (interested) users of a specified user in a user group, and synthesizing the interest of these similar users in a certain multimedia data to form a system prediction of the interest of the specified user in the media data. Multimedia data-based collaborative filtering refers to finding similar multimedia data B based on the user's preference for multimedia data A, and predicting the user's interest in multimedia data B based on the user's interest in multimedia data A.

[0051] It is understood that the sample objects and target objects in this application may refer to users. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of user information such as user media behavior data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, a computer device can only obtain user information such as user media behavior data when it obtains the user's authorization information for the above information.

[0052] For example, a computer device can display a permission prompt interface in the multimedia interface of a multimedia platform. The permission prompt interface is used to prompt the user that the user's media behavior data is currently being collected. After the user issues a confirmation operation on the permission prompt interface, the step of obtaining the user's media behavior data is started, otherwise it ends.

[0053] Further, see Figure 3 , is a flowchart of a multimedia data processing method provided by an embodiment of the present application. Figure 3 As shown, this method can be Figure 1 The terminal can also be executed by Figure 1 It can also be executed by the server in Figure 1 The terminal and server in the embodiment are executed together. In this application, the devices used to execute the method can be collectively referred to as computer devices. The multimedia data processing method may include the following steps S101 to S105:

[0054] S101. A computer device obtains media attribute information corresponding to M multimedia data, and media behavior data of N target objects regarding the M multimedia data within a first time period; one target object and one multimedia data together correspond to one media behavior data.

[0055] In the present application, a computer device can obtain media data information corresponding to M multimedia data from a multimedia platform, as well as media behavior data of N target objects with respect to the M multimedia data within a first time period. The M multimedia data can refer to all multimedia data in the multimedia platform, or the M multimedia data can be obtained by obtaining one multimedia data from each of M media categories in the multimedia platform. It is understood that the first time period can refer to the past week, the past 10 days, etc. The target objects can refer to registered users in the multimedia platform, or the target objects can refer to users with mutual relationships among registered users in the multimedia platform. The relationships here can include friend relationships, family and friends relationships, colleague relationships, etc. Friend relationships can refer to relationships between users who have shared interests in and collected the same type of multimedia data, or friend relationships can refer to relationships between users who belong to the same communication group. In particular, when a target object does not perform any operations on a certain multimedia data within the first time period, the missing data can be used as the target object's media behavior data with respect to the multimedia data.

[0056] S102: The computer device determines the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, and j is a positive integer less than or equal to M.

[0057] In this application, the above relevant information P ij Used to reflect the target object N i The corresponding M media behavior data and multimedia data M j The correlation between the media attribute information of the target object N to a certain extent can reflect the i For multimedia data M j For example, the greater the relevance, the more likely the target object N i For multimedia data M j The preference for performing media operations (such as clicking and playing) is higher; the lower the correlation, the higher the preference for performing media operations (such as clicking and playing); ... i For multimedia data M j The preference for performing media operations is low. Since the media attribute information of different multimedia data is different, the media behavior data of different target objects for the same multimedia data is different, and the media behavior preferences of different target objects for the same multimedia data are inconsistent. Therefore, the computer device can determine the target object N iThe corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , and execute this step repeatedly until the relevant information of each target object on M multimedia data is obtained.

[0058] Optionally, the computer device may determine the target object N in any one of the following two ways or a combination of the two ways: i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij :

[0059] Method 1: The relevant information P ij To reflect the multimedia data M j The media attribute information of the target object N i The covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij The computer device can obtain the multimedia data M j The expected value of the media attribute information is used as the attribute expected value to obtain the target object N i The expected value of the corresponding M media behavior data is used as the behavior expected value to obtain the multimedia data M j The media attribute information (media attribute information of multiple dimensions) of the target object N is obtained by calculating the difference between the media attribute information (media attribute information of multiple dimensions) and the attribute expected value, obtaining the first difference, and obtaining the target object N i The difference between the corresponding M media behavior data and the behavior expectation value is used to obtain the second difference. Further, the product of the first difference and the second difference is used to obtain the first value, and the covariance matrix Σ is generated according to the expectation of the first value. ij , that is, the covariance matrix includes the expectation of the first value, and the covariance matrix is ​​determined as the relevant information P ij Here the relevant information P ij To reflect the multimedia data M j The media attribute information of the target object N i The covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij , that is, the covariance matrix Σ ij The larger the value in, the more likely the target object N is. i The corresponding M media behavior data and multimedia data M j The higher the correlation between the media attribute information, the higher the covariance matrix Σ ij The smaller the value in, the more likely the target object N is. i The corresponding M media behavior data and multimedia data M j The lower the correlation between the media attribute information.

[0060] Method 2: The computer device can be based on the target object N i The corresponding M media behavior data statistical target objects N i Regarding the behavior values ​​of M multimedia data, according to the target object N i Determine the target object N based on the behavior values ​​of M multimedia data respectively i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij For example, the multimedia data is audio data, and the media behavior data of the target object with respect to multimedia data 1 is played 3 times, clicked 5 times, and exited 2 times. The computer device can set a behavior weight for each media behavior data, such as the behavior weights corresponding to play, click, and exit are the first behavior weight, the second behavior weight, and the third behavior weight, respectively. Among them, the first behavior weight is greater than the second behavior weight, and the second behavior weight is greater than the third behavior weight. The first behavior weight and the second behavior weight can be greater than 0, and the third behavior weight can be less than 0. The target object N is assigned a behavior weight based on the behavior weight corresponding to each behavior data. i About multimedia data M j The media behavior data is weighted to obtain the target object N i About multimedia data M j The behavior value of the target object is obtained by looping, and the behavior values ​​corresponding to the M multimedia data of the target object are obtained. The cumulative behavior values ​​corresponding to the M multimedia data of the target object are obtained, and the sum of the behavior values ​​is obtained. i About multimedia data M j The behavior ratio between the behavior value and the total behavior value determines the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij The larger the behavior ratio, the more likely the target object N is. i The corresponding M media behavior data and multimedia data M j The higher the correlation between the media attribute information, the lower the behavior ratio, indicating that the target object N i The corresponding M media behavior data and multimedia data M j The lower the correlation between the media attribute information of the target object N i About multimedia data M j The media behavior data of M multimedia data is missing data. In this case, the media attribute information and multimedia data M can be obtained. j The relevant information corresponding to the multimedia data matching the media attribute information is used as the target object N iThe corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij .

[0061] S103: If the relevant information corresponding to the M multimedia data is determined, the computer device generates the interaction weights of the N target objects with respect to the M multimedia data according to the P relevant information, the M media attribute information, and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the relevant information P is the interaction weight of the N target objects with respect to the M multimedia data. ij Belong to the P related information.

[0062] In this application, if the relevant information corresponding to the M multimedia data is determined, the computer device can generate the interaction weights of the N target objects with respect to the M multimedia data according to the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects; the target object N i About multimedia data M j The interaction weight can be used to reflect the target object N i and multimedia data M j The interaction parameters between the target objects (such as play time, click times), and one or more of the preferences of the multimedia data. The preference of the multimedia data can refer to the popularity of the multimedia data, that is, the popularity of the multimedia data is determined based on the media behavior data such as clicks and plays of N target objects. For example, if the target object N i About multimedia data M j The click rate is relatively low, the target object N i About multimedia data M j The interaction weight is set to a small value; if the target object N i About multimedia data M j The click rate is relatively high, and the target object N i About multimedia data M j The interaction weight is set to a larger value, which is conducive to making personalized recommendations to the target object based on the target object's media behavior preferences and improving the accuracy of multimedia data recommendations.

[0063] It is understandable that the computer device may generate the interaction weights of the N target objects with respect to the M multimedia data in any one of the following two ways or a combination of the two ways:

[0064] Method 1: The computer device can i About multimedia data M j Related information about the target object N i About multimedia data Mj The product of media behavior data is determined as the target object N i About multimedia data M j In particular, if the target object N i About multimedia data M j If the media behavior data is missing data, the computer device can obtain the media attribute information and M j The media attribute information matches the matching multimedia data, the target object N i The media behavior data of the matching multimedia data is non-missing data, and the target object N i Regarding the interaction weight of matching multimedia data, the target object N is determined i About multimedia data M j The interaction weight of .

[0065] Method 2: The relevant information P ij For reflecting the multimedia data M j The media attribute information of the target object N i The covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij The computer device may generate a signal reflecting the target object N i The corresponding media behavior feature matrix X of M media behavior data i , and for reflecting the multimedia data M j The media attribute feature matrix Y of the media attribute information j , according to the media behavior feature matrix X i , the covariance matrix Σ ij and the media attribute feature matrix Y j The product between them generates the target object N i Regarding the multimedia data M j The interaction weight equation is about the target object N i About multimedia data M j Furthermore, the computer device can generate the target object N according to the covariance matrix corresponding to the P related information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects. i Regarding the multimedia data M jBased on the interaction weight constraints of the N target objects, the interaction weights of the N target objects with respect to the M multimedia data are determined according to the interaction weight equations corresponding to the N target objects and the interest weight constraints corresponding to the N target objects. In other words, based on the interest weight constraints corresponding to the N target objects, the interaction weight equations corresponding to the N target objects are solved to obtain the interaction weights of the N target objects with respect to the M multimedia data. The interaction weights of the target objects with respect to each multimedia data are determined through the media attribute information, the media behavior data, and the correlation between the media attribute information and the media behavior data, which is conducive to improving the accuracy of obtaining the interaction weights.

[0066] It is understandable that in the above-mentioned second method, the covariance matrix corresponding to the P relevant information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects are generated. i Regarding the multimedia data M j The interaction weight constraint condition includes: determining the covariance matrix corresponding to the P relevant information and the target object N i The M covariance matrices associated with the multimedia data M j The N covariance matrices associated with the target object N i The associated M covariance matrices include: target object N i Regarding the covariance matrix of multimedia data M1, target object N i Regarding the covariance matrix of multimedia data M2, ..., target object N i About multimedia data M M The covariance matrix of the multimedia data M j The associated N covariance matrices include: target object N1 with respect to multimedia data M j The covariance matrix of the target object N2 about the multimedia data M j The covariance matrix of the target object N N About multimedia data M j Further, according to the media behavior feature matrix X i , the media attribute feature matrix corresponding to the M media attribute information, and the target object N i The associated M covariance matrices determine the first constraint condition; the first constraint condition is used to constrain the target object N i The cumulative sum of the interaction weights of the M multimedia data, where the target object N i The M multimedia data respectively include: target object N iRegarding the interaction weight of multimedia data M1, the target object N i Regarding the interaction weights of multimedia data M2, ..., target objects N i About multimedia data M M Then, according to the media attribute feature matrix Y j , the N media behavior feature matrices corresponding to the N target objects, and the multimedia data M j The associated N covariance matrices determine the second constraint condition; the second constraint condition is used to constrain the N target objects to be respectively related to the multimedia data M j The cumulative sum of the interaction weights of the N target objects, where the N target objects are respectively about the multimedia data M j The interaction weights include: target object N1 about multimedia data M j The interaction weight of the target object N2 on the multimedia data M j The interaction weight of , ..., target object N N About multimedia data M j The computer device may generate a third constraint condition, and determine the first constraint condition, the second constraint condition, and the third constraint condition as the target object N i Regarding the multimedia data M j interaction weight constraint; the third constraint is used to constrain the cumulative sum of the interaction weights of the N target objects with respect to the M multimedia data.

[0067] For example, the following formula (1) represents the interaction weight equation, which is used to reflect: target object N i Regarding the multimedia data M j The interaction weight W ij Equal to the media behavior feature matrix X i , the covariance matrix Σ ij and the media attribute feature matrix Y j The following formulas (2-1), (2-2), and (2-3) represent the target object N i Regarding the multimedia data M j The interaction weight constraints, formula (2-1), formula (2-2), and formula (2-3) represent the first constraint, the second constraint, and the third constraint, respectively. Formula (2-1), formula (2-2), and formula (2-3) are called a system of equations. The st in the system of equations represents the system of equations used to constrain the interaction weight W in formula (1). ij The value of . Among them, k in this equation group x To limit the constant, it can refer to an empirical value.

[0068] Wij =X i ∑ ij Y j (1)

[0069]

[0070] S104. The computer device collaboratively processes the M media attribute information and the M media behavior data corresponding to the N target objects based on the interaction weights of the N target objects with respect to the M multimedia data, and obtains the target interest levels of the N target objects with respect to the M multimedia data.

[0071] In the present application, the computer device can collaboratively process the M media attribute information and the M media behavior data corresponding to the N target objects based on the interaction weights of the N target objects with respect to the M multimedia data, and obtain the target interest of the N target objects with respect to the M multimedia data. In other words, by introducing the interaction weights, the differences between the media behavior data of the target objects with respect to different multimedia data are expanded, and the accuracy of obtaining the target interest is improved. Furthermore, by recommending multimedia data to the N target objects based on their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain target object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data.

[0072] It is understood that the above-mentioned collaborative processing may refer to at least one of multimedia data-based collaborative processing, user-based collaborative processing, and multimedia data and user-based collaborative processing. User-based collaborative processing is a method of analyzing the media behavior data of a target object, and performing the following operations on N-1 target objects (excluding target object N in N target objects). i Find a target object N i Similar (interested) target objects, the system forms the target object N by combining the media behavior data of these similar target objects on a certain multimedia data as well as their own media behavior data and media attribute information. iA method for predicting interest in this media data. Collaborative filtering based on multimedia data refers to finding similar multimedia data B based on the target object's media attribute information for multimedia data A, and predicting the target object's interest in multimedia data B based on the target object's media behavior data regarding multimedia data A and multimedia data B. Collaborative processing based on multimedia data and users is obtained based on the above-mentioned collaborative processing based on multimedia data and collaborative processing based on users. For example, two target objects with similar media behavior data are regarded as similar users, and similar users include target object 1 and target object 2. Two similar multimedia data are determined based on media attribute information, and similar multimedia data include multimedia data A and multimedia data B. The interest of target object 1 in multimedia data A and multimedia data B, respectively, can be determined based on target object 1's media behavior data with respect to multimedia data A and multimedia data B, target object 2's media behavior data with respect to multimedia data A and multimedia data B, and the media attribute information of multimedia data A and multimedia data B. Similarly, the interest of target object 2 in multimedia data A and multimedia data B can be determined based on the media behavior data of target object 1 regarding multimedia data A and multimedia data B, the media behavior data of target object 2 regarding multimedia data A and multimedia data B, and the media attribute information of multimedia data A and multimedia data B.

[0073] It can be understood that the above-mentioned collaborative processing of the M media attribute information and the M media behavior data corresponding to the N target objects respectively according to the interaction weights of the N target objects with respect to the M multimedia data to obtain the target interest of the N target objects respectively with respect to the M multimedia data includes: associating and identifying the M media attribute information and the M media behavior data corresponding to the N target objects respectively to obtain the initial interest of the N target objects respectively with respect to the M multimedia data; adjusting the initial interest of the N target objects respectively with respect to the M multimedia data according to the interaction weights of the N target objects respectively with respect to the M multimedia data to obtain the target interest of the N target objects respectively with respect to the M multimedia data.

[0074] The computer device can perform association recognition on the M media attribute information and the M media behavior data corresponding to the N target objects, and obtain the initial interest of the N target objects in the M multimedia data. Association recognition here refers to identifying the association relationship (such as similarity) between the media attribute information of the multimedia data and the association relationship between the media behavior data of the target objects; the target object N i Regarding the multimedia data M j The initial interest degree is used to reflect the target object N i Regarding the multimedia data M jThe rough interest of the target object N i Regarding the multimedia data M j The initial interest of the N target objects is missing and has a relatively low accuracy. Therefore, the computer device can adjust the initial interest of the N target objects with respect to the M multimedia data according to the interaction weights of the N target objects with respect to the M multimedia data, and obtain the target interest of the N target objects with respect to the M multimedia data. i Regarding the multimedia data M j The interaction weight is used to reflect the target object N i Regarding the media behavior data of M multimedia data and the multimedia data M j In other words, the target object N i Regarding the multimedia data M j The target interest degree is determined based on the following three factors: 1. The correlation between the media attribute information of M multimedia data (such as similarity); 2. The correlation between the media behavior data of N target objects; 3. The correlation between the target object N i The media behavior data of M multimedia data and the multimedia data M j It can be seen that by mining the target object N i The media behavior data of M multimedia data and the multimedia data M j The interactivity between media attribute information is conducive to obtaining the accuracy of target interest, that is, the target object N i Regarding the multimedia data M j The target interest degree is used to reflect the target object N i The precise interest level of the multimedia data Mj.

[0075] It is understandable that the above-mentioned association and identification of the M media attribute information and the M media behavior data corresponding to the N target objects, respectively, to obtain the initial interest of the N target objects in the M multimedia data, includes: the computer device calls the feature division layer of the first media recognition model to divide the M media attribute information to obtain first media attribute information with dense features and second media attribute information with sparse features; calls the feature division layer to divide the M media behavior data corresponding to the N target objects, respectively, to obtain first media behavior data with dense features and second media behavior data with sparse features. Then, the deep learning layer of the first media recognition model can be called to extract key media attribute information from the second media attribute information and key media behavior data from the second media behavior data; calls the first interest recognition layer of the first media recognition model to associate and identify the key media behavior data, the key media attribute information, the first media attribute information and the first media behavior data to obtain the initial interest of the N target objects in the M multimedia data.

[0076] like Figure 4 As shown, the computer device can call the feature partitioning layer of the first media recognition model to partition the M media attribute information into first media attribute information with dense features and second media attribute information with sparse features. The first media attribute information with dense features refers to the number of invalid media attribute feature values ​​used to reflect the first media attribute information being less than a first threshold value; the second media attribute information with sparse features refers to the number of invalid media attribute feature values ​​used to reflect the second media attribute information being greater than or equal to the first threshold value. Invalid values ​​may refer to feature values ​​without practical meaning, such as 0. For example, the computer device can call the feature partitioning layer of the first media recognition model to count the number of invalid values ​​in the media feature values ​​corresponding to each multimedia data item, classify the media attribute information of the multimedia data with a number less than the first threshold value as first media attribute information, and classify the media attribute information of the multimedia data with a number greater than or equal to the first threshold value as second media attribute information. Similarly, the first media behavior data with dense features may refer to the number of invalid values ​​in the media behavior feature values ​​corresponding to the first media behavior data being less than a second threshold value, and the second media behavior data with sparse features may refer to the number of invalid values ​​in the media behavior feature values ​​corresponding to the second media behavior data being greater than or equal to the second threshold value. For example, the computer device can call the feature division layer to count the target object N i Regarding the multimedia data M jThe number of media behavior feature values ​​corresponding to the media behavior data is an invalid value, and the media behavior data whose corresponding number is less than the second number threshold is classified into the first media behavior data, and the media behavior data whose corresponding number is greater than or equal to the second number threshold is classified into the second media behavior data.

[0077] Furthermore, the computer device can call the deep learning layer of the first media recognition model to extract key media attribute information from the second media attribute information and extract key media behavior data from the second media behavior data. Key media attribute information refers to the attribute information of primary significance in the second media attribute information, and the dimension of key media attribute information is smaller than the dimension of the second media attribute information; similarly, key media behavior data refers to the behavioral data of primary significance in the second media behavior data, and the dimension of key media behavior is smaller than the dimension of the second media behavior data. Furthermore, the computer device can call the first interest recognition layer of the first media recognition model to associate and identify the key media behavior data, the key media attribute information, the first media attribute information and the first media behavior data, and obtain the initial interest levels of the N target objects respectively with respect to the M multimedia data. By dividing the media behavior data and media attribute information, it is beneficial to save computing resources and improve the accuracy of obtaining the initial interest level.

[0078] It can be understood that the above-mentioned calling of the first interest recognition layer of the first media recognition model, associating and identifying the key media behavior data, the key media attribute information, the first media attribute information and the first media behavior data, and obtaining the initial interest levels of the N target objects with respect to the M multimedia data, includes: the computer device can call the first interest recognition layer, splice the key media behavior data with the first media behavior data, and obtain the spliced ​​media behavior data, call the first interest recognition layer, splice the key media attribute information with the first media attribute information, and obtain the spliced ​​media attribute information, and associate and identify the spliced ​​media behavior data and the spliced ​​media attribute information, and obtain the candidate interest levels of the N target objects with respect to the M multimedia data. Then, the candidate interest degrees of the N target objects with respect to the M multimedia data can be expanded to obtain the initial interest degrees of the N target objects with respect to the M multimedia data; for example, the candidate interest degrees of the N target objects with respect to the M multimedia data can be multiplied by a constant C to obtain the initial interest degrees of the N target objects with respect to the M multimedia data, which is conducive to expanding the distinctiveness of the initial interest degrees of the N target objects with respect to the M multimedia data.

[0079] It is understandable that the above adjustment of the initial interest of the N target objects with respect to the M multimedia data according to the interaction weights of the N target objects with respect to the M multimedia data, and obtaining the target interest of the N target objects with respect to the M multimedia data, includes: the computer device can call the average processing layer of the second media recognition model to average the initial interest of the N target objects with respect to the M multimedia data, and obtain the averaged initial interest, and the averaged initial interest includes the overall average interest corresponding to the initial interest of the N target objects with respect to the M multimedia data, the target object N i The average interest of the object of the M multimedia data, the average interest of the N target objects in the multimedia data M j One or more of the average media interest of the target object N i The average object interest of the M multimedia data is calculated based on the target object N i The initial interest of the M multimedia data is calculated, and the N target objects are related to the multimedia data M j The average media interest of the N target objects can be based on the multimedia data M j The initial interest degree is calculated. Further, the computer device can call the standard processing layer of the second media recognition model, and generate an object standard feature matrix and a media standard feature matrix according to the initial interest degrees of the N target objects with respect to the M multimedia data, and the interaction weights of the N target objects with respect to the M multimedia data. The object standard feature matrix is ​​used to reflect the standard media behavior characteristics of the N target objects with respect to the M multimedia data. The object standard feature matrix can also be called an object scoring matrix. The media standard feature matrix is ​​used to reflect the standard media attribute characteristics of the M multimedia data. The media standard feature matrix can also be called a media scoring matrix. Then, the second interest recognition layer of the second media recognition model is called to determine the target interest degrees of the N target objects with respect to the M multimedia data according to the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the initial interest degrees after averaging. By introducing the interaction weights, the accuracy of obtaining the target interest degrees is improved, and further, the accuracy of multimedia data recommendation is improved.

[0080] It is understandable that the above-mentioned calling the average processing layer of the second media recognition model to average the initial interest levels of the N target objects with respect to the M multimedia data to obtain the averaged initial interest levels includes: the computer device calling the average processing layer of the second media recognition model to determine the initial interest levels of the N target objects with respect to the multimedia data M from the initial interest levels of the N target objects with respect to the M multimedia data. j The initial interest of the target object N i The initial interest of the M multimedia data is respectively related to the N target objects. j The initial interest of the N target objects is averaged to obtain the interest of the multimedia data M j The average media interest is used to reflect the interest of N target objects in multimedia data M j Then, for the target object N i The initial interest of the M multimedia data is averaged to obtain the target object N i The average object interest of the M multimedia data is used to reflect the target object N. i The average rating of the M multimedia data. Furthermore, the initial interest of the N target objects with respect to the M multimedia data is averaged to obtain an overall average interest; the overall average interest is used to reflect the average rating of the N target objects for the M multimedia data. Then, the overall average interest, the object average interest corresponding to the N target objects, and the media average interest corresponding to the M multimedia data are determined as the initial interest after averaging. By calculating the media average interest, the object average interest, and the overall average interest, it is effectively avoided that a certain target object scores a certain multimedia data too high or too low, resulting in a relatively low accuracy in obtaining the target interest.

[0081] It is understandable that the above-mentioned calling of the second interest recognition layer of the second media recognition model determines the target interest of the N target objects with respect to the M multimedia data according to the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the initial interest after the averaging process, including: the computer device can call the second interest recognition layer, determine the target interest of the N target objects with respect to the M multimedia data from the object standard feature matrix, i Regarding the standard media behavior characteristics U of the M multimedia data i From the media standard feature matrix, determine the multimedia data M j Standard media attribute characteristics V jThen, establish the eigenvalue equation group corresponding to the interaction weight matrix, solve the eigenvalue equation group, obtain the eigenvalue vector of the interaction weight matrix, average the eigenvalues ​​in the eigenvalue vector, and obtain the average eigenvalue. The interaction weight matrix includes the interaction weights of the N target objects with respect to the M multimedia data. Then, obtain the standard media behavior feature U i The transpose of the target object N i Regarding the multimedia data M j The interest weight of the standard media attribute feature V j The product between them is used to obtain the first interest degree; the average characteristic root and the multimedia data M are obtained. j The corresponding media average interest, the target object N i The product of the corresponding object average interest is used to obtain the second interest. The target object N is determined based on the overall average interest, the first interest and the second interest. i Regarding the multimedia data M j Target interest.

[0082] For example, in the following formula (3), p(U i , V j , W ij ) represents the target object N i Regarding the multimedia data M j The target interest level, μ represents the overall average interest level, represents the mean characteristic root. Indicates the standard media behavior characteristics U i The transpose of the matrix formed. S in formula (4-1) and formula (4-2) ij refers to the target object N i About multimedia data M j The initial interest degree, k u 、k v They represent the user's restriction constant and the media's restriction constant, respectively, both of which are empirical values. Represents the target object N i Regarding the transpose of the matrix composed of the interaction weights of M multimedia data, W j Represents N target objects respectively about multimedia data M j The matrix composed of interaction weights. i Indicates the target object N i Regarding the average interest of the M multimedia data objects, b j Represents N target objects about the multimedia data M j Average media interest. They represent the first degree of interest and the second degree of interest respectively. The st in the equation group formed by formula (4-1) and formula (4-2) indicates that formula (3) is constrained by formula (4-1) and formula (4-2).

[0083]

[0084] For example, the eigenvalue vector of the interaction weight matrix can be expressed by the following formula (5), and the average eigenvalue of the eigenvalue vector can be expressed by the following formula (6):

[0085] Δ={λ l |l≤min(N,M)} (5)

[0086]

[0087] Among them, λ in formula (5) l is the characteristic root of the interaction weight matrix, Δ is the characteristic root vector of the interaction weight matrix, and the formula (6) represents the mean characteristic root.

[0088] It is understandable that the target object N is determined based on the overall average interest, the first interest and the second interest. i Regarding the multimedia data M j The target interest degree includes: as shown in the above formula (3), the computer device can obtain the cumulative sum of the overall average interest degree, the first interest degree and the second interest degree to obtain the target object N i Regarding the multimedia data M j The selected interest degree is obtained to obtain the initial interest degree matrix for reflecting the initial interest degrees of the N target objects respectively with respect to the M multimedia data. i About the media data M j If the initial interest degree is missing data, the target object N is used. i Regarding the multimedia data M j The initial interest matrix is ​​filled with the selected interest degrees to obtain the target interest matrix. i About the media data M j If the initial interest degree is non-missing data, the target object N is used. i Regarding the multimedia data M j The selected interest degree replaces the target object N in the initial interest matrix i Regarding the multimedia data M j The initial interest degree of the target object N in the target interest matrix is ​​obtained. i Regarding the multimedia data M jThe candidate interest degree is used as the target object N i Regarding the multimedia data M j By filling or replacing the initial interest matrix according to the selected interest, the completeness and accuracy of the target interest of the target object with respect to the multimedia data can be improved.

[0089] S105 : The computer device recommends multimedia data to the N target objects respectively according to their target interest levels with respect to the M multimedia data.

[0090] In this application, the computer can i Regarding the target interest of M multimedia data, the target multimedia data whose target interest is greater than the interest threshold is selected from the M multimedia data, and the target multimedia data is sent to the target object N. i Recommending target multimedia data avoids recommending all multimedia data to the target subject, improving the accuracy of multimedia data recommendations and saving computing resources. Furthermore, by introducing interaction weights, the differentiation between target subjects' preferences for different multimedia data can be increased, resolving the issue of poor multimedia data recommendation effectiveness caused by a relatively uniform distribution of click-through rates for target subjects' multimedia data. For example, if target subject 1 has a click-through rate of 0.6 for both multimedia data A and multimedia data B, and multimedia data A is more popular than multimedia data B, then target subject 1's interaction weight for multimedia data A is greater than its interaction weight for multimedia data B. Consequently, target subject 1's target interest in multimedia data A is higher than its target interest in multimedia data B, effectively increasing the differentiation between target subjects' preferences for multimedia data A and multimedia data B. Furthermore, by introducing interaction weights, the target subject's preference (i.e., interest) for multimedia data with low click-through rates or relatively low play rates can be reduced, eliminating the need to recommend such multimedia data to the target subject, saving computing resources. In other words, this solution is suitable for multimedia data recommendation scenarios where the scale of multimedia data is relatively large and the target objects' media behavior data such as click rate and play rate of multimedia data are relatively evenly distributed.

[0091] In this application, the computer device can obtain the media attribute information corresponding to M multimedia data, and the media behavior data of N target objects respectively about the M multimedia data in the first time period, and determine the target object N. i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , the relevant information P ij Used to reflect the target user N iMedia behavior data and multimedia data M j Further, if the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated based on the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects. i About multimedia data M j The interaction weight is mainly used to reflect the target object N i About multimedia data M j interaction parameters (such as clicks, playback frequency). Then, according to the interaction weights of the N target objects respectively with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest of the N target objects respectively with respect to the M multimedia data. It can be seen that in the process of obtaining the target interest of the target object with respect to multimedia data, by introducing the interaction weight, it is beneficial to expand the distinctiveness between the target object's media behavior data with respect to different multimedia data, and then expand the distinctiveness of the target object's preference with respect to different multimedia data, which can improve the accuracy of obtaining the target interest. By recommending multimedia data to the N target objects respectively based on their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data.

[0092] Further, see Figure 5 , is a flowchart of a multimedia data processing method provided by an embodiment of the present application. Figure 5 As shown, this method can be Figure 1 The terminal can also be executed by Figure 1 It can also be executed by the server in Figure 1 The terminal and server in the embodiment are executed together. In this application, the devices used to execute the method can be collectively referred to as computer devices. The multimedia data processing method may include the following steps S201 to S208:

[0093] S201. The computer device obtains media attribute information corresponding to H multimedia data, as well as media behavior data of L sample objects regarding the H multimedia data within a second time period, and historical media behavior data of the L sample objects regarding the H multimedia data within a third time period; the third time period is a time period after the second time period, and the first time period is a time period after the third time period.

[0094] In the present application, a computer device can train a first initial media recognition model based on the media behavior data of the sample objects regarding multimedia data and the media attribute information of the multimedia data to obtain a first media recognition model. Specifically, the computer device can obtain the media attribute information corresponding to H multimedia data respectively from the multimedia platform, as well as the media behavior data of L sample objects regarding the H multimedia data respectively during a second time period. The H multimedia data here include multimedia data of multiple media categories, the second time period can be the time period before the first time period, and the L sample objects can refer to all registered users or some registered users on the multimedia platform.

[0095] S202, the computer device calculates the sample object L e Regarding the multimedia data H r The historical media behavior data of the sample object L e Regarding the multimedia data H r The annotation interest of the sample object L e Belongs to the L sample objects, e is a positive integer less than or equal to L, the multimedia data H r Belonging to the H multimedia data, r is a positive integer less than or equal to H.

[0096] In this application, the computer device can use the sample object Le to determine the multimedia data H r The historical media behavior data of the sample object L e Regarding the multimedia data H r The annotated interest of each sample object with respect to each of the H multimedia data is obtained in this cycle. In other words, the annotated interest of each sample object with respect to the multimedia data is determined by the media behavior data of the sample object with respect to the multimedia data, without the need for manual annotation, thereby improving the accuracy and efficiency of obtaining the annotated interest.

[0097] For example, if the sample object Le is related to the multimedia data H r The media behavior data reflects the sample object Le's response to the multimedia data H r Execute the click operation and the play operation, and the first marked interest level is used as the sample object L e Regarding the multimedia data H r The annotation interest of the sample object Le is related to the multimedia data H r The media behavior data reflects the sample object Le's response to the multimedia data H r If the click operation is executed and the play operation is not executed, the second marked interest level is used as the sample object L e Regarding the multimedia data H rThe first annotation interest degree is greater than the second annotation interest degree.

[0098] S203. If the annotated interest levels corresponding to the L sample objects are determined, the computer device trains the first initial media recognition model based on the annotated interest levels corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain a first media recognition model.

[0099] In the present application, the first initial media recognition model may refer to a model used to identify a user's initial interest in multimedia data. The accuracy of the first initial media recognition model is relatively low. Therefore, once the labeled interest corresponding to the L sample objects is determined, the computer device trains the first initial media recognition model based on the labeled interest corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain a first media recognition model. By training the first initial media recognition model, the accuracy of the first media recognition model in identifying interest can be improved, that is, the first media recognition model can refer to a model with relatively high interest recognition accuracy.

[0100] It is understandable that the above step S203 includes: Figure 6 As shown, the computer device can divide the L sample objects according to the division ratio to obtain K training objects and G test objects; L is the sum of K and G, and the division ratio can refer to the ratio between the number of training objects and the number of test objects, such as 8:2. Then, the first initial media recognition model is called to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training objects, to obtain the first predicted interest levels of the K training objects with respect to the H multimedia data; if the first predicted interest levels of the K training objects with respect to the H multimedia data are not much different from the labeled interest levels of the K training objects with respect to the H multimedia data, it indicates that the interest recognition accuracy of the first initial recognition model is relatively high; if the first predicted interest levels of the K training objects with respect to the H multimedia data are significantly different from the labeled interest levels of the K training objects with respect to the H multimedia data, it indicates that the interest recognition accuracy of the first initial recognition model is relatively low. Therefore, the computer device can adjust the model parameters of the first initial media recognition model according to the labeled interest levels corresponding to the K training objects and the first predicted interest levels corresponding to the K training objects to obtain the adjusted first initial media recognition model.

[0101] Furthermore, if the adjusted first initial media recognition model is not in a converged state, indicating that the adjusted first initial media recognition model still has a relatively low accuracy in identifying the interest levels of the K training subjects, the adjusted first initial media recognition model continues to be trained until the adjusted first initial media recognition model is in a converged state. If the adjusted first initial media recognition model is in a converged state, indicating that the adjusted first initial media recognition model still has a relatively high accuracy in identifying the interest levels of the K training subjects, the computer device may use the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects to test the interest level identification performance of the adjusted first initial media recognition model and obtain a test result. If the test result indicates that the adjusted first initial media recognition model has failed the interest recognition performance test, indicating that the adjusted first initial media recognition model's interest recognition accuracy for the G test subjects is still relatively low, then the computer device can continue to iteratively train the adjusted first initial media recognition model based on the media behavior data of the K training subjects and the H media attribute information, or reacquire the test subjects, their media behavior data, and the H media attribute information to iteratively train the adjusted first initial media recognition model. The adjusted first initial media recognition model is then tested again until the test result indicates that the adjusted first initial media recognition model has passed the interest recognition performance test. The adjusted first initial media recognition model is then determined as the first media recognition model. Training the first initial recognition model based on the media behavior data and media attribute information corresponding to the test subjects and training subjects is beneficial for improving the interest recognition accuracy of the first media recognition model and enhancing the generalization capability of the first media recognition model.

[0102] It is understandable that the computer device calls the first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training objects, and obtains the first predicted interest level of the K training objects with respect to the H multimedia data. The implementation process can refer to the above-mentioned association recognition of the M media attribute information and the M media behavior data corresponding to the N target objects, and obtain the initial interest level of the N target objects with respect to the M multimedia data. The repetitive parts will not be repeated.

[0103] It is understood that if the adjusted first initial media recognition model is in a convergence state, the interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects, to obtain a test result, including: the computer device can call the adjusted first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training subjects, to obtain second predicted interest levels of the K training subjects with respect to the H multimedia data. If the second predicted interest levels of the K training subjects with respect to the H multimedia data differ significantly from the annotated interest levels of the K training subjects with respect to the H multimedia data, it indicates that the interest recognition accuracy of the adjusted first initial recognition model is relatively low. If the second predicted interest levels of the K training subjects with respect to the H multimedia data differ significantly from the annotated interest levels of the K training subjects with respect to the H multimedia data, it indicates that the interest recognition accuracy of the adjusted first initial recognition model is relatively high. Therefore, the computer device can measure the interest recognition accuracy of the adjusted first initial media recognition model based on the second predicted interest of the K training objects regarding the H multimedia data and the labeled interest of the K training objects regarding the H multimedia data.

[0104] For example, a computer device obtains the error function of the adjusted first initial recognition model, which is a function that reflects the relationship between the labeled interest, the predicted interest, and the model parameters. The partial derivative of the error function with respect to the model parameters can be calculated to obtain a partial derivative error function. The labeled interest corresponding to the K training objects and the second predicted interest corresponding to the K training objects are substituted into the partial derivative function to determine the gradient of the adjusted first initial media recognition model. The gradient of the adjusted first initial media recognition model is used to reflect the descent distance of the interest recognition error of the adjusted first initial media recognition model. If the gradient is smaller, it indicates that the descent distance is smaller, that is, the interest recognition error of the adjusted first initial media recognition model is closer to the minimum value; if the gradient is larger, it indicates that the descent distance is larger, that is, the interest recognition error of the adjusted first initial media recognition model is more different from the minimum value. Therefore, if the gradient of the adjusted first initial media recognition model is less than the gradient threshold, the adjusted first initial media recognition model is determined to be in a convergence state. The interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects to obtain a test result. The first initial media recognition model is trained using a gradient descent method to obtain a first media recognition model, thereby improving the interest recognition accuracy of the first media recognition model.

[0105] It is understood that if the adjusted first initial media recognition model is in a convergence state, the interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects, to obtain a test result, including: if the adjusted first initial media recognition model is in a convergence state, calling the adjusted first initial media recognition model to perform association prediction on the media behavior data corresponding to the G test subjects and the media attribute information of the H multimedia data, to obtain the test interest levels of the G test subjects with respect to the H multimedia data. If the difference between the test interest levels of the G test subjects with respect to the H multimedia data and the annotated interest levels of the G test subjects with respect to the H multimedia data is relatively large, it indicates that the interest recognition accuracy of the adjusted first initial recognition model for the K training subjects is relatively high, but the interest recognition accuracy for the G test subjects is relatively low, and the generalization ability of the adjusted first initial recognition model is relatively poor. If the difference between the second predicted interest levels of the K training subjects for the H multimedia data and the annotated interest levels of the K training subjects for the H multimedia data is relatively small, it indicates that the adjusted first initial recognition model has a relatively high accuracy in identifying the interest levels of the K training subjects and the G test subjects, and the adjusted first initial recognition model has a relatively strong generalization capability. Therefore, the computer device can determine the interest level recognition performance parameters of the adjusted first initial media recognition model based on the test interest levels corresponding to the G test subjects and the annotated interest levels corresponding to the G test subjects. The interest level recognition performance parameters include one or more of recall rate, precision rate, and AUC. Based on the interest level recognition performance parameters, a test result for the interest level recognition performance of the adjusted first initial media recognition model is generated. By testing the adjusted first initial media recognition model based on the media behavior data corresponding to the G test subjects and the media attribute information of the H multimedia data, the generalization capability of the first media recognition model is improved.

[0106] It is understandable that after obtaining the first media recognition model with relatively high interest recognition accuracy, the computer device can train the second initial media recognition model in the following manner to obtain the second media recognition model. Specifically, Figure 7 As shown, the computer device can call the first media recognition model to associate and identify the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the L sample objects, and obtain the initial interest of the L sample objects in relation to the H multimedia data. eThe corresponding H media behavior data and the multimedia data H r The related information between the media attribute information of the sample object L e The corresponding H media behavior data and the multimedia data H r The correlation of the media attribute information. Then, if the relevant information corresponding to the H multimedia data is determined, the interaction weights of the L sample objects with respect to the H multimedia data are generated according to the Q relevant information, the H media attribute information and the H media behavior data corresponding to the L sample objects; Q is the product of H and L; the second initial media recognition model is called, and the H media attribute information and the H media behavior data corresponding to the L sample objects are collaboratively predicted according to the interaction weights of the L sample objects with respect to the H multimedia data, so as to obtain the target interest of the L sample objects with respect to the H multimedia data. Further, the second initial media recognition model is trained according to the initial interest corresponding to the L sample objects and the target interest corresponding to the L sample objects to obtain the second media recognition model.

[0107] It is understandable that the computer device calls the second initial media recognition model, and according to the interaction weights of the L sample objects respectively with respect to the H multimedia data, collaboratively predicts the H media attribute information and the H media behavior data corresponding to the L sample objects, and obtains the target interest of the L sample objects respectively with respect to the H multimedia data. The realization process can refer to the above-mentioned computer device performing collaborative processing on the M media attribute information and the M media behavior data corresponding to the N target objects according to the interaction weights of the N target objects respectively with respect to the M multimedia data, and obtaining the target interest of the N target objects respectively with respect to the M multimedia data. The repeated parts will not be repeated.

[0108] It can be understood that the second initial media recognition model is trained according to the initial interest degrees corresponding to the L sample objects and the target interest degrees corresponding to the L sample objects to obtain the second media recognition model, including: the computer device can determine the interest degree recognition error of the second initial media recognition model according to the initial interest degrees corresponding to the L sample objects and the target interest degrees corresponding to the L sample objects; if the interest degree recognition error of the second initial media recognition model is not in a convergent state, the model parameters of the second initial media recognition model are adjusted according to the interest degree recognition error; and the second initial media recognition model after the model parameters are adjusted is determined as the second media recognition model.

[0109] For example, the computer device can obtain the initial interest degrees corresponding to the L sample objects respectively, and the mean square error between the target interest degrees corresponding to the L sample objects respectively, and determine the mean square error as the interest degree recognition error of the second initial media recognition model. If the interest degree recognition error of the second initial media recognition model is in a convergent state, and the interest degree recognition accuracy of the second initial media recognition is relatively high, then the second initial media recognition model is determined as the second media recognition model. If the interest degree recognition error of the second initial media recognition model is not in a convergent state, and the interest degree recognition accuracy of the second initial media recognition is relatively low, then the model parameters of the second initial media recognition model are adjusted according to the interest degree recognition error; and the second initial media recognition model after the model parameters are adjusted is determined as the second media recognition model.

[0110] S204. The computer device obtains media attribute information corresponding to M multimedia data, and media behavior data of N target objects regarding the M multimedia data within a first time period; one target object and one multimedia data together correspond to one media behavior data.

[0111] S205: The computer device determines the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, and j is a positive integer less than or equal to M.

[0112] S206. If the relevant information corresponding to the M multimedia data is determined, the computer device generates the interaction weights of the N target objects with respect to the M multimedia data according to the P relevant information, the M media attribute information, and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the relevant information P is the interaction weight of the N target objects with respect to the M multimedia data. ij Belong to the P related information.

[0113] S207. The computer device calls the first media recognition model, and collaboratively processes the M media attribute information and the M media behavior data corresponding to the N target objects according to the interaction weights of the N target objects with respect to the M multimedia data, to obtain the target interest levels of the N target objects with respect to the M multimedia data.

[0114] After obtaining the first media recognition model, the computer device can call the first media recognition model, and according to the interaction weights of the N target objects with respect to the M multimedia data, collaboratively process the M media attribute information and the M media behavior data corresponding to the N target objects, to obtain the target interest levels of the N target objects with respect to the M multimedia data, thereby improving the accuracy of obtaining the target interest levels.

[0115] S208: The computer device recommends multimedia data to the N target objects respectively according to their target interest levels with respect to the M multimedia data.

[0116] In this application, the computer device can obtain the media attribute information corresponding to M multimedia data, and the media behavior data of N target objects respectively about the M multimedia data in the first time period, and determine the target object N. i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , the relevant information P ij Used to reflect the target user N i Media behavior data and multimedia data M j Further, if the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated based on the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects. i About multimedia data M j The interaction weight is mainly used to reflect the target object N i About multimedia data M j interaction parameters (such as clicks, playback frequency). Then, according to the interaction weights of the N target objects with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest of the N target objects with respect to the M multimedia data. It can be seen that in the process of obtaining the target interest of the target object with respect to multimedia data, by introducing the interaction weight, it is beneficial to expand the distinction between the media behavior data of different target objects with respect to multimedia data, and the accuracy of obtaining the target interest can be improved. By recommending multimedia data to the N target objects respectively according to their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data.

[0117] See Figure 8 , is a structural diagram of a multimedia data processing device provided in an embodiment of the present application. The multimedia data processing device can be a computer program (including program code) running in a network device, for example, the multimedia data processing device is an application software; the device can be used to execute the corresponding steps of the method provided in an embodiment of the present application. Figure 8 As shown, the multimedia data processing device may include: an acquisition module 811, a determination module 812, a generation module 813, a recommendation module 814, and a training module 815.

[0118] an acquisition module, configured to acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; one target object and one multimedia data together correspond to one piece of media behavior data;

[0119] Determination module, used to determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M;

[0120] The generating module is configured to generate the interaction weights of the N target objects with respect to the M multimedia data respectively based on the P relevant information, the M media attribute information and the M media behavior data respectively corresponding to the N target objects after the relevant information corresponding to the M multimedia data is determined; P is the product of M and N, and the relevant information P ij Belonging to the P relevant information;

[0121] The recommendation module is used to collaboratively process the M media attribute information and the M media behavior data corresponding to the N target objects based on the interaction weights of the N target objects with respect to the M multimedia data, obtain the target interest levels of the N target objects with respect to the M multimedia data, and recommend multimedia data to the N target objects based on the target interest levels of the N target objects with respect to the M multimedia data.

[0122] Optionally, the relevant information P ij For reflecting the multimedia data M j The media attribute information of the target object N iThe covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij ;Generation module, including generation unit and determination unit:

[0123] A generating unit, configured to generate a signal reflecting the target object N i The corresponding media behavior feature matrix X of M media behavior data i , and for reflecting the multimedia data M j The media attribute feature matrix Y of the media attribute information j According to the media behavior feature matrix X i , the covariance matrix Σ ij With the media attribute feature matrix Y j The product between them generates the target object N i Regarding the multimedia data M j Generate the target object N according to the covariance matrix corresponding to the P relevant information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects. i Regarding the multimedia data M j The interaction weight constraints of

[0124] A determination unit is used to determine the interaction weights of the N target objects with respect to the M multimedia data according to the interaction weight equations corresponding to the N target objects and the interest weight constraints corresponding to the N target objects.

[0125] Optionally, the generating unit generates the target object N according to the covariance matrix corresponding to the P related information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects. i Regarding the multimedia data M j The interaction weight constraints include:

[0126] Determine the covariance matrix corresponding to the target object N from the covariance matrix corresponding to the P related information. i The M covariance matrices associated with the multimedia data M j The associated N covariance matrices;

[0127] According to the media behavior feature matrix X i , the media attribute feature matrix corresponding to the M media attribute information, and the target object N i The associated M covariance matrices determine the first constraint condition; the first constraint condition is used to constrain the target object Ni The cumulative sum of the interaction weights of the M multimedia data respectively;

[0128] According to the media attribute feature matrix Y j , N media behavior feature matrices corresponding to the N target objects, and the multimedia data M j The associated N covariance matrices determine the second constraint condition; the second constraint condition is used to constrain the N target objects to be respectively related to the multimedia data M j The cumulative sum of the interaction weights;

[0129] Generate a third constraint condition, determine the first constraint condition, the second constraint condition and the third constraint condition as the target object N i Regarding the multimedia data M j interaction weight constraint; the third constraint is used to constrain the cumulative sum of the interaction weights of the N target objects with respect to the M multimedia data.

[0130] Optionally, the recommendation module includes an identification unit and an adjustment unit:

[0131] an identification unit, configured to associate and identify the M media attribute information and the M media behavior data corresponding to the N target objects, and obtain initial interest levels of the N target objects with respect to the M multimedia data;

[0132] An adjustment unit is used to adjust the initial interest levels of the N target objects with respect to the M multimedia data according to the interaction weights of the N target objects with respect to the M multimedia data, so as to obtain the target interest levels of the N target objects with respect to the M multimedia data.

[0133] Optionally, the identification unit associates and identifies the M media attribute information and the M media behavior data corresponding to the N target objects to obtain initial interest levels of the N target objects with respect to the M multimedia data, including:

[0134] Invoking a feature partitioning layer of a first media recognition model to partition the M media attribute information to obtain first media attribute information having dense features and second media attribute information having sparse features;

[0135] Calling the feature segmentation layer to segment the M media behavior data corresponding to the N target objects respectively to obtain first media behavior data with dense features and second media behavior data with sparse features;

[0136] Invoking the deep learning layer of the first media recognition model to extract key media attribute information from the second media attribute information and to extract key media behavior data from the second media behavior data;

[0137] The first interest recognition layer of the first media recognition model is called to associate and identify the key media behavior data, the key media attribute information, the first media attribute information and the first media behavior data to obtain the initial interest levels of the N target objects with respect to the M multimedia data.

[0138] Optionally, the recognition unit calls a first interest recognition layer of the first media recognition model to perform association recognition on the key media behavior data, the key media attribute information, the first media attribute information, and the first media behavior data, and obtains initial interest levels of the N target objects with respect to the M multimedia data, including:

[0139] calling the first interest recognition layer to concatenate the key media behavior data with the first media behavior data to obtain concatenated media behavior data;

[0140] calling the first interest recognition layer to concatenate the key media attribute information and the first media attribute information to obtain concatenated media attribute information;

[0141] The media behavior data after the splicing process and the media attribute information after the splicing process are associated and identified to obtain candidate interest levels of the N target objects with respect to the M multimedia data.

[0142] An expansion process is performed on the candidate interest degrees of the N target objects with respect to the M multimedia data to obtain initial interest degrees of the N target objects with respect to the M multimedia data.

[0143] Optionally, the adjusting unit adjusts the initial interest levels of the N target objects with respect to the M multimedia data according to the interaction weights of the N target objects with respect to the M multimedia data, to obtain target interest levels of the N target objects with respect to the M multimedia data, including:

[0144] Invoking an averaging layer of the second media recognition model to perform averaging processing on the initial interest levels of the N target objects with respect to the M multimedia data, thereby obtaining averaged initial interest levels;

[0145] Invoking the standard processing layer of the second media recognition model to generate an object standard feature matrix and a media standard feature matrix based on the initial interest levels of the N target objects with respect to the M multimedia data and the interaction weights of the N target objects with respect to the M multimedia data; the object standard feature matrix is ​​used to reflect the standard media behavior characteristics of the N target objects with respect to the M multimedia data, and the media standard feature matrix is ​​used to reflect the standard media attribute characteristics of the M multimedia data;

[0146] The second interest recognition layer of the second media recognition model is called to determine the target interest levels of the N target objects with respect to the M multimedia data based on the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the initial interest levels after averaging.

[0147] Optionally, the adjustment unit calls an averaging layer of the second media recognition model to perform averaging processing on the initial interest levels of the N target objects with respect to the M multimedia data, respectively, to obtain the averaged initial interest levels, including:

[0148] The average processing layer of the second media recognition model is called to determine the initial interest of the N target objects with respect to the M multimedia data from the initial interest of the N target objects with respect to the M multimedia data. j The initial interest level of the target object N i Initial interest levels of the M multimedia data respectively;

[0149] The N target objects are respectively about the multimedia data M j The initial interest of the N target objects is averaged to obtain the interest of the N target objects on the multimedia data M j Average media interest level;

[0150] For the target object N i The initial interest levels of the M multimedia data are averaged to obtain the target object N i Average object interest levels of the M multimedia data;

[0151] averaging the initial interest levels of the N target objects with respect to the M multimedia data to obtain an overall average interest level;

[0152] The overall average interest degree, the object average interest degrees corresponding to the N target objects, and the media average interest degrees corresponding to the M multimedia data are determined as initial interest degrees after averaging processing.

[0153] Optionally, the adjustment unit calls the second interest recognition layer of the second media recognition model to determine the target interest of the N target objects with respect to the M multimedia data based on the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the initial interest after averaging, including:

[0154] Call the second interest recognition layer to determine the target object N from the object standard feature matrix i Regarding the standard media behavior characteristics U of the M multimedia data i ; From the media standard feature matrix, determine the multimedia data M j Standard media attribute characteristics V j ;

[0155] Determining an eigenvalue vector of an interaction weight matrix, and averaging the eigenvalues ​​in the eigenvalue vector to obtain an average eigenvalue; the interaction weight matrix includes interaction weights of the N target objects with respect to the M multimedia data;

[0156] Get the standard media behavior feature U i The transpose of the target object N i Regarding the multimedia data M j The interest weight of the standard media attribute feature V j The product between them is the first interest degree;

[0157] Get the average characteristic root and the multimedia data M j The corresponding media average interest, the target object N i The product of the corresponding average interest levels of the objects is used to obtain the second interest level;

[0158] Determine the target object N according to the overall average interest, the first interest and the second interest i Regarding the multimedia data M j Target interest.

[0159] Optionally, the adjustment unit determines the target object N according to the overall average interest, the first interest and the second interest. i Regarding the multimedia data M j Target interest, including:

[0160] Obtain the cumulative sum of the overall average interest, the first interest, and the second interest to obtain the target object N i Regarding the multimedia data Mj The degree of interest of the candidate;

[0161] Acquire an initial interest matrix for reflecting the initial interest of the N target objects with respect to the M multimedia data;

[0162] If the target object N i Regarding the multimedia data M j If the initial interest degree of the target object N is missing data, the target object N i Regarding the multimedia data M j Fill the initial interest matrix with the selected interest degrees to obtain the target interest matrix;

[0163] If the target object N i Regarding the multimedia data M j If the initial interest degree is non-missing data, the target object N is used. i Regarding the multimedia data M j The selected interest degree replaces the target object N in the initial interest matrix i Regarding the multimedia data M j The initial interest degree of , and the target interest degree matrix is ​​obtained;

[0164] The target object N in the target interest matrix i Regarding the multimedia data M j The selected interest degree is used as the target object N i Regarding the multimedia data M j Target interest.

[0165] The apparatus further includes a training module, the training module being configured to obtain media attribute information corresponding to each of H multimedia data items, media behavior data of each of L sample objects with respect to the H multimedia data items within a second time period, and historical media behavior data of each of the L sample objects with respect to the H multimedia data items within a third time period; the third time period being a time period after the second time period;

[0166] According to the sample object L e Regarding the multimedia data H r The historical media behavior data of the sample object L is determined e Regarding the multimedia data H r The labeled interest degree of the sample object L e belongs to the L sample objects, e is a positive integer less than or equal to L, the multimedia data H r Belonging to the H multimedia data, r is a positive integer less than or equal to H;

[0167] If the annotated interest levels corresponding to the L sample objects are determined, the first initial media recognition model is trained based on the annotated interest levels corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain a first media recognition model.

[0168] Optionally, the training module trains a first initial media recognition model based on the annotated interest levels corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain a first media recognition model, including:

[0169] Divide the L sample objects into K training objects and G test objects; L is the sum of K and G;

[0170] Invoking the first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training subjects, thereby obtaining first predicted interest levels of the K training subjects with respect to the H multimedia data.

[0171] Adjusting the first initial media recognition model according to the labeled interest levels corresponding to the K training objects and the first predicted interest levels corresponding to the K training objects to obtain an adjusted first initial media recognition model;

[0172] If the adjusted first initial media recognition model is in a converged state, then using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects, the interest level recognition performance of the adjusted first initial media recognition model is tested to obtain a test result;

[0173] If the test result indicates that the adjusted first initial media recognition model passes the interest recognition performance test, the adjusted first initial media recognition model is determined as the first media recognition model.

[0174] Optionally, if the adjusted first initial media recognition model is in a converged state, the training module uses the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects to test the interest level recognition performance of the adjusted first initial media recognition model, and obtains test results including:

[0175] Calling the adjusted first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training subjects, to obtain second predicted interest levels of the K training subjects with respect to the H multimedia data;

[0176] determining a gradient of the adjusted first initial media recognition model according to the annotated interest levels corresponding to the K training objects and the second predicted interest levels corresponding to the K training objects;

[0177] If the gradient of the adjusted first initial media identification model is less than the gradient threshold, determining that the adjusted first initial media identification model is in a convergence state;

[0178] The interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects to obtain test results.

[0179] Optionally, if the adjusted first initial media recognition model is in a converged state, the training module uses the media behavior data corresponding to the G test subjects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test subjects to test the interest level recognition performance of the adjusted first initial media recognition model, and obtains test results including:

[0180] If the adjusted first initial media recognition model is in a converged state, calling the adjusted first initial media recognition model to perform association prediction on the media behavior data corresponding to the G test subjects and the media attribute information of the H multimedia data, thereby obtaining the test interest levels of the G test subjects with respect to the H multimedia data;

[0181] determining, according to the test interest levels corresponding to the G test objects and the annotated interest levels corresponding to the G test objects, an interest level recognition performance parameter of the adjusted first initial media recognition model;

[0182] A test result of the interest recognition performance of the adjusted first initial media recognition model is generated according to the interest recognition performance parameter.

[0183] Optionally, the training module is further configured to call the first media recognition model to perform association recognition on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the L sample objects, thereby obtaining initial interest levels of the L sample objects with respect to the H multimedia data.

[0184] Determine the sample object L e The corresponding H media behavior data and the multimedia data H r Related information between media attribute information;

[0185] If the relevant information corresponding to the H multimedia data is determined, then based on the Q relevant information, the H media attribute information, and the H media behavior data corresponding to the L sample objects, the interaction weights of the L sample objects with respect to the H multimedia data are generated; Q is the product of H and L;

[0186] Invoking a second initial media recognition model to collaboratively predict the H media attribute information and the H media behavior data corresponding to the L sample objects based on their interaction weights with respect to the H multimedia data, thereby obtaining target interest levels of the L sample objects with respect to the H multimedia data.

[0187] The second initial media recognition model is trained according to the initial interest levels corresponding to the L sample objects and the target interest levels corresponding to the L sample objects to obtain a second media recognition model.

[0188] Optionally, the training module is further configured to train the second initial media recognition model based on the initial interest levels corresponding to the L sample objects and the target interest levels corresponding to the L sample objects to obtain a second media recognition model, including:

[0189] Determining an interest degree recognition error of the second initial media recognition model according to the initial interest degrees corresponding to the L sample objects and the target interest degrees corresponding to the L sample objects;

[0190] If the interest recognition error of the second initial media recognition model is not in a converged state, adjusting the model parameters of the second initial media recognition model according to the interest recognition error;

[0191] The second initial media recognition model after the model parameters are adjusted is determined as the second media recognition model.

[0192] According to one embodiment of the present application, Figure 3 The steps involved in the multimedia data processing method shown can be represented by Figure 8 The various modules in the multimedia data processing device shown are executed. For example, Figure 3 The step S101 shown in FIG. Figure 8 The acquisition module 811 in is executed, Figure 3 The step S102 shown in FIG. Figure 8 The determination module 812 is executed, Figure 3 The step S103 shown in FIG. Figure 8 The generation module 813 in is executed, Figure 3 Steps S104 and S105 shown in FIG can be replaced by Figure 8 The recommendation module 814 in is executed.

[0193] According to one embodiment of the present application, Figure 8 The various modules in the multimedia data processing device shown can be individually or all combined into one or several units to form, or one (or some) of the units can be further divided into at least two functionally smaller sub-units to achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the functions of one module can also be implemented by at least two units, or the functions of at least two modules can be implemented by one unit. In other embodiments of the present application, the multimedia data processing device can also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of at least two units.

[0194] According to one embodiment of the present application, the program can be executed by running on a general computer device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing components and storage components. Figure 3 A computer program (including program code) for each step involved in the corresponding method shown in Figure 8 The multimedia data processing device shown in and the multimedia data processing method of the embodiment of the present application are implemented. The above computer program can be recorded on, for example, a computer readable recording medium, and loaded into the above computing device through the computer readable recording medium and run therein.

[0195] In this application, the computer device can obtain the media attribute information corresponding to M multimedia data, and the media behavior data of N target objects respectively about the M multimedia data in the first time period, and determine the target object N. i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , the relevant information P ij Used to reflect the target user N i Media behavior data and multimedia data M jFurther, if the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated based on the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects. i About multimedia data M j The interaction weight is mainly used to reflect the target object N i About multimedia data M j interaction parameters (such as clicks, playback frequency). Then, according to the interaction weights of the N target objects with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest of the N target objects with respect to the M multimedia data. It can be seen that in the process of obtaining the target interest of the target object with respect to multimedia data, by introducing the interaction weight, it is beneficial to expand the distinction between the media behavior data of different target objects with respect to multimedia data, and the accuracy of obtaining the target interest can be improved. By recommending multimedia data to the N target objects respectively according to their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data.

[0196] See Figure 9 , is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 9 As shown, the above-mentioned computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may optionally also be at least one storage device away from the aforementioned processor 1001. As Figure 9 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0197] exist Figure 9 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0198] Acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; one target object and one multimedia data together correspond to one piece of media behavior data;

[0199] Determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M;

[0200] If the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated according to the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the relevant information P is the interaction weight of the N target objects with respect to the M multimedia data. ij Belonging to the P relevant information;

[0201] According to the interaction weights of the N target objects with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest levels of the N target objects with respect to the M multimedia data. According to the target interest levels of the N target objects with respect to the M multimedia data, multimedia data are recommended to the N target objects respectively.

[0202] In this application, the computer device can obtain the media attribute information corresponding to M multimedia data, and the media behavior data of N target objects respectively about the M multimedia data in the first time period, and determine the target object N. i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij , the relevant information P ij Used to reflect the target user N i Media behavior data and multimedia data Mj Further, if the relevant information corresponding to the M multimedia data is determined, then the interaction weights of the N target objects with respect to the M multimedia data are generated based on the P relevant information, the M media attribute information and the M media behavior data corresponding to the N target objects. i About multimedia data M j The interaction weight is mainly used to reflect the target object N i About multimedia data M j interaction parameters (such as clicks, playback frequency). Then, according to the interaction weights of the N target objects with respect to the M multimedia data, the M media attribute information and the M media behavior data corresponding to the N target objects are collaboratively processed to obtain the target interest of the N target objects with respect to the M multimedia data. It can be seen that in the process of obtaining the target interest of the target object with respect to multimedia data, by introducing the interaction weight, it is beneficial to expand the distinction between the media behavior data of different target objects with respect to multimedia data, and the accuracy of obtaining the target interest can be improved. By recommending multimedia data to the N target objects respectively according to their target interest with respect to the M multimedia data, the problem of low push accuracy of multimedia data when the media behavior data of a certain object with respect to multimedia data is relatively evenly distributed can be avoided. This solution can improve the push accuracy of multimedia data and achieve accurate push of multimedia data.

[0203] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 3 or Figure 5 The description of the multimedia data processing method in the corresponding embodiment can also be performed as described above. Figure 8 The description of the multimedia data processing device in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0204] In addition, it should be noted that: the embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the multimedia data processing device mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the multimedia data processing device mentioned above. Figure 3 And the previous article Figure 5 The description of the multimedia data processing method described in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0205] As an example, the above program instructions may be deployed on a computer device for execution, or deployed on at least two computer devices at one location for execution, or executed on at least two computer devices distributed at at least two locations and interconnected via a communication network. The at least two computer devices distributed at at least two locations and interconnected via a communication network may constitute a blockchain network.

[0206] The computer-readable storage medium may be the multimedia data processing device provided in any of the aforementioned embodiments or the central storage unit of the computer device, such as a hard disk or a memory card of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may include both the central storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0207] The terms "first," "second," and the like in the description, claims, and drawings of the embodiments of the present application are used to distinguish between contents in different media, rather than to describe a specific order. In addition, the terms "including" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules that are not listed, or may optionally include other steps and units inherent to these processes, methods, apparatuses, products, or devices.

[0208] The present application also provides a computer program product, including a computer program / instruction, which implements the above-mentioned Figure 3 and Figure 5 The description of the multimedia data processing method described in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiment of this application.

[0209] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0210] The methods and related devices provided by the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided by the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable network-connected device to generate a machine, so that the instructions executed by the processor of the computer or other programmable network-connected device generate instructions for implementing the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable network-connected device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the function specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable network connected device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide the functions for implementing the process. Figure 1 The flow or flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.

[0211] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A multimedia data processing method, characterized in that: include: Acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; A target object and a multimedia data together correspond to a piece of media behavior data; Determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M; The relevant information P ij For reflecting the multimedia data M j The media attribute information of the target object N i The covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij ; If the relevant information corresponding to the M multimedia data is determined, then generating interaction weights of the N target objects with respect to the M multimedia data according to the P relevant information, the M media attribute information, and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the related information P ij Belonging to the P relevant information; Performing association identification on the M media attribute information and the M media behavior data corresponding to the N target objects, respectively, to obtain initial interest levels of the N target objects with respect to the M multimedia data; According to the interaction weights of the N target objects respectively with respect to the M multimedia data, the initial interest levels of the N target objects respectively with respect to the M multimedia data are adjusted to obtain the target interest levels of the N target objects respectively with respect to the M multimedia data; and according to the target interest levels of the N target objects respectively with respect to the M multimedia data, multimedia data are recommended to the N target objects respectively.

2. The method according to claim 1, wherein Generating interaction weights of the N target objects with respect to the M multimedia data based on the P relevant information, the M media attribute information, and the M media behavior data corresponding to the N target objects, respectively, includes: Generate a signal to reflect the target object N i The corresponding media behavior feature matrix X of M media behavior data i , and for reflecting the multimedia data M j The media attribute feature matrix Y of the media attribute information j ; According to the media behavior feature matrix X i , the covariance matrix Σ ij With the media attribute feature matrix Y j The product between them generates the target object N i Regarding the multimedia data M j The interaction weight equation; Generate the target object N according to the covariance matrix corresponding to the P related information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects. i Regarding the multimedia data M j The interaction weight constraints of The interaction weights of the N target objects with respect to the M multimedia data are determined according to the interaction weight equations corresponding to the N target objects and the interest weight constraints corresponding to the N target objects.

3. The method according to claim 2, wherein The target object N is generated according to the covariance matrix corresponding to the P related information, the media attribute feature matrix corresponding to the M media attribute information, and the media behavior feature matrix corresponding to the N target objects. i Regarding the multimedia data M j The interaction weight constraints include: From the covariance matrices corresponding to the P relevant information, determine the N i The M covariance matrices associated with the multimedia data M j The associated N covariance matrices; According to the media behavior feature matrix X i , the media attribute feature matrix corresponding to the M media attribute information, and the target object N i The associated M covariance matrices determine the first constraint condition; the first constraint condition is used to constrain the target object N i The cumulative sum of the interaction weights of the M multimedia data respectively; According to the media attribute feature matrix Y j , N media behavior feature matrices corresponding to the N target objects, and the multimedia data M j The associated N covariance matrices determine the second constraint condition; the second constraint condition is used to constrain the N target objects to be respectively related to the multimedia data M j The cumulative sum of the interaction weights; Generate a third constraint condition, determine the first constraint condition, the second constraint condition and the third constraint condition as the target object N i Regarding the multimedia data M j interaction weight constraint; the third constraint is used to constrain the cumulative sum of the interaction weights of the N target objects with respect to the M multimedia data.

4. The method according to claim 1, wherein The associating and identifying the M media attribute information and the M media behavior data corresponding to the N target objects to obtain initial interest levels of the N target objects with respect to the M multimedia data includes: Invoking a feature partitioning layer of a first media recognition model to partition the M media attribute information to obtain first media attribute information having dense features and second media attribute information having sparse features; Calling the feature segmentation layer to segment the M media behavior data corresponding to the N target objects respectively to obtain first media behavior data with dense features and second media behavior data with sparse features; Invoking the deep learning layer of the first media recognition model to extract key media attribute information from the second media attribute information and to extract key media behavior data from the second media behavior data; The first interest recognition layer of the first media recognition model is called to associate and identify the key media behavior data, the key media attribute information, the first media attribute information and the first media behavior data to obtain the initial interest levels of the N target objects with respect to the M multimedia data.

5. The method according to claim 4, wherein The calling of the first interest recognition layer of the first media recognition model to associate and recognize the key media behavior data, the key media attribute information, the first media attribute information, and the first media behavior data to obtain initial interest levels of the N target objects with respect to the M multimedia data, respectively, includes: calling the first interest recognition layer to concatenate the key media behavior data with the first media behavior data to obtain concatenated media behavior data; calling the first interest recognition layer to concatenate the key media attribute information and the first media attribute information to obtain concatenated media attribute information; Performing association identification on the spliced ​​media behavior data and the spliced ​​media attribute information to obtain candidate interest levels of the N target objects with respect to the M multimedia data; An expansion process is performed on the candidate interest degrees of the N target objects with respect to the M multimedia data to obtain initial interest degrees of the N target objects with respect to the M multimedia data.

6. The method according to claim 1, wherein The adjusting, based on the interaction weights of the N target objects with respect to the M multimedia data, the initial interest levels of the N target objects with respect to the M multimedia data to obtain the target interest levels of the N target objects with respect to the M multimedia data, includes: Invoking an averaging layer of the second media recognition model to perform averaging processing on the initial interest levels of the N target objects with respect to the M multimedia data, thereby obtaining averaged initial interest levels; Invoking the standard processing layer of the second media recognition model to generate an object standard feature matrix and a media standard feature matrix based on the initial interest levels of the N target objects with respect to the M multimedia data and the interaction weights of the N target objects with respect to the M multimedia data; the object standard feature matrix is ​​used to reflect the standard media behavior characteristics of the N target objects with respect to the M multimedia data, and the media standard feature matrix is ​​used to reflect the standard media attribute characteristics of the M multimedia data; the object standard feature matrix is ​​an object scoring matrix, and the media standard feature matrix is ​​a media scoring matrix; The second interest recognition layer of the second media recognition model is called to determine the target interest levels of the N target objects with respect to the M multimedia data based on the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the initial interest levels after averaging.

7. The method according to claim 6, wherein The calling of the averaging layer of the second media recognition model to perform averaging processing on the initial interest levels of the N target objects with respect to the M multimedia data to obtain the averaged initial interest levels includes: The average processing layer of the second media recognition model is called to determine the initial interest of the N target objects with respect to the M multimedia data from the initial interest of the N target objects with respect to the M multimedia data. j The initial interest level of the target object N i Initial interest levels of the M multimedia data respectively; The N target objects are respectively about the multimedia data M j The initial interest of the N target objects is averaged to obtain the interest of the N target objects on the multimedia data M j Average media interest level; For the target object N i The initial interest levels of the M multimedia data are averaged to obtain the target object N i Average object interest levels of the M multimedia data; averaging the initial interest levels of the N target objects with respect to the M multimedia data to obtain an overall average interest level; The overall average interest degree, the object average interest degrees corresponding to the N target objects, and the media average interest degrees corresponding to the M multimedia data are determined as initial interest degrees after averaging processing.

8. The method according to claim 7, wherein The calling of the second interest recognition layer of the second media recognition model to determine the target interest of the N target objects with respect to the M multimedia data according to the object standard feature matrix, the media standard feature matrix, the interaction weights of the N target objects with respect to the M multimedia data, and the averaged initial interest includes: Call the second interest recognition layer to determine the target object N from the object standard feature matrix i Regarding the standard media behavior characteristics U of the M multimedia data i ; From the media standard feature matrix, determine the multimedia data M j Standard media attribute characteristics V j ; Determining an eigenvalue vector of an interaction weight matrix, and averaging the eigenvalues ​​in the eigenvalue vector to obtain an average eigenvalue; the interaction weight matrix includes interaction weights of the N target objects with respect to the M multimedia data; Get the standard media behavior feature U i The transpose of the target object N i Regarding the multimedia data M j The interest weight of the standard media attribute feature V j The product between them is the first interest degree; Get the average characteristic root and the multimedia data M j The corresponding media average interest, the target object N i The product of the corresponding average interest levels of the objects is used to obtain the second interest level; Determine the target object N according to the overall average interest, the first interest and the second interest i Regarding the multimedia data M j Target interest.

9. The method according to claim 8, wherein The target object N is determined based on the overall average interest, the first interest, and the second interest. i Regarding the multimedia data M j Target interest, including: Obtain the cumulative sum of the overall average interest, the first interest, and the second interest to obtain the target object N i Regarding the multimedia data M j The degree of interest of the candidate; Acquire an initial interest matrix for reflecting the initial interest of the N target objects with respect to the M multimedia data; If the target object N i Regarding the multimedia data M j If the initial interest degree of the target object N is missing data, the target object N i Regarding the multimedia data M j Fill the initial interest matrix with the selected interest degrees to obtain the target interest matrix; If the target object N i Regarding the multimedia data M j If the initial interest degree is non-missing data, the target object N is used. i Regarding the multimedia data M j The selected interest degree replaces the target object N in the initial interest matrix i Regarding the multimedia data M j The initial interest degree of , and the target interest degree matrix is ​​obtained; The target object N in the target interest matrix i Regarding the multimedia data M j The selected interest degree is used as the target object N i Regarding the multimedia data M j Target interest.

10. The method according to claim 4, wherein The method further comprises: Obtaining media attribute information corresponding to H multimedia data, media behavior data of L sample objects with respect to the H multimedia data within a second time period, and historical media behavior data of the L sample objects with respect to the H multimedia data within a third time period; the third time period being a time period after the second time period; According to the sample object L e Regarding the multimedia data H r The historical media behavior data of the sample object L is determined e Regarding the multimedia data H r The labeled interest degree of the sample object L e belongs to the L sample objects, e is a positive integer less than or equal to L, the multimedia data H r Belonging to the H multimedia data, r is a positive integer less than or equal to H; If the annotated interest levels corresponding to the L sample objects are determined, the first initial media recognition model is trained based on the annotated interest levels corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain a first media recognition model.

11. The method according to claim 10, wherein The first initial media recognition model is trained based on the annotated interest levels corresponding to the L sample objects, the media attribute information corresponding to the H multimedia data, and the media behavior data corresponding to the L sample objects to obtain the first media recognition model, including: Divide the L sample objects into K training objects and G test objects; L is the sum of K and G; Invoking the first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training subjects, thereby obtaining first predicted interest levels of the K training subjects with respect to the H multimedia data. Adjusting the first initial media recognition model according to the labeled interest levels corresponding to the K training objects and the first predicted interest levels corresponding to the K training objects to obtain an adjusted first initial media recognition model; If the adjusted first initial media recognition model is in a converged state, then using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects, the interest level recognition performance of the adjusted first initial media recognition model is tested to obtain a test result; If the test result indicates that the adjusted first initial media recognition model passes the interest recognition performance test, the adjusted first initial media recognition model is determined as the first media recognition model.

12. The method according to claim 11, wherein If the adjusted first initial media recognition model is in a converged state, the interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects, to obtain test results, including: Calling the adjusted first initial media recognition model to perform association prediction on the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the K training subjects, to obtain second predicted interest levels of the K training subjects with respect to the H multimedia data; determining a gradient of the adjusted first initial media recognition model according to the annotated interest levels corresponding to the K training objects and the second predicted interest levels corresponding to the K training objects; If the gradient of the adjusted first initial media identification model is less than the gradient threshold, determining that the adjusted first initial media identification model is in a convergence state; The interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects to obtain test results.

13. The method according to claim 11, wherein If the adjusted first initial media recognition model is in a converged state, the interest recognition performance of the adjusted first initial media recognition model is tested using the media behavior data corresponding to the G test objects, the media attribute information of the H multimedia data, and the annotated interest levels corresponding to the G test objects, to obtain test results, including: If the adjusted first initial media recognition model is in a converged state, calling the adjusted first initial media recognition model to perform association prediction on the media behavior data corresponding to the G test subjects and the media attribute information of the H multimedia data, thereby obtaining the test interest levels of the G test subjects with respect to the H multimedia data; determining, according to the test interest levels corresponding to the G test objects and the annotated interest levels corresponding to the G test objects, an interest level recognition performance parameter of the adjusted first initial media recognition model; A test result of the interest recognition performance of the adjusted first initial media recognition model is generated according to the interest recognition performance parameter.

14. The method according to claim 10, wherein The method further comprises: Invoking the first media recognition model to associate and identify the media attribute information corresponding to the H multimedia data and the media behavior data corresponding to the L sample objects, thereby obtaining initial interest levels of the L sample objects with respect to the H multimedia data. Determine the sample object L e The corresponding H media behavior data and the multimedia data H r Related information between media attribute information; If the relevant information corresponding to the H multimedia data is determined, then based on the Q relevant information, the H media attribute information, and the H media behavior data corresponding to the L sample objects, the interaction weights of the L sample objects with respect to the H multimedia data are generated; Q is the product of H and L; Invoking a second initial media recognition model to collaboratively predict the H media attribute information and the H media behavior data corresponding to the L sample objects based on their interaction weights with respect to the H multimedia data, thereby obtaining target interest levels of the L sample objects with respect to the H multimedia data. The second initial media recognition model is trained according to the initial interest levels corresponding to the L sample objects and the target interest levels corresponding to the L sample objects to obtain a second media recognition model.

15. The method according to claim 14, wherein The training of the second initial media recognition model according to the initial interest levels corresponding to the L sample objects and the target interest levels corresponding to the L sample objects to obtain the second media recognition model includes: Determining an interest degree recognition error of the second initial media recognition model according to the initial interest degrees corresponding to the L sample objects and the target interest degrees corresponding to the L sample objects; If the interest recognition error of the second initial media recognition model is not in a converged state, adjusting the model parameters of the second initial media recognition model according to the interest recognition error; The second initial media recognition model after the model parameters are adjusted is determined as the second media recognition model.

16. A multimedia data processing device, characterized in that: include: an acquisition module, configured to acquire media attribute information corresponding to M multimedia data, and media behavior data of N target objects with respect to the M multimedia data within a first time period; A target object and a multimedia data together correspond to a piece of media behavior data; Determination module, used to determine the target object N i The corresponding M media behavior data and multimedia data M j The related information P between the media attribute information ij ; The target object N i Belonging to the N target objects, the multimedia data M j Belonging to the M multimedia data, i is a positive integer less than or equal to N, j is a positive integer less than or equal to M; The relevant information P ij For reflecting the multimedia data M j The media attribute information of the target object N i The covariance matrix Σ of the distribution correlation between the corresponding M media behavior data ij ; a generating module configured to generate, upon completion of determining the relevant information corresponding to the M multimedia data, interaction weights of the N target objects with respect to the M multimedia data based on the P relevant information, the M media attribute information, and the M media behavior data corresponding to the N target objects; P is the product of M and N, and the related information P ij Belonging to the P relevant information; The recommendation module includes an identification unit and an adjustment unit; an identification unit, configured to associate and identify the M media attribute information and the M media behavior data corresponding to the N target objects, and obtain initial interest levels of the N target objects with respect to the M multimedia data; an adjusting unit, configured to adjust initial interest levels of the N target objects with respect to the M multimedia data according to interaction weights of the N target objects with respect to the M multimedia data, to obtain target interest levels of the N target objects with respect to the M multimedia data; The recommendation module is configured to recommend multimedia data to the N target objects respectively according to their target interest levels with respect to the M multimedia data.

17. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Data pushing method and device

    CN106055617A