Media data processing method and device, equipment and storage medium
By introducing knowledge graphs and media triples into the multimedia data recommendation system for iterative training, the aggregation of media features in long-tail multimedia data is explicitly constrained, thus solving the problems of high model training complexity and low efficiency caused by the sparsity of long-tail data, and achieving more efficient model training and accurate recommendation results.
Patent Information
- Application Number
- CN202410632836.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-28
AI Technical Summary
In existing multimedia data recommendation systems, the sparsity of long-tail multimedia data leads to an imbalance in training samples, resulting in biased model recommendation performance. Existing methods increase training complexity and are inefficient by optimizing the model structure.
By introducing knowledge graphs, effective and ineffective media triples are generated. The initial media recommendation model is iteratively trained, and the aggregation of media features under the same attribute profile label is explicitly constrained to avoid ineffective features. The model training data includes long-tail and popular multimedia data.
It reduces model training complexity, improves training efficiency and accuracy, enhances the learning effect of long-tail multimedia data, and solves the problem of sample sparsity.
Smart Images

Figure CN121029979A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a media data processing method and device, equipment and a storage medium. BACKGROUND
[0002] In the current multimedia data recommendation system, the multimedia data to be processed generally has a large scale, and the number of multimedia data is usually in the order of millions or tens of millions. In addition to some popular multimedia data (which often has good investment effect and sufficient accumulated samples), there are also a large number of long-tail multimedia data (unpopular, newly invested, etc.). The historical conversion amount of these long-tail multimedia data is small, and the accumulated sample amount is sparse.
[0003] Therefore, when constructing samples based on historical multimedia data and training a new round of media recommendation model, the relevant samples of the sampled popular multimedia data are relatively more, while the relevant samples of the long-tail multimedia data are relatively sparse. The imbalance of the training samples will cause the recommendation effect of the media recommendation model to be biased, such as more sufficient learning of popular multimedia data and more accurate recommendation; insufficient training of long-tail multimedia data will make it difficult for the multimedia data recommendation system to accurately recommend the long-tail multimedia data to the interested population. Therefore, when modeling the multimedia data recommendation system, the sparsity problem of the long-tail multimedia data needs to be considered. At present, the model structure is mainly optimized to process the samples of the popular multimedia data and the long-tail multimedia data by using different network models, which will double the model parameters, increase the training complexity, and result in low training efficiency. SUMMARY
[0004] The embodiments of the present application provide a media data processing method, device, equipment and storage medium, which reduce the training complexity and improve the training efficiency.
[0005] The embodiments of the present application provide a media data processing method, device, equipment and storage medium, which reduce the training complexity and improve the training efficiency.
[0006] The method comprises the following steps:
[0007] According to the first media feature and the knowledge graph associated with the sample multimedia data, an effective media triple associated with the sample multimedia data is generated. The effective media triple includes the first media feature, the first attribute portrait label and the first association semantic. The first association semantic reflects that the first media feature and the first attribute portrait label have an association relationship.
[0008] Based on the above valid media triples, N invalid media triples that are not associated with the above sample multimedia data are generated. Each invalid media triple includes a second media feature, a second attribute profile label, and the above first association semantics. There is no association relationship between the above second media feature and the above second attribute profile label. N is a positive integer.
[0009] Based on the first association semantics within the effective media triples, the first distance between the first media feature and the first attribute profile label is determined using the initial media recommendation model. Based on the first association semantics within each of the invalid media triples, the second distance between the second media feature and the second attribute profile label within each of the invalid media triples is determined.
[0010] The initial media recommendation model is used to identify the first media feature and the object feature to obtain the predicted interest degree of the sample object for the sample multimedia data. The initial media recommendation model is iteratively trained based on the predicted interest degree, the labeled interest degree, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping condition for iterative training, thus obtaining the target media recommendation model for recommending multimedia data.
[0011] One embodiment of this application provides a media data processing apparatus, including:
[0012] The acquisition module is used to acquire the object features of the sample object, the first media features of the sample multimedia data, and the labeled interest degree of the sample object in relation to the sample multimedia data.
[0013] The first generation module is used to generate a valid media triple associated with the sample multimedia data based on the knowledge graph associated with the first media feature and the sample multimedia data. The valid media triple includes the first media feature, the first attribute profile label, and the first association semantics, and the first association semantics reflects the association relationship between the first media feature and the first attribute profile label.
[0014] The second generation module is used to generate N invalid media triples that are not associated with the above sample multimedia data based on the above valid media triples. Each of the above invalid media triples includes a second media feature, a second attribute profile label and the above first association semantics. There is no association relationship between the above second media feature and the above second attribute profile label. N is a positive integer.
[0015] The determination module is used to determine, through the initial media recommendation model, the first distance between the first media feature and the first attribute profile label based on the first association semantics within the effective media triples, and the second distance between the second media feature and the second attribute profile label within each of the invalid media triples based on the first association semantics within each of the invalid media triples.
[0016] The training module is used to identify the first media feature and the object feature through the initial media recommendation model, obtain the predicted interest degree of the sample object for the sample multimedia data, and iteratively train the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping condition for iterative training, and obtain the target media recommendation model for recommending multimedia data.
[0017] One embodiment of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0018] One embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0019] This application includes, but is not limited to, the following beneficial effects: (1) introducing a knowledge graph to characterize the association graph between sample multimedia data and its attribute profile tags. The sample multimedia data includes long-tail multimedia data and popular multimedia data. When long-tail multimedia data and popular multimedia data have the same attribute profile tags, both long-tail multimedia data and popular multimedia data are connected to the same attribute profile tag. That is, through the knowledge graph, the connection between long-tail multimedia data and popular multimedia data under the same attribute profile tag can be realized.
[0020] (2) Iteratively training the initial media recommendation model using the first distance corresponding to the effective media triplet helps to explicitly constrain the media features of sample multimedia data under the same attribute profile label to be as similar as possible. This ensures that the media features of long-tail multimedia data and popular multimedia data under the same attribute profile label (i.e., the first attribute profile label) cluster towards the same central point, which can be the first attribute profile label. This also makes the media features of long-tail multimedia data under the same attribute profile label closer to the media features of popular multimedia data, increasing the number of samples under the same profile attribute label and reducing the training complexity of the initial media recommendation model. At the same time, it helps long-tail multimedia data learn information from popular multimedia data under the same attribute profile label, thereby improving the learning effect of the initial media recommendation model on the media features of long-tail multimedia data and solving the sparsity problem of long-tail multimedia data. (3) Invalid media triplets refer to incorrect triplets. Iteratively training the initial media recommendation model using the second distance corresponding to the invalid media triplet helps to make the second media features and second attribute profile labels within the invalid media triplet move away from the central point corresponding to the first attribute profile label, which helps to improve the training accuracy of the model. (4) During the training process, a single initial media recommendation model is shared for both long-tail multimedia data and popular multimedia data. This eliminates the need to create separate network models for long-tail multimedia data and popular multimedia data, thereby reducing the training complexity of the model and improving the training efficiency. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of a media data processing system provided in this application;
[0023] Figure 2 This is a flowchart illustrating a media data processing method provided in this application;
[0024] Figure 3 This is a schematic diagram of a knowledge graph provided in this application;
[0025] Figure 4 This is a flowchart illustrating a media data processing method provided in this application;
[0026] Figure 5 This is a schematic diagram of the network structure of an initial media recommendation model provided in this application;
[0027] Figure 6 This is a schematic diagram of the structure of a media data processing device provided in an embodiment of this application;
[0028] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0030] This application's embodiments may relate to artificial intelligence (AI) technology, as well as fields such as autonomous driving and intelligent transportation. AI refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0031] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0032] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0033] This application primarily involves machine learning techniques within artificial intelligence. It utilizes machine learning to construct an initial media recommendation model. During the iterative training of this initial model, the first distance corresponding to effective media triples is used for iterative training. This helps to explicitly constrain the media features of sample multimedia data under the same attribute profile label to be as similar as possible. This ensures that the media features of long-tail multimedia data and popular multimedia data under the same attribute profile label (i.e., the first attribute profile label) cluster towards the same central point, which can be the first attribute profile label. This also makes the media features of long-tail multimedia data under the same attribute profile label closer to the media features of popular multimedia data, increasing the number of samples under the same profile attribute label and improving the learning effect of the initial media recommendation model for the media features of long-tail multimedia data, thus solving the sparsity problem of long-tail multimedia data. The second distance corresponding to invalid media triples is used for iterative training of the initial media recommendation model. This helps to move the second media features and second attribute profile labels within invalid media triples away from the central point corresponding to the first attribute profile label, thereby improving the training accuracy of the model.
[0034] To facilitate a clearer understanding of this application, the media data processing system implementing this application will be introduced first, such as... Figure 1 As shown, this media data processing system includes a server 10 and a terminal cluster. The terminal cluster can include one or more terminals; the number of terminals is not limited here. Figure 1 As shown, taking a terminal cluster consisting of 4 terminals as an example, the terminal cluster can specifically include terminal 11a, terminal 12a, terminal 13a, and terminal 14a. It can be understood that terminal 11a, terminal 12a, terminal 13a, and terminal 14a can all connect to server 10 via the network so that each terminal can interact with server 10 through the network connection.
[0035] Understandably, a server can be a single physical server, a server cluster or distributed system consisting of at least two physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, ContentDelivery Network (CDN), and big data and artificial intelligence platforms. A terminal can specifically refer to in-vehicle terminals, smartphones, tablets, laptops, desktop computers, smart speakers, speakers with screens, smart TVs, smartwatches, etc., but is not limited to these. Various terminals and servers can be directly or indirectly connected via wired or wireless communication. Furthermore, the number of terminals and servers can be one or at least two; this application does not impose any restrictions.
[0036] Each of the aforementioned terminals may have one or more target applications installed. These target applications can refer to applications with data processing capabilities (such as multimedia data recommendation), and may include standalone applications, web applications, or mini-programs within a host application. Specifically, target applications may include browser applications, shopping applications, application download applications, video playback applications, social applications, content publishing applications, and game applications.
[0037] In this context, server 10 refers to a device that provides backend services for the target application in the terminal. Server 10 may include an initial media recommendation model, etc. Server 10 can be used to iteratively train the initial media recommendation model based on the method provided in this application to obtain the target media recommendation model; the terminal and server 10 can provide multimedia data recommendation services to users through the target application and the target media recommendation model.
[0038] For example, taking terminal 11a as an example, terminal 11a may include a target application. Terminal 11a can log in to the target application based on the user's basic information and display the target application's recommendation interface. Here, the recommendation page may refer to a multimedia data display interface, etc. The user can input basic media features such as media type through the recommendation interface displayed on terminal 11a. Terminal 11a can obtain the basic media features input by the user and send them to server 10. Server 10 can be used to obtain candidate multimedia data to be recommended based on the basic media features, call the target media recommendation model, and identify the target object's interest in the candidate multimedia data based on the object features of the target object corresponding to terminal 11a and the media features of the candidate multimedia data. It returns candidate multimedia data with an interest level greater than the interest threshold to terminal 11a so that terminal 11a can display the returned candidate multimedia data to the target object on the recommendation interface.
[0039] In one embodiment, after obtaining the target media recommendation model, the server 10 can send the target media recommendation model to each terminal (taking terminal 11a as an example). Terminal 11a can store the target media recommendation model locally and call the target media recommendation model from the local device through the target application to recommend multimedia data to the user.
[0040] In one embodiment, after obtaining the target media recommendation model, the server 10 can send the target media recommendation model to the cloud server. Each terminal (taking terminal 11a as an example) can call the target media recommendation model in the cloud server through the target application to recommend multimedia data to the user.
[0041] Multimedia data can refer to advertisements, items, applications, audio and video data, game items, online media activities, etc. Sample multimedia data can refer to multimedia data used for iterative training of the initial media recommendation model, and candidate multimedia data can refer to multimedia data to be recommended. Sample multimedia data can include candidate multimedia data. Both sample multimedia data and candidate multimedia data can include long-tail multimedia data and popular multimedia data.
[0042] Long-tail multimedia data can refer to newly launched multimedia data, or it can be multimedia data whose popularity is below a popularity threshold within a historical time period. Popular multimedia data refers to multimedia data whose popularity is greater than or equal to a popularity threshold within a historical time period. The popularity of multimedia data can be determined based on user media interaction data within a historical time period. Media interaction data can include the number of clicks, favorites, purchases, trials, and views. For example, the more clicks multimedia data receives within a historical time period, the higher its popularity; conversely, the fewer clicks multimedia data receives within a historical time period, the lower its popularity.
[0043] The media characteristics of multimedia data can include basic media characteristics and non-basic media characteristics. Basic media characteristics include media identifiers and names of multimedia data, while non-basic media characteristics can include the release time, media interaction data, main content, attribute profile tags, etc. of multimedia data.
[0044] In this context, the attribute profile tags for multimedia data can refer to the media category to which the multimedia data belongs. Multimedia data typically includes multi-level attribute profile tags, which include parent attribute profile tags and child attribute profile tags. The child attribute profile tags are sub-tags of the parent attribute profile tags. For example, the attribute profile tags for a milk advertisement can include food, dairy products, etc. Food is the parent attribute profile tag (i.e., the first-level attribute profile tag, also known as the first-level functional attribute category), and dairy products are the child attribute profile tags (i.e., the second-level attribute profile tag, also known as the second-level functional attribute category).
[0045] In this context, "object" can refer to a user. The object's characteristics can include basic user information (such as user ID) and interests / preferences. Interests / preferences can include multimedia data related to the user's interactions (likes, favorites, plays, purchases, etc.) within a historical time period. In other words, interests / preferences reflect the multimedia data that the user was interested in during a specific historical period. The historical time period can be the past week, month, or year. "Sample object" can refer to the users used to train the initial media model. For a given sample multimedia data, the sample object can include positive and negative sample objects. A positive sample object can be the target user group of the sample multimedia data, meaning users who are interested in it. The positive sample object's interest level for the sample multimedia data is 1. A negative sample object can be the non-target user group of the sample multimedia data, meaning users who are not interested in it. The negative sample object's interest level for the sample multimedia data is 0.
[0046] Further, please see Figure 2 This is a flowchart illustrating a media data processing method provided in an embodiment of this application. Figure 2 As shown, this method can be derived from... Figure 1 It can be executed by any terminal in the terminal cluster, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1 The terminal cluster in the application uses terminals and servers to collaboratively execute the media data processing method. The device used to execute this method can be collectively referred to as a computer device. The method may include the following steps:
[0047] S101. Obtain the object features of the sample object, the first media features of the sample multimedia data, and the labeled interest degree of the sample object in relation to the sample multimedia data.
[0048] In this application, a computer device can obtain the first media features of the sample multimedia data, the object features of the sample object, and the labeled interest level of the sample object towards the sample multimedia data from local storage or other devices. The labeled interest level can be obtained based on the sample object's historical interaction data with the sample multimedia data. The historical interaction data can include the number of interactions such as liking, collecting, playing, and purchasing by the sample object towards the sample multimedia data. The labeled interest level can be 1 or 0. A labeled interest level of 1 for the sample multimedia data indicates that the sample object is interested in the sample multimedia data, that is, the sample object is the target user group of the sample multimedia data. A labeled interest level of 0 for the sample multimedia data indicates that the sample object is not interested in the sample multimedia data, that is, the sample object is a non-target user group of the sample multimedia data.
[0049] It should be noted that there can be multiple sample multimedia data sets, multiple sample objects, and multiple media types. These media types include various categories; for example, if the sample multimedia data is advertising, the media types could include car advertisements, cosmetics advertisements, household goods advertisements, etc. When the sample multimedia data is audio-visual data, the media types could include short videos, films, live videos, video conferencing, music, etc. The sample objects include positive and negative sample objects for each media type.
[0050] S102. Based on the knowledge graph associated with the first media feature and the sample multimedia data, generate an effective media triple associated with the sample multimedia data; the effective media triple includes the first media feature, the first attribute profile label, and the first association semantics, wherein the first association semantics reflects the association relationship between the first media feature and the first attribute profile label.
[0051] The knowledge graph associated with the sample multimedia data reflects the relationship between the sample multimedia data and attribute profile tags. It also reflects the relationship between multi-level attribute profile tags within the sample multimedia data. The knowledge graph can be a knowledge network with entities as nodes and directed edges representing relationships between entities. Entities can include head entities and tail entities. The knowledge graph includes multiple initial triples, each consisting of (head entity, relation, tail entity). The head entity can refer to either the sample multimedia data or the attribute profile tag, and the tail entity can refer to the attribute profile tag of the sample multimedia data. The relation indicates the relationship between the head and tail entities. When the head entity in the initial triple is sample multimedia data, the initial triple can be called the initial media triple. The relation within the initial media triple can be called the first association semantic, reflecting the relationship between the sample multimedia data and the attribute profile tag. This relationship indicates which level of attribute profile tag the attribute profile tag belongs to within the sample multimedia data. When the head entity in the initial triple is an attribute profile label, the initial triple can be called the initial label triple. The relation in the initial label triple can be called the second association semantics, which reflects the association relationship between different attribute profile labels. In other words, the second association semantics reflects the subordinate relationship between different attribute profile labels.
[0052] In this application, a computer device can query an initial media triplet belonging to the sample multimedia data based on a first media feature. The initial media triplet is (sample multimedia data, first association semantics, first attribute profile label), where the first association semantics reflects the association between the sample multimedia data and the first attribute profile label. The sample multimedia data in the initial media triplet is replaced with the first media feature to obtain a valid media triplet, which is (first media feature, first association semantics, first attribute profile label). Here, the first association semantics also reflects the association between the first media feature and the first attribute profile label. The first association semantics can serve as translation information between the first media feature and the first attribute profile label, i.e., as a mapping from the first media feature to the first attribute profile label.
[0053] It should be noted that sample multimedia data with the same type of head entity and the same type of tail entity have the same relationship. Different relationships will be assigned when the types of head entity and tail entity are different. The relationship between head entity and tail entity can be determined according to the labeling method of the scenario corresponding to the sample multimedia data.
[0054] For example, such as Figure 3As shown, taking an advertising scenario as an example, the primary attribute profile tags include food, cosmetics, and daily necessities; the sub-attribute profile tags for food (secondary attribute profile tags in the advertising scenario) include dairy products, vegetables, etc. The initial media triplet for milk advertising includes (milk advertising, primary functional attribute category, food), (milk advertising, secondary functional attribute category, dairy products), etc.; the initial tag triplet includes (food, sub-functional attribute category, dairy products), etc., that is, the sub-functional attribute category is used to indicate that dairy products are sub-attribute profile tags for food.
[0055] S103. Based on the above valid media triples, generate N invalid media triples that are not associated with the above sample multimedia data. Each invalid media triple includes a second media feature, a second attribute profile label, and the above first association semantics. There is no association relationship between the above second media feature and the above second attribute profile label. N is a positive integer.
[0056] In this application, the computer device can randomly generate N invalid media triples that are not associated with the sample multimedia data based on the valid media triples. The second media feature and the second attribute profile label in the invalid media triples are not associated, that is, the first association semantics cannot reflect the association between the second media feature and the second attribute profile label, that is, the first association semantics cannot be used as a mapping between the second media feature and the second attribute profile label.
[0057] For example, taking milk advertising as an example, the effective media triplet for milk advertising includes (media characteristics of the milk advertisement, primary functional attribute category, food), while the ineffective media triplet for milk advertising includes (media characteristics of the milk advertisement, primary functional attribute category, cosmetics) and (media characteristics of lipstick, primary functional attribute category, food). Here, the media characteristics of milk advertising can include the number of times the milk advertisement is clicked, the number of times it is viewed, the publication time of the milk advertisement, and the brand, etc. The media characteristics of lipstick can include the lipstick color, the target audience, the publication time, the number of copies sold, and the number of times it is favorited, etc.
[0058] It should be noted that valid media triples can be actual triples from the sample multimedia data; that is, valid media triples can be generated based on the first media features and the knowledge graph, and valid media triples are correct. Invalid media triples, on the other hand, can refer to triples that are irrelevant to the sample multimedia data. Invalid media triples can be randomly generated based on valid media triples; that is, invalid media triples are incorrect triples.
[0059] It should be noted that a sample multimedia data set has one or more valid media triples, and a valid media triple can correspond to one or more invalid media triples. This application mainly uses the example of a sample multimedia data set having one valid media triple as an example. In this application, all N invalid media triples correspond to the same valid media triple. For implementation methods where the sample multimedia data set has multiple valid media triples, please refer to the implementation methods illustrated in this application.
[0060] S104. Using the initial media recommendation model, based on the first association semantics within the effective media triples, determine the first distance between the first media feature and the first attribute profile label. Based on the first association semantics within each invalid media triple, determine the second distance between the second media feature and the second attribute profile label within each invalid media triple.
[0061] In this application, a computer device can input valid media triples and invalid media triples into an initial media recommendation model, and output the media recommendation model to determine a first distance between a first media feature and a first attribute profile label based on the first association semantics within the valid media triples, and to determine a second distance between a second media feature and a second attribute profile label within each invalid media triple based on the first association semantics within the invalid media triples.
[0062] The first distance reflects the distance between the first media feature and the first attribute profile label in the vector space, i.e., it reflects the clustering between the first media feature and the first attribute profile label. A shorter first distance indicates a greater clustering between the first media feature and the first attribute profile label. To cluster more multimedia data samples under the same attribute profile label, and to bring long-tail and popular multimedia data under the same attribute profile label to the same central point, the initial media recommendation model can be trained to shorten the first distance between the first media feature and the first attribute profile label.
[0063] Similarly, the second distance reflects the distance between the second media feature and the second attribute profile label in the vector space, that is, it reflects the clustering between the second media feature and the second attribute profile label. A shorter second distance indicates a greater clustering between the second media feature and the second attribute profile label. Since the second media feature and the second attribute profile label belong to invalid media triples, meaning there is no correlation between them, the initial media recommendation model can be trained to increase the second distance between them.
[0064] S105. The first media feature and the object feature are identified by the initial media recommendation model to obtain the predicted interest degree of the sample object for the sample multimedia data. The initial media recommendation model is iteratively trained according to the predicted interest degree, the labeled interest degree, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping iterative training condition, and the target media recommendation model for recommending multimedia data is obtained.
[0065] In this application, a computer device can input the first media feature and the object feature into an initial media recommendation model. The initial media recommendation model then identifies the first media feature and the object feature to obtain the predicted interest level of the sample object in the sample multimedia data. There is a positive correlation between the first distance corresponding to the effective media triple and the fit between the first associated semantics, the first media feature, and the first attribute profile label. That is, the larger the first distance, the lower the fit between the first associated semantics, the first media feature, and the first attribute profile label. This indicates a poor clustering effect between the first media feature and the first attribute profile label. The initial media recommendation model fails to accurately cluster long-tail multimedia data and popular multimedia data under the same attribute profile label towards the same central point. In other words, the initial media recommendation model fails to enable long-tail multimedia data to learn more information from popular multimedia data. Conversely, the smaller the first distance, the higher the fit between the first associated semantics, the first media feature, and the first attribute profile label. In other words, the aggregation effect between the first media feature and the first attribute profile label is better. The initial media recommendation model can accurately aggregate long-tail multimedia data and popular multimedia data under the same attribute profile label to the same central point. It can also be said that the initial media recommendation model can enable long-tail multimedia data to learn more information from popular multimedia data.
[0066] Similarly, there is a positive correlation between the second distance corresponding to the invalid media triple and the fit between the first associated semantics, the second media feature, and the second attribute profile label. That is, the larger the second distance, the lower the fit between these three elements, indicating a poorer clustering effect between the second media feature and the second attribute profile label. In this case, the initial media recommendation model can accurately move the second media feature away from the second attribute profile label and away from the first attribute profile label. Conversely, the smaller the second distance, the higher the fit between these elements, indicating a better clustering effect between the second media feature and the second attribute profile label. In this case, the initial media recommendation model cannot accurately move the second media feature away from the second attribute profile label and away from the first attribute profile label.
[0067] Furthermore, the computer device can iteratively train the initial media recommendation model based on the predicted interest, the labeled interest, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping condition for iterative training, thus obtaining the target media recommendation model for recommending multimedia data. This helps long-tail multimedia data learn information about popular multimedia data under the same attribute profile label, thereby improving the learning effect of the initial media recommendation model on the media features of long-tail multimedia data and solving the sparsity problem of long-tail multimedia data.
[0068] In one embodiment, obtaining the target media recommendation model for recommending multimedia data until the initial media recommendation model after iterative training meets the stopping iterative training condition includes: the computer device can determine whether the convergence state of the initial media recommendation model after iterative training is converged; when the convergence state of the initial media recommendation model after iterative training is converged, it indicates that the recommendation accuracy of the initial media recommendation model after iterative training is relatively high, and it can be determined that the initial media recommendation model after iterative training meets the stopping iterative training condition. Alternatively, the computer device can count the number of iterations of the initial media recommendation model; when the number of iterations of the initial media recommendation model is greater than a threshold, it is determined that the initial media recommendation model after iterative training meets the stopping iterative training condition; the initial media recommendation model after iterative training that meets the stopping iterative training condition is determined as the target media recommendation model for recommending multimedia data.
[0069] In this context, a converged state for the initial media recommendation model after iterative training can be defined as the total predicted loss of the initial media recommendation model being less than a loss threshold; a non-converged state can be defined as the total predicted loss of the initial media recommendation model being greater than or equal to the loss threshold. The loss threshold can be the minimum value of the loss function of the initial media recommendation model, or it can be determined based on the requirements of the application scenario of the trained target media recommendation model. The aforementioned number of iterations threshold can be determined based on the requirements of the application scenario of the target media recommendation model, or it can be determined based on the equipment resources of the computer device.
[0070] In this model, there can be P samples of multimedia data. These P samples can be divided into m batches, with each batch containing M = P / m samples. Updating the parameters of the initial media recommendation model based on one batch of multimedia data completes one iteration, also known as one iteration of training. When the parameters of the initial media recommendation model are updated based on m batches of multimedia data, it is considered that the initial media recommendation model has completed one epoch.
[0071] This application increases the number of samples under the same profile attribute label, reducing the training complexity of the initial media recommendation model. Simultaneously, it facilitates long-tail multimedia data learning information from popular multimedia data under the same attribute profile label, thereby improving the initial media recommendation model's learning performance on media features of long-tail multimedia data and addressing the sparsity problem of long-tail multimedia data. Furthermore, during training, a single initial media recommendation model is shared for both long-tail and popular multimedia data, eliminating the need to create separate network models for each type of data, thus reducing training complexity and improving training efficiency.
[0072] Further, please see Figure 4 This is a flowchart illustrating a media data processing method provided in an embodiment of this application. Figure 4 As shown, this method can be derived from... Figure 1 It can be executed by any terminal in the terminal cluster, or by... Figure 1 The server in the middle can be used to execute it, or it can be executed by... Figure 1 The terminal cluster in the application uses terminals and servers to collaboratively execute the media data processing method. The device used to execute this method can be collectively referred to as a computer device. The method may include the following steps:
[0073] S201. Obtain the object features of the sample object, the first media features of the sample multimedia data, and the labeled interest degree of the sample object in relation to the sample multimedia data.
[0074] S202. Based on the knowledge graph associated with the first media feature and the sample multimedia data, generate a valid media triple associated with the sample multimedia data; the valid media triple includes the first media feature, the first attribute profile label, and the first association semantics, the first association semantics reflecting the association relationship between the first media feature and the first attribute profile label.
[0075] S203. Based on the above valid media triples, generate N invalid media triples that are not associated with the above sample multimedia data. Each invalid media triple includes a second media feature, a second attribute profile label, and the above first association semantics. There is no association relationship between the above second media feature and the above second attribute profile label. N is a positive integer.
[0076] In one embodiment, the aforementioned N invalid media triples include a first invalid media triplet and a second invalid media triplet. Generating N invalid media triples unrelated to the aforementioned sample multimedia data based on the aforementioned valid media triplets includes: a computer device replacing the first media feature within the aforementioned valid media triplets with a second media feature unrelated to the aforementioned sample multimedia data to obtain a first invalid media triplet; and replacing the first attribute profile label within the aforementioned valid media triplets with a second attribute profile label unrelated to the aforementioned sample multimedia data to obtain a second invalid media triplet; wherein the second attribute profile label within the aforementioned first invalid media triplet is the same as the first attribute profile label; and the second media feature within the aforementioned second invalid media triplet is the same as the first media feature.
[0077] It should be noted that a valid media triple can be represented as: (first media feature, first related semantic, first attribute profile label); a first invalid media triple can be represented as: (second media feature, first related semantic, first attribute profile label); and a second invalid media triple can be represented as: (first media feature, first related semantic, second attribute profile label).
[0078] In this context, the second media feature in the first invalid media triplet is erroneous information. Therefore, the first invalid media triplet can be used to control the initial media recommendation model to keep the second media feature away from the first attribute profile label. In other words, the first invalid media triplet is used to control the initial media recommendation model to keep the erroneous second media feature away from the first attribute profile label. That is, the first invalid media triplet is used to control the initial media recommendation model to keep the second media feature of the sample multimedia data that does not have the first media feature away from the first attribute profile label, and to keep the second media feature away from the first attribute profile that is not associated with it.
[0079] Similarly, the second attribute profile label in the second invalid media triple is incorrect information. Therefore, the second invalid media triple can be used to control the initial media recommendation model to keep the first media feature away from the second attribute profile label. That is, the second invalid media triple is used to control the initial media recommendation model to keep the first media feature away from the incorrect second attribute profile label. In other words, the second invalid media triple is used to control the initial media recommendation model to keep the first media feature of the sample multimedia data that does not have the first attribute profile label away from the second attribute profile label, and to keep the first media feature away from the second attribute profile label that is not associated with it.
[0080] S204. Using the aforementioned main recommendation network, a first media embedding vector reflecting the first media feature and a second media embedding vector reflecting the second media feature are generated. This initial media recommendation model includes a main recommendation network and an auxiliary recommendation network;
[0081] In this application, the initial media recommendation model may include a main recommendation network and an auxiliary recommendation network. The main recommendation network may be a network used for online prediction, and the auxiliary recommendation network may be a network used to assist in training the main recommendation network. During the training process of the initial media recommendation model, the computer device can use the main recommendation network to identify the first media feature to obtain a first media embedding vector reflecting the first media feature, and identify the second media feature to obtain a second media embedding vector reflecting the second media feature.
[0082] S205. Through the above-mentioned auxiliary recommendation network, a first association embedding vector reflecting the first association semantics within the above-mentioned effective media triples is generated, and a second association embedding vector reflecting the first association semantics within each of the above-mentioned invalid media triples is generated.
[0083] In this application, a computer device can use the aforementioned auxiliary recommendation network to query a first association embedding vector reflecting the first association semantics within the effective media triplet from the association semantic table, based on the first association semantics, and generate a second association embedding vector reflecting the first association semantics within each of the aforementioned invalid media triplets.
[0084] The association semantic table includes association embedding vectors corresponding to various association semantics. Each association embedding vector includes d-dimensional association feature values, which are real numbers belonging to the range [-1, 1].
[0085] S206. Through the above-mentioned auxiliary recommendation network, a first tag embedding vector reflecting the first attribute profile tag is generated, and a second tag embedding vector reflecting the second attribute profile tag in each of the above-mentioned invalid media triples is generated.
[0086] In this application, a computer device can use the auxiliary recommendation network to query a first tag embedding vector reflecting the first attribute profile tag from the attribute profile tag table based on the first attribute profile tag, and generate a second tag embedding vector reflecting the second attribute profile tag in each of the invalid media triples.
[0087] The attribute profile label table includes label embedding vectors corresponding to various attribute profile labels. Each label embedding vector includes d-dimensional label feature values, which are real numbers belonging to the range [-1, 1].
[0088] S207. Based on the first association embedding vector, the first media embedding vector, and the first tag embedding vector, determine the first distance between the first media feature and the first attribute profile tag. Based on the second association embedding vector, the second media embedding vector corresponding to each invalid media triplet, and the second tag embedding vector, determine the second distance between the second media feature and the second attribute profile tag within each invalid media triplet.
[0089] In one embodiment, the computer device can obtain the first distance between the first media feature and the first attribute profile label, and the second distance between the second media feature and the second attribute profile label within each invalid media triplet, through either of the following two implementation methods. In implementation method one, step S207 may include the following steps S11 to S14:
[0090] S11. Summing the first media embedding vector and the first association embedding vector to obtain the first distribution embedding vector from the first media feature to the first attribute portrait label.
[0091] Specifically, the first distribution embedding vector is used to reflect the distribution of the first media embedding vector in the space centered on the first attribute portrait label, such as the distribution position.
[0092] S12. Based on the difference between the first distribution embedding vector and the first tag embedding vector, determine the first distance between the first media feature and the first attribute portrait tag.
[0093] Specifically, the first distance can be used to reflect whether the first media embedding vectors are clustered near the first center point, which is the first attribute profile label. The smaller the first distance, the more clustered the first media embedding vectors (first media features) are near the first center point; conversely, the larger the first distance, the further away the first media embedding vectors are from the first center point. Therefore, iteratively training the initial media recommendation model using the first distance helps it cluster the media embedding vectors of sample multimedia data under the same attribute profile label near the same center point. This also helps the media embedding vectors of long-tail multimedia data under the same attribute profile label learn information from the media embedding vectors of popular multimedia data, improving the initial media recommendation model's learning effect on the media embedding vectors of long-tail multimedia data and solving the sparsity problem of long-tail multimedia data.
[0094] S13. Summing the second associated embedding vector with the second media embedding vector corresponding to each invalid media triple to obtain the second distribution embedding vector from the second media feature to the second attribute profile label in each invalid media triple.
[0095] Specifically, the second distribution embedding vector is used to reflect the distribution of the second media embedding vector in the space centered on the second attribute portrait label, such as the distribution location.
[0096] S14. Based on the difference between the second distribution embedding vector and the second label embedding vector corresponding to each invalid media triple, determine the second distance between the second media feature and the second attribute profile label in each invalid media triple.
[0097] Specifically, this second distance can be used to reflect whether the second media embedding vectors are clustered near the second center point, which is either the first attribute profile label or the second attribute profile label. That is, the smaller the second distance, the more clustered the second media embedding vectors (second media features) are near the second center point; conversely, the larger the second distance, the further away the second media embedding vectors are from the second center point. Therefore, iteratively training the initial media recommendation model using the second distance helps to move the media embedding vectors of multimedia samples without a certain attribute profile label away from the center point corresponding to that attribute profile label. This increases the distinguishability between the media embedding vectors of multimedia samples with different attribute profile labels, avoids clustering the media embedding vectors of multimedia samples with different attribute profile labels towards the same center point, and improves the training accuracy of the initial recommendation model.
[0098] It should be noted that after the embedding vectors corresponding to the first media feature, the second media feature, the first attribute portrait label, the second attribute portrait label, and the first associated semantics are respectively, the computer device can perform knowledge graph reasoning. Knowledge graph reasoning includes the above step S207. The purpose of knowledge graph reasoning is to constrain the vector distribution of entities and relationships in the knowledge graph to evolve towards the expected set constraint conditions, thereby achieving the effect of optimizing the media embedding vectors of the sample multimedia data.
[0099] For different sample multimedia data, the following situation may occur: When different sample media data are used as header entities, the effective media triples of the sample multimedia data are embedded to obtain embedded effective media triples. Here, embedding can refer to replacing the corresponding parameters with the embedding vectors corresponding to each parameter in the effective media triples. The embedded effective media triples can be represented as: (h i ,r, ij ), h i Let r be the media embedding vector of the i-th sample multimedia data, r be the first association embedding vector corresponding to the first association semantic, and t be the media embedding vector of the i-th sample multimedia data. ij This is the label embedding vector corresponding to the j-th attribute profile label of the i-th sample multimedia data. When multiple different samples of multimedia data share the same attribute profile label t... ij When an association occurs, the goal of knowledge graph reasoning is: in the vector space, to identify entity t with attribute profiles and labels. ij Centered on this point, the media embedding vectors corresponding to all associated multimedia data samples should cluster around this center point, while not intersecting with the attribute profile label t. ij The media embedding vectors corresponding to other associated sample multimedia data should be far away from the attribute profile label entity t. ij The corresponding center point.
[0100] During the modeling process, popular multimedia data, due to its sufficient sample size, will dominate the attribute profile label entity t. ij The corresponding center point's position in the vector space is adjusted. During the iterative training of the initial recommendation model, the media embedding vector (i.e., the first media embedding vector) of the long-tail sample multimedia data will also continuously move towards the attribute profile label entity t. ij The proximity of corresponding center points is beneficial to the media embedding vectors of long-tail multimedia data under the same attribute profile label. Information can be learned from the media embedding vectors of popular multimedia data, injecting more information into the media embedding vectors of long-tail multimedia data, thereby achieving optimization.
[0101] As described above, the reasoning process of knowledge graphs is based on the effective media triples of constrained embedding (h) i ,r, ijAs the distance between two entities in the vector space decreases, how can we quantify this process? We can use r as a mapping from the first media feature (i.e., the first media embedding vector) to the first attribute profile label (the first label embedding vector), so that h... i +r should be as close to t as possible ij They are equal, which can be expressed by formula (1) (h) i ,,t ij The relationship between the three:
[0102] h i +r≈t ij (1)
[0103] Based on formula (1), the computer device can use the following formula (2) to calculate the first distance between the first media feature and the first attribute portrait label:
[0104] E(h i ,r, ij ) = ||h i + rt ij || (2)
[0105] In formula (2), E(h) i ,r, ij E(h) represents the first distance between the first media feature and the first attribute profile label of the i-th sample multimedia data, and ∥∥ represents the calculation of the modulus length. i ,,t j The larger ) is, the more h i ,r, ij The higher the compatibility between them; conversely, the lower the compatibility, the better. i ,r, ij The smaller h is i ,,t ij The lower the compatibility between them.
[0106] Wherein, the embedded first invalid media triplet corresponding to the first invalid media triplet of the i-th sample multimedia data is: (h i ′ ,r, ij The computer device can use the following formula (3) to calculate the second distance between the second media feature in the first invalid media triplet and the second attribute portrait label:
[0107] E(h i i ,r, ij ) = ||h i ′ + rt ij || (3)
[0108] In formula (3), E(h) i ′ ,,t ij ) represents the second distance between the second media feature within the first invalid media triplet and the second attribute profile label. The embedded second invalid media triplet corresponding to the second invalid media triplet of the i-th sample multimedia data is: (h i ,,t i ′ j The computer device can use the following formula (4) to calculate the second distance between the second media feature in the second invalid media triplet and the second attribute portrait label:
[0109] E(h i , ,t i ′ j ) = ||h i + rr i ′ j || (4)
[0110] In formula (4), E(h) i ,, i ′ j ) represents the second distance between the second media feature within the second invalid media triple and the second attribute profile label.
[0111] In implementation method two, step S207 above may include the following steps S21 to S23:
[0112] S21. Multiply the transformation embedding vector corresponding to the first associated semantic with the first media embedding vector to obtain the first media transformation vector. Multiply the first tag embedding vector with the transformation embedding vector to obtain the first tag transformation vector.
[0113] Specifically, the auxiliary recommendation network can generate a transformed embedding vector corresponding to the first associated semantics based on the first associated semantics. This transformed embedding vector is used to map the media embedding vector and the label vector into the same vector space so that the distance between them can be calculated subsequently. The computer device can multiply the transformed embedding vector corresponding to the first associated semantics with the first media embedding vector to obtain the first media transformed vector, and multiply the first label embedding vector with the transformed embedding vector to obtain the first label transformed vector. That is, the first media transformed vector and the first label transformed vector belong to the same vector space.
[0114] S22. Multiply the above-mentioned transformation embedding vector with the above-mentioned second media embedding vector corresponding to each of the above-mentioned invalid media triplets to obtain the second media transformation vector corresponding to each of the above-mentioned invalid media triplets. Multiply the above-mentioned second tag embedding vector corresponding to each of the above-mentioned invalid media triplets with the above-mentioned transformation embedding vector to obtain the second tag transformation vector of each of the above-mentioned invalid media triplets.
[0115] In this context, the second media transformation vector corresponding to each invalid media triplet belongs to the same vector space as the second tag transformation vector.
[0116] S23. Based on the first association embedding vector, the first media transformation vector, and the first tag transformation vector, determine the first distance between the first media feature and the first attribute profile tag. Based on the second association embedding vector, the second media transformation vector corresponding to each invalid media triplet, and the second tag transformation vector, determine the second distance between the second media feature and the second attribute profile tag within each invalid media triplet.
[0117] Specifically, the computer device determines the first distance between the first media feature and the first attribute portrait label using the converted vector, and determines the second distance between the second media feature and the second attribute portrait label within each invalid media triplet, which helps to improve the accuracy of the calculation.
[0118] In one embodiment, step S23 may include the following steps S231 to S234:
[0119] S231. The first association embedding vector and the first media transformation vector are summed to obtain the third distribution embedding vector from the first media feature to the first attribute portrait label.
[0120] Specifically, the third distribution embedding vector is used to reflect the distribution of the first media embedding vector in the space centered on the first attribute portrait label.
[0121] S232. Based on the difference between the third distribution embedding vector and the first tag transformation vector, the first distance between the first media feature and the first attribute portrait tag is determined.
[0122] Specifically, the computer device can perform a modulo operation on the difference between the third distribution embedding vector and the first tag transformation vector to obtain the first distance between the first media feature and the first attribute portrait tag. The first distance reflects the distribution distance between the first media feature and the first attribute portrait tag in the vector space.
[0123] S233. Summing the second association embedding vector and the second media transformation vector corresponding to each invalid media triplet, to obtain the fourth distribution embedding vector from the second media feature to the second attribute portrait label in each invalid media triplet.
[0124] Specifically, the fourth distribution embedding vector is used to reflect the distribution of the second media embedding vector in the space centered on the second attribute portrait label.
[0125] S234. Based on the difference between the fourth distribution embedding vector and the second label transformation vector corresponding to each of the above invalid media triples, determine the second distance between the second media feature and the second attribute profile label in each of the above invalid media triples.
[0126] Specifically, the computer device can perform a modulo operation on the difference between the fourth distribution embedding vector and the second label transformation vector corresponding to each invalid media triple to obtain the second distance between the second media feature and the second attribute portrait label in each invalid media triple; the second distance reflects the distribution distance of the second media feature and the second attribute portrait label in the vector space.
[0127] It should be noted that for the same sample multimedia data, the following situation may occur: the number of valid multimedia triples associated with the same sample multimedia data is multiple. Among these valid media triples, different valid media triples have different first attribute profile labels and first association semantics. The sample multimedia data has multiple different first attribute profile labels, and the differences between the different first attribute profile labels are relatively large. This is reflected in the modeling process, which will roughly show that the distribution of the first media features and first association semantics of the sample multimedia data in the vector space is relatively different, which increases the modeling difficulty.
[0128] Based on this, computer devices can use fully connected network layers to obtain different transformed embedding vectors according to different first association semantics. By transforming the embedding vectors, the first media embedding vector and the corresponding first tag embedding vector are mapped to a dedicated vector space to measure the distance. Different first association semantics will map the corresponding first media embedding vector and the corresponding first tag embedding vector to different vector spaces.
[0129] Specifically, the computer device can calculate the first distance between the first media feature and the first attribute portrait label based on the transformed embedding vector and the following formula (5):
[0130] E(h i , , t ij ) = || r h i+ r - r t ij || (5)
[0131] Among them, T r It can be the transformed embedding vector corresponding to the association semantics reflecting the relationship between the i-th sample multimedia data and the j-th attribute portrait label. This transformed embedding vector can be calculated using the following formula (6):
[0132] T r = Tanh(W r +b) (6)
[0133] Where Tanh represents the activation function, and W and b are the trainable parameter matrices of the fully connected network layer.
[0134] Optionally, the computer device may use the following formula (7) to calculate the second distance between the second media feature within the first invalid media triplet and the second attribute profile label:
[0135] E(h i ′ ,r, ij ) = || r h i ′ + rT r t ij || (7)
[0136] Optionally, the computer device may use the following formula (8) to calculate the second distance between the second media feature within the second invalid media triple and the second attribute profile label:
[0137] E(h i , ,t i ′ j ) = || r h i + rT r t i ′ j || (8)
[0138] S208. The first media feature and the object feature are identified by the initial media recommendation model to obtain the predicted interest degree of the sample object for the sample multimedia data. The initial media recommendation model is iteratively trained according to the predicted interest degree, the labeled interest degree, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping iterative training condition, and the target media recommendation model for recommending multimedia data is obtained.
[0139] In one example, the iterative training of the initial media recommendation model based on the predicted interest, the labeled interest, the first distance, and the second distances corresponding to the N invalid media triples includes: determining the interest prediction loss of the initial media recommendation model based on the predicted interest and the labeled interest; determining the feature recognition loss of the initial media recommendation model based on the first distance and the second distances corresponding to the N invalid media triples; and iteratively training the initial media recommendation model based on the interest prediction loss and the feature recognition loss.
[0140] Specifically, the computer device can determine the interest prediction loss of the initial media recommendation model based on the predicted interest level and the labeled interest level. This interest prediction loss reflects the accuracy of the initial media recommendation model in predicting the interest level of the sample objects. That is, the larger the interest prediction loss, the lower the accuracy of the initial media recommendation model in predicting the interest level of the sample objects; conversely, the smaller the interest prediction loss, the higher the accuracy of the initial media recommendation model in predicting the interest level of the sample objects. The computer device can determine the feature recognition loss of the initial media recommendation model based on the first distance and the second distance corresponding to the N invalid media triples. The feature recognition loss constrains the initial multimedia recommendation model to reduce the distance between the media features (i.e., media embedding vectors) of sample multimedia data under the same attribute profile label and their associated attribute profile labels, and to increase the distance between the media features of sample multimedia data and their unassociated attribute profile labels. Therefore, the computer device can iteratively train the initial media recommendation model based on the interest prediction loss and the aforementioned feature recognition loss to improve the recommendation accuracy of the initial media recommendation model.
[0141] For example, the labeled and predicted interest levels of the sample multimedia data can be substituted into the loss function of the initial media recommendation model to obtain the interest prediction loss of the initial media recommendation model. This loss function can be any of the following: cross-entropy loss, mean squared error loss, divergence loss, etc. Taking the cross-entropy loss function as an example, during the model training phase, the cross-entropy loss function is used to measure the difference between the predicted and labeled interest levels of the initial media recommendation model, guiding the parameter optimization of the initial media recommendation model.
[0142] Specifically, the computer device can use the following formula (9) to calculate the interest prediction loss of the initial media recommendation model:
[0143]
[0144] Among them, in formula (9) Let M be the interest prediction loss of the initial media recommendation model during one iteration of training, and M be the number of multimedia data samples in one iteration of training for the initial media recommendation model. i Let the labeled interest level be the multimedia data of the i-th sample. Let be the predicted interest level of the multimedia data for the i-th sample.
[0145] In one example, determining the feature recognition loss of the initial media recommendation model based on the first distance and the second distances corresponding to the N invalid media triples includes: subtracting the first distance from the second distances corresponding to the N invalid media triples to obtain distance differences for each of the N invalid media triples; summing the distance differences for each of the N invalid media triples to obtain a total distance difference. The feature recognition loss of the initial media recommendation model is then determined based on this total distance difference.
[0146] Specifically, the computer device can use the following formula (10) to calculate the feature recognition loss of the initial media recommendation model:
[0147]
[0148] In formula (10), the feature recognition loss of the initial media recommendation model during one iteration of training is defined. That is, the initial media recommendation model needs to reduce the E(h) corresponding to the effective media triples. i ,r,t′ ij ), expanding the E(h) corresponding to the invalid media triple. i ,r,t ij ), E(h i ′,r,t ijIn other words, formula (10) is used to constrain the initial media recommendation model to reduce the distance between the media features (media embedding vectors) of sample multimedia data under the same attribute profile label and their associated attribute profile labels. For example, it can reduce the distance between the media features of long-tail media data and the media features of popular media data under the same attribute profile label. This is equivalent to injecting the media features of popular media data under the same attribute profile label into the media features of long-tail media data. That is, it is equivalent to the media features of long-tail media data learning information from the media features of popular media data under the same attribute profile label, improving the learning effect of the initial media recommendation model for long-tail multimedia data and solving the sparsity problem of long-tail multimedia data. Formula (10) is also used to constrain the initial media recommendation model to expand the distance between media features and their unrelated attribute profile labels, avoiding the aggregation of media features of sample multimedia data under different attribute profile labels to the same center point, improving the representational distinctiveness of media features of sample multimedia data under different attribute profile labels, and improving the accuracy of model training.
[0149] In one embodiment, the iterative training of the initial media recommendation model based on the interest prediction loss and the feature recognition loss includes: summing the interest prediction loss and the feature recognition loss to obtain the total prediction loss of the initial media recommendation model; determining the convergence state of the initial media recommendation model based on the total prediction loss. When the convergence state of the initial media recommendation model is non-converged, adjusting the parameters of the initial media recommendation model based on the total prediction loss to obtain the iteratively trained initial media recommendation model.
[0150] Specifically, the computer device can sum the aforementioned interest prediction loss and feature recognition loss to obtain the total predicted loss of the initial media recommendation model. This summation can include cumulative addition or weighted summation. When the total predicted loss is less than a loss threshold, the initial media recommendation model is determined to be converged, and it can be designated as the target media recommendation model. When the total predicted loss is greater than or equal to the loss threshold, the initial media recommendation model is determined to be non-converged. Based on the total predicted loss, the parameters of the initial media recommendation model are adjusted to obtain the iteratively trained initial media recommendation model. Here, the computer device can adjust the parameters of the main recommendation network and auxiliary recommendation network of the initial media recommendation model according to the gradient direction propagation based on the total predicted loss to obtain the iteratively trained initial media recommendation model.
[0151] For example, a computer device can use the following formula (11) to calculate the total prediction loss of the initial media recommendation model:
[0152]
[0153] Among them, in formula (11) β is the total predicted loss of the initial media recommendation model during one iteration of training, and β is a hyperparameter used to control the contribution of the two losses. The value of β can be an empirical value.
[0154] In one embodiment, a computer device can receive a multimedia data recommendation request for a target object. This request includes the object features of the target object, and may also include media type and media keywords. Based on the multimedia data recommendation request, the computer device can retrieve the media features of candidate multimedia data to be recommended from a database. Through the main recommendation network of the target media model, it identifies the media features of the candidate multimedia data and the object features of the target object, obtaining the target object's interest level in the candidate multimedia data. Based on this interest level, the candidate multimedia data is recommended to the target object. The target media recommendation model trained as described above has the ability to learn information from the media features of popular multimedia data, specifically targeting long-tail multimedia data. Therefore, when the candidate multimedia data is long-tail multimedia data, the target media recommendation model can also recommend it to suitable users, improving the accuracy of multimedia data recommendations.
[0155] In one embodiment, the method may further include the following step S31
[0156] S31. From the knowledge graph, query the valid tag triples associated with the multimedia data of the sample; the valid tag triples include the third attribute profile tag, the fourth attribute profile tag and the second association semantics, the second association semantics reflecting the association relationship between the third attribute profile tag and the fourth attribute profile tag.
[0157] Here, the third attribute portrait tag is a sub-attribute portrait tag of the fourth attribute portrait tag. The third attribute portrait tag here can be the first attribute portrait tag mentioned earlier, or it can be any attribute portrait tag mentioned earlier other than the first attribute portrait tag. The second association semantic here reflects the subordinate relationship between the third attribute portrait tag and the fourth attribute portrait tag.
[0158] Specifically, computer devices can query the knowledge graph for valid tag triples associated with the sample multimedia data based on the first media features of the sample multimedia data.
[0159] S32. Based on the valid label triplet, generate M invalid label triplets that are not associated with the sample multimedia data. Each invalid label triplet includes a fifth attribute profile label, a sixth attribute profile label, and the second association semantics. There is no association relationship between the fifth attribute profile label and the sixth attribute profile label. M is a positive integer.
[0160] Specifically, taking M=2 as an example, the M invalid label triplets include the first invalid label triplet and the second invalid label triplet. The computer device can use a fifth attribute profile label that is not associated with the sample multimedia data and the third attribute profile label in the valid label triplet to obtain the first invalid label triplet. Then, it can use a sixth attribute profile label that is not associated with the sample multimedia data to replace the fourth attribute profile label in the valid label triplet to obtain the second invalid label triplet. That is, the sixth attribute profile label in the first invalid label triplet is the same as the fourth attribute profile label in the valid label triplet, and the fifth attribute profile label in the second invalid label triplet is the same as the third attribute profile label in the valid label triplet.
[0161] It should be noted that a valid tag triple can be represented as: (third attribute profile tag, second association semantic, fourth attribute profile tag), a first invalid tag triple can be represented as: (fifth attribute profile tag, second association semantic, fourth attribute profile tag), and a second invalid tag triple can be represented as: (third attribute profile tag, second association semantic, sixth attribute profile tag).
[0162] S33. Using the initial media recommendation model, based on the second association semantics within the effective tag triplet, determine the third distance between the third attribute profile tag and the fourth attribute profile tag, and based on the second association semantics within each invalid tag triplet, determine the fourth distance between the fifth attribute profile tag and the sixth attribute profile tag within each invalid tag triplet.
[0163] Specifically, the computer device can use an auxiliary recommendation network to generate embedding vectors reflecting the third attribute profile label, the fourth attribute profile label, the fifth attribute profile label, the sixth attribute profile label, and the second related semantics, respectively. In formula (2) or formula (5), the embedding vector corresponding to the third attribute profile label is used to replace h. i The fourth attribute portrait tag corresponding to the embedding vector replacement t ij The embedding vector r corresponding to the second associated semantic is replaced, and the third distance between the third attribute portrait label and the fourth attribute portrait label is calculated.
[0164] Similarly, in formula (3) or formula (7) above, the embedding vector corresponding to the fifth attribute portrait tag in the first invalid tag triplet is used to replace h. i ′ The embedding vector corresponding to the fourth attribute portrait label is replaced with t. ij The embedding vector r corresponding to the second associated semantics is replaced, and the fourth distance between the fifth attribute portrait label and the sixth attribute portrait label in the first invalid label triplet is calculated.
[0165] Similarly, in formula (4) or formula (8) above, the embedding vector corresponding to the third attribute portrait tag is used to replace h. i The embedding vector corresponding to the sixth attribute portrait tag is replaced with t. i ′ j The embedding vector r corresponding to the second associated semantics is replaced, and the fourth distance between the fifth attribute portrait label and the sixth attribute portrait label in the first invalid label triplet is calculated.
[0166] S34. Based on the predicted interest, the labeled interest, the first distance, the second distance and the third distance corresponding to the N invalid media triples, and the fourth distance corresponding to the M invalid label triples, the initial media recommendation model is iteratively trained.
[0167] Specifically, the computer device can calculate the interest prediction loss of the initial media recommendation model according to the above formula (9), sum the second distance and the first distance corresponding to the N invalid media triples respectively to obtain the distance difference corresponding to the N invalid media triples respectively, and subtract the fourth distance and the third distance corresponding to the M invalid label triples respectively to obtain the distance difference corresponding to the M invalid label triples respectively. The distance differences corresponding to the N invalid media triples and the distance differences corresponding to the M invalid label triples are summed to obtain the total distance difference. Based on the total distance difference, the feature recognition loss of the initial media recommendation model is determined.
[0168] It should be noted that the initial media recommendation model in this application includes a main recommendation network and an auxiliary recommendation network, such as... Figure 5 As shown, the main recommendation network includes a feature embedding layer 51a, an object-side key feature extraction layer 52a, a media-side key feature extraction layer 53a, and a fully connected layer 56a; the auxiliary recommendation network includes a feature embedding layer 57a and a knowledge graph reasoning layer 58a.
[0169] Specifically, the feature embedding layer 51a can be used to identify the object features of the sample object to obtain an initial object embedding vector reflecting the object features, and to identify the media features of the sample multimedia data to obtain an initial media embedding vector reflecting the media features. The object layer key feature extraction layer 52a is used to extract key features from the initial object embedding vector to obtain the object embedding vector 54a of the sample object. The media layer key feature extraction layer 53a is used to extract key features from the initial media embedding vector to obtain the media embedding vector 55a of the sample multimedia data (the first media embedding vector, the second media embedding vector, etc. mentioned above).
[0170] Among them, the fully connected layer 56a can be used to predict the predicted interest degree of the sample object for the sample multimedia data based on the media embedding vector (i.e., the first media embedding vector and the second media embedding vector) and the object embedding vector, and determine the interest degree prediction loss of the initial media recommendation model based on the predicted interest degree and the labeled interest degree according to the main task objective function. The main task objective function can be any one of the cross-entropy loss function, mean squared error loss function, divergence loss function, etc. mentioned above.
[0171] Among them, the feature embedding layer 57a can be used to generate label embedding vectors (first label embedding vector, second label embedding vector) reflecting attribute profile labels, and to generate association embedding vectors (such as the first association embedding vector mentioned above) reflecting association semantics. The knowledge graph inference layer 58a can be used to determine the feature recognition loss of the initial media recommendation model based on the media embedding vector, association embedding vector, label embedding vector and auxiliary objective function. The auxiliary task objective function can refer to the formula (10) mentioned above.
[0172] It should be noted that the knowledge graph-driven auxiliary recommendation network proposed in this application is used to implement an auxiliary task, which is jointly modeled end-to-end with the main task of the recommendation system. The main task of the recommendation system can be predicting the predicted interest level of sample objects in relation to sample multimedia data, while the auxiliary task can be assisting the main recommendation network in generating media embedding vectors for the sample multimedia data. Each task has its own independent objective function, and the sum of the two objective functions serves as the learning objective of the entire initial media recommendation model. Furthermore, the method proposed in this application has good universality and can be combined with various main tasks of recommendation systems, such as click-through rate prediction models and conversion rate prediction models. The method proposed in this application optimizes the learning effect of media features of multimedia data in the main task through auxiliary modeling, and plays a particularly positive role in improving the recommendation effect for long-tail multimedia data.
[0173] Please see Figure 6 This is a schematic diagram of the structure of a media data processing device provided in an embodiment of this application.Figure 6 As shown, the media data processing device may include:
[0174] The acquisition module 611 is used to acquire the object features of the sample object, the first media features of the sample multimedia data, and the labeled interest degree of the sample object in relation to the sample multimedia data.
[0175] The first generation module 612 is used to generate a valid media triple associated with the sample multimedia data based on the knowledge graph associated with the first media feature and the sample multimedia data. The valid media triple includes the first media feature, the first attribute profile label and the first association semantics, and the first association semantics reflects the association relationship between the first media feature and the first attribute profile label.
[0176] The second generation module 613 is used to generate N invalid media triples that are not associated with the above sample multimedia data based on the above valid media triples. Each of the above invalid media triples includes a second media feature, a second attribute profile label and the above first association semantics. There is no association relationship between the above second media feature and the above second attribute profile label. N is a positive integer.
[0177] The determination module 614 is used to determine, through the initial media recommendation model, the first distance between the first media feature and the first attribute profile label based on the first association semantics within the effective media triplet, and the second distance between the second media feature and the second attribute profile label within each of the invalid media triplets based on the first association semantics within each of the invalid media triplets.
[0178] The training module 615 is used to identify the first media feature and the object feature through the initial media recommendation model, obtain the predicted interest degree of the sample object for the sample multimedia data, and iteratively train the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance, and the second distance corresponding to the N invalid media triples, until the initial media recommendation model after iterative training meets the stopping iterative training condition, and obtain the target media recommendation model for recommending multimedia data.
[0179] Optionally, the initial media recommendation model described above includes a main recommendation network and an auxiliary recommendation network;
[0180] The aforementioned determining module 614 is specifically used for:
[0181] Through the aforementioned main recommendation network, a first media embedding vector reflecting the first media feature is generated, and a second media embedding vector reflecting the second media feature is generated.
[0182] Through the above-mentioned auxiliary recommendation network, a first association embedding vector reflecting the first association semantics within the above-mentioned effective media triples is generated, and a second association embedding vector reflecting the first association semantics within each of the above-mentioned invalid media triples is generated.
[0183] Through the aforementioned auxiliary recommendation network, a first tag embedding vector reflecting the first attribute profile tag is generated, and a second tag embedding vector reflecting the second attribute profile tag within each of the aforementioned invalid media triples is generated.
[0184] Based on the first association embedding vector, the first media embedding vector, and the first tag embedding vector, a first distance between the first media feature and the first attribute profile tag is determined. Based on the second association embedding vector, the second media embedding vector corresponding to each invalid media triple, and the second tag embedding vector, a second distance between the second media feature and the second attribute profile tag within each invalid media triple is determined.
[0185] Optionally, the aforementioned determining module 614 is specifically used for:
[0186] Summing the first media embedding vector and the first association embedding vector yields the first distribution embedding vector from the first media feature to the first attribute profile label.
[0187] Based on the difference between the first distribution embedding vector and the first label embedding vector, the first distance between the first media feature and the first attribute profile label is determined.
[0188] Summing the second association embedding vector with the second media embedding vector corresponding to each invalid media triplet yields a second distribution embedding vector from the second media feature to the second attribute profile label within each invalid media triplet.
[0189] Based on the difference between the second distribution embedding vector and the second label embedding vector corresponding to each of the above invalid media triples, the second distance between the second media feature and the second attribute profile label in each of the above invalid media triples is determined.
[0190] Optionally, the aforementioned determining module 614 is specifically used for:
[0191] The first media embedding vector is multiplied by the first associated semantics to obtain the first media embedding vector, and the first tag embedding vector is multiplied by the first media embedding vector to obtain the first tag embedding vector.
[0192] The above-mentioned transformation embedding vector is multiplied with the above-mentioned second media embedding vector corresponding to each of the above-mentioned invalid media triplets to obtain the second media transformation vector corresponding to each of the above-mentioned invalid media triplets. The above-mentioned second tag embedding vector corresponding to each of the above-mentioned invalid media triplets is multiplied with the above-mentioned transformation embedding vector to obtain the second tag transformation vector corresponding to each of the above-mentioned invalid media triplets.
[0193] Based on the first association embedding vector, the first media transformation vector, and the first tag transformation vector, a first distance between the first media feature and the first attribute profile tag is determined. Based on the second association embedding vector, the second media transformation vector corresponding to each invalid media triple, and the second tag transformation vector, a second distance between the second media feature and the second attribute profile tag within each invalid media triple is determined.
[0194] Optionally, the aforementioned determining module 614 is specifically used for:
[0195] Summing the first association embedding vector and the first media transformation vector yields the third distribution embedding vector from the first media feature to the first attribute portrait label.
[0196] Based on the difference between the third distribution embedding vector and the first label transformation vector, the first distance between the first media feature and the first attribute portrait label is determined.
[0197] The second association embedding vector and the second media transformation vector corresponding to each invalid media triple are summed to obtain the fourth distribution embedding vector from the second media feature to the second attribute portrait label in each invalid media triple.
[0198] Based on the difference between the fourth distribution embedding vector and the second label transformation vector corresponding to each of the above invalid media triples, the second distance between the second media feature and the second attribute profile label in each of the above invalid media triples is determined.
[0199] Optionally, the aforementioned training module 615 is specifically used for:
[0200] Based on the predicted interest and the labeled interest, the interest prediction loss of the initial media recommendation model is determined.
[0201] Based on the first distance and the second distance corresponding to the N invalid media triples, the feature recognition loss of the initial media recommendation model is determined.
[0202] Based on the interest prediction loss and the feature recognition loss described above, the initial media recommendation model is iteratively trained.
[0203] Optionally, the aforementioned training module 615 is specifically used for:
[0204] The difference between the second distance and the first distance corresponding to the above N invalid media triples is calculated to obtain the distance difference values corresponding to the above N invalid media triples.
[0205] The total distance difference is obtained by summing the distance differences corresponding to the above N invalid media triples.
[0206] Based on the total distance difference mentioned above, the feature recognition loss of the initial media recommendation model is determined.
[0207] Optionally, the aforementioned training module 615 is specifically used for:
[0208] The summation of the interest prediction loss and the feature recognition loss yields the total prediction loss of the initial media recommendation model.
[0209] Based on the predicted total loss, the convergence state of the initial media recommendation model is determined.
[0210] When the initial media recommendation model is in a non-convergent state, the parameters of the initial media recommendation model are adjusted according to the predicted total loss to obtain the iteratively trained initial media recommendation model.
[0211] Optionally, the above N invalid media triples include the first invalid media triple and the second invalid media triple;
[0212] The first generation module 612 is specifically used to replace the first media feature in the effective media triplet with a second media feature that is not associated with the above-mentioned sample multimedia data to obtain a first invalid media triplet.
[0213] The first attribute profile tag in the above valid media triplet is replaced with the second attribute profile tag that is not associated with the above sample multimedia data to obtain the second invalid media triplet.
[0214] Among them, the second attribute profile tag in the first invalid media triplet is the same as the first attribute profile tag; the second media feature in the second invalid media triplet is the same as the first media feature.
[0215] Optional, training module 615, specifically used for:
[0216] From the knowledge graph above, query the valid tag triples associated with the multimedia data of the above samples; the valid tag triples include the third attribute profile tag, the fourth attribute profile tag and the second association semantics, the second association semantics reflecting the association relationship between the third attribute profile tag and the fourth attribute profile tag.
[0217] Based on the above valid label triplets, M invalid label triplets that are not associated with the above sample multimedia data are generated. Each invalid label triplet includes a fifth attribute profile label, a sixth attribute profile label, and the above second association semantics. There is no association relationship between the fifth attribute profile label and the sixth attribute profile label. M is a positive integer.
[0218] Based on the second association semantics within each of the effective tag triples, the third distance between the third attribute profile tag and the fourth attribute profile tag is determined using the initial media recommendation model. Based on the second association semantics within each of the invalid tag triples, the fourth distance between the fifth attribute profile tag and the sixth attribute profile tag within each of the invalid tag triples is determined.
[0219] The above-mentioned initial media recommendation model is iteratively trained based on the predicted interest, the labeled interest, the first distance, and the second distance corresponding to the N invalid media triples, including:
[0220] Based on the predicted interest, the labeled interest, the first distance, the second distance and the third distance corresponding to the N invalid media triples, and the fourth distance corresponding to the M invalid label triples, the initial media recommendation model is iteratively trained.
[0221] Optionally, the device further includes a recommendation module 616, which is used for:
[0222] Receive a multimedia data recommendation request for the target object; the multimedia data recommendation request includes the object characteristics of the target object.
[0223] Based on the aforementioned multimedia data recommendation request, the media features of the candidate multimedia data to be recommended are obtained from the database;
[0224] Through the main recommendation network of the target media model, the media features of the candidate multimedia data and the object features of the target object are identified to obtain the identification interest degree of the target object in the candidate multimedia data.
[0225] Based on the aforementioned interest levels, the aforementioned candidate multimedia data are recommended to the aforementioned target objects.
[0226] This application increases the number of samples under the same profile attribute label, reducing the training complexity of the initial media recommendation model. Simultaneously, it facilitates long-tail multimedia data learning information from popular multimedia data under the same attribute profile label, thereby improving the initial media recommendation model's learning performance on media features of long-tail multimedia data and addressing the sparsity problem of long-tail multimedia data. Furthermore, during training, a single initial media recommendation model is shared for both long-tail and popular multimedia data, eliminating the need to create separate network models for each type of data, thus reducing training complexity and improving training efficiency.
[0227] Please see Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 7 As shown, the aforementioned computer device 1000 can refer to a terminal or server, including: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the aforementioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. In some embodiments, the user interface 1003 may include a display screen (DcSPlay) and a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WC-FC interface). The memory 1005 may be a high-speed RAM or a non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 7 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and computer programs.
[0228] exist Figure 7 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the computer program stored in the memory 1005 to implement the steps in the various method embodiments of this application.
[0229] This application includes, but is not limited to, the following beneficial effects: (1) introducing a knowledge graph to characterize the association graph between sample multimedia data and its attribute profile tags. The sample multimedia data includes long-tail multimedia data and popular multimedia data. When long-tail multimedia data and popular multimedia data have the same attribute profile tags, both long-tail multimedia data and popular multimedia data are connected to the same attribute profile tag. That is, through the knowledge graph, the connection between long-tail multimedia data and popular multimedia data under the same attribute profile tag can be realized.
[0230] (2) Iteratively training the initial media recommendation model using the first distance corresponding to the effective media triplet helps to explicitly constrain the media features of sample multimedia data under the same attribute profile label to be as similar as possible. This ensures that the media features of long-tail multimedia data and popular multimedia data under the same attribute profile label (i.e., the first attribute profile label) cluster towards the same central point, which can be the first attribute profile label. This also makes the media features of long-tail multimedia data under the same attribute profile label closer to the media features of popular multimedia data, increasing the number of samples under the same profile attribute label and reducing the training complexity of the initial media recommendation model. At the same time, it helps long-tail multimedia data learn information from popular multimedia data under the same attribute profile label, thereby improving the learning effect of the initial media recommendation model on the media features of long-tail multimedia data and solving the sparsity problem of long-tail multimedia data. (3) Invalid media triplets refer to incorrect triplets. Iteratively training the initial media recommendation model using the second distance corresponding to the invalid media triplet helps to make the second media features and second attribute profile labels within the invalid media triplet move away from the central point corresponding to the first attribute profile label, which helps to improve the training accuracy of the model. (4) During the training process, a single initial media recommendation model is shared for both long-tail multimedia data and popular multimedia data. This eliminates the need to create separate network models for long-tail multimedia data and popular multimedia data, thereby reducing the training complexity of the model and improving the training efficiency.
[0231] It should be understood that the computer device described in the embodiments of this application can execute the media data processing method described in the corresponding embodiments above, and can also execute the media data processing apparatus described in the corresponding embodiments above, which will not be repeated here. In addition, the beneficial effects of using the same method will not be repeated here either.
[0232] In practice, the collection and processing of data in this application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the data subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the data subject.
[0233] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program executed by the aforementioned media data processing apparatus. This computer program includes program instructions, which, when executed by the processor, enable the execution of the media data processing method described in the corresponding embodiments above. Therefore, these descriptions will not be repeated here. Additionally, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments of this application, please refer to the description of the method embodiments of this application.
[0234] As an example, the above program instructions can be deployed and executed on a computer device, or deployed and executed on at least two computer devices in one location, or executed on at least two computer devices distributed in at least two locations and interconnected by a communication network. At least two computer devices distributed in at least two locations and interconnected by a communication network can form a blockchain network.
[0235] The aforementioned computer-readable storage medium may be a media data processing apparatus provided in any of the foregoing embodiments or a central storage unit of the aforementioned computer device, such as a hard disk or central storage unit of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium may include both the central storage unit and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0236] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish content in different media, rather than to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0237] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0238] This application also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the media data processing method and decoding method described in the preceding embodiments, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiments of this application.
[0239] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0240] The methods and related apparatus provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable network-connected device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable network-connected device, generate instructions for implementing the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable network-connected device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable network-connected device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.
[0241] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method of processing media data, the method comprising: The method comprises: obtaining an object feature of a sample object, a first media feature of sample multimedia data, and a labeled interest degree of the sample object for the sample multimedia data; generating an effective media triple associated with the sample multimedia data according to a knowledge graph associated with the first media feature and the sample multimedia data, wherein the effective media triple comprises the first media feature, a first attribute portrait label, and a first association semantic, and the first association semantic reflects that the first media feature and the first attribute portrait label have an association relationship; generating N ineffective media triples not associated with the sample multimedia data according to the effective media triple, wherein each ineffective media triple comprises a second media feature, a second attribute portrait label, and the first association semantic, the second media feature and the second attribute portrait label do not have the association relationship, and N is a positive integer; determining, by an initial media recommendation model, a first distance between the first media feature and the first attribute portrait label according to the first association semantic in the effective media triple, and determining a second distance between the second media feature and the second attribute portrait label in each ineffective media triple according to the first association semantic in each ineffective media triple; identifying, by the initial media recommendation model, the first media feature and the object feature to obtain a predicted interest degree of the sample object for the sample multimedia data, and iteratively training the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance, and the second distances corresponding to the N ineffective media triples, until the initial media recommendation model after the iterative training meets a stop iteration training condition, to obtain a target media recommendation model for recommending multimedia data.
2. The method of claim 1, wherein, The initial media recommendation model comprises a main recommendation network and an auxiliary recommendation network. The method comprises: generating, by the main recommendation network, a first media embedding vector reflecting the first media feature, and a second media embedding vector reflecting the second media feature; generating, by the auxiliary recommendation network, a first association embedding vector reflecting the first association semantic in the effective media triple, and a second association embedding vector reflecting the first association semantic in each ineffective media triple; generating, by the auxiliary recommendation network, a first label embedding vector reflecting the first attribute portrait label, and a second label embedding vector reflecting the second attribute portrait label in each ineffective media triple; and determine a first distance between the first media feature and the first attribute portrait label according to the first correlation embedding vector, the first media embedding vector and the first label embedding vector, and determine a second distance between the second media feature in each of the invalid media triplets and the second attribute portrait label according to the second correlation embedding vector, the second media embedding vector corresponding to each of the invalid media triplets and the second label embedding vector.
3. The method of claim 2, wherein, The determining the first distance between the first media feature and the first attribute portrait label according to the first correlation embedding vector, the first media embedding vector and the first label embedding vector, and the determining the second distance between the second media feature in each of the invalid media triplets and the second attribute portrait label according to the second correlation embedding vector, the second media embedding vector corresponding to each of the invalid media triplets and the second label embedding vector, comprises: performing summation operation on the first media embedding vector and the first correlation embedding vector to obtain a first distribution embedding vector of the first media feature to the first attribute portrait label; determining the first distance between the first media feature and the first attribute portrait label according to a difference between the first distribution embedding vector and the first label embedding vector; performing summation operation on the second correlation embedding vector and the second media embedding vector corresponding to each of the invalid media triplets to obtain a second distribution embedding vector of the second media feature in each of the invalid media triplets to the second attribute portrait label; determining the second distance between the second media feature in each of the invalid media triplets and the second attribute portrait label according to a difference between the second distribution embedding vector corresponding to each of the invalid media triplets and the second label embedding vector.
4. The method of claim 2, wherein, The determining the first distance between the first media feature and the first attribute portrait label according to the first correlation embedding vector, the first media embedding vector and the first label embedding vector, and the determining the second distance between the second media feature in each of the invalid media triplets and the second attribute portrait label according to the second correlation embedding vector, the second media embedding vector corresponding to each of the invalid media triplets and the second label embedding vector, comprises: performing product operation on the conversion embedding vector corresponding to the first correlation semantic and the first media embedding vector to obtain a first media conversion vector, and performing product operation on the first label embedding vector and the conversion embedding vector to obtain a first label conversion vector; performing product operation on the conversion embedding vector and the second media embedding vector corresponding to each of the invalid media triplets to obtain a second media conversion vector corresponding to each of the invalid media triplets, and performing product operation on the second label embedding vector corresponding to each of the invalid media triplets and the conversion embedding vector to obtain a second label conversion vector corresponding to each of the invalid media triplets; According to the first association embedding vector, the first media conversion vector and the first label conversion vector, a first distance between the first media feature and the first attribute portrait label is determined, and according to the second association embedding vector, the second media conversion vector corresponding to each invalid media triple and the second label conversion vector, a second distance between the second media feature in each invalid media triple and the second attribute portrait label is determined.
5. The method of claim 4, wherein, According to the first association embedding vector, the first media conversion vector and the first label conversion vector, a first distance between the first media feature and the first attribute portrait label is determined, and according to the second association embedding vector, the second media conversion vector corresponding to each invalid media triple and the second label conversion vector, a second distance between the second media feature in each invalid media triple and the second attribute portrait label is determined. The first media feature is summed with the first media conversion vector to obtain a third distribution embedding vector of the first media feature to the first attribute portrait label; The first distance between the first media feature and the first attribute portrait label is determined according to the difference between the third distribution embedding vector and the first label conversion vector; The second media feature in each invalid media triple is summed with the second media conversion vector corresponding to each invalid media triple to obtain a fourth distribution embedding vector of the second media feature in each invalid media triple to the second attribute portrait label; The second distance between the second media feature in each invalid media triple and the second attribute portrait label is determined according to the difference between the fourth distribution embedding vector corresponding to each invalid media triple and the second label conversion vector.
6. The method of claim 1, wherein, The iteration training of the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance and the second distance corresponding to the N invalid media triples includes: According to the predicted interest degree and the labeled interest degree, an interest degree prediction loss of the initial media recommendation model is determined; According to the first distance and the second distance corresponding to the N invalid media triples, a feature recognition loss of the initial media recommendation model is determined; According to the interest degree prediction loss and the feature recognition loss, the initial media recommendation model is iteratively trained.
7. The method of claim 6, wherein, According to the first distance and the second distance corresponding to the N invalid media triples, the feature recognition loss of the initial media recommendation model is determined, including: The second distance corresponding to the N invalid media triples and the first distance are subtracted to obtain distance difference values corresponding to the N invalid media triples; The distance difference values corresponding to the N invalid media triples are summed to obtain a total distance difference value; According to the total distance difference value, the feature recognition loss of the initial media recommendation model is determined.
8. The method of claim 6, wherein, The iterative training of the initial media recommendation model according to the interest degree prediction loss and the feature recognition loss comprises: The interest degree prediction loss and the feature recognition loss are summed to obtain a prediction total loss of the initial media recommendation model; According to the prediction total loss, a convergence state of the initial media recommendation model is determined; When the convergence state of the initial media recommendation model is a non-converged state, parameters of the initial media recommendation model are adjusted according to the prediction total loss to obtain an iteratively trained initial media recommendation model.
9. The method of claim 1, wherein, The N invalid media triplets comprise a first invalid media triplet and a second invalid media triplet; The generating of the N invalid media triplets that are not associated with the sample multimedia data according to the valid media triplet comprises: The first media feature in the valid media triplet is replaced by a second media feature that is not associated with the sample multimedia data to obtain the first invalid media triplet; The first attribute portrait label in the valid media triplet is replaced by a second attribute portrait label that is not associated with the sample multimedia data to obtain the second invalid media triplet; The second attribute portrait label in the first invalid media triplet is the same as the first attribute portrait label, and the second media feature in the second invalid media triplet is the same as the first media feature.
10. The method of claim 1, wherein, The method further comprises: An effective label triplet associated with the sample multimedia data is queried from the knowledge graph, the effective label triplet comprising a third attribute portrait label, a fourth attribute portrait label, and a second association semantic reflecting that the third attribute portrait label and the fourth attribute portrait label have an association relationship; M invalid label triplets that are not associated with the sample multimedia data are generated according to the effective label triplet, each invalid label triplet comprising a fifth attribute portrait label, a sixth attribute portrait label, and the second association semantic, the fifth attribute portrait label and the sixth attribute portrait label not having the association relationship, and M being a positive integer; A third distance between the third attribute portrait label and the fourth attribute portrait label is determined according to the second association semantic in the effective label triplet by the initial media recommendation model, and a fourth distance between the fifth attribute portrait label and the sixth attribute portrait label in each invalid label triplet is determined according to the second association semantic in each invalid label triplet; The iterative training of the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance, the second distances of the N invalid media triplets, the third distance, and the fourth distances of the M invalid label triplets comprises: The initial media recommendation model is iteratively trained according to the predicted interest degree, the labeled interest degree, the first distance, the second distances of the N invalid media triplets, the third distance, and the fourth distances of the M invalid label triplets.
11. The method of claim 1, wherein, The method further comprises: receiving a multimedia data recommendation request of a target object; the multimedia data recommendation request comprising object features of the target object; obtaining media features of candidate multimedia data to be recommended from a database according to the multimedia data recommendation request; identifying the media features of the candidate multimedia data and the object features of the target object through a main recommendation network of the target media model to obtain an identified interest degree of the target object for the candidate multimedia data; recommending the candidate multimedia data to the target object according to the identified interest degree.
12. A media data processing apparatus, characterized by comprising: comprise: an obtaining module, configured to obtain object features of a sample object, first media features of sample multimedia data, and a labeled interest degree of the sample object for the sample multimedia data; a first generating module, configured to generate valid media triplets associated with the sample multimedia data according to a knowledge graph associated with the first media features and the sample multimedia data; the valid media triplets comprising first media features, first attribute portrait labels, and first associated semantics, the first associated semantics reflecting that the first media features and the first attribute portrait labels have an associated relationship; a second generating module, configured to generate N invalid media triplets not associated with the sample multimedia data according to the valid media triplets, each of the invalid media triplets comprising second media features, second attribute portrait labels, and the first associated semantics, the second media features and the second attribute portrait labels not having the associated relationship, N being a positive integer; a determining module, configured to determine, through an initial media recommendation model, a first distance between the first media features and the first attribute portrait labels according to the first associated semantics in the valid media triplets, and determine a second distance between the second media features and the second attribute portrait labels in each of the invalid media triplets according to the first associated semantics in each of the invalid media triplets; a training module, configured to identify the first media features and the object features through the initial media recommendation model to obtain a predicted interest degree of the sample object for the sample multimedia data, and iteratively train the initial media recommendation model according to the predicted interest degree, the labeled interest degree, the first distance, and the second distances corresponding to the N invalid media triplets respectively until the iteratively trained initial media recommendation model meets a stop iteration training condition to obtain a target media recommendation model for recommending multimedia data.
13. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 11.
14. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 11. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 11.