A data processing method and device, computer equipment and readable storage medium

By constructing a multi-task recognition model and combining shared and dedicated networks, the problem of low media recognition accuracy caused by the sparsity of training data is solved, thereby improving the accuracy and efficiency of multimedia data delivery.

CN117093769BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-07-21
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing media recognition models suffer from low accuracy in media recognition and multimedia data delivery due to the sparsity of training data.

Method used

By constructing a multi-task recognition model that combines a shared network and a dedicated network, shared media features and dedicated media features are extracted from the spliced ​​media features of samples to identify multi-dimensional interactive labels. The shared network is used for multiple media recognition tasks, while the dedicated network is used for a single task, reducing training resource overhead and improving recognition accuracy.

Benefits of technology

It improves the accuracy and efficiency of multimedia data delivery, avoids the problem of low recognition accuracy caused by the sparsity of training data, and enhances the training efficiency and accuracy of multi-task recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117093769B_ABST
    Figure CN117093769B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method and device, computer equipment and a readable storage medium. The method comprises: obtaining sample spliced media features and N-dimensional labeled interaction labels of a sample object for sample multimedia data; calling a shared network in an initial recognition model to extract shared media features associated with N media recognition tasks from the sample spliced media features; calling a dedicated network i in the initial recognition model to extract dedicated media features i from the sample spliced media features; calling a recognition network i in the initial recognition model to identify an i-dimensional predicted interaction label according to the shared media features and the dedicated media features i; and training the initial recognition model according to the N-dimensional labeled interaction labels and the N-dimensional predicted interaction labels to obtain a multi-task recognition model. The media recognition accuracy of the trained initial recognition model can be improved, and the push accuracy of the multimedia data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology, and in particular to a data processing method, apparatus, computer equipment, and readable storage medium. Background Technology

[0002] With the development of internet technology and the increasing scale of network data, the needs of objects (such as users) are becoming more and more diversified and personalized. Push systems have become a common solution for the public to filter out multimedia data (such as advertising data, news data, and game data) of interest to objects when faced with massive amounts of internet information. By pushing multimedia data, the content presented to the object can be reasonably controlled, thereby better realizing the value of multimedia data.

[0003] Currently, in push systems, multimedia data is mainly pushed to objects through trained media recognition models. In practice, it has been found that due to factors such as the sparsity of training data, the media recognition accuracy of the trained media recognition model is relatively low, which in turn leads to low accuracy in pushing multimedia data. Summary of the Invention

[0004] This application provides a data processing method, apparatus, computer device, and readable storage medium to improve the media recognition accuracy of the initial recognition model after training, thereby improving the accuracy of multimedia data delivery.

[0005] One embodiment of this application provides a data processing method, including:

[0006] Obtain the sample splicing media features and the N-dimensional annotation interaction labels of the sample objects for the sample multimedia data; the sample splicing media features are obtained by splicing the media attribute features of the sample multimedia data and the historical media interaction features of the sample objects; the N-dimensional annotation interaction labels correspond to N media recognition tasks; N is an integer greater than 1;

[0007] The shared network in the initial recognition model is invoked to extract shared media features associated with N media recognition tasks from the sample splicing media features;

[0008] Call the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample splicing media features; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks in the initial recognition model, and i is a positive integer less than or equal to N;

[0009] Call the recognition network i in the initial recognition model for media recognition task i, and based on the shared media features and the specific media features i, identify the sample object and predict the interactive label for the i-th dimension of the sample multimedia data;

[0010] The initial recognition model is trained based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping condition, thus obtaining the multi-task recognition model.

[0011] One embodiment of this application provides a data processing apparatus, including:

[0012] The first acquisition module is used to acquire sample splicing media features and N-dimensional annotation interaction labels of sample objects for sample multimedia data; the sample splicing media features are obtained by splicing the media attribute features of sample multimedia data and the historical media interaction features of sample objects; the N-dimensional annotation interaction labels correspond to N media recognition tasks; N is an integer greater than 1;

[0013] The first extraction module is used to call the shared network in the initial recognition model to extract shared media features associated with N media recognition tasks from the sample splicing media features;

[0014] The second extraction module is used to call the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample splicing media features; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks of the initial recognition model, and i is a positive integer less than or equal to N.

[0015] The first recognition module is used to call the recognition network i in the initial recognition model for media recognition task i, and to identify the i-th dimension of the sample object for the sample multimedia data based on the shared media features and the exclusive media features i.

[0016] The training module is used to train the initial recognition model based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping training condition, thus obtaining the multi-task recognition model.

[0017] One embodiment of this application provides a computer device, including: a processor and a memory;

[0018] The processor is connected to a memory, which stores a computer program. When the computer program is executed by the processor, it causes the computer device to perform the method provided in the embodiments of this application.

[0019] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the method provided in this application.

[0020] One embodiment of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in this application embodiment.

[0021] In this embodiment, a multi-task recognition model is constructed based on sample-stitched media features. This model can identify multi-dimensional interactive tags of objects in multimedia data, with each interactive tag corresponding to a media recognition task. Therefore, the multi-task recognition model can handle multiple media recognition tasks without requiring a separate model for each task, reducing resource overhead during training and improving training efficiency. During training, an initial recognition model is first constructed, comprising a shared network, multiple dedicated networks, and multiple recognition networks. Each dedicated network and recognition network corresponds to a media recognition task. When identifying the i-th predicted interactive tag for media recognition task i, the i-th predicted interactive tag of the sample object in multimedia data is identified by combining the recognition network i in the initial model with the shared media features and the dedicated media features i associated with media recognition task i. The initial recognition model is then trained based on the N-dimensional labeled interactive tags and the N-dimensional predicted interactive tags to obtain the multi-task recognition model. The shared media features here are extracted from the sample-stitched media features by the shared network of the initial recognition model, while the specific media feature i is extracted from the sample-stitched media features by the specific network in the initial recognition model associated with media recognition task i. In other words, when recognizing the i-th predicted interaction label under media recognition task i, not only the specific media feature i directly associated with media recognition task i is combined, but also the shared media features indirectly associated with media recognition task i are combined, providing more information for recognizing the i-th predicted interaction label under media recognition task i. That is, different media recognition tasks can share shared media features, which provides more training data for the training process for different media recognition tasks. This avoids the problem of low media recognition accuracy of the initial recognition model (i.e., task recognition model) after training due to the sparsity of training data, improves the media recognition accuracy of the initial recognition model after training, and thus improves the accuracy of multimedia data push. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application;

[0024] Figure 2 This is a schematic diagram illustrating an application scenario of data processing provided in an embodiment of this application;

[0025] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0026] Figure 4 This is a schematic diagram of the model structure of an initial recognition model provided in an embodiment of this application;

[0027] Figure 5 This is a schematic diagram of a network structure of a gated subnetwork provided in an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of the model structure of an initial recognition model provided in an embodiment of this application;

[0029] Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0030] Figure 8 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0031] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0033] This application relates to the field of artificial intelligence technology. Specifically, this application can train an initial recognition model by splicing media features from samples to obtain a multi-task recognition model for multi-task recognition. For example, the multi-task recognition model can be used to simultaneously recognize multi-dimensional recognition interaction tags (such as attention rate tags, purchase rate tags, sharing rate tags, etc.) under multiple media recognition tasks, thereby improving the media recognition accuracy of the initial recognition model after training and thus improving the accuracy of multimedia data push.

[0034] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, machine learning / deep learning, autonomous driving, and intelligent transportation.

[0035] Specifically, this application relates to machine learning, a branch of artificial intelligence technology. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0036] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application. For example... Figure 1 As shown, the data processing system may include server 10 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices; the number of terminal devices is not limited here. Figure 1As shown, it may specifically include terminal device 100a, terminal device 100b, terminal device 100c, ..., terminal device 100n. For example... Figure 1 As shown, terminal devices 100a, 100b, 100c, ..., 100n can each connect to the server 10 via a network, allowing each terminal device to interact with the server 10 through this network connection. Of course, terminal devices 100a, 100b, 100c, ..., 100n can communicate with each other via a direct network connection, enabling point-to-point communication. In other words, when two terminal devices need to exchange data, one terminal device (the sending terminal device) can directly send data to the other terminal device (the receiving terminal device).

[0037] Each terminal device in the terminal device cluster can include: smartphones, tablets, laptops, desktop computers, smart voice interaction devices, smart home appliances (e.g., smart TVs), wearable devices, in-vehicle terminals, and other smart terminals with data processing capabilities. It should be understood that, as... Figure 1 Each terminal device in the terminal device cluster shown can be equipped with an application that has data processing capabilities. When the application runs on each terminal device, it can interact with the aforementioned applications. Figure 1 Data interaction occurs between the servers 10 shown. Specific applications may include media recognition applications, multimedia data push applications, etc. The applications in this embodiment can be integrated into another application, or they can be independent applications (e.g., news applications). This embodiment does not limit the type of application. For ease of understanding, this embodiment can... Figure 1 From the multiple terminal devices shown, one terminal device is selected as the target terminal device. For example, in the embodiments of this application, a terminal device can be selected as the target terminal device. Figure 1 The terminal device 100a shown is the target terminal device. The target terminal device can have an application with data processing function installed. At this time, the target terminal device can realize data interaction with the server 10 through the application.

[0038] Among them, such as Figure 1 As shown, server 10 is a device that can provide backend services for applications in terminal devices. Server 10 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0039] It should be understood that, based on Figure 1 One data processing system described herein is applicable to multimedia data push scenarios. It is understood that this application can train an initial recognition model based on sample-concatenated media features. After training the initial recognition model, a multi-task recognition model is obtained. This multi-task recognition model can be used to identify multi-dimensional recognition interaction tags of business objects (such as users) for multimedia data to be pushed (such as advertising data) under multiple media recognition tasks. One media recognition task corresponds to one recognition interaction tag. Media recognition tasks can include click-through rate recognition tasks, shallow conversion rate recognition tasks, and deep conversion rate recognition tasks, etc. Multi-dimensional interaction tags under multiple media recognition tasks can include click-through rate tags under click-through rate recognition tasks, shallow conversion rate tags under shallow conversion rate recognition tasks, and deep conversion rate tags under deep conversion rate recognition tasks, etc. Shallow conversion rate refers to the probability that an object performs a shallow operation after clicking on multimedia data, such as follow rate and like rate. Deep conversion rate refers to the probability that an object performs a deep operation after clicking on multimedia data, such as purchase rate and share rate. In this way, multi-dimensional recognition interaction tags can be identified within a single model, and data from multiple media recognition tasks can be shared, improving the accuracy and efficiency of multi-task recognition. Furthermore, based on the multi-dimensional recognition interaction tags output by the multi-task recognition model, multimedia data to be pushed can be sent to business objects. For example, the push score of the multimedia data to be pushed can be determined based on click-through rate tags and purchase rate tags. Multimedia data with a push score greater than a threshold can be pushed to business objects, improving the efficiency and accuracy of multimedia data push.

[0040] In this application embodiment, multimedia data can refer to advertising data (such as advertising video data, advertising audio data, advertising text data, etc.), or news data, game data, etc. The sample object in this application embodiment can refer to a user who has been pushed multimedia data within a historical time period. The sample multimedia data can refer to the multimedia data that currently needs to be identified using N-dimensional interactive tags, such as identifying click-through rate tags, conversion rate tags, etc., for the sample multimedia data, where N is an integer greater than 1. The N-dimensional labeled interactive tags correspond to N media identification tasks, with one-dimensional labeled interactive tags corresponding to one media identification task. The N media identification tasks can include click-through rate identification tasks, deep conversion rate identification tasks, shallow conversion rate identification tasks, etc. The N-dimensional interactive tags (such as labeled interactive tags, predicted interactive tags, and identified interactive tags) can include click-through rate tags, shallow conversion rate tags, deep conversion rate tags, etc. The media attribute features of the sample multimedia data can include identifier features (such as sparse features) and statistical features (such as dense features). For example, when the sample multimedia data is advertising data, the media attribute characteristics of the sample multimedia data can include attributes such as price, purpose, theme, and appearance of the advertising objects (such as items or virtual characters) in the advertising data. When the sample multimedia data is news data or game data, the media attribute characteristics of the sample multimedia data can include attributes such as content theme and content viewing time. Taking advertising data as an example, identification features can include object identifiers, identifiers of sample objects (such as game items or game characters) in the sample multimedia data, and identifiers of objects (such as game items or game characters) that have been pushed to the multimedia data. Statistical features can include the price of the sample objects and the price of the pushed objects.

[0041] Similarly, the historical media interaction characteristics of a sample object can also include identifier-type characteristics and statistical characteristics. Specifically, the historical media interaction characteristics of a sample object can include its historical interaction characteristics with previously pushed multimedia data, its object attribute characteristics, and the media attribute characteristics of the pushed multimedia data. Previously pushed multimedia data can refer to one or more multimedia data items pushed to the sample object within a historical time period. The historical interaction characteristics of a sample object with respect to previously pushed multimedia data can include the sample object's interaction behavior with the pushed multimedia data, and the number of times different interaction behaviors occurred. For example, whether the sample object clicked on the pushed multimedia data and the number of times it clicked; whether the sample object made a purchase after clicking on the pushed multimedia data and the number of times it made a purchase; whether the sample object commented after clicking on the pushed multimedia data and the number of times it commented, etc.

[0042] The object attribute features of the sample objects can include the object identifier and age of the object. The media attribute features of the pushed multimedia data can include the media content features of the pushed multimedia data. For example, if the pushed multimedia data is advertising data, the media attribute features of the pushed multimedia data can include the identifier, price, purpose, theme, and appearance of the advertising object (such as an item or virtual character) in the advertising data. If the pushed multimedia data is news data or game data, the media attribute features of the pushed multimedia data can include the content theme and viewing time. The N-dimensional annotation interactive tags of the sample objects for the sample multimedia data can refer to the interactive tags manually annotated under N media recognition tasks. For example, the manually annotated click-through rate tag of the sample objects for the sample multimedia data is 100% under the click-through rate recognition task, and the manually annotated purchase rate tag of the sample objects for the sample multimedia data is 50% under the purchase rate recognition task, etc.

[0043] Specifically, the initial recognition model in this application includes a shared network, a dedicated network, and a recognition network. The shared network is used to extract shared media features associated with N media recognition tasks from the sample spliced ​​media features. It can be understood that shared media features can be used for label recognition in N media recognition tasks. Shared media features can include the media features required for label recognition in N media recognition tasks, as well as the correlation features between media features. In other words, different media recognition tasks can share features, thus extracting more features and improving the accuracy of multi-task recognition. For example, taking multimedia data as advertising data, when the N media recognition tasks include click-through rate (CTR) recognition and purchase rate (PCR) recognition tasks, the shared media features can be a fusion of features such as object attribute features, media attribute features of the sample multimedia data, click features required for the CTR recognition task, purchase features required for the PCR recognition task, and correlation features between click features and purchase features. Click features can include the media attribute features of the pushed multimedia data clicked by the sample object and the number of clicks, while purchase features can include the media attribute features of the pushed multimedia data clicked by the sample object and the number of clicks. For example, if the sample objects perform actions such as clicking, purchasing, and sharing on the pushed multimedia data, it can reflect that the sample objects have a high level of interest in the pushed multimedia data.

[0044] Therefore, feature weights can be assigned to the media attribute features of pushed multimedia data with high interest levels to accurately identify click-through rate tags, purchase rate tags, etc., for sample objects targeting the multimedia data. In other words, by sharing media features, different media recognition tasks can utilize features from other media recognition tasks, thereby increasing feature diversity. Furthermore, when using purchase rate recognition, not only purchase features but also click features can be obtained from shared media features. For example, when purchase features are scarce, pushed multimedia data with a high number of clicks can be obtained from click features. A high number of clicks indicates a high level of interest in the sample object and a high probability of purchase; purchase rate tags can be predicted from pushed multimedia data with a high number of clicks. This addresses the problem of insufficient training data for some media recognition tasks.

[0045] It is understandable that different media recognition tasks can share the network parameters in the shared network mentioned above, and perform feature extraction through these parameters. Simultaneously, training data corresponding to N media recognition tasks can be shared, providing the correlations between the training data for each of the N media recognition tasks. This avoids building a model for each media recognition task, and also prevents poor model performance due to insufficient training data for each media recognition task, thereby improving the efficiency and accuracy of multi-task recognition (i.e., improving the efficiency and accuracy of multi-dimensional interactive label recognition).

[0046] The initial recognition model includes dedicated networks associated with N media recognition tasks, one for each task. These dedicated networks extract specific media features relevant to each task. For example, the network for click-through rate (CTR) recognition extracts features reflecting whether a user clicked on multimedia data, while the network for purchase rate recognition extracts features reflecting whether a user purchased multimedia data. Since the network parameters in the shared network are shared across the N tasks, they are affected by the recognition losses of each task. This means that while the trained shared network can extract features required by all N tasks, it may not reach its optimal performance for all tasks. Consequently, the trained shared network may extract more features for some tasks and fewer features for others, resulting in a seesaw effect.

[0047] Therefore, this application, by configuring a dedicated network for each identification task, can extract dedicated media features directly associated with the corresponding media identification task. These dedicated media features are only used for the corresponding media identification task, so the network parameters in the dedicated network are only affected by the identification loss of the corresponding media identification task and not by the identification loss of other media identification tasks, thus effectively solving the seesaw effect. For example, in a click-through rate (CTR) identification task, the dedicated media features corresponding to the CTR identification task can only include features associated with the CTR identification task, such as the media attribute features of the multimedia data pushed by the sample object, the number of clicks, the object attribute features of the sample object, and the media attribute features of the sample multimedia data. In this way, by using a dedicated network for each media identification task, the seesaw effect problem that occurs when all media identification tasks share a common network can be avoided. This seesaw effect refers to the situation where some media identification tasks have high accuracy while others have low accuracy.

[0048] Simultaneously, the initial recognition model includes recognition networks associated with N media recognition tasks, one for each task. Through the recognition network corresponding to each task, predicted interaction labels for that task can be extracted. For example, the recognition network for the click-through rate (CTR) recognition task identifies the CTR label of a sample object for the sample multimedia data, while the recognition network for the purchase rate recognition task identifies the purchase rate label of a sample object for the sample multimedia data. Since each media recognition task has its own dedicated network and recognition network, the training of the network parameters in each task's dedicated network and recognition network is not affected by other media recognition tasks. This weakens the seesaw effect between tasks, thereby improving the accuracy of multi-task recognition and ultimately improving the accuracy of multimedia data delivery.

[0049] For better understanding, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram illustrating an application scenario of digital processing provided in an embodiment of this application. For example... Figure 2 Terminal devices 201a, 202a, 203a, and 204a in the terminal device cluster 20a shown can be the aforementioned Figure 1 The terminal devices in the terminal device cluster in the corresponding embodiment, such as Figure 2 The terminal device 20c shown can be the above-mentioned Figure 1 The terminal device 100b in the corresponding embodiment, such as Figure 2 The server 20d shown can be the one described above. Figure 1The corresponding embodiment uses server 10. Terminal devices 20c and 20a in the terminal device cluster 20a are connected to server 20d via a network, allowing them to interact with server 20d. It is understood that the terminal devices in terminal device cluster 20a can send media attribute characteristics of sample multimedia data and historical media interaction characteristics of sample objects to server 20d. Taking advertising data as an example, the media attribute characteristics of the sample multimedia data can include the identifier, price, purpose, theme, and appearance of the advertising object (such as an item or virtual character). The historical media interaction characteristics of the sample object can include the historical interaction characteristics of the sample object with respect to previously pushed multimedia data, the object attribute characteristics of the sample object, and the media attribute characteristics of the previously pushed multimedia data, specifically as described above regarding historical media interaction characteristics. Previously pushed multimedia data can refer to advertising data, news data, game data, etc. It is understood that previously pushed multimedia data includes not only advertising data of the same type as the sample multimedia data but also other types of data besides advertising data.

[0050] In this process, the managed object 20b can manually input N-dimensional annotation interactive tags for the sample multimedia data via the terminal device 20c. Examples of tags include click-through rate and purchase rate. When the terminal device 20c receives these N-dimensional annotation interactive tags from the managed object 20b, it can send them to the server 20d. Further, after receiving the media attribute features of the sample multimedia data and the historical media interaction features of the sample object, the server 20d can perform feature concatenation to obtain the concatenated media features. The server 20d can then input these concatenated media features into the initial recognition model, call the shared network within the initial recognition model, and extract features from the concatenated media features to extract shared media features associated with N media recognition tasks. Understandably, N media recognition tasks can share the network parameters and structure of a shared network in the initial recognition model. This shared network allows for the extraction of shared media features required by all N media recognition tasks. Furthermore, the sample-stitched media features can be generated based on training data associated with the N media recognition tasks, enabling training data sharing between different tasks. In this way, the shared network avoids the problem of poor model performance when training data for a media recognition task is insufficient, thus improving the efficiency and accuracy of multi-task recognition (i.e., improving the efficiency and accuracy of multi-dimensional interactive label recognition).

[0051] Simultaneously, server 20d can call the dedicated network i in the initial recognition model to extract features from the sample spliced ​​media features, extracting the dedicated media features i associated with media recognition task i. The initial recognition model includes N dedicated networks corresponding to N media recognition tasks, with each media recognition task having its own dedicated network (i.e., one dedicated network per media recognition task). Dedicated network i is the dedicated network associated with media recognition task i among the N dedicated networks in the initial recognition model. In other words, dedicated network i is a network specifically configured for media recognition task i, used only to extract media features associated with media recognition task i, and is not affected by other media recognition tasks. This mitigates the seesaw effect caused by sharing network layers among the N media recognition tasks (e.g., some media recognition tasks have higher accuracy, some have lower accuracy, and the N media recognition tasks cannot reach the optimal level together), thereby improving the recognition accuracy of the media recognition tasks.

[0052] Furthermore, server 20d can invoke the recognition network i corresponding to media recognition task i in the initial recognition model. Based on the exclusive media feature i and the shared media feature, it identifies the i-th dimension predicted interaction label of the sample object for the sample multimedia data. The initial recognition model includes N recognition networks corresponding to N media recognition tasks, each media recognition task having its own configured recognition network (i.e., one media recognition task corresponds to one recognition network). Recognition network i is the recognition network associated with media recognition task i among the N exclusive networks in the initial recognition model. In other words, recognition network i is a network exclusively configured for media recognition task i, used only to identify the interaction label corresponding to media recognition task i, and is not affected by other media recognition tasks. Specifically, recognition network i includes a gating mechanism. This gating mechanism controls the importance of the exclusive media feature i and the shared media feature, respectively. Based on the importance of the exclusive media feature i and the shared media feature, feature fusion is performed on the exclusive media feature i and the shared media feature. Then, the fused business-integrated media feature is identified to obtain the i-th dimension predicted interaction label of the sample object for the sample multimedia data. In this way, the gating mechanism can alleviate feature conflicts between different media features and effectively solve the seesaw effect between N media recognition tasks, thereby improving the accuracy of multi-task recognition.

[0053] Similarly, referring to the method of obtaining the i-th dimension predicted interaction label, server 20d can call the dedicated network corresponding to the remaining media recognition task to obtain the dedicated media features of the remaining media recognition task. The remaining media recognition task refers to the task other than the media recognition task among the N media recognition tasks. Further, server 20d can call the recognition network corresponding to the remaining media recognition task, and based on the dedicated media features and shared media features of the remaining media recognition task, identify the N-1 dimension predicted interaction label of the sample object for the sample multimedia data, until N dimension predicted interaction labels corresponding to the N media recognition tasks are obtained. Further, server 20d can train the initial recognition model based on the N dimension labeled interaction label and the N dimension predicted interaction label, until the trained initial recognition model meets the stopping training condition, thus obtaining the multi-task recognition model. It can be seen that, by sharing the network parameters and network structure in the shared network for N media recognition tasks, this embodiment of the application can realize the establishment of a model for N media recognition tasks, avoiding the problems caused by modeling N media recognition tasks separately, and reducing the cost and efficiency of multi-task recognition. Meanwhile, by sharing a network, training data can be shared among the N media recognition tasks. This avoids the problem of insufficient model training caused by limited training data for each media recognition task (such as limited training data for a purchase rate recognition task), thus improving the performance of the multi-task recognition model and consequently increasing the accuracy of multi-task recognition. Furthermore, this embodiment utilizes a dedicated network for each media recognition task, effectively addressing the seesaw effect among the N media recognition tasks and improving the accuracy of multi-task recognition, thereby enhancing the accuracy of multimedia data delivery.

[0054] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 3 As shown, this method can be derived from... Figure 1 It can be executed by any terminal device in the system, or by... Figure 1 Server 10 in the middle can be used to execute it, and it can also be executed by Figure 1 The terminal device and server in this application work together to execute the method. The device used to execute this method can be collectively referred to as a computer device. The data processing method may include, but is not limited to, the following steps:

[0055] S101, obtain the media features of the sample splicing, and the N-dimensional annotation interaction labels of the sample object for the sample multimedia data.

[0056] Specifically, computer equipment can acquire sample spliced ​​media features, which are obtained by splicing the media attribute features of the sample multimedia data and the historical media interaction features of the sample object. The sample multimedia data refers to multimedia data that requires the initial recognition model to perform N-dimensional interaction tag recognition. The sample object refers to a user who has been pushed multimedia data within a historical time period. The N-dimensional interaction tags used by the initial recognition model to identify the sample object's interaction with the sample multimedia data are N, where N is a positive integer. The N-dimensional labeled interaction tags correspond to N media recognition tasks, and the one-dimensional labeled interaction tags correspond to one media recognition task, such as the i-th dimension labeled interaction tag corresponding to media recognition task i. The N media recognition tasks can include click-through rate recognition tasks, shallow conversion rate recognition tasks, deep conversion rate recognition tasks, etc. The shallow conversion rate can represent the probability of an object performing a shallow operation after clicking on multimedia data, such as the follow rate and like rate. The deep conversion rate represents the probability of an object performing a deep operation after clicking on multimedia data, such as the purchase rate and share rate. N-dimensional interactive labels can include click-through rate labels, shallow conversion rate labels, and deep conversion rate labels.

[0057] Here, N-dimensional annotation interactive labels refer to the N-dimensional annotation interactive labels that the management object annotates under N media recognition tasks based on the media features of the sample. In other words, N-dimensional annotation interactive labels are the interactive labels manually annotated by the management object for each of the N media recognition tasks. For example, based on the media features of the sample, the management object determines that the click-through rate label for the sample object in the click-through rate recognition task is 100%, and the purchase rate label for the sample object in the purchase rate recognition task is 50%, etc.

[0058] Specifically, the media attribute features of sample multimedia data can include identifier-based features (such as sparse features) and statistical features (such as dense features). For example, when the sample multimedia data is advertising data, its media attribute features can include the price, purpose, theme, and appearance of the advertising objects (such as items or virtual characters). When the sample multimedia data is news data or game data, its media attribute features can include content theme and viewing duration. Similarly, the historical media interaction features of sample objects can also include identifier-based features (such as sparse features) and statistical features (such as dense features). The historical media interaction features of sample objects can include the sample object's historical interaction features with previously pushed multimedia data, the object attribute features of the sample object, and the media attribute features of the previously pushed multimedia data. The multimedia data that has been pushed can refer to one or more multimedia data that have been pushed to the sample object within a historical time period. The historical interaction characteristics of the sample object with respect to the multimedia data that has been pushed can include the sample object's interaction behavior with the multimedia data that has been pushed, as well as the number of times different interaction behaviors are performed. For example, whether the sample object clicked on the multimedia data that has been pushed and the number of times the click behavior was performed; whether the sample object made a purchase after clicking on the multimedia data that has been pushed and the number of times the purchase behavior was performed; whether the sample object made a comment after clicking on the multimedia data that has been pushed and the number of times the comment behavior was performed, etc.

[0059] The object attribute characteristics of the sample object can include the object identifier and age of the object. The media attribute characteristics of the pushed multimedia data can include the media content characteristics of the pushed multimedia data. For example, if the pushed multimedia data is advertising data, the media attribute characteristics of the pushed multimedia data can include the identifier, price, purpose, theme, appearance, and other attributes of the advertising object (such as an item or virtual character) in the advertising data. If the pushed multimedia data is news data or game data, the media attribute characteristics of the pushed multimedia data can include the content theme, content viewing time, and other attributes.

[0060] Specifically, the computer device can obtain historical media interaction data of sample objects from the user side, such as object attribute data of sample objects and interaction data of sample objects with pushed multimedia data. Simultaneously, the computer device can obtain media attribute data of pushed multimedia data and media attribute data of sample multimedia data from the multimedia data push side (e.g., the advertising side). Further, the computer device can use the vector transformation network in the initial recognition model to perform vector transformation on the media attribute data of historical media interaction data and sample multimedia data to obtain historical media interaction features and media attribute features of sample multimedia data. Specifically, the computer device can query a 64-dimensional embedding vector for each key in the sparse features (e.g., sparse features) of the media attribute data of historical media interaction data and sample multimedia data, and concatenate the vectors obtained from each key query. This concatenation is then further combined with the statistical features (e.g., dense features) of the media attribute data of historical media interaction data and sample multimedia data to obtain the sample concatenated media features. The historical media interaction data of the sample objects may include training data associated with N media recognition tasks. For example, historical media interaction data may include click behavior data associated with the click-through rate recognition task, purchase behavior data associated with the purchase rate recognition task, and attention behavior data associated with the attention rate recognition task.

[0061] S102, invoke the shared network in the initial recognition model to extract shared media features associated with N media recognition tasks from the sample splicing media features.

[0062] Specifically, the initial recognition model can be used to identify N media recognition tasks, obtaining N-dimensional predicted interaction labels for each task, such as the click-through rate (CTR) label for the CTR recognition task and the purchase rate (PCR) label for the purchase rate recognition task. Specifically, the initial recognition model includes a shared network. Computer devices can access this shared network to extract shared media features associated with the N media recognition tasks from the sample-stitched media features. Understandably, the N media recognition tasks can share the network parameters and structure of the shared network. Thus, a single initial recognition model can identify all N media recognition tasks, instead of building a model for each task, reducing the efficiency and cost of multi-task recognition. Furthermore, since the sample-stitched media features are obtained based on training data associated with the N media recognition tasks, these tasks can share training data. For example, the purchase rate recognition task can utilize the training data corresponding to the CTR recognition task. This avoids the problem of insufficient training data for a particular media recognition task (e.g., insufficient training data for the purchase rate recognition task), leading to inadequate model training and improving the accuracy of multi-task recognition. It should be understood that the network parameters in the shared network are influenced by the recognition losses of the N media recognition tasks. The trained shared network can extract shared media features required by all N media recognition tasks. In other words, the shared media features extracted by the shared network are shared by the N media recognition tasks, and each media recognition task can use the shared media features for interactive label prediction.

[0063] Optionally, the computer device may invoke the shared network in the initial recognition model to extract shared media features associated with N media recognition tasks from the sample spliced ​​media features. Specifically, this may involve: if there is only one shared network, then invoking the M expert sub-networks included in the shared network to extract M expert media features associated with the N media recognition tasks from the sample spliced ​​media features, where M is a positive integer. The extracted M expert media features are then identified as the shared media features associated with the N media recognition tasks.

[0064] Specifically, the initial recognition model can have one or more shared networks (i.e., network layers). Each shared network can include M expert subnetworks. When M is an integer greater than 1, the network parameters of different expert subnetworks are different. For example, the network parameters of expert subnetwork z001 and expert subnetwork z002 in the M expert subnetworks are different. Because the network parameters of different expert subnetworks are different, different expert subnetworks can be used to extract features in different spaces. In other words, each expert subnetwork has a preferred feature extraction domain. Through the network structure of M expert subnetworks, more complex data features can be processed, thereby improving model performance. When there is only one shared network, the computer device can input the sample spliced ​​media features into the M expert subnetworks in the shared network, and call the M expert subnetworks included in the shared network to extract M expert media features associated with N media recognition tasks from the sample spliced ​​media features.

[0065] Understandably, an expert subnetwork can extract an expert media feature associated with N media recognition tasks from the sample splicing media features. Specifically, expert subnetwork z001 among the M expert subnetworks can extract expert media features t001 associated with N media recognition tasks from the sample splicing media features, expert subnetwork z002 among the M expert subnetworks can extract expert media features t002 associated with N media recognition tasks from the sample splicing media features, and so on, until the expert media features extracted by the M expert subnetworks are obtained, resulting in M ​​expert media features. The computer device can determine the extracted M expert media features as shared media features associated with the N media recognition tasks. Since the N media recognition tasks share the shared media features for interactive label prediction and adjust the network parameters in the shared network based on the prediction loss of the N media recognition tasks, the media features extracted by the shared network are associated with the N media recognition tasks; that is, the media features extracted by the shared network are needed by all N media recognition tasks.

[0066] Specifically, each of the M expert sub-networks has different initial network parameters, and the network structure within each expert sub-network can be the same or different. Each expert sub-network can refer to a neural network, and its structure may include a fully connected sub-network and an activation function. The fully connected sub-network is used to perform convolutional processing on the sample splicing media features, and the activation function is used to activate the convolutional media features input to the fully connected sub-network (e.g., nonlinear feature transformation). It can be understood that the fully connected sub-network can map the features learned from the sample splicing media features to a sample label space, and perform classification and recognition through this sample label space. The activation sub-network can transform the low-order features input to the fully connected sub-network into high-order features to extract richer features. In this embodiment, the network layer data of the fully connected sub-network and activation function in each expert sub-network are not limited.

[0067] Optionally, if there are multiple shared networks in the initial recognition model, these shared networks have a hierarchical relationship. The specific method by which the computer device calls the shared networks in the initial recognition model to extract shared media features associated with N media recognition tasks from the sample splicing media features may include: if the shared network includes a first shared network and a second shared network, then the M expert sub-networks included in the first shared network are called to extract M expert media features associated with the N media recognition tasks from the sample splicing media features, where M is a positive integer, such as 1, 2, 3, etc. The N initial proprietary media features and the M expert media features corresponding to the first shared network are fused to obtain the first fused media features; the N initial proprietary media features are extracted from the N first proprietary networks in the initial recognition model, and the N first proprietary networks and the first shared network are located in the same network layer. The M expert sub-networks included in the second shared network are called to extract shared media features associated with the N media recognition tasks from the first fused media features.

[0068] Specifically, if the shared network in the initial recognition model includes a first shared network and a second shared network, and both the first and second shared networks include M expert sub-networks, with the first shared network having a lower network level than the second shared network, meaning the media features output by the first shared network can be used as input to the second shared network. The computer device can call upon the M expert sub-networks included in the first shared network to extract M expert media features associated with N media recognition tasks from the sample splicing media features. Similarly, each expert sub-network in the first shared network can extract one expert media feature from the sample splicing media features, until each of the M expert sub-networks in the first shared network has extracted its corresponding expert media feature, thus obtaining the M expert media features extracted by the first shared network. Further, the computer device can call upon the gated sub-networks in the shared network to fuse the N initial proprietary media features and the M expert media features extracted by the first shared network to obtain the first fused media features. Here, the N initial proprietary media features are extracted by the N first proprietary networks in the initial recognition model, and the N first proprietary networks are located at the same network layer as the first shared network.

[0069] The gated subnetwork in the shared network can be a self-attention network. Through this subnetwork, it can automatically learn the commonalities and characteristics between different media recognition tasks, thereby effectively determining the fusion weights corresponding to the N initial specific media features and the M expert media features extracted by the first shared network. That is, one feature corresponds to one fusion weight. Further, the computer device can fuse the N initial specific media features and the M expert media features extracted by the first shared network using the fusion weights corresponding to these weights, obtaining the first fused media feature. Specifically, the computer device can invoke the gated subnetwork in the shared network to weight each initial specific media feature according to its corresponding fusion weight, obtaining M weighted specific media features. The computer device can also invoke the gated subnetwork in the shared network to weight the M expert media features extracted by the first shared network according to their corresponding fusion weights, obtaining M weighted expert media features. Furthermore, the computer device can invoke the gated subnetwork in the shared network to add the M weighted specific media features and the M weighted expert media features to obtain the first fused media feature. This avoids feature conflicts between N media recognition tasks, improves model performance, and consequently enhances the accuracy of multi-task recognition.

[0070] In this system, an initial specific media feature is extracted from the sample-stitched media features by a first specific network in the initial recognition model, and is associated with the corresponding media recognition task. For example, initial specific media feature i in N initial specific media features is extracted from the sample-stitched media features by the first specific network i, and is associated with media recognition task i. Further, the computer device can input the first fused media features into the M expert sub-networks included in the second shared network, and call the M expert sub-networks included in the second shared network to extract M expert media features associated with the N media recognition tasks from the first fused media features. Further, the computer device can determine the M expert media features extracted by the second shared network as shared media features associated with the N media recognition tasks. The M expert sub-networks included in the first shared network and the M expert sub-networks included in the second shared network have different network parameters, but their corresponding network structures can be the same, both including fully connected sub-networks and activation functions. The fully connected subnetwork is used to perform convolutional processing on the media features, and the activation function is used to activate the convolutional media features input to the fully connected subnetwork (such as nonlinear feature transformation).

[0071] Optionally, the number of shared networks can be three or more. When the number of shared networks is three or more, the feature extraction content of the second shared network can be referenced. The input of the third shared network can be the fusion of the output of the second shared network and the outputs of N second dedicated networks. The input of the fourth shared network can be the fusion of the output of the third shared network and the outputs of N third dedicated networks, and so on, until the output of the last shared network is obtained. The output of the last shared network is used as the shared media feature associated with N media recognition tasks.

[0072] S103, invoke the dedicated network i in the initial recognition model to extract the dedicated media features i associated with the media recognition task i from the sample splicing media features.

[0073] Specifically, the initial recognition model can include a dedicated network for each of the N media recognition tasks. The number of dedicated networks for each media recognition task can be one or more. The computer device can call upon dedicated network i in the initial recognition model to extract dedicated media features i associated with media recognition task i from the sample spliced ​​media features. Dedicated network i is the dedicated network associated with media recognition task i among the N dedicated networks in the initial recognition model. In other words, the computer device can input the sample spliced ​​media features into the dedicated network corresponding to each media recognition task, call upon the dedicated network corresponding to each media recognition task, and extract the dedicated media features associated with the corresponding media recognition task from the sample spliced ​​media features. For example, the dedicated network corresponding to the click-through rate recognition task is used to extract features reflecting whether the sample object clicked on the sample multimedia data, and the dedicated network corresponding to the purchase rate recognition task is used to extract features reflecting whether the sample object purchased the sample multimedia data, etc. In this way, by extracting the dedicated media features associated with the corresponding media recognition task through the dedicated network corresponding to each media recognition task, the seesaw effect problem that occurs when all media recognition tasks share a common network can be avoided, thus improving the accuracy of multi-task recognition. The seesaw effect refers to a situation where some media recognition tasks have high accuracy, while others have low accuracy.

[0074] Optionally, when there is only one dedicated network for each media recognition task, the computer device calls dedicated network i in the initial recognition model to extract the dedicated media feature i associated with media recognition task i from the sample spliced ​​media features. The specific method may include: if there is only one dedicated network i, then the fully connected sub-network in dedicated network i is called to extract the associated media features from the sample spliced ​​media features. The activation sub-network in dedicated network i is then called to perform a non-linear feature transformation on the associated media features to obtain the dedicated media feature i associated with media recognition task i.

[0075] Specifically, the network structure of the dedicated network corresponding to each media recognition task can include a fully connected subnetwork and an activation subnetwork. If there is only one dedicated network i corresponding to media recognition task i, the computer device can input the sample-assembled media features into the fully connected subnetwork of dedicated network i, and then use the fully connected subnetwork of dedicated network i to extract the associated media features related to media recognition task i from the sample-assembled media features. For example, the dedicated network corresponding to the click-through rate recognition task can extract click features reflecting whether the sample object clicked on the sample multimedia data from the sample-assembled media features. It is understandable that since the network parameters of the dedicated network i corresponding to media recognition task i are adjusted by the recognition loss of media recognition task i, the media features extracted by the dedicated network i corresponding to media recognition task i are specific to media recognition task i.

[0076] Furthermore, it can be understood that the fully connected subnetwork in dedicated network i can map the features learned from the sample splicing media features to the sample label space, obtaining associated media features. The computer device can input the associated media features into the activation subnetwork in dedicated network i, and call the activation subnetwork in dedicated network i to perform nonlinear feature transformation on the associated media features, obtaining dedicated media features i associated with media recognition task i. It can be understood that since the associated media features output by the fully connected subnetwork in dedicated network i are low-order features, the activation subnetwork (i.e., the activation function) in dedicated network i can transform the low-order features input by the fully connected subnetwork into high-order features, thereby extracting richer features. Therefore, each media recognition task is configured with a dedicated network, through which dedicated media features associated with the corresponding media recognition task can be extracted. This avoids the seesaw effect caused by the shared media features extracted by the shared network being biased towards some media recognition tasks, thus improving model performance and the accuracy of multi-task recognition.

[0077] It is understandable that when there are multiple shared networks, the number of dedicated networks corresponding to each media recognition task can also be multiple, and the number of shared networks can be the same as the number of dedicated networks. Similarly, when there are multiple dedicated networks, the number of dedicated networks corresponding to each media recognition task can also be multiple and the same, and these dedicated networks for each media recognition task have a hierarchical relationship. For example, media recognition task R001 in N media recognition tasks includes three dedicated network layers: dedicated network layer C001, dedicated network layer C002, and dedicated network layer C003. The hierarchical relationship between these three dedicated network layers can be: dedicated network layer C001 -> dedicated network layer C002 -> dedicated network layer C003, meaning that the network level of dedicated network layer C001 is lower than that of dedicated network layer C002, and the network level of dedicated network layer C002 is lower than that of dedicated network layer C003. Similarly, media recognition task R002 in N media recognition tasks also includes three dedicated network layers.

[0078] Optionally, taking the dedicated network i corresponding to media recognition task i as an example, when dedicated network i includes a first dedicated network i and a second dedicated network i, the specific method by which the computer device calls dedicated network i in the initial recognition model to extract the dedicated media feature i associated with media recognition task i from the sample splicing media features may include: if dedicated network i includes a first dedicated network i and a second dedicated network i, calling the first dedicated network i to extract the initial dedicated media feature associated with media recognition task i from the sample splicing media features; fusing the initial dedicated media feature extracted by the first dedicated network i with the M expert media features extracted by the first shared network to obtain the second fused media feature; and calling the second dedicated network i to extract the dedicated media feature i associated with media recognition task i from the second fused media feature.

[0079] Specifically, the computer device can input the sample-stitched media features into a first dedicated network i, and invoke the first dedicated network i to extract initial dedicated media features associated with media recognition task i from the sample-stitched media features. Specifically, the first dedicated network i may also include fully connected subnetworks and activation subnetworks. The computer device can invoke the fully connected subnetwork in the first dedicated network i to extract associated features from the sample-stitched media features, obtaining associated media features associated with media recognition task i. Further, the activation subnetwork in the first dedicated network i is invoked to perform nonlinear feature transformation on the associated media features extracted by the fully connected subnetwork in the first dedicated network i, obtaining the initial dedicated media features associated with media recognition task i. Since the computer device inputs the sample-stitched media features into a first shared network, it invokes the M expert subnetworks included in the first shared network to extract M expert media features.

[0080] Furthermore, the computer device can invoke the gated subnetwork corresponding to the first dedicated network i to fuse the initial dedicated media features extracted by the first dedicated network i and the M expert media features extracted by the first shared network to obtain the second fused media features. The gated subnetwork corresponding to the first dedicated network i can be a self-attention network, which can automatically learn the commonalities and characteristics between different media recognition tasks. Based on the commonalities and characteristics between different media recognition tasks, it determines the fusion weights corresponding to the initial dedicated media features extracted by the first dedicated network i and the M expert media features extracted by the first shared network. Furthermore, the computer device can invoke the gated subnetwork corresponding to the first dedicated network i to perform weighted processing on the initial dedicated media features extracted by the first dedicated network i according to the fusion weights corresponding to the initial dedicated media features extracted by the first dedicated network i, to obtain weighted initial dedicated media features. The computer device can invoke the gated subnetwork corresponding to the first dedicated network i and the fusion weights corresponding to the M expert media features extracted by the first shared network to perform weighted processing on the M expert media features extracted by the first shared network, to obtain M weighted expert media features. Furthermore, the computer device can invoke the gated subnetwork corresponding to the first dedicated network i to add the weighted initial dedicated media features and the M weighted expert media features to obtain the second fused media features. This avoids feature conflicts between N media recognition tasks, improves model performance, and consequently enhances the accuracy of multi-task recognition.

[0081] Furthermore, the computer device can input the second fused media features into the second dedicated network i, and invoke the second dedicated network i to extract the dedicated media features i associated with the media recognition task i from the second fused media features. Similarly, the second dedicated network i may also include a fully connected subnetwork and an activation subnetwork. The computer device can invoke the fully connected subnetwork in the second dedicated network i to perform associated feature extraction on the second fused media features, obtaining the associated fused media features associated with the media recognition task i. Further, the activation subnetwork in the second dedicated network i is invoked to perform nonlinear feature transformation on the associated fused media features extracted by the fully connected subnetwork in the second dedicated network i, obtaining the dedicated media features i associated with the media recognition task i.

[0082] Optionally, there can be three or more dedicated networks i. When there are three or more dedicated networks i, the feature extraction content of the second dedicated network i can be referenced. The input of the third dedicated network i can be the fusion of the output of the second dedicated network i and the output of the second shared network. The input of the fourth dedicated network i can be the fusion of the output of the third shared network and the output of the third dedicated network i. This process continues until the output of the last dedicated network i is obtained, and the output of the last dedicated network i is used as the dedicated media feature i associated with the media recognition task i. In this way, through multiple dedicated networks, media features can be extracted from samples and concatenated for multi-dimensional extraction, resulting in dedicated media features i associated with the media recognition task i. This makes the dedicated media features i richer and improves the recognition accuracy of the media recognition task i. Optionally, when there are three or more dedicated networks i, the input of the third dedicated network i can be the output of the second dedicated network i, and the input of the fourth dedicated network i can be the fusion of the outputs of the third dedicated network i. This process continues until the output of the last dedicated network i is obtained, and the output of the last dedicated network i is used as the dedicated media feature i associated with the media recognition task i.

[0083] S104, call the recognition network i corresponding to media recognition task i in the initial recognition model, and based on the shared media features and the exclusive media features i, identify the sample object and predict the interactive label for the i-th dimension of the sample multimedia data.

[0084] Specifically, the initial recognition model can include recognition networks corresponding to N media recognition tasks. Each media recognition task has its own dedicated recognition network. These networks identify the predicted interaction labels for the corresponding media recognition task based on the specific media features and shared media features. Specifically, the computer device can input the specific media feature *i* and the shared media feature into the recognition network *i* corresponding to media recognition task *i* in the initial recognition model. The recognition network *i* is then invoked to identify the i-th dimension of the predicted interaction label for the sample multimedia data based on the shared media features and the specific media feature *i*. The i-th dimension of the predicted interaction label corresponds to media recognition task *i*. For example, the recognition network corresponding to the click-through rate (CTR) recognition task can be invoked to identify the CTR label for the sample multimedia data based on the specific media features and shared media features corresponding to the CTR recognition task.

[0085] Optionally, the computer device invokes the recognition network i in the initial recognition model to identify the i-th dimension of the sample object's predicted interactive label for the sample multimedia data based on shared media features and specific media features i. The specific method may include: invoking the gated sub-network included in recognition network i to determine the feature weights corresponding to the shared media features and specific media features i, respectively; invoking the gated sub-network included in recognition network i to perform feature fusion on the shared media features and specific media features i based on their respective feature weights, thereby obtaining the task-fused media features associated with the media recognition task i; and invoking the label prediction sub-network included in recognition network i to predict the label on the task-fused media features, thereby obtaining the i-th dimension predicted interactive label of the sample object for the sample multimedia data.

[0086] Specifically, the computer device can invoke the gating subnetwork included in recognition network i to determine the feature weights corresponding to the shared media features and the specific media features i, respectively. Specifically, the gating subnetwork included in recognition network i can also be a self-attention network. Through the gating subnetwork included in recognition network i, the correlation between the shared media features, the specific media features i, and the media recognition task i can be learned. Further, based on the correlation between the shared media features, the specific media features i, and the media recognition task i, the feature weights corresponding to the shared media features and the specific media features i are determined. For example, the higher the correlation, the higher the corresponding feature weight; conversely, the lower the correlation, the lower the corresponding feature weight. Further, the computer device can invoke the gating subnetwork corresponding to recognition network i to perform weighted processing on the shared media features according to the feature weights corresponding to the shared media features, obtaining the weighted shared media features. Simultaneously, the computer device can invoke the gating subnetwork corresponding to recognition network i to perform weighted processing on the specific media features i according to the feature weights corresponding to the specific media features i, obtaining the weighted specific media features i. The computer device can connect the weighted shared media features and the weighted exclusive media features i to obtain the task-fused media features associated with the media recognition task i.

[0087] Furthermore, the computer device can invoke the label prediction sub-network included in recognition network i to predict labels for the task-fused media features, thereby obtaining the i-th dimension predicted interactive label of the sample object for the sample multimedia data. It can be understood that the label prediction sub-network included in recognition network i is equivalent to a dedicated tower corresponding to media recognition task i. Through the dedicated tower corresponding to media recognition task i, feature classification can be performed on the task-fused media features to obtain the i-th dimension predicted interactive label of the sample object for the sample multimedia data. Therefore, by using the recognition network corresponding to each media recognition task, the predicted interactive label of the sample object for the sample multimedia data under the corresponding media recognition task can be identified.

[0088] Optionally, the specific method by which the computer device determines the feature weights corresponding to the shared media feature and the specific media feature i through the gating sub-network may include: calling the gating sub-network included in the recognition network i to determine a first correlation degree between the shared media feature and the media recognition task i, and to determine a second correlation degree between the specific media feature and the media recognition task i. Based on the first correlation degree, the feature weights corresponding to the shared media feature are generated, and based on the second correlation degree, the feature weights corresponding to the specific media feature i are generated.

[0089] Specifically, the gating subnetwork included in recognition network i can also be a self-attention network. Through the gating subnetwork included in recognition network i, the commonalities and characteristics between different media recognition tasks are automatically learned, thereby determining the first correlation degree between shared media features and media recognition task i, and the second correlation degree between specific media features and media recognition task i. Specifically, since the shared media features include M expert media features, the gating subnetwork included in recognition network i can identify the correlation degree corresponding to each expert media feature. Further, the computer device can call the gating subnetwork included in recognition network i to generate feature weights corresponding to the shared media features based on the first correlation degree, and generate feature weights corresponding to the specific media features i based on the second correlation degree. Specifically, when generating feature weights, the computer device can sum the first and second correlation degrees to obtain a correlation sum, and determine the feature weight corresponding to the shared media feature as the ratio between the first correlation degree and the correlation sum, and determine the feature weight corresponding to the specific media feature i as the ratio between the second correlation degree and the correlation sum. In this way, by controlling the importance of shared media features and specific media features i respectively through the gating subnetworks included in the identification network i, the conflict between different media identification tasks can be alleviated, thus resolving the seesaw effect among N media identification tasks, improving model performance, and consequently enhancing the accuracy of multi-task identification. The identification network i can also include several fully connected subnetworks, which together perform feature identification on the service-integrated media features, outputting the i-th dimension predicted interaction label of the sample object for the sample multimedia data under media identification task i.

[0090] S105, train the initial recognition model based on the N-dimensional labeled interactive labels and the N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping training condition, and obtain the multi-task recognition model.

[0091] Specifically, the computer equipment can train the initial recognition model based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping condition, thus obtaining a multi-task recognition model. The stopping condition can be that the total loss of the initial recognition model is less than a loss threshold, or that the initial recognition model has reached the target number of training iterations.

[0092] Optionally, the computer device trains the initial recognition model based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping condition, obtaining the multi-task recognition model. Specific methods may include: determining the recognition loss value i for media recognition task i based on the i-th dimension labeled interactive label and the i-th dimension predicted interactive label corresponding to the media recognition task; obtaining the initial loss influence weight of the recognition loss value i on the initial recognition model; determining the total recognition loss value for the initial recognition model based on the recognition loss values ​​and initial loss influence weights corresponding to the N media recognition tasks, and the total recognition loss function for the initial recognition model; and training the initial recognition model based on the total recognition loss value and the recognition loss values ​​corresponding to the N media recognition tasks until the trained initial recognition model meets the stopping condition, thus obtaining the multi-task recognition model.

[0093] Specifically, the computer device can obtain the difference between the i-th dimension labeled interactive label and the i-th dimension predicted interactive label corresponding to media recognition task i, as the recognition loss value i for media recognition task i. In this way, the computer device can obtain the recognition loss values ​​corresponding to N media recognition tasks. Further, the computer device can obtain the initial loss influence weight of the recognition loss value i for the initial recognition model. Specifically, the computer device can randomly determine a loss influence weight from the loss influence weight range, as the initial loss influence weight of the recognition loss value i for the initial recognition model; the loss influence weight range can be a range greater than 0. The initial loss influence weight corresponding to each media recognition task can be the same or different. Further, the computer device can determine the total recognition loss value for the initial recognition model based on the recognition loss values ​​and initial loss influence weights corresponding to the N media recognition tasks, as well as the total recognition loss function for the initial recognition model.

[0094] Specifically, when N equals 2, that is, when the N media recognition tasks are the first media recognition task and the second media recognition task, the total recognition loss function of the initial recognition model can be found in the following formula (1).

[0095] (1)

[0096] Among them, in formula (1) The weights of the loss effect in the recognition loss function for the first media recognition task. The weights of the loss effect in the recognition loss function for the second media recognition task. Let be the recognition loss function for the first media recognition task. Let be the recognition loss function for the first media recognition task. The recognition loss function for both the first and second media recognition tasks can be any one or a combination of several of the following: cross-entropy loss function, mean squared error loss function, Euclidean distance loss function, KL (Kullback-Leibler divergence) loss function, etc.

[0097] Furthermore, the computer equipment can train the initial recognition model based on the total recognition loss value and the recognition loss values ​​corresponding to each of the N media recognition tasks, until the trained initial recognition model meets the stopping condition, thus obtaining a multi-task recognition model. Since the initial recognition model is a multi-task model, the recognition loss values ​​of different media recognition tasks will affect the final result of the initial recognition model. Therefore, by setting the loss influence weights corresponding to the recognition loss values ​​of different media recognition tasks, the optimal state of the initial recognition model during training can be accurately determined, thereby improving the model's performance. In other words, by setting the loss influence weights corresponding to the recognition loss values ​​of different media recognition tasks, the initial recognition model can achieve higher performance when the stopping condition is met.

[0098] Optionally, the computer device trains the initial recognition model based on the total recognition loss value and the recognition loss values ​​corresponding to each of the N media recognition tasks until the trained initial recognition model meets the stopping condition. The specific method for obtaining the multi-task recognition model may include: if the total recognition loss value is greater than a loss threshold, adjusting the weights of the initial losses corresponding to each of the N media recognition tasks; adjusting the network parameters in the initial recognition model based on the recognition loss values ​​corresponding to each of the N media recognition tasks to obtain the trained initial recognition model; if the total recognition loss value of the trained initial recognition model is less than or equal to the loss threshold, then the trained initial recognition model is determined to meet the stopping condition, and the trained initial recognition model is identified as the multi-task recognition model.

[0099] Specifically, after obtaining the total recognition loss value of the initial recognition model, the computer device compares the total recognition loss value with a loss threshold. If the total recognition loss value is less than or equal to the loss threshold, the initial recognition model meets the stopping training condition and is designated as a multi-task recognition model. If the total recognition loss value is greater than the loss threshold, it indicates that the initial recognition model does not meet the stopping training condition. In this case, the computer device can adjust the initial loss influence weights corresponding to the N media recognition tasks. Specifically, the computer device can use the gradient descent algorithm to adjust the initial loss influence weights corresponding to the N media recognition tasks. Simultaneously, the computer device can adjust the network parameters in the initial recognition model based on the recognition loss values ​​corresponding to the N media recognition tasks to obtain the trained initial recognition model.

[0100] Furthermore, the computer device can continue to train the initial recognition model using sample-stitched media features, referring to the training method of the initial recognition model, until the total recognition loss value of the trained initial recognition model is obtained. The computer device can compare the total recognition loss value of the trained initial recognition model with a loss threshold. If the total recognition loss value of the trained initial recognition model is greater than the loss threshold, the number of training iterations of the trained initial recognition model is obtained. If the number of training iterations of the trained initial recognition model is less than the target number, it means that the trained initial recognition model has not met the stopping condition, and the initial recognition model continues to be trained using sample-stitched media features. If the number of training iterations of the trained initial recognition model is greater than or equal to the target number, it means that the trained initial recognition model meets the stopping condition, and the trained initial recognition model is determined to be a multi-task recognition model. Specifically, if the total recognition loss value of the trained initial recognition model is less than or equal to the loss threshold, it is determined that the trained initial recognition model meets the stopping condition, and the trained initial recognition model is determined to be a multi-task recognition model.

[0101] Optionally, the computer device adjusts the network parameters in the initial recognition model based on the recognition loss values ​​corresponding to N media recognition tasks to obtain the trained initial recognition model. Specific methods may include: obtaining the descent gradient of recognition network i based on the recognition loss value corresponding to media recognition task i; adjusting the network parameters in recognition network i based on the descent gradient of recognition network i; performing gradient backpropagation on the descent gradient of recognition network i to obtain the descent gradient of dedicated network i; adjusting the network parameters in dedicated network i based on the descent gradient of dedicated network i; and performing gradient backpropagation on the descent gradients of the dedicated networks corresponding to the N media recognition tasks to obtain the descent gradient of the shared network; adjusting the network parameters in the shared network based on the descent gradient of the shared network to obtain the trained initial recognition model.

[0102] Specifically, the computer device can differentiate the loss function of the recognition network i based on the recognition loss value corresponding to media recognition task i, obtaining the loss derivative of recognition network i. Further, the computer device can determine the descent gradient of recognition network i based on the loss derivative of recognition network i and the current network parameters in recognition network i. The computer device can obtain the product between the descent gradient of recognition network i and the learning step size to obtain the parameter adjustment threshold of recognition network i. Furthermore, the computer device can obtain the difference between the current network parameters in recognition network i and the parameter adjustment threshold to obtain the adjusted network parameters of recognition network i. The computer device can then update the current network parameters in recognition network i to the adjusted network parameters corresponding to recognition network i. It is evident that the adjustment of the network parameters in recognition network i is only affected by the recognition loss of media recognition task i and is not affected by the recognition losses of other media recognition tasks, enabling the trained recognition network i to accurately recognize interactive labels related to media recognition task i.

[0103] Furthermore, since the input to recognition network i is the output of dedicated network i, the computer device can perform gradient backpropagation on the descent gradient of recognition network i to obtain the descent gradient of dedicated network i. Specifically, the computer device can determine the recognition loss value of dedicated network i based on the recognition loss value of recognition network i, and then determine the descent gradient of dedicated network i based on the recognition loss value of dedicated network i. The computer device can obtain the product between the descent gradient of dedicated network i and the learning step size to obtain the parameter adjustment threshold corresponding to dedicated network i. The computer device can obtain the difference between the current network parameters in dedicated network i and the parameter adjustment threshold corresponding to dedicated network i to obtain the adjusted network parameters corresponding to dedicated network i, and update the current network parameters in dedicated network i to the adjusted network parameters corresponding to dedicated network i. It can be seen that, since the input to recognition network i is the output of dedicated network i, during gradient direction propagation, the parameter adjustment in dedicated network i is only affected by the recognition loss of media recognition task i, and is not affected by the recognition loss of other media recognition tasks, which enables the trained dedicated network i to accurately extract the media features required by media recognition task i. In other words, the dedicated network i is exclusive to media recognition task i and will not be affected by other media recognition tasks. This can effectively solve the seesaw effect between N media recognition tasks, improve the performance of the model, and thus improve the accuracy of multi-task recognition.

[0104] Furthermore, since the shared media output by the shared network serves as the input to N dedicated networks corresponding to N media recognition tasks, the computer device can backpropagate the descent gradients of the dedicated networks corresponding to each of the N media recognition tasks to obtain the descent gradient of the shared network. Specifically, the computer device can determine the recognition loss value of the shared network based on the recognition loss values ​​corresponding to the N dedicated networks, and then determine the descent gradient of the shared network based on the recognition loss value of the shared network. Further, the computer device can obtain the product between the descent gradient of the shared network and the learning step size to obtain the parameter adjustment threshold corresponding to the shared network. The computer device can obtain the difference between the current network parameters in the shared network and the parameter adjustment threshold corresponding to the shared network to obtain the adjusted network parameters corresponding to the shared network, and update the current network parameters in the shared network to the adjusted network parameters corresponding to the shared network to obtain the trained initial recognition model. It can be seen that since the network parameters in the shared network are adjusted based on the recognition losses of the N media recognition tasks, the trained shared network can extract shared media features associated with the N media recognition tasks. In this way, multiple media recognition tasks can be built into one model, which can improve the efficiency of multi-task recognition. At the same time, through the shared network, N media recognition tasks can share sample data (i.e. training data), which can solve the problem of insufficient model training due to insufficient sample data for some media recognition tasks.

[0105] Optionally, the specific method by which the computer device adjusts the initial loss influence weights corresponding to the N media recognition tasks may include: Differentiating the total recognition loss function based on the initial loss influence weight i to obtain the loss derivative with respect to the initial loss influence weight i; obtaining the descent gradient with respect to the initial loss influence weight i based on the loss derivative and the recognition loss value corresponding to media recognition task i; obtaining the product between the descent gradient of the initial loss influence weight i and the learning step size to obtain the parameter adjustment threshold; obtaining the difference between the initial loss influence weight corresponding to the initial loss influence weight i and the parameter adjustment threshold to obtain the adjusted loss influence weight; and updating the initial loss influence weight corresponding to the initial loss influence weight i to the adjusted loss influence weight.

[0106] Specifically, the computer device differentiates the total recognition loss function based on the initial loss influence weight *i*, obtaining the loss derivative with respect to the initial loss influence weight *i*. Substituting the recognition loss value corresponding to media recognition task *i* into the loss derivative with respect to the initial loss influence weight *i*, the computer device obtains the descent gradient with respect to the initial loss influence weight *i*. Furthermore, the computer device can obtain the product between the descent gradient of the initial loss influence weight *i* and the learning step size, thus obtaining the parameter adjustment threshold corresponding to the initial loss influence weight *i*. The learning step size can be adjusted according to specific circumstances; this application does not impose any restrictions on the learning step size. Further, the computer device can obtain the adjusted loss influence weight by taking the difference between the initial loss influence weight corresponding to the initial loss influence weight *i* and the parameter adjustment threshold, and update the initial loss influence weight corresponding to the initial loss influence weight *i* to the adjusted loss influence weight. It is evident that the initial recognition model can automatically learn the loss influence weights while minimizing the overall loss, enabling the initial recognition model to reach its optimal state after training, thereby improving the performance of the multi-task recognition model and ultimately increasing the accuracy of multi-task recognition.

[0107] like Figure 4 As shown, Figure 4 This is a schematic diagram of the model structure of an initial recognition model provided in an embodiment of this application, as shown below. Figure 4 As shown, the computer device can acquire a sample data set 40a, which includes media attribute data of sample multimedia data and historical media interaction data of sample objects. This data includes sparse data 1, ..., sparse data r, and statistical data. The computer device can perform vector queries on the sparse feature data to obtain sparse feature 1, ..., sparse feature r, and simultaneously perform vector transformation on the statistical data to obtain statistical features. Further, the computer device can input the sparse features and statistical features into a feature concatenation network 40b, concatenating the sparse features and statistical features to obtain sample concatenated media features. Figure 4 As shown, taking one shared network and one dedicated network for each media recognition task as an example, and N media recognition tasks as the first media recognition task and the second media recognition task, the computer device can input the sample splicing features into the shared network 40c in the initial recognition model, and input them into the dedicated network corresponding to each media recognition task in the initial recognition model respectively.

[0108] Specifically, the computer device invokes the shared network 40c to extract shared media features associated with N media recognition tasks (i.e., the first media recognition task and the second media recognition task) from the sample splicing media features. The shared network 40c may include M expert subnetworks. The extraction process of the shared media features can be found in step S102 above, and will not be repeated here. Simultaneously, the computer device can invoke the dedicated network 40d corresponding to the first media recognition task to extract dedicated media features associated with the first media recognition task from the sample splicing media features, and invoke the dedicated network 40e corresponding to the second media recognition task to extract dedicated media features associated with the second media recognition task from the sample splicing media features. The extraction of dedicated media features can be found in step S103 above.

[0109] Specifically, the feature extraction logic of the dedicated network 40d corresponding to the first media recognition task can be found in the following formula (2).

[0110] R1_FC=MatMul( (embedding)(2)

[0111] In formula (2), R1_FC represents the dedicated network 40d corresponding to the first media recognition task, and MatMul represents matrix multiplication. The parameters are defined in the dedicated network 40d for the first media recognition task, and the embedding is the media feature of the sample splicing.

[0112] Specifically, the feature extraction logic of the dedicated network 40e corresponding to the second media recognition task is similar to the feature extraction logic of the dedicated network 40d corresponding to the first media recognition task, as shown in the following formula (3).

[0113] R2_FC=MatMul( (embedding)(3)

[0114] In formula (3), R2_FC represents the dedicated network 40e corresponding to the second media recognition task, and MatMul represents matrix multiplication. For the dedicated network 40e for the second media recognition task, embedding represents the media features of the sample splicing.

[0115] Specifically, the feature extraction logic of the shared network 40c can be found in the following formula (4).

[0116] SHARE_FC=MatMul( (embedding)(4)

[0117] In formula (4), SHARE_FC represents the shared network 40c, and MatMul represents matrix multiplication. The network parameters in the shared network 40c are defined, and the embedding is the feature of the sample splicing media.

[0118] Furthermore, the computer device can input the sample-assembled media features into the gated subnetwork 40f corresponding to the first media recognition task, and call the gated subnetwork 40f to determine the feature weights corresponding to the specific media features and shared media features corresponding to the first media recognition task. Based on the feature weights corresponding to the specific media features and shared media features corresponding to the first media recognition task, the specific media features and shared media features corresponding to the first media recognition task are fused to obtain the task-fused media features of the first media recognition task. Similarly, the computer device can input the sample-assembled media features into the gated subnetwork 40g corresponding to the second media recognition task, and call the gated subnetwork 40g to determine the feature weights corresponding to the specific media features and shared media features corresponding to the second media recognition task. Based on the feature weights corresponding to the specific media features and shared media features corresponding to the second media recognition task, the specific media features and shared media features corresponding to the second media recognition task are fused to obtain the task-fused media features of the second media recognition task.

[0119] Specifically, the feature fusion logic of the gating subnetwork 40f corresponding to the first media recognition task can be found in the following formula (5).

[0120] MK1_FC= R1_FC+ SHARE_FC(5)

[0121] Wherein, MK1_FC is the gated sub-network 40f corresponding to the first media recognition task, and R1_FC is the specific media feature corresponding to the first media recognition task. Here, SHRAE_FC represents the feature weights of the specific media features corresponding to the first media recognition task, and SHRAE_FC represents the shared media features input to the shared network. The feature weights are for shared media features. and Equals softmax(MatMul( ,embedding)) The network parameters are those in the gated subnetwork 40f corresponding to the first media recognition task.

[0122] Specifically, the feature fusion logic of the gated subnetwork 40g corresponding to the second media recognition task is similar to the feature fusion logic of the gated subnetwork 40f corresponding to the first media recognition task, as shown in the following formula (6).

[0123] MK2_FC= R1_FC+ SHARE_FC(6)

[0124] Wherein, MK2_FC is the gated sub-network 40g corresponding to the second media recognition task, and R2_FC is the specific media feature corresponding to the second media recognition task. Here, SHARE_FC represents the feature weights of the specific media features corresponding to the second media identification task, and SHARE_FC represents the shared media features output by the shared network. The feature weights are for shared media features. and Equals softmax(MatMul( ,embedding)) The network parameters in the gated subnetwork 40g corresponding to the second media recognition task.

[0125] Furthermore, the computer device invokes the fully connected network 40h and the prediction network 40i corresponding to the first media recognition task. Based on the business-converged media features corresponding to the first media recognition task, it identifies the predicted interaction tags of the sample object for the sample multimedia data under the first media recognition task. The recognition network corresponding to the first media recognition task includes a gating sub-network 40f, a fully connected network 40h, and a prediction network 40i. For example, the first media recognition task could be a click-through rate (CTR) recognition task. The computer device invokes the recognition network corresponding to the CTR recognition task and, based on the business-converged media features corresponding to the CTR recognition task, identifies the predicted CTR tags of the sample object for the sample multimedia data under the CTR recognition task. Similarly, the computer device invokes the fully connected network 40k and the prediction network 40l corresponding to the second media recognition task. Based on the business-converged media features corresponding to the second media recognition task, it identifies the predicted interaction tags of the sample object for the sample multimedia data under the second media recognition task. The recognition network corresponding to the second media recognition task includes a gating sub-network 40g, a fully connected network 40k, and a prediction network 40l. For example, the second media identification task could be the purchase rate identification task, which calls the identification network corresponding to the purchase rate identification task and identifies the sample object's predicted purchase rate label for the sample multimedia data under the purchase rate identification task based on the business converged media characteristics corresponding to the purchase rate identification task.

[0126] Furthermore, the computer device can determine the recognition loss 40j corresponding to the first media recognition task based on the predicted interaction labels and labeled interaction labels (which can be labeled by the administrator). Similarly, the computer device can determine the recognition loss 40m corresponding to the second media recognition task based on the predicted interaction labels and labeled interaction labels (which can be labeled by the administrator). Then, based on the recognition losses corresponding to the first and second media recognition tasks, the initial recognition model is trained to obtain a multi-task recognition model.

[0127] like Figure 5 As shown, Figure 5 This is a schematic diagram of a network structure of a gated subnetwork provided in an embodiment of this application, such as... Figure 5 As shown, the gated subnetwork includes fully connected layers and a normalization function (such as a softmax function). Specifically, taking the gated subnetwork 40f corresponding to the first media recognition task as an example to obtain the service fusion media features corresponding to the first media recognition task, the computer device can input the sample spliced ​​media features, the exclusive media features corresponding to the first media recognition task, and the shared media features into the gated subnetwork 40f. Through the fully connected layers in the gated subnetwork 40f, feature extraction is performed on the sample spliced ​​media features to learn the commonalities and characteristics between different media recognition tasks. Furthermore, the computer device can input the media features output by the fully connected layers in the gated subnetwork 40f into the normalization function in the gated subnetwork 40f. The normalization function layer normalizes the media features output from the fully connected layer. It takes as input the feature weight w1 corresponding to the specific media features of the first media recognition task and outputs the feature weight w2 corresponding to the shared media features. Further, the computer device can invoke the gated sub-network 40f to obtain the product between the feature weight w1 and the specific media features, resulting in weighted specific media features; and obtain the product between the feature weight w2 and the shared media features, resulting in weighted shared media features. The weighted specific media features and the weighted shared media features are then fused to obtain the service-fused media features corresponding to the first media recognition task.

[0128] like Figure 6 As shown, Figure 6 This is a schematic diagram of the model structure of an initial recognition model provided in an embodiment of this application, as shown below. Figure 6 As shown, relative to Figure 4 The initial recognition model in Figure 6In the initial recognition model, each media recognition task corresponds to two dedicated networks, and there are also two shared networks. Taking N media recognition tasks as the first and second media recognition tasks as an example, the computer device can acquire a sample data set 60a. This sample data set 60a includes media attribute data of sample multimedia data and historical media interaction data of sample objects. The media attribute data of sample multimedia data and the historical media interaction data of sample objects include sparse data 1, ..., sparse data r, and statistical data. The computer device can perform vector queries on the sparse feature data to obtain sparse feature 1, ..., sparse feature r, and simultaneously perform vector transformation on the statistical data to obtain statistical features. Furthermore, the computer device can input the sparse features and statistical features into the splicing network 60b, and perform feature splicing on the sparse features and statistical features to obtain sample spliced ​​media features.

[0129] It is understood that both the first shared network 60c and the second shared network 60i include M expert subnetworks. The computer device can input the sample spliced ​​media features into the first shared network 60c, call the M expert subnetworks included in the first shared network 60c, and extract M expert media features associated with N media recognition tasks from the sample spliced ​​media features. The extraction process of the M expert media features extracted by the first shared network 60c can be found in step S102 above, and will not be repeated here. Simultaneously, the computer device can input the sample spliced ​​media features into the first dedicated network 60d corresponding to the first media recognition task, call the first dedicated network 60d, and extract initial dedicated media features associated with the first media recognition task from the sample spliced ​​media features. The computer device can input the sample spliced ​​media features into the first dedicated network 60d corresponding to the second media recognition task, call the first dedicated network 60e, and extract initial dedicated media features associated with the second media recognition task from the sample spliced ​​media features. The extraction content of the initial dedicated media features can be found in step S103 above.

[0130] Furthermore, the computer device can invoke the gated sub-network 60g corresponding to the first shared network 60c to fuse the N initial proprietary media features extracted by the N first proprietary networks and the M expert media features extracted by the first shared network 60c to obtain the first fused media features. Further, the computer device can invoke the M expert sub-networks included in the second shared network 60i to extract shared media features associated with the N media recognition tasks from the first fused media features. For details, please refer to step S102 above; this embodiment will not be repeated here.

[0131] Specifically, the feature extraction logic of the gated sub-network 60g corresponding to the first shared network 60c can be found in the following formula (7).

[0132] SHARE_MK= R1_FC1+ SHARE_FC1+ R2_FC1(7)

[0133] In formula (7), SHARE_MK is the gated sub-network 60g corresponding to the first shared network 60c, R1_FC1 is the initial exclusive media feature output by the first exclusive network 60d, and R2_FC is the initial exclusive media feature output by the first exclusive network 60e. The feature weights are the initial proprietary media features output by the first proprietary network 60d. Here, SHAER_FC1 represents the feature weights corresponding to the initial proprietary media features output by the first proprietary network 60e, and SHAER_FC1 represents the M expert media features extracted by the first shared network 60c. The feature weights are the M expert media features extracted from the first shared network 60c. Among them, , as well as Equals softmax(MatMul( ,embedding)) These are the network parameters in the gated subnet 60g.

[0134] Furthermore, the computer device can invoke the gated sub-network 60f corresponding to the first dedicated network 60d to fuse the initial dedicated media features extracted by the first dedicated network 60d and the M expert media features extracted by the first shared network 60c to obtain a second fused media feature. The computer device can further invoke the second dedicated network 60j corresponding to the first media recognition task to extract the dedicated media features associated with the first media recognition task from the second fused media feature. Similarly, the computer device can invoke the gated sub-network 60h corresponding to the first dedicated network 60e to fuse the initial dedicated media features extracted by the first dedicated network 60e and the M expert media features extracted by the first shared network 60c to obtain a third fused media feature. The computer device can further invoke the second dedicated network 60k corresponding to the second media recognition task to extract the dedicated media features associated with the second media recognition task from the third fused media feature. The feature extraction logic of the first dedicated network 60d, the second dedicated network 60j, the first dedicated network 60e, and the second dedicated network 60k are similar and can be referred to in formula (2) or formula (3) above. This application will not repeat them. The feature extraction logic of the gated subnetwork 60f and the gated subnetwork 60h can be referred to in formula (5) or formula (6) above. This application will not repeat them.

[0135] Furthermore, the computer device can input the second converged media feature output by the gated subnetwork 60f into the gated subnetwork 60l, and invoke the gated subnetwork 60l to fuse the exclusive media feature output by the second dedicated network 60j and the shared media feature output by the second shared network 60i to obtain the service converged media feature corresponding to the first media identification task. Similarly, the computer device can input the second converged media feature output by the gated subnetwork 60h into the gated subnetwork 60m, and invoke the gated subnetwork 60m to fuse the exclusive media feature output by the second dedicated network 60k and the shared media feature output by the second shared network 60i to obtain the service converged media feature corresponding to the second media identification task.

[0136] The feature fusion logic of the gated subnetwork 60l can be found in the following formula (8).

[0137] MK3_FC= R1_FC2+ SHAER_FC2(8)

[0138] In formula (8), MK3_FC is the gated sub-network 60l, and R1_FC2 is the exclusive media feature output by the second exclusive network 60j. SHAER_FC2 represents the feature weights of the proprietary media features output by the second proprietary network 60j, and SHAER_FC2 represents the shared media features output by the second shared network 60i. The feature weights are for shared media features. and Equals softmax(MatMul( ,embedding)) These are the network parameters in the gated subnetwork 60l. The feature fusion logic of the gated subnetwork 60m is similar to that of the gated subnetwork 60l, and will not be described further here.

[0139] Furthermore, the computer device invokes the fully connected network 60n and the prediction network 60p corresponding to the first media recognition task. Based on the business-integrated media features corresponding to the first media recognition task, it identifies the predicted interaction tags of the sample objects for the sample multimedia data under the first media recognition task. For example, the first media recognition task could be a click-through rate (CTR) recognition task. The computer device invokes the recognition network corresponding to the CTR recognition task and, based on the business-integrated media features corresponding to the CTR recognition task, identifies the predicted CTR tags of the sample objects for the sample multimedia data under the CTR recognition task. Similarly, the computer device invokes the fully connected network 60o and the prediction network 60q corresponding to the second media recognition task. Based on the business-integrated media features corresponding to the second media recognition task, it identifies the predicted interaction tags of the sample objects for the sample multimedia data under the second media recognition task. For example, the second media recognition task could be a purchase rate recognition task. The computer device invokes the recognition network corresponding to the purchase rate recognition task and, based on the business-integrated media features corresponding to the purchase rate recognition task, identifies the predicted purchase rate tags of the sample objects for the sample multimedia data under the purchase rate recognition task. Furthermore, the computer device can determine the recognition loss 60r corresponding to the first media recognition task based on the predicted interaction tags and labeled interaction tags (which can be labeled by administrators). Similarly, the computer device can determine the recognition loss 60s corresponding to the second media recognition task based on the predicted interaction label and the labeled interaction label (which can be labeled by the administrator). Then, based on the recognition loss corresponding to the first media recognition task and the recognition loss corresponding to the second media recognition task, the initial recognition model is trained to obtain a multi-task recognition model.

[0140] In this embodiment, a multi-task recognition model is constructed based on sample-stitched media features. This model can identify multi-dimensional interactive tags of objects in multimedia data, with each interactive tag corresponding to a media recognition task. Therefore, the multi-task recognition model can handle multiple media recognition tasks without requiring a separate model for each task, reducing resource overhead during training and improving training efficiency. During training, an initial recognition model is first constructed, comprising a shared network, multiple dedicated networks, and multiple recognition networks. Each dedicated network and recognition network corresponds to a media recognition task. When identifying the i-th predicted interactive tag for media recognition task i, the i-th predicted interactive tag of the sample object in multimedia data is identified by combining the recognition network i in the initial model with the shared media features and the dedicated media features i associated with media recognition task i. The initial recognition model is then trained based on the N-dimensional labeled interactive tags and the N-dimensional predicted interactive tags to obtain the multi-task recognition model. The shared media features here are extracted from the sample splicing media features by the shared network of the initial recognition model, while the specific media feature i is extracted from the sample splicing media features by the specific network in the initial recognition model associated with media recognition task i. In other words, when recognizing the i-th predicted interaction label under media recognition task i, not only the specific media feature i directly associated with media recognition task i is combined, but also the shared media features indirectly associated with media recognition task i are combined, providing more information for recognizing the i-th predicted interaction label under media recognition task i. That is, different media recognition tasks can share shared media features, which provides more training data for the training process of different media recognition tasks. This avoids the problem of low media recognition accuracy of the initial recognition model (i.e., task recognition model) after training due to the sparsity of training data, thus improving the media recognition accuracy of the initial recognition model after training and thereby improving the accuracy of multimedia data push. In addition, this application can improve the recognition accuracy of the multi-task recognition model by setting the loss influence weight of the loss function of different media recognition tasks.

[0141] Further, please see Figure 7 , Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 7 As shown, this method can be derived from... Figure 1 It can be executed by any terminal device in the system, or by... Figure 1 Server 10 in the middle can be used to execute it, and it can also be executed by Figure 1 The terminal device and server in this application work together to execute the method. The device used to execute this method can be collectively referred to as a computer device. The data processing method may include, but is not limited to, the following steps:

[0142] S201, Obtain the characteristics of the media used for business splicing.

[0143] Specifically, when it is necessary to push a target object to a business entity, the computer device can generate corresponding multimedia data to be pushed. For example, in a game scenario, if a new game item is developed, advertising data corresponding to the new game item can be generated and accurately pushed to the business entity to promote the new game item. Specifically, after training the initial recognition model to obtain a multi-task recognition model, the computer device can use the multi-task recognition model to identify the N-dimensional recognition interaction tags of the business entity for the multimedia data to be pushed. The N-dimensional recognition interaction tags correspond to N media recognition tasks. These N media recognition tasks can include click-through rate (CTR) recognition tasks, shallow conversion rate (SCR) recognition tasks, and deep conversion rate (DCR) recognition tasks, etc. The N-dimensional recognition interaction tags can include CTR tags under the CTR recognition task, shallow conversion rate tags under the shallow conversion rate recognition task, and deep conversion rate tags under the DCR recognition task, etc. Specifically, the computer device can obtain the historical media interaction characteristics of the business entity currently requiring multimedia data push, as well as the media attribute characteristics of the multimedia data to be pushed. The historical media interaction features of the business object can include the media attribute features of multimedia data that has been pushed to the business object within a historical time period (such as the previous week, the previous month, etc.), the historical interaction features of the business object with the pushed multimedia data, and the object attribute features of the business object. The computer device can perform feature concatenation on the historical media interaction features of the business object and the media attribute features of the multimedia data to be pushed to obtain the business concatenated media features. Then, it can call a multi-task recognition model to identify the N-dimensional recognition interaction tags of the business object with respect to the multimedia data to be pushed based on the business concatenated media features.

[0144] Specifically, the media attribute characteristics of the multimedia data to be pushed can include identifying features (such as sparse features) and statistical features (such as dense features). For example, when the multimedia data to be pushed is advertising data, its media attribute characteristics can include the price, purpose, theme, and appearance of the advertising object (such as an item or virtual character). When the multimedia data to be pushed is news data or game data, its media attribute characteristics can include the content theme and viewing duration. Similarly, the historical media interaction characteristics of the business object can also include identifying features (such as sparse features) and statistical features (such as dense features). The historical interaction characteristics of the business object with the pushed multimedia data can include the interaction behavior of the business object with the pushed multimedia data, and the number of times different interaction behaviors occur. For example, whether the business object clicked on the pushed multimedia data and the number of times it clicked; whether the business object made a purchase after clicking on the pushed multimedia data and the number of times it made a purchase; whether the business object commented after clicking on the pushed multimedia data and the number of times it commented, etc. The object attribute characteristics of the business object can include the object identifier and object age, etc. The media attribute characteristics of the pushed multimedia data can include the media content characteristics of the pushed multimedia data. For example, if the pushed multimedia data is advertising data, the media attribute characteristics of the pushed multimedia data can include the identifier, price, purpose, theme, appearance, and other attributes of the advertising object (such as an item or virtual character) in the advertising data. If the pushed multimedia data is news data or game data, the media attribute characteristics of the pushed multimedia data can include the content theme, content viewing duration, and other attributes.

[0145] S202, invoke the shared network in the multi-task recognition model to extract the shared media features associated with N media recognition tasks from the media features of the service splicing.

[0146] Specifically, the computer device inputs the service splicing media features into the multi-task recognition model, calls the shared network in the multi-task recognition model, and extracts the service shared media features associated with N media recognition tasks from the service splicing media features. Specifically, the number of shared networks can be one or more, and the shared network can include M expert sub-networks. The computer device can call one or more of the M expert sub-networks included in the shared network to extract the service shared media features associated with N media recognition tasks from the service splicing media features. The feature extraction method of the shared network in the multi-task recognition model is the same as that of the shared network in the initial recognition model. The acquisition of service shared media features can refer to the acquisition of shared media features in step S102 above, which will not be repeated here in this embodiment. It is understood that different media recognition tasks can share the network parameters of the shared network in the multi-task recognition model and perform feature extraction through the network parameters in the shared network. Furthermore, since the network parameters of the shared network in the multi-task recognition model are obtained based on joint training of N media recognition tasks, the service shared media features mentioned by the shared network in the multi-task recognition model are required by N media recognition tasks.

[0147] S203, invoke the dedicated network i in the multi-task recognition model to extract the business-specific media feature i associated with the media recognition task i from the business splicing media features.

[0148] Specifically, the computer device inputs the service splicing media features into the multi-task recognition model, calls the dedicated network i in the multi-task recognition model, and extracts the service-specific media feature i associated with media recognition task i from the service splicing media features. In this way, the computer device can call N dedicated networks in the multi-task recognition model to extract N service-specific media features associated with N media recognition tasks from the service splicing media features. In this way, by extracting the service-specific media features associated with the corresponding media recognition task through the dedicated network corresponding to each media recognition task, the seesaw effect problem that occurs when all media recognition tasks share a network can be avoided, and the accuracy of multi-task recognition can be improved. The content of the extraction of service-specific media feature i can be referred to the content of the extraction of dedicated media feature i in step S103 above, and will not be repeated here in this embodiment.

[0149] S204, invoke the recognition network i for media recognition task i in the multi-task recognition model, and identify the i-th dimension recognition interaction label of the business object for the multimedia data to be pushed based on the business-shared media features and business-specific media features i.

[0150] Specifically, the computer device can input the service-specific media feature i into the recognition network i of the multi-task recognition model for media recognition task i, and call the gating sub-network in the recognition network i to determine the feature weights corresponding to the service-shared media feature and the service-specific media feature i, respectively. Based on the feature weights corresponding to the service-shared media feature and the service-specific media feature i, feature fusion is performed on the service-shared media feature and the service-specific media feature i to obtain the service-fused media feature. In this way, the gating mechanism can control the importance of the service-shared media feature and the service-specific media feature i, effectively mitigating the feature conflict problem between N media recognition tasks. Further, the computer device can use the label prediction sub-network in the recognition network i to perform label prediction on the service-fused media feature, obtaining the i-th dimension recognition interaction label of the service object for the multimedia data to be pushed. This i-th dimension recognition interaction label corresponds to the media recognition task i, that is, the i-th dimension recognition interaction label is the interaction label of the service object for the multimedia data to be pushed under the media recognition task i. The recognition content of the i-th dimension recognition interaction label can refer to the content of step S104 above, and will not be repeated here in this embodiment.

[0151] S205: Based on the N-dimensional recognition interaction tags, push the multimedia data to be pushed to the business object.

[0152] Specifically, N-dimensional interactive tags can include click-through rate tags, shallow conversion rate tags, and deep conversion rate tags. Computer devices can use these N-dimensional interactive tags to push multimedia data to target users. This improves the accuracy and efficiency of multimedia data delivery.

[0153] Optionally, the specific method by which the computer device pushes multimedia data to be pushed to the business object based on the N-dimensional recognition interaction tag may include: determining the push score of the multimedia data to be pushed based on the N-dimensional recognition interaction tag; and pushing the multimedia data to be pushed to the business object with a recommendation score greater than the score threshold.

[0154] Computer devices can determine the push score of multimedia data to be pushed based on N-dimensional recognition interaction tags. Specifically, the computer device can perform a weighted summation of the N-dimensional recognition interaction tags to obtain a total interaction tag, which is then used as the push score of the multimedia data to be pushed. Furthermore, the computer device can push multimedia data with a recommendation score greater than a threshold value to the target business object; this threshold value can be set according to specific circumstances.

[0155] In this embodiment, after obtaining the service splicing media features, these features can be input into a multi-task recognition model. The shared network within the model is invoked to extract shared media features associated with N media recognition tasks. Further, a dedicated network i within the model is invoked to extract dedicated media features i associated with media recognition task i. Then, the recognition network i for media recognition task i is invoked, and based on the shared and dedicated media features i, the i-th dimension of the interactive label for the multimedia data to be pushed is identified. Based on the N-dimensional interactive label, the multimedia data is pushed to the business object. Thus, by using the N-dimensional interactive label identified by the highly accurate multi-task recognition model to push multimedia data to the business object, the accuracy of multimedia data push can be improved.

[0156] Further, please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. The data processing device may include: a first acquisition module 11, a first extraction module 12, a second extraction module 13, a first identification module 14, a training module 15, a second acquisition module 16, a third extraction module 17, a fourth extraction module 18, a second identification module 19, and a push module 20.

[0157] The first acquisition module 11 is used to acquire sample splicing media features and N-dimensional annotation interaction labels of sample objects for sample multimedia data; the sample splicing media features are obtained by splicing the media attribute features of sample multimedia data and the historical media interaction features of sample objects; the N-dimensional annotation interaction labels correspond to N media recognition tasks; N is an integer greater than 1;

[0158] The first extraction module 12 is used to call the shared network in the initial recognition model to extract shared media features associated with N media recognition tasks from the sample splicing media features;

[0159] The second extraction module 13 is used to call the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample splicing media features; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks of the initial recognition model, and i is a positive integer less than or equal to N.

[0160] The first identification module 14 is used to call the identification network i in the initial identification model for media identification task i, and to identify the i-th dimension of the sample object for the sample multimedia data based on the shared media features and the exclusive media features i.

[0161] Training module 15 is used to train the initial recognition model based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping training condition, thus obtaining a multi-task recognition model.

[0162] The first extraction module 12 includes a first extraction unit 1201 and a first determination unit 1202.

[0163] The first extraction unit 1201 is used to call the M expert sub-networks included in the shared network if the number of shared networks is one, and extract M expert media features associated with N media recognition tasks from the sample splicing media features, where M is a positive integer;

[0164] The first determining unit 1202 is used to determine the extracted M expert media features as shared media features associated with N media recognition tasks.

[0165] The second extraction module 13 includes a second extraction unit 1301 and a feature conversion unit 1302.

[0166] The second extraction unit 1301 is used to call the fully connected sub-network in the dedicated network i if the number of dedicated networks i is one, and extract the associated media features related to the media recognition task i from the sample splicing media features.

[0167] The feature transformation unit 1302 is used to call the activation sub-network in the dedicated network i to perform nonlinear feature transformation on the associated media features to obtain the dedicated media features i associated with the media recognition task i.

[0168] The first extraction module 12 further includes: a third extraction unit 1203, a first fusion unit 1204, and a fourth extraction unit 1205.

[0169] The third extraction unit 1203 is used to call the M expert subnetworks included in the first shared network if the shared network includes the first shared network and the second shared network, and extract M expert media features associated with N media recognition tasks from the sample splicing media features, where M is a positive integer;

[0170] The first fusion unit 1204 is used to fuse N initial proprietary media features and M expert media features extracted by the first shared network to obtain the first fused media feature; the N initial proprietary media features are extracted by the N first proprietary networks in the initial recognition model, and the N first proprietary networks and the first shared network are located in the same network layer.

[0171] The fourth extraction unit 1205 is used to call the M expert sub-networks included in the second shared network to extract the shared media features associated with the N media recognition tasks from the first fused media features.

[0172] The second extraction module 13 further includes a fifth extraction unit 1303, a second fusion unit 1304, and a sixth extraction unit 1305.

[0173] The fifth extraction unit 1303 is used to call the first exclusive network i if the exclusive network i includes the first exclusive network i and the second exclusive network i, and extract the initial exclusive media features associated with the media recognition task i from the sample splicing media features.

[0174] The second fusion unit 1304 is used to fuse the initial exclusive media features extracted by the first exclusive network i and the M expert media features extracted by the first shared network to obtain the second fused media features.

[0175] The sixth extraction unit 1305 is used to call the second dedicated network i to extract the dedicated media feature i associated with the media recognition task i from the second fused media features.

[0176] The initial recognition model includes a gating subnetwork and a tag prediction subnetwork for media recognition task i; the first recognition module 14 includes a second determination unit 1401, a third fusion unit 1402 and a tag prediction unit 1403.

[0177] The second determining unit 1401 is used to call the gated sub-network included in the identification network i to determine the feature weights corresponding to the shared media feature and the exclusive media feature i respectively.

[0178] The third fusion unit 1402 is used to call the gated sub-network included in the recognition network i, and perform feature fusion on the shared media feature and the exclusive media feature i according to the feature weights corresponding to the shared media feature and the exclusive media feature i respectively, to obtain the task fused media feature associated with the media recognition task i.

[0179] The label prediction unit 1403 is used to call the label prediction sub-network included in the recognition network i to perform label prediction on the task fusion media features and obtain the i-th dimension predicted interactive label of the sample object for the sample multimedia data.

[0180] Specifically, the second determining unit 1401 is used for:

[0181] Invoke the gated subnetworks included in the identification network i to determine the first degree of correlation between shared media features and media identification task i, and to determine the second degree of correlation between exclusive media features and media identification task i;

[0182] Based on the first degree of correlation, generate feature weights corresponding to the shared media features;

[0183] Based on the second degree of correlation, the feature weights corresponding to the specific media feature i are generated.

[0184] The training module 15 includes: a third determination unit 1501, an acquisition unit 1502, a fourth determination unit 1503, and a training unit 1504.

[0185] The third determining unit 1501 is used to determine the recognition loss value i for media recognition task i based on the i-th dimension labeled interaction label and the i-th dimension predicted interaction label corresponding to media recognition task i.

[0186] The acquisition unit 1502 is used to acquire the initial loss influence weight i of the recognition loss value i on the initial recognition model;

[0187] The fourth determining unit 1503 is used to determine the total recognition loss value of the initial recognition model based on the recognition loss value and initial loss influence weight corresponding to the N media recognition tasks, as well as the total recognition loss function of the initial recognition model.

[0188] Training unit 1504 is used to train the initial recognition model based on the total recognition loss value and the recognition loss values ​​corresponding to the N media recognition tasks, until the trained initial recognition model meets the stopping training condition, thus obtaining the multi-task recognition model.

[0189] Specifically, training unit 1504 is used for:

[0190] If the total recognition loss value is greater than the loss threshold, the initial loss impact weights corresponding to the N media recognition tasks will be adjusted.

[0191] Based on the recognition loss values ​​corresponding to the N media recognition tasks, the network parameters in the initial recognition model are adjusted to obtain the trained initial recognition model.

[0192] If the total recognition loss of the initial recognition model after training is less than or equal to the loss threshold, then the initial recognition model after training is determined to meet the stop training condition.

[0193] The initial recognition model after training was determined to be a multi-task recognition model.

[0194] Specifically, based on the recognition loss values ​​corresponding to the N media recognition tasks, the network parameters in the initial recognition model are adjusted to obtain the trained initial recognition model, including:

[0195] The descent gradient of recognition network i is obtained based on the recognition loss value corresponding to media recognition task i, and the network parameters in recognition network i are adjusted based on the descent gradient of recognition network i.

[0196] Gradient backpropagation is performed on the descent gradient of the recognition network i to obtain the descent gradient of the dedicated network i. Based on the descent gradient of the dedicated network i, the network parameters in the dedicated network i are adjusted.

[0197] Gradient backpropagation is performed on the descent gradients of the dedicated networks corresponding to the N media recognition tasks to obtain the descent gradient of the shared network. Based on the descent gradient of the shared network, the network parameters in the shared network are adjusted to obtain the initial recognition model after training.

[0198] Specifically, the initial loss influence weights corresponding to the N media recognition tasks are adjusted, including:

[0199] Based on the initial loss influence weight i, the derivative of the total recognition loss function is obtained by taking the derivative of the loss with respect to the initial loss influence weight i.

[0200] Based on the loss derivative and the recognition loss value corresponding to media recognition task i, obtain the descent gradient with respect to the influence weight i of the initial loss;

[0201] The product of the descent gradient of the initial loss influence weight i and the learning step size is obtained to get the parameter adjustment threshold. The difference between the initial loss influence weight i and the parameter adjustment threshold is obtained to get the adjusted loss influence weight.

[0202] Update the initial loss impact weight corresponding to the initial loss impact weight i to the adjusted loss impact weight.

[0203] The data processing device also includes:

[0204] The second acquisition module 16 is used to acquire the business splicing media features; the business splicing media features are obtained by splicing the media attribute features of the multimedia data to be pushed and the historical media interaction features of the business object.

[0205] The third extraction module 17 is used to call the shared network in the multi-task recognition model to extract the shared media features associated with N media recognition tasks from the media features of the spliced ​​media.

[0206] The fourth extraction module 18 is used to call the dedicated network i in the multi-task recognition model to extract the business-specific media feature i associated with the media recognition task i from the business splicing media features;

[0207] The second identification module 19 is used to call the identification network i of the media identification task i in the multi-task identification model, and identify the i-th dimension identification interaction label of the business object for the multimedia data to be pushed based on the business-shared media features and the business-specific media features i.

[0208] The push module 20 is used to push multimedia data to be pushed to business objects based on N-dimensional recognition interaction tags.

[0209] The push module 20 includes:

[0210] The fifth determining unit 2001 is used to determine the push score of the multimedia data to be pushed based on the N-dimensional identification interaction tags.

[0211] The push unit 2002 is used to push multimedia data to be pushed to business objects when the recommendation score is greater than the score threshold.

[0212] According to one embodiment of this application, Figure 8 The modules in the data processing apparatus shown can be individually or entirely combined into one or more units, or one or more of these units can be further divided into at least two functionally smaller sub-units to achieve the same operation without affecting the technical effects of the embodiments of this application. The above modules are based on logical functional division. In practical applications, the function of one module can also be implemented by at least two units, or the function of at least two modules can be implemented by one unit. In other embodiments of this application, the data processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by at least two units.

[0213] According to one embodiment of this application, a general-purpose computer device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), can perform operations such as... Figure 3 The computer program (including program code) involved in each step of the corresponding method shown, to construct such... Figure 8 The data processing apparatus shown herein, and the data processing method for implementing the embodiments of this application, are described. The computer program described above may be recorded on, for example, a computer-readable recording medium, loaded onto the computer device via the computer-readable recording medium, and run therein.

[0214] In this embodiment, a multi-task recognition model is constructed based on sample-stitched media features. This model can identify multi-dimensional interactive tags of objects in multimedia data, with each interactive tag corresponding to a media recognition task. Therefore, the multi-task recognition model can handle multiple media recognition tasks without requiring a separate model for each task, reducing resource overhead during training and improving training efficiency. During training, an initial recognition model is first constructed, comprising a shared network, multiple dedicated networks, and multiple recognition networks. Each dedicated network and recognition network corresponds to a media recognition task. When identifying the i-th predicted interactive tag for media recognition task i, the i-th predicted interactive tag of the sample object in multimedia data is identified by combining the recognition network i in the initial model with the shared media features and the dedicated media features i associated with media recognition task i. The initial recognition model is then trained based on the N-dimensional labeled interactive tags and the N-dimensional predicted interactive tags to obtain the multi-task recognition model. The shared media features here are extracted from the sample splicing media features by the shared network of the initial recognition model, while the specific media feature i is extracted from the sample splicing media features by the specific network in the initial recognition model associated with media recognition task i. In other words, when recognizing the i-th predicted interaction label under media recognition task i, not only the specific media feature i directly associated with media recognition task i is combined, but also the shared media features indirectly associated with media recognition task i are combined, providing more information for recognizing the i-th predicted interaction label under media recognition task i. That is, different media recognition tasks can share shared media features, which provides more training data for the training process of different media recognition tasks. This avoids the problem of low media recognition accuracy of the initial recognition model (i.e., task recognition model) after training due to the sparsity of training data, thus improving the media recognition accuracy of the initial recognition model after training and thereby improving the accuracy of multimedia data push. In addition, this application can improve the recognition accuracy of the multi-task recognition model by setting the loss influence weight of the loss function of different media recognition tasks.

[0215] Further, please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device may be a terminal device or a server. Figure 9As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. In some embodiments, the user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. Optionally, the network interface 1004 may include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 9 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0216] In such Figure 9 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0217] Obtain the sample splicing media features and the N-dimensional annotation interaction labels of the sample objects for the sample multimedia data; the sample splicing media features are obtained by splicing the media attribute features of the sample multimedia data and the historical media interaction features of the sample objects; the N-dimensional annotation interaction labels correspond to N media recognition tasks; N is an integer greater than 1;

[0218] The shared network in the initial recognition model is invoked to extract shared media features associated with N media recognition tasks from the sample splicing media features;

[0219] Call the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample splicing media features; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks in the initial recognition model, and i is a positive integer less than or equal to N;

[0220] Call the recognition network i in the initial recognition model for media recognition task i, and based on the shared media features and the specific media features i, identify the sample object and predict the interactive label for the i-th dimension of the sample multimedia data;

[0221] The initial recognition model is trained based on N-dimensional labeled interactive labels and N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping training condition, thus obtaining the multi-task recognition model.

[0222] It should be understood that the computer device 1000 described in the embodiments of this application can perform the foregoing... Figure 7 The description of the data processing method in the corresponding embodiments can also be performed as described above. Figure 8 The description of the data processing apparatus in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0223] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned data processing device. When the processor executes the computer program, it can execute the aforementioned... Figure 3 or Figure 7 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.

[0224] Furthermore, it should be noted that this application also provides a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, causing the computer device to perform the aforementioned... Figure 3 or Figure 7 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product embodiments related to this application, please refer to the description of the method embodiments of this application.

[0225] It should be noted that the data collection and processing described in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the data subject (or have a legal basis), and conduct subsequent data use and processing within the scope of laws and regulations and the authorization of the data subject. For example, this application needs to obtain the informed consent or separate consent of the business object or sample object when obtaining the historical media interaction characteristics of the business object or sample object.

[0226] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0227] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A data processing method, characterized in that, include: Obtain the features of the spliced ​​media of the samples, as well as the N-dimensional annotation and interactive tags of the sample objects for the sample multimedia data; The sample splicing media feature is obtained by splicing the media attribute features of the sample multimedia data and the historical media interaction features of the sample object; the N-dimensional labeled interaction label corresponds to N media recognition tasks; N is an integer greater than 1; The shared network in the initial recognition model is invoked to extract the shared media features associated with the N media recognition tasks from the spliced ​​media features of the samples; The dedicated network i in the initial recognition model is invoked to extract the dedicated media feature i associated with the media recognition task i from the media features spliced ​​from the sample; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks of the initial recognition model, and i is a positive integer less than or equal to N; Invoke the recognition network i in the initial recognition model for the media recognition task i, and based on the shared media features and the specific media features i, identify the i-th dimension predicted interaction label of the sample object for the sample multimedia data; The initial recognition model is trained based on the N-dimensional labeled interactive labels and the N-dimensional predicted interactive labels until the trained initial recognition model meets the stopping training condition, thus obtaining a multi-task recognition model.

2. The method according to claim 1, characterized in that, The step of invoking the shared network in the initial recognition model to extract shared media features associated with the N media recognition tasks from the sample spliced ​​media features includes: If the number of the shared network is one, then the M expert sub-networks included in the shared network are invoked to extract the M expert media features associated with the N media recognition tasks from the sample splicing media features, where M is a positive integer; The extracted M expert media features are identified as shared media features associated with the N media recognition tasks.

3. The method according to claim 1, characterized in that, The step of calling the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample spliced ​​media features includes: If there is only one dedicated network i, then the fully connected sub-network in the dedicated network i is invoked to extract the associated media features related to the media recognition task i from the sample spliced ​​media features; The activation subnetwork in the dedicated network i is invoked to perform nonlinear feature transformation on the associated media features to obtain the dedicated media features i associated with the media recognition task i.

4. The method according to claim 1, characterized in that, The step of invoking the shared network in the initial recognition model to extract shared media features associated with the N media recognition tasks from the sample spliced ​​media features includes: If the shared network includes a first shared network and a second shared network, then the M expert sub-networks included in the first shared network are invoked to extract M expert media features associated with the N media recognition tasks from the sample splicing media features, where M is a positive integer; The first fused media feature is obtained by fusing N initial proprietary media features and M expert media features extracted by the first shared network; the N initial proprietary media features are extracted by the N first proprietary networks in the initial recognition model, and the N first proprietary networks are located in the same network layer as the first shared network. The M expert subnetworks included in the second shared network are invoked to extract shared media features associated with the N media identification tasks from the first fused media features.

5. The method according to claim 4, characterized in that, The step of calling the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample spliced ​​media features includes: If the dedicated network i includes a first dedicated network i and a second dedicated network i, the first dedicated network i is invoked to extract the initial dedicated media features associated with the media recognition task i from the sample splicing media features; The initial exclusive media features extracted from the first exclusive network i and the M expert media features extracted from the first shared network are fused to obtain the second fused media features; The second dedicated network i is invoked to extract the dedicated media feature i associated with the media identification task i from the second fused media features.

6. The method according to claim 1, characterized in that, The initial recognition model includes a gating subnetwork and a tag prediction subnetwork for the media recognition task i. The step of calling the recognition network i in the initial recognition model, and recognizing the i-th dimension predicted interaction label of the sample object for the sample multimedia data based on the shared media features and the specific media features i, includes: The gating subnetworks included in the identification network i are invoked to determine the feature weights corresponding to the shared media feature and the exclusive media feature i, respectively. The gating subnetwork included in the recognition network i is invoked, and feature fusion is performed on the shared media feature and the exclusive media feature i according to the feature weights corresponding to the shared media feature and the exclusive media feature i respectively, to obtain the task fused media feature associated with the media recognition task i; The label prediction subnetwork included in the recognition network i is invoked to perform label prediction on the task-fused media features, thereby obtaining the i-th dimension predicted interactive label of the sample object for the sample multimedia data.

7. The method according to claim 6, characterized in that, The step of determining the feature weights corresponding to the shared media feature and the specific media feature i through the gated sub-network includes: The gating subnetwork included in the identification network i is invoked to determine the first correlation degree between the shared media feature and the media identification task i, and to determine the second correlation degree between the exclusive media feature and the media identification task i; Based on the first correlation degree, generate the feature weights corresponding to the shared media features; Based on the second correlation degree, the feature weights corresponding to the exclusive media feature i are generated.

8. The method according to claim 1, characterized in that, The step of training the initial recognition model based on the N-dimensional labeled interaction labels and the N-dimensional predicted interaction labels until the trained initial recognition model meets the stopping training condition, thereby obtaining a multi-task recognition model, includes: Based on the i-th dimension labeled interaction label and the i-th dimension predicted interaction label corresponding to the media recognition task i, determine the recognition loss value i for the media recognition task i; Obtain the initial loss influence weight i of the recognition loss value i on the initial recognition model; Based on the recognition loss value and initial loss influence weight corresponding to the N media recognition tasks respectively, and the total recognition loss function of the initial recognition model, determine the total recognition loss value of the initial recognition model; The initial recognition model is trained based on the total recognition loss value and the recognition loss values ​​corresponding to the N media recognition tasks until the trained initial recognition model meets the stopping training condition, thus obtaining the multi-task recognition model.

9. The method according to claim 8, characterized in that, The step of training the initial recognition model based on the total recognition loss value and the recognition loss values ​​corresponding to the N media recognition tasks, until the trained initial recognition model meets the stopping training condition, to obtain a multi-task recognition model, includes: If the total recognition loss value is greater than the loss threshold, the initial loss impact weights corresponding to the N media recognition tasks are adjusted respectively; Based on the recognition loss values ​​corresponding to the N media recognition tasks, the network parameters in the initial recognition model are adjusted to obtain the trained initial recognition model. If the total recognition loss value of the trained initial recognition model is less than or equal to the loss threshold, then the trained initial recognition model is determined to meet the stop training condition. The trained initial recognition model is determined as a multi-task recognition model.

10. The method according to claim 9, characterized in that, The step of adjusting the network parameters in the initial recognition model based on the recognition loss values ​​corresponding to the N media recognition tasks to obtain the trained initial recognition model includes: The descent gradient of the recognition network i is obtained based on the recognition loss value corresponding to the media recognition task i, and the network parameters in the recognition network i are adjusted based on the descent gradient of the recognition network i. Gradient backpropagation is performed on the descent gradient of the recognition network i to obtain the descent gradient of the dedicated network i. Based on the descent gradient of the dedicated network i, the network parameters in the dedicated network i are adjusted. The descent gradients of the dedicated networks corresponding to the N media recognition tasks are backpropagated to obtain the descent gradient of the shared network. Based on the descent gradient of the shared network, the network parameters in the shared network are adjusted to obtain the initial recognition model after training.

11. The method according to claim 9, characterized in that, The adjustment of the initial loss influence weights corresponding to the N media recognition tasks includes: Based on the initial loss influence weight i, the total recognition loss function is differentiated to obtain the loss derivative with respect to the initial loss influence weight i; Based on the loss derivative and the recognition loss value corresponding to the media recognition task i, obtain the descent gradient with respect to the initial loss influence weight i; Obtain the product between the descent gradient of the initial loss influence weight i and the learning step size to obtain the parameter adjustment threshold; obtain the difference between the initial loss influence weight i and the parameter adjustment threshold to obtain the adjusted loss influence weight. Update the initial loss influence weight corresponding to the initial loss influence weight i to the adjusted loss influence weight.

12. The method according to claim 1, characterized in that, The method further includes: Obtain the media splicing features of the business; the media splicing features of the business are obtained by splicing the media attribute features of the multimedia data to be pushed and the historical media interaction features of the business object. The shared network in the multi-task recognition model is invoked to extract the shared media features associated with the N media recognition tasks from the service splicing media features; Call the dedicated network i in the multi-task recognition model to extract the service-specific media feature i associated with the media recognition task i from the service splicing media features; The identification network i in the multi-task identification model for the media identification task i is invoked, and the i-th dimension identification interaction tag of the business object for the multimedia data to be pushed is identified based on the shared media features of the business and the business-specific media features i. Based on the N-dimensional recognition interaction tags, the multimedia data to be pushed is pushed to the business object.

13. The method according to claim 12, characterized in that, The step of pushing the multimedia data to be pushed to the business object based on the N-dimensional recognition interaction tags includes: Based on the N-dimensional recognition interaction tags, determine the push score of the multimedia data to be pushed; Multimedia data with a recommendation score greater than the score threshold will be pushed to the business object.

14. A data processing apparatus, characterized in that, include: The first acquisition module is used to acquire the features of the sample splicing media and the N-dimensional annotation interaction tags of the sample object for the sample multimedia data. The sample splicing media feature is obtained by splicing the media attribute features of the sample multimedia data and the historical media interaction features of the sample object; the N-dimensional labeled interaction label corresponds to N media recognition tasks; N is an integer greater than 1; The first extraction module is used to call the shared network in the initial recognition model to extract the shared media features associated with the N media recognition tasks from the sample splicing media features; The second extraction module is used to call the dedicated network i in the initial recognition model to extract the dedicated media feature i associated with the media recognition task i from the sample splicing media features; the dedicated network i is the dedicated network associated with the media recognition task i among the N dedicated networks of the initial recognition model, and i is a positive integer less than or equal to N; The first identification module is used to call the identification network i in the initial identification model for the media identification task i, and identify the i-th dimension predicted interaction label of the sample object for the sample multimedia data based on the shared media features and the specific media features i. The training module is used to train the initial recognition model based on the N-dimensional labeled interaction labels and the N-dimensional predicted interaction labels until the trained initial recognition model meets the stop training condition, thus obtaining a multi-task recognition model.

15. A computer device, characterized in that, include: Processor and memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to cause the computer device to perform the method according to any one of claims 1-13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-13.

17. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium and adapted to be read and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-13.