Model training method and device based on artificial intelligence, equipment and storage medium

By enhancing training of unexposed samples, a target enhanced recommendation model is formed, which solves the problem of inconsistent training and inference data space in the recommendation system and improves the accuracy of the recommended model.

CN120336842APending Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410075268.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When training recommended models, the existing recommendation system lacks user feedback due to the lack of user feedback on the unexposed data, resulting in the model training in the exposed data space and inference in the mixed space, resulting in inconsistent training and inference data space, and it is impossible to effectively use unexposed data for model training, affecting the accuracy of recommendation.

Method used

By enhancing training for each unexposed sample, an enhanced recommendation model is obtained, and the unexposed sample with the first loss less than the second loss is used as the enhancement sample. The basic recommendation model is trained together with the second exposure sample to form a target enhanced recommendation model to ensure the comprehensiveness of the training data.

Benefits of technology

The recommendation accuracy of the recommended model is improved, ensuring that the model's training data distribution in the unexposed data space is consistent with the actual inference data distribution, and improving the recommendation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336842A_ABST
    Figure CN120336842A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device based on artificial intelligence, electronic equipment, a computer program product and a computer readable storage medium, and the method can be applied to an artificial intelligence scene, and comprises the steps: carrying out the training of a basic recommendation model based on each unexposed sample, obtaining an enhanced recommendation model corresponding to each unexposed sample; performing forward propagation on the first exposure sample in each enhanced recommendation model to obtain a first loss corresponding to each enhanced recommendation model, and performing forward propagation on the first exposure sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model; when the first loss is smaller than the second loss, taking the unexposed sample corresponding to the first loss as an enhanced sample; and training the basic recommendation model based on the enhanced sample and the second exposure sample to obtain a target enhanced recommendation model. Through the method, the recommendation accuracy of the recommendation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a model training method based on artificial intelligence, a recommendation processing method based on artificial intelligence, an apparatus, an electronic device, a computer program product, and a computer-readable storage medium. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, such as natural language processing technology and machine learning / deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0003] In the recommendation ranking link, for the candidate information set to reach the user's personal client display interface, the following links are required: recall, rough ranking, fine ranking, mixed ranking, distribution, and exposure. In the recall and rough ranking links, items preferred by the user are quickly screened out from the information set, and then the information to be processed is mixed and ranked through fine ranking and mixed ranking, and finally the final distribution sequence is obtained. In the related technology, the recommendation system retrieves all exposure information, processes it through feature engineering, and uses it to train the recommendation model. The recommendation model is trained using exposure data to perform recommendation ranking on the distributed data, resulting in low recommendation accuracy of the recommendation model. Summary of the Invention

[0004] Embodiments of this application provide a model training method based on artificial intelligence, a recommendation processing method based on artificial intelligence, an apparatus, an electronic device, a computer program product, and a computer-readable storage medium, which can improve the recommendation accuracy of the recommendation model.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a model training method based on artificial intelligence, and the method includes:

[0007] Training a basic recommendation model based on each unexposed sample respectively to obtain an enhanced recommendation model corresponding to each unexposed sample;

[0008] Performing forward propagation of the first exposure sample in each enhanced recommendation model to obtain a first loss corresponding to each enhanced recommendation model, and performing forward propagation of the first exposure sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model;

[0009] When the first loss is less than the second loss, use the unexposed sample corresponding to the first loss as the enhanced sample;

[0010] Train the basic recommendation model based on the enhanced sample and the second exposed sample to obtain a target enhanced recommendation model.

[0011] An embodiment of the present application provides a recommendation processing method based on artificial intelligence, and the method includes:

[0012] Obtain information to be recommended;

[0013] Call the target enhanced recommendation model to perform prediction processing on the recommendation metrics of the information to be recommended, and obtain the predicted recommendation metrics of the information to be recommended, where the target enhanced recommendation model is trained by the model training method based on artificial intelligence provided by the embodiment of the present application;

[0014] Based on the predicted recommendation metrics of the information to be recommended, perform recommendation processing on the information to be recommended.

[0015] An embodiment of the present application provides a model training device based on artificial intelligence, including:

[0016] A basic training module, configured to train a basic recommendation model based on each unexposed sample respectively to obtain an enhanced recommendation model corresponding to each unexposed sample;

[0017] A loss calculation module, configured to perform forward propagation of the first exposed sample in each enhanced recommendation model to obtain a first loss corresponding to each enhanced recommendation model, and perform forward propagation of the first exposed sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model;

[0018] A sample screening module, configured to use the unexposed sample corresponding to the first loss as the enhanced sample when the first loss is less than the second loss;

[0019] An enhanced training module, configured to train the basic recommendation model based on the enhanced sample and the second exposed sample to obtain a target enhanced recommendation model.

[0020] In the above solution, the basic training module is further configured to obtain a third exposed sample different from the first exposed sample, and train an initialized recommendation model based on the third exposed sample to obtain the basic recommendation model.

[0021] In the above solution, the basic training module is further configured to perform the following processing on each of the unexposed samples: Propagate the unexposed sample forward in the basic recommendation model to obtain a first predicted recommendation metric corresponding to the unexposed sample. Based on the true label of the unexposed sample and the first predicted recommendation metric corresponding to the unexposed sample, determine a third loss corresponding to the unexposed sample. Based on the third loss corresponding to the unexposed sample, perform an update process on the basic recommendation model to obtain an enhanced recommendation model corresponding to the unexposed sample.

[0022] In the above solution, the loss calculation module is further configured to: Propagate the first exposed sample forward in each of the enhanced recommendation models to obtain a second predicted recommendation metric corresponding to each of the enhanced recommendation models. Based on the true label of the first exposed sample and the second predicted recommendation metric corresponding to each of the enhanced recommendation models, determine a first single-sample loss of the first exposed sample corresponding to each of the enhanced recommendation models. For each of the enhanced recommendation models, fuse the first single-sample losses of the multiple first exposed samples corresponding to the enhanced recommendation model to obtain a first loss corresponding to the enhanced recommendation model.

[0023] In the above solution, the loss calculation module is further configured to: Propagate the first exposed sample forward in the basic recommendation model to obtain a third predicted recommendation metric corresponding to the first exposed sample. Based on the true label of the first exposed sample and the third predicted recommendation metric corresponding to the first exposed sample, determine a second single-sample loss corresponding to the first exposed sample. Fuse the second single-sample losses of the multiple first exposed samples to obtain a second loss corresponding to the basic recommendation model.

[0024] In the above solution, the enhanced training module is further configured to propagate each of the enhanced samples forward in the basic recommendation model to obtain a fourth loss corresponding to each of the enhanced samples. Based on the enhancement value of each of the enhanced samples, perform a weighted process on the fourth losses of the multiple enhanced samples to obtain an enhanced sample loss. Propagate each of the second exposed samples forward in the basic recommendation model to obtain a fifth loss corresponding to each of the second exposed samples, and fuse the fifth losses of the multiple second exposed samples to obtain an exposed sample loss. Based on the fusion result of the enhanced sample loss and the exposed sample loss, perform an update process on the basic recommendation model to obtain the target enhanced recommendation model.

[0025] In the above solution, the enhancement training module is further configured to perform the following processing on each of the enhanced samples: Propagate the enhanced sample forward in the basic recommendation model to obtain a fourth predicted recommendation metric corresponding to the enhanced sample; Determine a fourth loss corresponding to the enhanced sample based on the true label of the enhanced sample and the fourth predicted recommendation metric corresponding to the enhanced sample.

[0026] In the above solution, the enhancement training module is further configured to, when the true label of the enhanced sample is set to not clicked, use the difference between the first loss and the second loss as the enhancement value of the enhanced sample.

[0027] In the above solution, the enhancement training module is further configured to perform an absolute value processing on the enhancement value corresponding to each of the enhanced samples to obtain an absolute value result corresponding to each of the enhanced samples, perform a normalization processing on the absolute value result corresponding to each of the enhanced samples to obtain a weight corresponding to each of the enhanced samples, and perform a weighted processing on the fourth losses of the multiple enhanced samples based on the weight corresponding to each of the enhanced samples to obtain the enhanced sample loss.

[0028] In the above solution, the basic training module is further configured to obtain the exposure data of the exposure information sample and the non-exposure data of the non-exposure information sample, perform a first embedding process on the exposure data based on the first embedding parameter to obtain the exposure sample, and perform a second embedding process on the non-exposure data based on the second embedding parameter to obtain the non-exposure sample.

[0029] An embodiment of the present application provides a recommendation processing device based on artificial intelligence, including:

[0030] An acquisition module, configured to acquire information to be recommended;

[0031] A prediction processing module, configured to call a target enhanced recommendation model to perform a recommendation metric prediction process on the information to be recommended to obtain a predicted recommendation metric of the information to be recommended, where the target enhanced recommendation model is trained by the model training method based on artificial intelligence provided by the embodiment of the present application;

[0032] A recommendation processing module, configured to perform a recommendation process on the information to be recommended based on the predicted recommendation metric of the information to be recommended.

[0033] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0034] A memory, configured to store computer-executable instructions;

[0035] A processor, when executing computer-executable instructions stored in the memory, implements the model training method based on artificial intelligence or the recommendation processing method based on artificial intelligence provided by the embodiments of the present application.

[0036] The embodiments of the present application provide a computer-readable storage medium storing computer-executable instructions, which are used to implement the model training method based on artificial intelligence or the recommendation processing method based on artificial intelligence provided by the embodiments of the present application when being executed by a processor.

[0037] The embodiments of the present application provide a computer program product including computer-executable instructions, which implement the model training method based on artificial intelligence or the recommendation processing method based on artificial intelligence provided by the embodiments of the present application when being executed by a processor.

[0038] The embodiments of the present application have the following beneficial effects:

[0039] By training the basic recommendation model respectively based on each unexposed sample, an enhanced recommendation model corresponding to each unexposed sample is obtained, so that the model parameters of the enhanced recommendation model corresponding to each unexposed sample are only affected by the training of the corresponding unexposed sample. By performing forward propagation of the first exposure sample in each enhanced recommendation model, the first loss corresponding to each enhanced recommendation model is obtained, and by performing forward propagation of the first exposure sample in the basic recommendation model, the second loss corresponding to the basic recommendation model is obtained. The loss impact of each unexposed sample on model training is measured by the first loss and the second loss. When the first loss is less than the second loss, the unexposed sample corresponding to the first loss is used as an enhanced sample, and the unexposed samples that play an enhancing role in the training of the basic recommendation model are screened out as enhanced samples. The basic recommendation model is trained based on the enhanced samples and the second exposure samples to obtain a target enhanced recommendation model, thereby ensuring that the unexposed samples used as enhanced samples and the second exposure samples can be used to train the basic recommendation model, ensuring the comprehensiveness of the training data of the basic recommendation model, and further improving the recommendation accuracy of the target enhanced recommendation model. Description of the Drawings

[0040] Figure 1 is a schematic structural diagram of the model training system provided by the embodiments of the present application;

[0041] Figure 2A is a schematic structural diagram of the server for executing the model training method based on artificial intelligence provided by the embodiments of the present application;

[0042] Figure 2B is a schematic structural diagram of the server for executing the recommendation processing method based on artificial intelligence provided by the embodiments of the present application;

[0043] Figure 3AIt is a schematic flowchart of a model training method based on artificial intelligence provided by an embodiment of the present application;

[0044] Figure 3B It is an optional flowchart of a model training method based on artificial intelligence provided by an embodiment of the present application Figure 1 ;

[0045] Figure 3C It is the second optional flowchart of a model training method based on artificial intelligence provided by an embodiment of the present application;

[0046] Figure 3D It is the third optional flowchart of a model training method based on artificial intelligence provided by an embodiment of the present application;

[0047] Figure 3E It is an optional flowchart of a model training method based on artificial intelligence provided by an embodiment of the present application Figure 4 ;

[0048] Figure 4 It is a schematic flowchart of a recommendation processing method based on artificial intelligence provided by an embodiment of the present application;

[0049] Figure 5 It is a schematic diagram of a subscription number recommendation interface provided by an embodiment of the present application;

[0050] Figure 6 It is a schematic diagram of a recommendation ranking link provided by an embodiment of the present application;

[0051] Figure 7 It is a pre-training process framework of a recommendation ranking model provided by an embodiment of the present application;

[0052] Figure 8 It is an application-side schematic diagram of a target enhanced recommendation model provided by an embodiment of the present application;

[0053] Figure 9 It is an application feedback data diagram of a target enhanced recommendation model provided by an embodiment of the present application. Specific embodiments

[0054] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0055] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0056] In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0057] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0058] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0059] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0060] 1) Data Augmentation: A method of expanding the training data set by generating more similar data to the target task through prior knowledge, which usually helps to improve the generalization ability and accuracy of the model. When applied to a recommendation system, data augmentation can provide a more complete characterization of the sample distribution in aspects such as users, items, and the interactions between them without significantly increasing the query and storage costs of the system. Since the quantity and quality of data determine the upper limit of the performance of the recommendation ranking model, researching better data augmentation methods is of great help for improving the personalized recommendation ability of the recommendation system.

[0061] 2) Negative Sampling: A data augmentation method. In a recommendation system, negative samples are usually much more numerous than positive samples. Negative sampling is a method of randomly selecting a part of the samples from the negative samples of the exposure data to make the ratio of positive and negative samples more balanced. This can improve the training efficiency of the model and avoid overfitting at the same time.

[0062] 3) Label: The labeled data relied on when training a deep neural network, such as "0 / 1" representing "belonging to / not belonging to" a certain class. For example, whether a click is generated.

[0063] 4) Embedding Vector: A numerical vector composed of multiple floating-point numbers, which describes various attributes and properties of items or users in a high-dimensional space.

[0064] 5) Receiver Operating Characteristic (ROC) Curve: A graphical tool used to represent the performance of a classification model, which depicts the performance of a classifier at different thresholds by taking the true positive rate and the false positive rate as the horizontal and vertical coordinates.

[0065] 6) Area Under the Curve (AUC) of the ROC Curve: Used to evaluate the performance of a click-through rate prediction model, and the higher the value, the better.

[0066] In the recommendation ranking link, the candidate information set needs to go through the following links to reach the user's personal client display interface: recall, rough ranking, fine ranking, mixed ranking, distribution, and exposure. The approximate order of magnitude of the number of information processed in each link is as follows: The approximate order of magnitude of the number of information processed in the recall link is 10 6 , the approximate order of magnitude of the number of information processed in the rough ranking link is 10 4 , the approximate order of magnitude of the number of information processed in the fine ranking link is 10 2 , and the approximate order of magnitude of the number of information processed in the mixed ranking link, distribution link, and exposure link is all 10. The recall and rough ranking links quickly screen out user-preferred items at the hundred level from an information set ranging from one hundred thousand to one million levels, and then through the precise personalized modeling of the fine ranking and the mixed ranking of multiple sources and multiple modalities of items, and finally obtain the final distribution sequence after being scattered by some business rules. Usually, the recommendation system retrieves all exposure information, processes it through feature engineering, and uses it to train the ranking model.

[0067] When the applicant implemented the embodiments of the present application, it was found that the following defects existed in the related art: The recommendation system retrieves all exposure information, processes it through feature engineering, and uses it to train the ranking model. In this case, there is an inconsistency in the data spaces of the offline training and online serving inference of the model: The model is trained in the exposure data space, that is, the model uses exposure data for training, and performs inference in the mixed arrangement space (for convenience, it can be approximately considered equal to the distribution space), that is, the model performs recommendation ranking on the distributed data. Therefore, the information items that have passed through the mixed arrangement and distribution links but have not been exposed can be used to expand the training data set of the recommendation model. Moreover, training on the data space after sample expansion will not cause a large deviation in the training-recommendation data distribution. The problem is that in the data space of unexposed distribution, the system does not collect the user's feedback on each piece of information, and the samples with unknown labels cannot be directly used to evaluate the goodness of the model's click-through rate prediction, so the model cannot be trained.

[0068] The embodiments of the present application provide an artificial intelligence-based model training method, an artificial intelligence-based recommendation processing method, an apparatus, a device, a computer-readable storage medium, and a computer program product, which can improve the recommendation accuracy of the recommendation model.

[0069] The following describes the exemplary applications of the electronic device provided in the embodiments of the present application. The device provided in the embodiments of the present application can be implemented as various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (for example, a mobile phone, a cellular phone, a portable music player, a personal digital assistant, a dedicated information device, a portable game device), a smart device (for example, a smart speaker, a smart watch, a smart TV, a smart home appliance, a smart voice interaction device), a vehicle-mounted terminal, an aircraft, etc., or can be implemented as a server. Next, the exemplary application when the device is implemented as a server will be described.

[0070] See Figure 1 , Figure 1 FIG. is a schematic diagram of the architecture of the model training system 100 provided in the embodiments of the present application. To support a model training application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0071] The terminal 400 is used to generate a model training request. For example, a user generates a model training request through the graphical interface 410 of the terminal 400. The server 200 is used to train the basic recommendation model respectively based on each unexposed sample according to the model training request, obtain an enhanced recommendation model corresponding to each unexposed sample, perform forward propagation of the first exposure sample in each enhanced recommendation model, obtain a first loss corresponding to each enhanced recommendation model, and perform forward propagation of the first exposure sample in the basic recommendation model, obtain a second loss corresponding to the basic recommendation model. When the first loss is less than the second loss, use the unexposed sample corresponding to the first loss as an enhanced sample, and train the basic recommendation model based on the enhanced sample and the second exposure sample to obtain a target enhanced recommendation model.

[0072] In some embodiments, the target enhanced recommendation model is deployed on the server 200. The terminal 400 generates a recommendation processing request. The server 200 obtains the information to be recommended from the database based on the recommendation processing request, calls the target enhanced recommendation model to perform a recommendation index prediction process on the information to be recommended, obtains a predicted recommendation index of the information to be recommended, performs a recommendation process on the information to be recommended based on the predicted recommendation index of the information to be recommended, obtains a recommendation result, and feeds back the recommendation result to the terminal 400.

[0073] In some embodiments, the target enhanced recommendation model is deployed on the terminal 400. The terminal 400 generates a recommendation processing request, and based on the recommendation processing request, obtains the information to be recommended from the database, calls the target enhanced recommendation model to perform a recommendation index prediction process on the information to be recommended, obtains a predicted recommendation index of the information to be recommended, and performs a recommendation process on the information to be recommended based on the predicted recommendation index of the information to be recommended to obtain a recommendation result.

[0074] In some embodiments, the information to be recommended may be product information, news information, video information, etc. In the application scenario where the information to be recommended is product information, the unexposed samples may be product information that is not displayed on the product display interface, and the exposed samples may be product information that is displayed on the product display interface. The predicted recommendation metric for the information to be recommended may be the probability of predicting that a user clicks on a product to enter the details interface, and the recommendation result may be the product information recommendation result displayed on the user's product recommendation interface; in the application scenario where the information to be recommended is news information, the unexposed samples may be news information that is not displayed on the news push page, and the exposed samples may be news information that is displayed on the news push page. The predicted recommendation metric for the information to be recommended may be the probability of predicting that a user clicks on news information for effective reading, and the recommendation result may be the news information recommendation result displayed on the user's news push page; in the application scenario where the information to be recommended is video information, the unexposed samples may be video information that is not displayed on the video recommendation interface, and the exposed samples may be video information that is displayed on the video recommendation interface. The predicted recommendation metric for the information to be recommended may be the probability of predicting that a user clicks on a video for effective viewing, and the recommendation result may be the video information recommendation result displayed on the user's video recommendation interface.

[0075] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present invention.

[0076] The embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as artificial intelligence. Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0077] With the research and progress of artificial intelligence technology, artificial intelligence technology is being studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interactions, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0078] See Figure 2A , Figure 2A is a schematic structural diagram of a server 200-1 that executes a model training method based on artificial intelligence provided by an embodiment of the present application. Figure 2A The server 200-1 shown includes: at least one processor 210, a memory 230, and at least one network interface 220. Each component in the server 200-1 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2A all kinds of buses are labeled as the bus system 240.

[0079] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0080] The memory 230 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 230 optionally includes one or more storage devices that are physically located far from the processor 210.

[0081] The memory 230 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (Random Access Memory, RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.

[0082] In some embodiments, the memory 230 is capable of storing data to support various operations. Examples of these data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0083] An operating system 231, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0084] A network communication module 232, for reaching other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;

[0085] In some embodiments, the model training device based on artificial intelligence provided by the embodiments of the present application can be implemented in software. Figure 2A Shown is a model training device 233 based on artificial intelligence stored in a memory 230, which can be software in the form of a program and a plug-in, etc., including the following software modules: a basic training module 2331, a loss calculation module 2332, a sample screening module 2333, and an enhanced training module 2334. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0086] In some embodiments, a terminal or a server can implement the model training method based on artificial intelligence provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be commands at the microprogram level, machine instructions, or software instructions. The computer program can be a native program or a software module in an operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a model training APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to a browser environment to run. In short, the above computer-executable instructions can be instructions in any form, and the above computer program can be an application program, a module, or a plug-in in any form.

[0087] See Figure 2B , Figure 2B is a schematic structural diagram of a server 200-2 that executes the recommendation processing method based on artificial intelligence provided by the embodiments of the present application. Figure 2BThe server 200-2 shown includes: at least one processor 250, a memory 260, and at least one network interface 270. Each component in the server 200-2 is coupled together through a bus system 280. It can be understood that the bus system 280 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 280 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2B all kinds of buses are labeled as the bus system 280.

[0088] The processor 250 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0089] The memory 260 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. Optionally, the memory 260 includes one or more storage devices that are physically located away from the processor 250.

[0090] The memory 260 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 260 described in the embodiments of the present application is intended to include any suitable type of memory.

[0091] In some embodiments, the memory 260 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.

[0092] An operating system 261, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0093] A network communication module 262, for reaching other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and Universal Serial Bus (USB), etc.;

[0094] In some embodiments, the recommendation processing apparatus based on artificial intelligence provided by the embodiments of the present application can be implemented in software. Figure 2B FIG. Figure 2B shows a recommendation processing apparatus 263 based on artificial intelligence stored in the memory 260, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 2631, a prediction processing module 2632, and a recommendation processing module 2633. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0095] In some embodiments, a terminal or a server can implement the recommendation processing apparatus based on artificial intelligence provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be commands at the microprogram level, machine instructions, or software instructions. The computer program can be a native program or a software module in an operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a recommendation sorting APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer-executable instructions can be instructions in any form, and the above computer programs can be application programs, modules, or plug-ins in any form.

[0096] Next, the model training method based on artificial intelligence provided by the embodiments of the present application will be described in conjunction with the exemplary applications and implementations of the server provided by the embodiments of the present application.

[0097] It should be noted that in the examples of recommendation processing below, information recommendation is used as an example for illustration. Those skilled in the art can apply the model training method based on artificial intelligence provided by the embodiments of the present application to the model training processing including other types of recommendation objects according to the understanding of the following text.

[0098] See Figure 3A , Figure 3A FIG. Figure 3A is a schematic flowchart of the model training method based on artificial intelligence provided by the embodiments of the present application, which will be described in conjunction with Figure 3A the steps 101 to 104 shown in FIG. Figure 3A .

[0099] In step 101, each unexposed sample is used to train the basic recommendation model respectively, and an enhanced recommendation model corresponding to each unexposed sample is obtained.

[0100] As an example, taking the product recommendation scenario as an example, the unexposed samples can be product information that is not displayed on the product display interface, the exposed samples can be product information that is displayed on the product display interface, and the predicted recommendation metric can be the probability of predicting that a user clicks on a product to enter the detail interface; taking the news recommendation scenario as an example, the unexposed samples can be news information that is not displayed on the news push page, the exposed samples can be news information that is displayed on the news push page, and the predicted recommendation metric can be the probability of predicting that a user clicks on the news information for effective reading, and the recommendation result can be the news information recommendation result displayed on the user's news push page; taking the video recommendation scenario as an example, the unexposed samples can be video information that is not displayed on the video recommendation interface, the exposed samples can be video information that is displayed on the video recommendation interface, and the predicted recommendation metric can be the probability of predicting that a user clicks on a video for effective browsing, and the recommendation result can be the video information recommendation result displayed on the user's video recommendation interface.

[0101] See Figure 3B , Figure 3B which is an optional process schematic of the artificial intelligence-based model training method provided by the embodiments of the present application Figure 1 . In some embodiments, Figure 3A Step 101 in Figure 3B can be implemented by steps 1011 to 1013 shown in

[0102] The following processing is performed on each unexposed sample:

[0103] In step 1011, the unexposed sample is propagated forward in the basic recommendation model to obtain the first predicted recommendation metric corresponding to the unexposed sample.

[0104] As an example, an unexposed sample is the embedded representation of the relevant basic features of the sample information that enters the distribution link obtained from the information data stream but does not enter the exposure link due to the low probability of being clicked by the user. The relevant basic features can be abstractly represented as: <target user-related features, subscription account-related features of the information to be recommended, features of the information to be recommended itself, statistical features of cross relationships>. Among them, the target user-related features can include: for example, user account identifier, user gender, etc., the number of exposed information of the user in the past 1 day, the number of exposed information of the user in the past 7 days, etc.; the subscription account-related features of the information to be recommended can include: for example, subscription account identifier, the number of users subscribing to the subscription account, the number of created information of the subscription account in the past 7 days, the number of click-through readings of the subscription account in the past 7 days, etc.; the features of the information to be recommended itself can include: for example, information identifier, the number of hours since the information was sent, the number of exposures of the information in the past 1 hour, the number of clicks of the information in the past 1 hour, etc.; the statistical features of cross relationships can include: for example, the number of exposures of the user to the subscription account in the past 28 days, the number of clicks of the user on the subscription account in the past 28 days, etc. Taking a single unexposed sample as the input, call the basic recommendation model to predict the recommendation metrics for the single unexposed sample, and obtain the first predicted recommendation metric corresponding to the single unexposed sample. The first predicted recommendation metric can be the predicted click-through rate, collection rate, forwarding rate, etc. The predicted click-through rate is the probability of predicting that the unexposed sample will be clicked by the user.

[0105] In step 1012, based on the true label of the unexposed sample and the first predicted recommendation metric corresponding to the unexposed sample, determine the third loss corresponding to the unexposed sample.

[0106] As an example, since the probability of an unexposed sample being clicked by the user is predicted to be low, it does not enter the exposure link. Therefore, the unexposed sample can be used as a negative sample to train the model. The true label of the unexposed sample can be set to 0 to indicate that it is a negative sample. The label values of positive and negative samples can be set according to actual needs. After obtaining the first predicted recommendation metric corresponding to the unexposed sample, the third loss of the basic recommendation model corresponding to the unexposed sample can be calculated based on the true label and the first predicted recommendation metric of the unexposed sample. For example, if the first predicted recommendation metric here is 0.7, it means that the probability of this unexposed sample being predicted to be clicked is 0.7. However, the probability of this unexposed sample being actually clicked is the true label 0. That is, the third loss is determined based on the difference between the first predicted recommendation metric and the true label. Specifically, the third loss here can be a cross-entropy loss function, etc.

[0107] In step 1013, based on the third loss corresponding to the unexposed sample, perform an update process on the basic recommendation model to obtain an enhanced recommendation model corresponding to the unexposed sample.

[0108] As an example, according to the third loss of the unexposed samples corresponding to the basic recommendation model, the model parameters of the basic recommendation model are updated to obtain an updated enhanced recommendation model. Since this enhanced recommendation model is trained using a single unexposed sample, the difference between its model parameters and those of the basic recommendation model is caused by using a single unexposed sample for training. Therefore, for multiple unexposed samples, a corresponding number of enhanced recommendation models can be obtained one by one.

[0109] By using unexposed samples to train the basic recommendation model and updating the basic recommendation model, an enhanced recommendation model corresponding to the unexposed sample is obtained. The influence result of the unexposed sample itself on the basic recommendation model is reflected through the change of the model parameters of the enhanced recommendation model, which is used to compare the differences between the basic recommendation model and the enhanced recommendation model in the subsequent process, so as to evaluate the influence degree of the unexposed sample on model training, and then screen out the unexposed samples that can enhance the recommendation accuracy of the recommendation model, and train the basic recommendation model to improve the recommendation accuracy of the target enhanced recommendation model.

[0110] Continue to refer to Figure 3A In step 102, the first exposure sample is propagated forward in each enhanced recommendation model to obtain the first loss corresponding to each enhanced recommendation model, and the first exposure sample is propagated forward in the basic recommendation model to obtain the second loss corresponding to the basic recommendation model.

[0111] Refer to Figure 3C , Figure 3C FIG. is the second optional flowchart of the model training method based on artificial intelligence provided by the embodiments of the present application. In some embodiments, Figure 3A The step of propagating the first exposure sample forward in each enhanced recommendation model in step 102 to obtain the first loss corresponding to each enhanced recommendation model can be implemented through Figure 3C The steps 1021 to 1023 shown in, and the following is a detailed description.

[0112] In step 1021, the first exposure sample is propagated forward in each enhanced recommendation model to obtain the second predicted recommendation index corresponding to each enhanced recommendation model.

[0113] As an example, the first exposure sample is an embedded representation of the relevant basic features of the sample information that enters the distribution link from the information data stream and enters the exposure link because the probability of being clicked by the user is predicted to be high. The relevant basic features here are similar to those of the unexposed samples. Taking the first exposure sample as the input, the enhanced recommendation model is called to predict the recommendation metrics for the first exposure sample, and the second predicted recommendation metrics corresponding to each enhanced recommendation model are obtained. The second predicted recommendation metrics are similar to the first predicted recommendation metrics and can be click-through rate, collection rate, forwarding rate, etc. The predicted click-through rate is the probability of predicting that the first exposure sample will be clicked by the user.

[0114] In step 1022, based on the true label of the first exposure sample and the second predicted recommendation metrics corresponding to each enhanced recommendation model, the first single-sample loss of the first exposure sample corresponding to each enhanced recommendation model is determined.

[0115] As an example, the first exposure sample is a sample that enters the exposure link because the probability of being clicked by the user is predicted to be high. Here, since the first exposure sample enters the exposure link, whether the first exposure sample is clicked by the user is clear. When the first exposure sample is clicked by the user, the true label of the first exposure sample can be set to 1. When the first exposure sample is not clicked by the user, the true label of the first exposure sample can be set to 0. According to the true label of the first exposure sample and the second predicted recommendation metrics, the first single-sample loss of the enhanced recommendation model corresponding to the first exposure sample is calculated. For example, the second predicted recommendation metric here is 0.7, indicating that the probability of predicting that the first exposure sample will be clicked is 0.7. However, the true probability of the first exposure sample being clicked is the true label 1. That is, the first single-sample loss is determined based on the difference between the second predicted recommendation metric and the true label. Specifically, the first single-sample loss here can be a cross-entropy loss function, etc.

[0116] In step 1023, for each enhanced recommendation model, the first single-sample losses of multiple first exposure samples corresponding to the enhanced recommendation model are fused to obtain the first loss corresponding to the enhanced recommendation model.

[0117] As an example, for the first single-sample losses of multiple first exposure samples, the fusion process can be an averaging process or other fusion methods, which are not limited in this application. The first single-sample losses of multiple first exposure samples are fused into the first loss corresponding to the enhanced recommendation model. Since the enhanced recommendation model is trained by using a certain unexposed sample, the first loss can be regarded as the loss of the enhanced recommendation model corresponding to that unexposed sample.

[0118] After training a basic recommendation model with a single unexposed sample to obtain an enhanced recommendation model corresponding to the unexposed sample, the first exposed sample (with known label) is input into the enhanced recommendation model corresponding to the unexposed sample, and the first loss of the enhanced recommendation model corresponding to the unexposed sample is calculated. The influence degree of the unexposed sample on the enhanced recommendation model is reflected by the first loss.

[0119] See Figure 3D , Figure 3D is the optional flowchart three of the model training method based on artificial intelligence provided by the embodiments of the present application. In some embodiments, Figure 3A The forward propagation of the first exposed sample in the basic recommendation model shown in step 102 to obtain the second loss corresponding to the basic recommendation model can be implemented through Figure 3D steps 1024 to 1026 shown in

[0120] In step 1024, the first exposed sample is forward propagated in the basic recommendation model to obtain the third predicted recommendation index corresponding to the first exposed sample.

[0121] As an example, the first exposed sample is input into the basic recommendation model, and the basic recommendation model performs forward propagation on the first exposed sample to obtain the third predicted recommendation index corresponding to the first exposed sample. Since the basic recommendation model is not trained with unexposed samples, the difference between the third predicted recommendation index corresponding to the first exposed sample and the second predicted recommendation index corresponding to each enhanced recommendation model is caused by the enhanced recommendation model being trained with unexposed samples.

[0122] In step 1025, based on the true label of the first exposed sample and the third predicted recommendation index corresponding to the first exposed sample, the second single-sample loss corresponding to the first exposed sample is determined.

[0123] As an example, taking the first exposed sample as a positive sample and setting its true label as 1, according to the true label of the first exposed sample and the third predicted recommendation index, the second single-sample loss of the basic recommendation model corresponding to the first exposed sample is calculated. For example, here the third predicted recommendation index is 0.6, indicating that the probability of the first exposed sample being predicted to be clicked is 0.6. However, the probability of the first exposed sample being truly clicked is the true label 1. That is, the second single-sample loss is determined based on the difference between the third predicted recommendation index and the true label. Specifically, the second single-sample loss here can be a cross-entropy loss function, etc.

[0124] In step 1026, the second single-sample losses of multiple first exposed samples are fused to obtain the second loss corresponding to the basic recommendation model.

[0125] As an example, for the second single-sample loss of multiple first exposure samples, a fusion process is performed. The fusion process can be an averaging process or other fusion methods, which are not limited in this application. The second single-sample losses of multiple first exposure samples are fused into the second loss corresponding to the basic recommendation model. Since the basic recommendation model is not trained using a single unexposed sample and is equivalent to a control model corresponding to an enhanced recommendation model trained using unexposed samples, the second loss can be regarded as the loss function of the control model corresponding to the first loss.

[0126] By using the first exposure samples as positive samples, inputting them into the basic recommendation model, and calculating the second loss of the basic recommendation model, a second loss is obtained as a control for the first loss of the enhanced recommendation model trained using unexposed samples, which is used for subsequent comparison with the first loss to evaluate the influence degree of the unexposed samples corresponding to the enhanced recommendation model on model enhancement.

[0127] Continue to refer to Figure 3A , in step 103, when the first loss is less than the second loss, the unexposed sample corresponding to the first loss is used as an enhanced sample.

[0128] As an example, since the model parameters of the enhanced recommendation model and the basic recommendation model are the same before training the enhanced recommendation model using a single unexposed sample, and after training the enhanced recommendation model using a single unexposed sample, the difference between the parameters of the enhanced recommendation model and the basic recommendation model is caused by the enhanced recommendation model being trained and updated using a single unexposed sample. Therefore, for the same or the same batch of first exposure samples, the difference between the first loss of the enhanced recommendation model and the second loss of the basic recommendation model is caused by the enhanced recommendation model being trained using a single unexposed sample. Thus, the difference value vi of the two losses = first loss - second loss, which can be used to measure whether it is helpful to introduce a single unexposed sample as a negative sample into the training dataset for the enhanced recommendation model. When vi is less than 0, it means that after training the enhanced recommendation model using this single unexposed sample, the loss function value of the first exposure sample is reduced, that is, it is considered that this single unexposed sample is an enhanced sample that helps to improve the prediction accuracy of the enhanced recommendation model.

[0129] In step 104, the basic recommendation model is trained based on the enhanced samples and the second exposure samples to obtain the target enhanced recommendation model.

[0130] Refer to Figure 3E , Figure 3E is an optional process schematic of the model training method based on artificial intelligence provided by the embodiments of this application Figure 4 . In some embodiments, Figure 3A The step 104 shown can be throughFigure 3E The implementation of steps 1041 to 1044 shown will be described in detail below.

[0131] In step 1041, each enhanced sample is propagated forward in the basic recommendation model to obtain a fourth loss corresponding to each enhanced sample.

[0132] In some embodiments, step 1041 can be implemented in the following manner: perform the following processing on each enhanced sample: propagate the enhanced sample forward in the basic recommendation model to obtain a fourth predicted recommendation metric corresponding to the enhanced sample, and determine the fourth loss corresponding to the enhanced sample based on the true label of the enhanced sample and the fourth predicted recommendation metric corresponding to the enhanced sample.

[0133] As an example, the enhanced sample is a sample screened from unexposed samples with a difference vi = first loss - second loss less than 0. Therefore, the enhanced sample is also a negative sample, and its true label is the same as that of the unexposed sample. Let its true label be 0. Taking the enhanced sample as a negative sample, propagate it forward in the basic recommendation model to obtain a fourth predicted recommendation metric corresponding to the enhanced sample. The fourth predicted recommendation metric can be a probability value used to characterize the probability that the enhanced sample is clicked by the user. Based on the true label of the enhanced sample and the fourth predicted recommendation metric, calculate the fourth loss corresponding to the enhanced sample. For example, here the fourth predicted recommendation metric is 0.2, indicating that the probability that the enhanced sample is predicted to be clicked is 0.2. However, the probability that the enhanced sample is actually clicked is the true label 0. That is, determine the fourth single-sample loss based on the difference between the fourth predicted recommendation metric and the true label. Specifically, the fourth loss here can be a cross-entropy loss function, etc.

[0134] By using the enhanced samples screened from unexposed samples as negative samples to train the basic recommendation model, on the one hand, the basic recommendation model can be trained with negative samples, and on the other hand, it is ensured that the negative samples used to train the basic recommendation model are negative samples that are helpful for improving the prediction accuracy of the target enhanced recommendation model, thereby improving the prediction accuracy of the target enhanced recommendation model.

[0135] In step 1042, based on the enhancement value of each enhanced sample, the fourth losses of multiple enhanced samples are weighted to obtain an enhanced sample loss.

[0136] In some embodiments, the enhancement value in step 1042 can be obtained in the following manner: when the true label of the enhanced sample is set to not clicked, the difference between the first loss and the second loss is used as the enhancement value of the enhanced sample.

[0137] As an example, the difference between the first loss and the second loss is caused by training the enhanced recommendation model with a single unexposed sample. Therefore, the difference between the two losses vi = first loss - second loss can be used to measure whether the introduction of the enhanced sample as a negative sample into the training data set is helpful for the enhanced recommendation model. Therefore, for the enhanced sample with the true label of not clicked, the difference vi is used as the enhancement value to evaluate the degree of influence of the enhanced sample on improving the prediction accuracy of the enhanced recommendation model.

[0138] By using the difference between the first loss and the second loss as the enhancement value of the enhanced sample, quantitatively evaluating the degree of influence of the enhanced sample on improving the prediction accuracy of the enhanced recommendation model, and using the enhancement value of the enhanced sample as the basis for weighted processing of the loss function, training the basic recommendation model to obtain the target enhanced recommendation model, thereby improving the prediction accuracy of the target enhanced recommendation model.

[0139] In some embodiments, step 1042 can be implemented in the following manner: taking the absolute value of the enhancement value corresponding to each enhanced sample to obtain the absolute value result corresponding to each enhanced sample, normalizing the absolute value result corresponding to each enhanced sample to obtain the weight corresponding to each enhanced sample, and based on the weight corresponding to each enhanced sample, performing weighted processing on the fourth loss of multiple enhanced samples to obtain the enhanced sample loss.

[0140] As an example, the first loss of each enhanced sample is less than the second loss. Therefore, the enhancement value of the enhanced sample = first loss - second loss < 0. It is necessary to first take the absolute value of the enhancement value of each enhanced sample to ensure that the absolute value result corresponding to each enhanced sample is positive, and then normalize the absolute value result corresponding to each enhanced sample to obtain the weight corresponding to each enhanced sample, and use formula (1) to perform weighted processing on the fourth loss of each enhanced sample with the corresponding weight to obtain the weighted enhanced sample loss Weighted Loss tgt ′:

[0141]

[0142] Among them, is the weight after taking the absolute value and normalizing the negative number v i , the part in the brackets is the fourth loss of the corresponding enhanced sample. The fourth loss can be the cross-entropy loss function or other types of loss functions. y i is the true label of the corresponding enhanced sample i, and p i is the fourth prediction recommendation index of the corresponding enhanced sample i.

[0143] By calculating the weight of the fourth loss for each augmented sample based on the augmentation value of each augmented sample, the influence degree of the fourth loss corresponding to different augmented samples on the augmented sample loss is made different. The fourth loss of the augmented sample with a greater influence on improving the prediction accuracy of the augmented recommendation model also has a greater influence on the augmented sample loss, enabling the basic recommendation model to update parameters with a preference, thereby improving the prediction accuracy of the target augmented recommendation model.

[0144] In step 1043, each second exposure sample is propagated forward in the basic recommendation model to obtain the fifth loss corresponding to each second exposure sample, and the fifth losses of multiple second exposure samples are fused to obtain the exposure sample loss.

[0145] As an example, the second exposure sample is a sample that enters the distribution link and the exposure link obtained from the information data stream. The second exposure sample can be the same as the first exposure sample or the third exposure sample, or different from the first exposure sample and the third exposure sample. Multiple second exposure samples are used as positive samples and trained in the basic recommendation model to obtain the fifth loss corresponding to each second exposure sample, and the fifth losses of multiple second exposure samples are fused to obtain the exposure sample loss. The fusion processing here can be averaging processing or other fusion processing methods, which are not limited in this application.

[0146] In step 1044, based on the fusion result of the augmented sample loss and the exposure sample loss, the basic recommendation model is updated to obtain the target augmented recommendation model.

[0147] As an example, the augmented sample loss and the exposure sample loss are fused to obtain a fusion loss, which includes the loss corresponding to each positive sample (i.e., the second exposure sample) and the weighted loss corresponding to each negative sample (i.e., the augmented sample). Based on the fusion loss, the basic recommendation model is updated to obtain the target augmented recommendation model.

[0148] By using the augmented sample that helps improve the model prediction accuracy as the negative sample and the second exposure sample as the positive sample, calculating the loss corresponding to each sample, and performing weighted processing on the loss of the augmented sample based on the augmentation value, on the one hand, the basic recommendation model can be trained with positive and negative samples, and on the other hand, the loss of the augmented sample with a greater influence on improving the prediction accuracy of the augmented recommendation model also has a greater influence on the augmented sample loss, enabling the basic recommendation model to update parameters with a preference, thereby improving the prediction accuracy of the target augmented recommendation model.

[0149] In some embodiments, before step 101, the following operations may also be performed: obtaining a third exposure sample different from the first exposure sample, training the initialized recommendation model based on the third exposure sample, and obtaining a basic recommendation model.

[0150] As an example, the third exposure sample is a sample that enters the distribution link and the exposure link obtained from the information data stream. The third exposure sample is input into the initialized recommendation model for forward propagation to obtain the predicted click probability corresponding to the third exposure sample. Suppose the true label of the third exposure sample is 1. Based on the true label and the predicted click probability of the third exposure sample, the corresponding loss is calculated, and the initialized recommendation model is updated based on this loss to obtain a basic recommendation model.

[0151] By training the initialized recommendation model with the third exposure sample as the positive sample, the initialized recommendation model can perform basic prediction and recommendation processing.

[0152] In some embodiments, the following operations may also be performed: obtaining the exposure data of the exposure information sample and the non-exposure data of the non-exposure information sample, performing a first embedding process on the exposure data based on the first embedding parameter to obtain an exposure sample, and performing a second embedding process on the non-exposure data based on the second embedding parameter to obtain a non-exposure sample.

[0153] As an example, outside the basic recommendation model or the enhanced recommendation model, two embedding layers are pre-set, namely a basic embedding layer for performing the first embedding process based on the first embedding parameter, and an additional embedding layer for performing the second embedding process based on the second embedding parameter. The two embedding layers are parallel. The first embedding parameter of the basic embedding layer can be regarded as the basic embedding vector corresponding to the features of the exposure data, and the second embedding parameter of the additional embedding layer can be regarded as the additional embedding layer corresponding to the characteristics of the non-exposure data. When the data input into the model is exposure data, the basic embedding layer is called to perform the first embedding process on the exposure data based on the first embedding parameter to obtain an exposure sample. When the data input into the model is non-exposure data, the additional embedding layer is called to perform the second embedding process on the non-exposure data based on the second embedding parameter to obtain a non-exposure sample.

[0154] As an example, during the model training process, the numerical values of the basic embedding vectors of the basic embedding layer and the additional embedding vectors of the additional embedding layer will be updated as the model parameters of the model are updated. For example, when training the initialized recommendation model using the third exposure sample, the numerical values of the basic embedding vectors of the basic embedding layer that processes the exposure data into the third exposure sample will also be updated accordingly. When training the basic recommendation model using the enhanced sample and the second exposure sample, the numerical values of the additional embedding vectors of the additional embedding layer that processes the unexposed data into the enhanced sample will also be updated accordingly, and the numerical values of the basic embedding vectors of the basic embedding layer that processes the exposure data into the second exposure sample will also be updated accordingly.

[0155] By performing embedding processing on the exposure data and the unexposed data respectively, the features between the obtained exposure samples and unexposed samples are distinguished, thereby improving the prediction accuracy of the target enhanced recommendation model.

[0156] Next, the exemplary application and implementation of the server provided in the embodiments of the present application will be combined to illustrate the recommendation processing method based on artificial intelligence provided in the embodiments of the present application.

[0157] It should be noted that in the examples of recommendation processing below, information recommendation is used as an example for illustration. Those skilled in the art can apply the recommendation processing method based on artificial intelligence provided in the embodiments of the present application to the recommendation processing including other types of recommended objects according to the understanding of the following text.

[0158] See Figure 4 , Figure 4 is a schematic flowchart of the recommendation processing method based on artificial intelligence provided in the embodiments of the present application, and will be described in combination with Figure 4 the steps 201 to 203 shown.

[0159] In step 201, the information to be recommended is obtained.

[0160] As an example, the information to be recommended is obtained from the information data stream. Among them, for application scenarios corresponding to different types of recommended content, the information to be recommended can be commodity information, news information, video information, etc.

[0161] In step 202, the target enhanced recommendation model is called to perform prediction processing on the recommendation metrics of the information to be recommended, and the predicted recommendation metrics of the information to be recommended are obtained.

[0162] As an example, the target enhanced recommendation model is trained by the model training method based on artificial intelligence provided by the embodiments of the present application. First, the basic features of the information to be recommended are obtained. The basic features can be abstractly represented as: <features related to the target user, features related to the subscription number to which the information to be recommended belongs, features related to the information to be recommended itself, statistical features of cross relationships>. At this time, the basic features are discrete features with statistical significance. At this time, it is necessary to perform embedding processing on the basic features to obtain the data features of the information to be recommended. At this time, the data features are input into the target enhanced recommendation model, and the data features perform forward propagation in the target enhanced recommendation model, and then the predicted recommendation index of the information to be recommended can be obtained. In the application scenario where the information to be recommended is commodity information, the predicted recommendation index of the information to be recommended can be the probability of predicting that the user clicks on the commodity to enter the details interface; in the application scenario where the information to be recommended is news information, the predicted recommendation index of the information to be recommended can be the probability of predicting that the user clicks on the news information for effective reading; in the application scenario where the information to be recommended is video information, the predicted recommendation index of the information to be recommended can be the probability of predicting that the user clicks on the video for effective browsing.

[0163] As an example, before inputting the data features into the target enhanced recommendation model, the data features can also be subjected to a first embedding process based on the first embedding parameter through the basic embedding layer. The reason is that on the application side, the information to be recommended can be regarded as information that may be exposed.

[0164] In step 203, based on the predicted recommendation index of the information to be recommended, recommendation processing is performed on the information to be recommended.

[0165] As an example, when performing recommendation index prediction processing on the information to be recommended, it can be to process multiple pieces of information to be recommended simultaneously, or to process a single piece of information to be recommended. When performing recommendation processing, it can be sorted in descending order according to the predicted recommendation indexes of multiple pieces of information to be recommended, and the information to be recommended with a higher ranking is subjected to recommendation processing. It can also be to preset a recommendation index threshold, and perform recommendation processing on the information to be recommended whose predicted recommendation index meets the recommendation index threshold.

[0166] By using the target enhanced recommendation model trained with enhanced samples and second exposure samples, prediction recommendation index processing is performed on the information to be recommended, and based on the predicted recommendation index of the information to be recommended, recommendation processing is performed on the information to be recommended, ensuring that the information to be recommended obtained by recommendation has a high probability of being clicked by the user.

[0167] Next, an exemplary application of the embodiments of the present application in an actual recommendation ranking model application scenario will be described.

[0168] The input of the recommended sorting model is the relevant basic features of the subscription account information, and the output is the click probability of the subscription account information. The subscription account information is sorted according to the click probability, and the top n subscription account information is used as the exposure result. Specifically, each sample processed by the recommended sorting model represents the relevant basic features of a user for a piece of subscription account information exposed to it. The basic features can be abstractly represented as: <target user-related features, subscription account-related features of the information to be recommended, information-related features of the information to be recommended itself, statistical features of cross relationships>. Among them, the target user-related features can include: for example, user account identifier, user age, user gender, user location, the number of exposed information of the user in the past 1 day, the number of exposed information of the user in the past 7 days, and so on; the subscription account-related features of the information to be recommended can include: for example, subscription account identifier, number of fans of the subscription account, number of created information of the subscription account in the past 7 days, number of click-through readings of the subscription account in the past 7 days, and so on; the information-related features of the information to be recommended itself can include: for example, information account identifier, number of hours since the information was sent, number of exposures of the information in the past 1 hour, number of clicks of the information in the past 1 hour, and so on. Among them, the subscription account identifier combined with the information account identifier obtains the account identifier of each item (i.e., subscription account information); the statistical features of cross relationships can include: for example, the number of exposures of the user to the subscription account in the past 28 days, the number of clicks of the user on the subscription account in the past 28 days, and so on. See Figure 5 , Figure 5 is a schematic diagram of the subscription account recommendation interface provided by an embodiment of the present application. As Figure 5 shown, the subscription account interface displays the recommendation sections of subscription account 1, subscription account 2, and subscription account 3 respectively. At least one piece of subscription account information is displayed in the recommendation section of each subscription account. The corresponding relationship between the subscription account and the subscription account information is: multiple pieces of subscription account information are included under the subscription account, and each piece of subscription account information has its own belonging subscription account.

[0169] See Figure 6 , Figure 6 is a schematic diagram of the recommended sorting link provided by an embodiment of the present application. As Figure 6 shown, in a recommended sorting link, for the candidate information set to reach the user's personal client display interface, the following links are required: recall, rough sorting, fine sorting, mixed sorting, distribution, and exposure. And the approximate order of magnitude of the number of information processed in each link is as follows: the approximate order of magnitude of the number of information processed in the recall link is 10 6 , the approximate order of magnitude of the number of information processed in the rough sorting link is 10 4 , the approximate order of magnitude of the number of information processed in the fine sorting link is 10 2, the approximate order of magnitude of the number of processed information in the mixed sorting stage, the distribution stage, and the exposure stage is all 10. In the recall and rough sorting stages, hundreds of user-preferred items are quickly selected from an information set ranging from one hundred thousand to one million levels. Then, through the precise personalized modeling of the fine sorting and the mixed sorting of various sources and various modal items, finally, after being scattered by some business rules, the final distribution sequence is obtained, that is, an information list sent to the user. Multiple pieces of information are sequentially displayed in the information list. For example, information a corresponding to the distribution position 1, information b corresponding to the distribution position 2, information c corresponding to the distribution position 3, and information d corresponding to the distribution position 4. Since not all the information in the information list can be exposed to the recommendation interface, as Figure 6 shown, the distribution position 3 is the exposure cut-off position. Then information a, information b, and information c are exposed, while the information after information d corresponding to the distribution position 4 is in the unexposed state of distribution. For the exposed information, the feedback result of whether the user clicks can be obtained. For example, for the exposed information a, the feedback result that the user has clicked is obtained, while for the exposed information b, the feedback result that the user has not clicked is obtained. For the other information in the unexposed state of distribution, it is unknown whether the user clicks. Usually, the recommendation system retrieves all the exposed information and, after feature engineering processing, uses it to train the sorting model. In this case, there is an inconsistency in the data space where the offline training of the model and the inference of the online service are located: the model is trained in the exposure data space, that is, the model uses the exposure data for training, and the inference is performed in the mixed sorting space (for convenience, it can be approximately considered equal to the distribution space), that is, the model performs recommendation sorting on the distribution data. Therefore, the part of the information items that have passed through the mixed sorting stage and the distribution stage but have not been exposed can be used to expand the training data set of the recommendation sorting model. Moreover, training on the data space after sample expansion will not cause a large deviation in the training-recommendation data distribution. The problem is that since in the data space of the unexposed distribution, the system has not collected the feedback of the user on each piece of information, the samples with unknown labels cannot be directly used to evaluate the quality of the model's click-through rate prediction, and thus the model cannot be trained.

[0170] See Figure 7 , Figure 7 is the pre-training process framework of the recommendation sorting model provided by the embodiment of the present application. As Figure 7As shown in the figure, the framework mainly consists of two parts: a basic recommendation model trained using exposure samples, and an enhanced recommendation model jointly trained using exposure samples and unexposed samples sent down (i.e., unexposed samples). In addition, there are a basic embedding layer (BaseEmbedding layer) and a subsidiary embedding layer (Subsidiary Embedding layer). The basic embedding layer stores the features of the corresponding exposure data and the basic embedding vectors of item account identifiers (base embedding vectors, corresponding to the first embedding parameter). When the input data is the exposure data of the corresponding exposure sample, the basic embedding layer is called to perform the first embedding process on the exposure data to obtain the exposure sample, and the basic embedding vectors are updated along with the training of the basic recommendation model (calculating gradients for the model parameters based on the loss function and using the gradients to update the base embedding vectors). The subsidiary embedding layer stores the features of the corresponding unexposed data sent down and the subsidiary embedding vectors of item account identifiers (Subsidiary embedding vectors, corresponding to the second embedding parameter). When the input data is the unexposed data of the corresponding unexposed sample sent down, the subsidiary embedding layer is called to perform the second embedding process on the unexposed data to obtain the unexposed sample sent down, and the subsidiary embedding vectors are updated along with the training of the enhanced recommendation model. The two parts of the model (including the basic embedding layer and the subsidiary embedding layer) use the same network layer structure and parameter scale. During the training process, the two parts of the model are updated alternately, and their parameters are synchronized after a certain number of steps. The specific process is as follows.

[0171] First step, train the basic recommendation model using exposure samples: continue to refer to Figure 7 , for the basic recommendation model, after N I iterations, perform gradient updates (forward propagation and backward propagation). In each iteration, obtain a certain number of exposure samples (i.e., the first exposure samples) from the data stream, and train and update the parameters of the basic recommendation model. Subsequently, synchronize the parameters of the updated basic recommendation model to Figure 7 the intermediate enhanced recommendation model.

[0172] Second step, continue to refer to Figure 7 , obtain the unexposed samples sent down (i.e., unexposed samples) from the data stream, and sample an unexposed sample set with a sample size of N scr . Then, the enhanced recommendation model traverses each unexposed sample sent down in this enhanced sample set one by one. For example, first use unexposed sample 1 to train the enhanced recommendation model (at this time, the model parameters of the enhanced recommendation model are the same as those of the basic recommendation model) to obtain the enhanced recommendation model corresponding to unexposed sample #N scr1 , and then use unexposed sample 2 to train the enhanced recommendation model (at this time, the model parameters of the enhanced recommendation model are the same as those of the basic recommendation model) to obtain the enhanced recommendation model corresponding to unexposed sample #Nscr2 The enhanced recommendation model, and so on. Among them, in one training for a single unexposed sample sent down, the unexposed sample sent down is regarded as a click negative sample (i.e., the label is 0). According to the output result and the true label of the sample, the loss function value of a single unexposed sample sent down is calculated. Based on the loss function value of a single unexposed sample sent down, the enhanced recommendation model is trained and updated. Denote the third loss of this single unexposed sample sent down as Loss src . Then, sample a batch of a fixed number of exposed samples from the remaining exposed samples that have not been used to train the basic recommendation model as the anchor data (i.e., the third exposed sample). After performing one round of gradient update on the enhanced recommendation model based on the loss function value of a single enhanced sample, the anchor data is processed in the enhanced recommendation model and the basic recommendation model respectively, and the first loss Loss tgt of the enhanced recommendation model and the second loss Loss base of the basic recommendation model are calculated. Since the model parameters of the enhanced recommendation model and the basic recommendation model are the same before one training for a single unexposed sample sent down, and after one training for a single unexposed sample sent down, the difference between the model parameters of the enhanced recommendation model and the basic recommendation model is caused by the enhanced recommendation model being trained and updated with a single unexposed sample sent down. Therefore, for the same anchor data, the difference between Loss tgt of the enhanced recommendation model and Loss base of the basic recommendation model is due to the training of the enhanced recommendation model with a single unexposed sample i. Therefore, the difference v i between the two loss functions = the first loss of the corresponding unexposed sample i - the second loss = Loss tgt (i) - Loss base can be used to measure whether the enhancement training of the enhanced recommendation model is helpful after introducing the unexposed sample i sent down as a negative sample into the training dataset. When v i <0, it means that after training the enhanced recommendation model with the unexposed sample i sent down, the loss function value of the anchor data is reduced, and it can be considered that the unexposed sample i sent down is a negative sample that is helpful for improving the prediction accuracy of the enhanced recommendation model. It should be noted that after obtaining Loss tgt and Loss base , the enhanced recommendation model and the basic recommendation model are not updated according to Loss tgt and Loss base , but gradient clipping is performed.

[0173] After completing the training of the enhanced recommendation model with a single unexposed sample sent down, the enhanced recommendation model is re-initialized as the updated basic recommendation model in the first step, that is, according to Loss srcAfter one round of gradient update on the enhanced recommendation model, the enhanced recommendation model is restored to ensure that the enhanced recommendation model can continue to evaluate the enhancement effect of subsequent unexposed samples sent down on the enhanced recommendation model. When traversing N scr unexposed samples sent down, N scr v i values are obtained, and each v i value corresponds to the enhancement value of each unexposed sample sent down.

[0174] In the third step, select the unexposed samples sent down with v i value < 0. For example, the enhancement value v1 of the unexposed sample sent down 1, the enhancement value v2 of the unexposed sample sent down 2, and the enhancement value v3 of the unexposed sample sent down 3. Screen the v i values less than 0 from v1, v2, and v3. The unexposed samples sent down with v i value < 0 (i.e., enhanced samples) and the exposed samples obtained from the data stream (i.e., the second exposed samples) form the expanded enhanced dataset. For the enhanced recommendation model synchronized with the parameters of the basic recommendation model, perform N U iterations, and sample a certain number of training samples from the enhanced dataset each time. Calculate the corresponding fifth loss Loss tgt ″ using the second exposed samples. When calculating the fourth loss Loss tgt ' of the enhanced samples, use the enhancement value v i obtained in the second step, and weight the loss function according to formula (2). Here, only negative samples, that is, enhanced samples, are weighted to obtain the enhanced sample loss Weighted Loss tgt ″ (the weighted result of the fourth loss):

[0175]

[0176] where is the weighted weight after taking the negative of the negative v i and performing min-max normalization. The part in the parentheses is the cross-entropy loss function calculated for negative samples. y i is the true label corresponding to the enhanced sample i, and p i is the predicted recommendation index corresponding to the enhanced sample i. Perform gradient update on the enhanced recommendation model based on Weighted Loss tgt ″.

[0177] In the fourth step, synchronize the model parameters of the trained enhanced recommendation model back to the basic recommendation model. On the continuously updated data stream, repeat the above first to third steps, and finally obtain the target enhanced recommendation model.

[0178] See Figure 8 for details.Figure 8 This is a schematic diagram of the application side of the target enhancement recommendation model provided in the embodiment of the present application. Due to continuous parameter synchronization, after the data enhancement training mode, to be applied to the reasoning link of the online service, only one of the basic recommendation model and the enhanced recommendation model needs to be deployed as the target enhancement recommendation model. Figure 8 In the example, the basic recommendation model after parameter synchronization is used as the target enhanced recommendation model to obtain the recommended information from the online data stream. Since the embedding vector values of the basic embedding layer and the additional embedding layer are different, the recommended information obtained by the default application side may be exposed. Figure 8 As shown, the basic embedding layer is selected as the embedding layer on the application side, and the basic embedding layer is called to perform the first embedding processing on the recommended information to obtain the corresponding data features, and the data features are input into the target enhanced recommendation model to perform predictive recommendation index processing to obtain the predicted recommendation index corresponding to the information to be recommended, and the recommendation processing is performed on the recommended information, and finally the recommended information is output.

[0179] See also Figure 9 , Figure 9 : is a target enhancement recommendation model application feedback data diagram provided by the embodiment of the present application. Figure 9 As shown, when the target enhanced recommendation model provided in the embodiment of the present application is applied to the recommendation sorting of subscription account information, compared with the control group A1+A2, the B3 group using the target enhanced recommendation model has an average number of click quadruplets increased by 0.616%, the proportion of valid click users increased by 0.505%, the proportion of users with consumption increased by 0.029%, and the ratio of the number of recommended flow click users to the number of hit users increased by 0.693%. The B4 group using the target enhanced recommendation model has an average number of click quadruplets increased by 0.228%, the proportion of valid click users increased by 0.970%, the proportion of users with consumption increased by 0.147%, and the ratio of the number of recommended flow click users to the number of hit users increased by 0.520%. It can be seen that the application of the target enhanced recommendation model provided in the embodiment of the present application to the recommendation sorting of subscription account information increases the utilization rate of log data. Compared with the baseline model, the average number of information readings and reading time are significantly improved. It can be seen that the target enhanced recommendation model provided in the embodiment of the present application has better sorting ability.

[0180] It is understandable that in the embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0181] The following is a description of an exemplary structure of the artificial intelligence-based model training device 233 provided in the embodiment of the present application implemented as a software module. In some embodiments,Figure 2A As shown, the software modules stored in the artificial intelligence-based model training device 233 of the memory 230 may include: a basic training module 2331, which is used to train the basic recommendation model respectively based on each unexposed sample to obtain an enhanced recommendation model corresponding to each unexposed sample; a loss calculation module 2332, which is used to perform forward propagation of the first exposure sample in each enhanced recommendation model to obtain a first loss corresponding to each enhanced recommendation model, and perform forward propagation of the first exposure sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model; a sample screening module 2333, which is used to, when the first loss is less than the second loss, use the unexposed sample corresponding to the first loss as an enhanced sample; an enhanced training module 2334, which is used to train the basic recommendation model based on the enhanced sample and the second exposure sample to obtain a target enhanced recommendation model.

[0182] In some embodiments, the basic training module 2331 is further used to obtain a third exposure sample different from the first exposure sample, and train the initialized recommendation model based on the third exposure sample to obtain a basic recommendation model.

[0183] In some embodiments, the basic training module 2331 is further used to perform the following processing on each unexposed sample: perform forward propagation of the unexposed sample in the basic recommendation model to obtain a first predicted recommendation metric corresponding to the unexposed sample, determine a third loss corresponding to the unexposed sample based on the true label of the unexposed sample and the first predicted recommendation metric corresponding to the unexposed sample, and perform an update process on the basic recommendation model based on the third loss corresponding to the unexposed sample to obtain an enhanced recommendation model corresponding to the unexposed sample.

[0184] In some embodiments, the loss calculation module 2332 is further used to perform the following processing on each first exposure sample: perform forward propagation of the first exposure sample in each enhanced recommendation model to obtain a second predicted recommendation metric corresponding to each enhanced recommendation model, determine a first single-sample loss of the first exposure sample corresponding to each enhanced recommendation model based on the true label of the first exposure sample and the second predicted recommendation metric corresponding to each enhanced recommendation model, and for each enhanced recommendation model, perform a fusion process on the first single-sample losses of the multiple first exposure samples corresponding to the enhanced recommendation model to obtain a first loss corresponding to the enhanced recommendation model.

[0185] In some embodiments, the loss calculation module 2332 is further configured to perform the following processing on each first exposure sample: perform forward propagation of the first exposure sample in the basic recommendation model to obtain a third predicted recommendation metric corresponding to the first exposure sample; determine a second single-sample loss corresponding to the first exposure sample based on the true label of the first exposure sample and the third predicted recommendation metric corresponding to the first exposure sample; and perform a fusion process on the second single-sample losses of multiple first exposure samples to obtain a second loss corresponding to the basic recommendation model.

[0186] In some embodiments, the enhanced training module 2334 is further configured to perform forward propagation of each enhanced sample in the basic recommendation model to obtain a fourth loss corresponding to each enhanced sample; perform a weighted process on the fourth losses of multiple enhanced samples based on the enhancement value of each enhanced sample to obtain an enhanced sample loss; perform forward propagation of each second exposure sample in the basic recommendation model to obtain a fifth loss corresponding to each second exposure sample; perform a fusion process on the fifth losses of multiple second exposure samples to obtain an exposure sample loss; and perform an update process on the basic recommendation model based on the fusion result of the enhanced sample loss and the exposure sample loss to obtain a target enhanced recommendation model.

[0187] In some embodiments, the enhanced training module 2334 is further configured to perform the following processing on each enhanced sample: perform forward propagation of the enhanced sample in the basic recommendation model to obtain a fourth predicted recommendation metric corresponding to the enhanced sample; and determine a fourth loss corresponding to the enhanced sample based on the true label of the enhanced sample and the fourth predicted recommendation metric corresponding to the enhanced sample.

[0188] In some embodiments, when the true label of the enhanced sample is set to not clicked, the enhanced training module 2334 is further configured to use the difference between the first loss and the second loss as the enhancement value of the enhanced sample.

[0189] In some embodiments, the enhanced training module 2334 is further configured to perform an absolute value process on the enhancement value corresponding to each enhanced sample to obtain an absolute value result corresponding to each enhanced sample; perform a normalization process on the absolute value result corresponding to each enhanced sample to obtain a weight corresponding to each enhanced sample; and perform a weighted process on the fourth losses of multiple enhanced samples based on the weight corresponding to each enhanced sample to obtain an enhanced sample loss.

[0190] In some embodiments, the basic training module 2331 is further configured to obtain the exposure data of the exposure information sample and the non-exposure data of the non-exposure information sample, perform a first embedding process on the exposure data based on a first embedding parameter to obtain an exposure sample, and perform a second embedding process on the non-exposure data based on a second embedding parameter to obtain a non-exposure sample.

[0191] Next, the exemplary structure of the software module implementation of the recommendation processing device 263 based on artificial intelligence provided in the embodiments of the present application will be further described. In some embodiments, as Figure 2B shown, the software module stored in the recommendation processing device 263 based on artificial intelligence in the memory 260 may include: an acquisition module 2631 for acquiring information to be recommended; a prediction processing module 2632 for calling a target enhanced recommendation model to perform recommendation index prediction processing on the information to be recommended to obtain a predicted recommendation index of the information to be recommended, where the target enhanced recommendation model is trained by the model training method based on artificial intelligence provided in the embodiments of the present application; and a recommendation processing module 2633 for performing recommendation processing on the information to be recommended based on the predicted recommendation index of the information to be recommended.

[0192] The embodiments of the present application provide a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the model training method based on artificial intelligence and the recommendation processing method based on artificial intelligence described above in the embodiments of the present application.

[0193] The embodiments of the present application provide a computer-readable storage medium storing computer executable instructions, where computer executable instructions or a computer program are stored, and when the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the model training method based on artificial intelligence and the recommendation processing method based on artificial intelligence provided in the embodiments of the present application. For example, as Figure 3A shown in the model training method based on artificial intelligence, Figure 4 shown in the recommendation processing method based on artificial intelligence.

[0194] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0195] In some embodiments, the computer executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0196] By way of example, the computer-executable instructions may or may not correspond to files in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program under discussion, or in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0197] By way of example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected by a communication network.

[0198] In summary, through the embodiments of the present application, negative sampling augmentation is performed on the "downstream unexposed data" that has not received research attention to achieve data enhancement. The information that is sent downstream by the recommendation system but not exposed is used to augment the training sample set of the recommendation ranking model, enabling the recommendation model to learn and fit personalized interests on a larger scale and more complete data. At the same time, a training mode suitable for using samples with unknown labels to assist model training is proposed to avoid bringing negative effects to the basic recommendation model after introducing augmented samples. When the embodiments of the present application are applied to the recommendation ranking of WeChat subscription information, the utilization rate of log data is increased. Compared with the baseline model, the average number of information readings per person and the reading duration are significantly improved, demonstrating that the model trained by the method provided by the embodiments of the present application has higher prediction accuracy and better ranking ability.

[0199] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A model training method based on artificial intelligence, characterized in that, The method includes: Training the basic recommendation model based on each unexposed sample respectively to obtain an enhanced recommendation model corresponding to each of the unexposed samples; Performing forward propagation of the first exposure sample in each of the enhanced recommendation models to obtain a first loss corresponding to each of the enhanced recommendation models, and performing forward propagation of the first exposure sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model; When the first loss is less than the second loss, using the unexposed sample corresponding to the first loss as an enhanced sample; Training the basic recommendation model based on the enhanced sample and the second exposure sample to obtain a target enhanced recommendation model.

2. The method according to claim 1, characterized in that, Before training the basic recommendation model based on each unexposed sample respectively, the method further includes: Obtaining a third exposure sample different from the first exposure sample; Training the initialized recommendation model based on the third exposure sample to obtain the basic recommendation model.

3. The method according to claim 1, wherein The training the basic recommendation model based on each unexposed sample respectively to obtain an enhanced recommendation model corresponding to each of the unexposed samples includes: Performing the following processing on each of the unexposed samples: Performing forward propagation of the unexposed sample in the basic recommendation model to obtain a first predicted recommendation metric corresponding to the unexposed sample; Determining a third loss corresponding to the unexposed sample based on the true label of the unexposed sample and the first predicted recommendation metric corresponding to the unexposed sample; Updating the basic recommendation model based on the third loss corresponding to the unexposed sample to obtain an enhanced recommendation model corresponding to the unexposed sample.

4. The method according to claim 1, wherein The performing forward propagation of the first exposure sample in each of the enhanced recommendation models to obtain a first loss corresponding to each of the enhanced recommendation models includes: Performing forward propagation of the first exposure sample in each of the enhanced recommendation models to obtain a second predicted recommendation metric corresponding to each of the enhanced recommendation models; Determining a first single-sample loss of the first exposure sample corresponding to each of the enhanced recommendation models based on the true label of the first exposure sample and the second predicted recommendation metric corresponding to each of the enhanced recommendation models; For each of the enhanced recommendation models, fusing the first single-sample losses of the multiple first exposure samples corresponding to the enhanced recommendation model to obtain a first loss corresponding to the enhanced recommendation model.

5. The method according to claim 1, characterized in that, The performing forward propagation of the first exposure sample in the basic recommendation model to obtain a second loss corresponding to the basic recommendation model includes: Performing forward propagation of the first exposure sample in the basic recommendation model to obtain a third predicted recommendation metric corresponding to the first exposure sample; Determining a second single-sample loss corresponding to the first exposure sample based on the true label of the first exposure sample and the third predicted recommendation metric corresponding to the first exposure sample; Fusing the second single-sample losses of the multiple first exposure samples to obtain a second loss corresponding to the basic recommendation model.

6. The method according to claim 1, wherein Training the basic recommendation model based on the enhanced samples and the second exposure samples to obtain a target enhanced recommendation model includes: Performing forward propagation of each of the enhanced samples in the basic recommendation model to obtain a fourth loss corresponding to each of the enhanced samples; Based on the enhancement value of each enhanced sample, performing weighted processing on the fourth losses of multiple enhanced samples to obtain an enhanced sample loss; Performing forward propagation of each of the second exposure samples in the basic recommendation model to obtain a fifth loss corresponding to each of the second exposure samples, and performing fusion processing on the fifth losses of multiple second exposure samples to obtain an exposure sample loss; Based on the fusion result of the enhanced sample loss and the exposure sample loss, performing update processing on the basic recommendation model to obtain the target enhanced recommendation model.

7. The method according to claim 6, wherein The performing forward propagation of each of the enhanced samples in the basic recommendation model to obtain a fourth loss corresponding to each of the enhanced samples includes: Performing the following processing on each of the enhanced samples: Performing forward propagation of the enhanced sample in the basic recommendation model to obtain a fourth predicted recommendation metric corresponding to the enhanced sample; Based on the true label of the enhanced sample and the fourth predicted recommendation metric corresponding to the enhanced sample, determining a fourth loss corresponding to the enhanced sample.

8. The method according to claim 6, wherein The method further includes: When the true label of the enhanced sample is set to not clicked, using the difference between the first loss and the second loss as the enhancement value of the enhanced sample.

9. The method according to claim 6, wherein The performing weighted processing on the fourth losses of multiple enhanced samples based on the enhancement value of each enhanced sample to obtain an enhanced sample loss includes: Performing absolute value processing on the enhancement value corresponding to each enhanced sample to obtain an absolute value result corresponding to each enhanced sample; Performing normalization processing on the absolute value result corresponding to each enhanced sample to obtain a weight corresponding to each enhanced sample; Based on the weight corresponding to each enhanced sample, performing weighted processing on the fourth losses of multiple enhanced samples to obtain the enhanced sample loss.

10. The method according to any one of claims 1 to 9, characterized in that The method further includes: Obtaining exposure data of an exposure information sample and unexposed data of an unexposed information sample; Performing first embedding processing on the exposure data based on a first embedding parameter to obtain the exposure sample, and performing second embedding processing on the unexposed data based on a second embedding parameter to obtain the unexposed sample.

11. A recommendation processing method based on artificial intelligence, characterized in that The method includes: Obtaining information to be recommended; Invoking the target enhanced recommendation model to perform recommendation metric prediction processing on the information to be recommended to obtain a predicted recommendation metric of the information to be recommended, where the target enhanced recommendation model is trained by the method according to any one of claims 1 to 10; Based on the predicted recommendation metric of the information to be recommended, performing a recommendation process on the information to be recommended.

12. An artificial intelligence-based model training device, characterized in that, The apparatus includes: A basic training module, configured to train a basic recommendation model based on each unexposed sample respectively to obtain an enhanced recommendation model corresponding to each unexposed sample; A loss calculation module, configured to perform forward propagation of the first exposure samples in each of the enhanced recommendation models to obtain first losses corresponding to each of the enhanced recommendation models, and perform forward propagation of the first exposure samples in the basic recommendation model to obtain second losses corresponding to the basic recommendation model; A sample screening module, configured to use the unexposed samples corresponding to the first losses as enhanced samples when the first losses are less than the second losses; An enhanced training module, configured to train the basic recommendation model based on the enhanced samples and second exposure samples to obtain a target enhanced recommendation model.

13. A recommendation processing device based on artificial intelligence, characterized in that, The apparatus includes: An acquisition module, configured to acquire information to be recommended; A prediction processing module, configured to call the target enhanced recommendation model to perform prediction processing on recommendation metrics of the information to be recommended, to obtain predicted recommendation metrics of the information to be recommended, where the target enhanced recommendation model is trained by the method according to any one of claims 1 to 10; A recommendation processing module, configured to perform recommendation processing on the information to be recommended based on the predicted recommendation metrics of the information to be recommended.

14. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the model training method according to any one of claims 1 to 10 or the recommendation processing method according to claim 11 when executing the computer-executable instructions stored in the memory.

15. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the model training method according to any one of claims 1 to 10 or the recommendation processing method according to claim 11.

16. A computer program product, comprising computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the model training method according to any one of claims 1 to 10 or the recommendation processing method according to claim 11.