Model training method and device, information recommendation method and device, electronic equipment, computer program product and computer readable storage medium

Through the recommendation model trained by inverse sampling and joint loss value, the problem of traditional algorithms being biased towards popular information is solved, more accurate and broader information recommendation is achieved, and the personalized recommendation effect of the model is improved.

CN120372073APending Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410107443.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional information recommendation algorithms are prone to favor popular recommendation information, resulting in low exposure and click-through rates of long-tail recommendation information, and high homogeneity of recommendation results, making it impossible to effectively utilize the fine-grained characteristics of long-tail recommendation information.

Method used

The initial training data set is sampled using inverse sampling probability, combined with the first sub-model and the second sub-model for prediction, and the recommended model is trained using joint loss values to balance the positive and negative sample ratio, reduce the risk of overfitting, and improve the generalization ability of the model.

Benefits of technology

The prediction accuracy and generalization ability of the recommendation model are improved, ensuring that the recommendation results are more in line with the personalized needs of the target object, and the recommendation scope is expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372073A_ABST
    Figure CN120372073A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and an information recommendation method. The method comprises the steps of obtaining a to-be-trained recommendation model, and obtaining an initial training data set; performing random sampling on the initial training data set to obtain a first training data set, and performing inverse sampling on the initial training data set based on an inverse sampling probability to obtain a second training data set; predicting each first training sample in the first training data set by using the first sub-model to obtain each first predicted value, and predicting each second training sample in the second training data set by using the second sub-model to obtain each second predicted value; determining a joint loss value by using the first sample label of each first training sample, each first prediction value and the second sample label of each second training sample; and training a to-be-trained recommendation model by using the joint loss value to obtain a trained recommendation model. According to the invention, the recommendation network model with a more accurate recommendation effect and a wider recommendation range can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a model training method, an information recommendation method, a device, an electronic device, a computer program product, and a computer-readable storage medium. Background Art

[0002] Information recommendation is a way of delivering information based on the interests of the recommended object, with the aim of increasing the click-through rate and conversion rate of the recommended information. Since popular recommended information gets more exposure and clicks, traditional sorting-based information recommendation algorithms tend to recommend these popular recommended information. This results in long-tail recommended information often being ignored, resulting in low exposure and click-through rates. Secondly, traditional recommendation algorithms usually tend to recommend recommended information that is similar to the recommended object's past behavior, which leads to the homogeneity of recommendation results. In addition, traditional recommended information classification is often simply divided according to the topic or tag of the recommended information, but there are more fine-grained features and related information in the long-tail recommended information. In the recommendation information system, the long-tail distribution phenomenon of recommended information categories is common, which means that only a few categories of recommended information get most of the exposure consumption, while most of the recommended information is difficult to get enough exposure, and thus difficult to be noticed and clicked by the recommended object, which is very disadvantageous to the recommended information owner. Summary of the invention

[0003] The embodiments of the present application provide a model training method, an information recommendation method, an apparatus, an electronic device, a computer program product, and a computer-readable storage medium, which can obtain a recommendation network model with more accurate recommendation effects and a wider recommendation range.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present application provides a model training method, the method comprising:

[0006] Acquire a recommendation model to be trained and acquire an initial training data set, wherein the recommendation model to be trained includes a first sub-model and a second sub-model;

[0007] Randomly sampling the initial training data set to obtain a first training data set, and inversely sampling the initial training data set based on an inverse sampling probability to obtain a second training data set, wherein the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training data set;

[0008] Using the first sub-model to predict each first training sample in the first training data set to obtain each first prediction value, and using the second sub-model to predict each second training sample in the second training data set to obtain each second prediction value;

[0009] Determine a combined loss value by using the first sample labels of the respective first training samples in the first training dataset, the respective first predicted values, and the second sample labels of the respective second training samples in the second training dataset;

[0010] Train the recommendation model to be trained by using the combined loss value to obtain a trained recommendation model.

[0011] An embodiment of the present application provides an information recommendation method, and the method includes:

[0012] Obtain the recommendation information to be released, the candidate object information of multiple candidate recommendation objects corresponding to the recommendation information to be released, and a trained recommendation model;

[0013] Determine the to-be-released recommendation features of the recommendation information to be released and the candidate object features of each candidate object information;

[0014] Use the feature processing module in the trained recommendation model to process the to-be-released recommendation features and the candidate object features of each candidate object information to obtain a second to-be-predicted feature vector;

[0015] Use the first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector to obtain the target object of the recommendation information to be released;

[0016] Send the recommendation information to be released to the terminal corresponding to the target object.

[0017] An embodiment of the present application provides a model training device, including:

[0018] A first acquisition module, configured to acquire a recommendation model to be trained and acquire an initial training dataset, where the recommendation model to be trained includes a first sub-model and a second sub-model;

[0019] A sampling module, configured to randomly sample the initial training dataset to obtain a first training dataset, and perform inverse sampling on the initial training dataset based on an inverse sampling probability, where the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training dataset;

[0020] A first prediction module, configured to use the first sub-model to predict each first training sample in the first training dataset to obtain respective first predicted values, and use the second sub-model to predict each second training sample in the second training dataset to obtain respective second predicted values;

[0021] A first determination module, configured to determine a joint loss value by using the first sample labels of the respective first training samples in the first training dataset, the respective first predicted values, and the second sample labels of the respective second training samples in the second training dataset;

[0022] A training module, configured to train the recommendation model to be trained by using the joint loss value, and obtain a trained recommendation model.

[0023] An information recommendation device provided by an embodiment of the present application includes:

[0024] A second acquisition module, configured to acquire to-be-published recommendation information, candidate object information of a plurality of candidate recommendation objects corresponding to the to-be-published recommendation information, and a trained recommendation model;

[0025] A second determination module, configured to determine to-be-published recommendation features of the to-be-published recommendation information and candidate object features of the respective candidate object information;

[0026] A processing module, configured to use a feature processing module in the trained recommendation model to process the to-be-published recommendation features and the candidate object features of the respective candidate object information, and obtain a second to-be-predicted feature vector;

[0027] A second prediction module, configured to use a first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector, and obtain a target object of the to-be-published recommendation information;

[0028] A sending module, configured to send the to-be-published recommendation information to a terminal corresponding to the target object.

[0029] An electronic device provided by an embodiment of the present application, the electronic device includes:

[0030] A memory, configured to store computer-executable instructions;

[0031] A processor, configured to implement the model training method provided by an embodiment of the present application, or implement the information recommendation method provided by an embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0032] A computer-readable storage medium provided by an embodiment of the present application, storing a computer program or computer-executable instructions, configured to implement the model training method provided by an embodiment of the present application, or implement the information recommendation method provided by an embodiment of the present application when being executed by a processor.

[0033] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the model training method provided by the embodiment of the present application is implemented, or the information recommendation method provided by the embodiment of the present application is implemented.

[0034] The embodiment of the present application has the following beneficial effects:

[0035] In the embodiment of the present application, an initial training data set is obtained and a recommendation model to be trained is prepared, including preparing a first sub-model and a second sub-model. Then, the initial training data set is randomly sampled to obtain a first training data set, and then the initial training data set is inversely sampled according to the inverse sampling probability to obtain a second training data set. Since the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training data set, therefore, using the second training data set can ensure that when training the recommendation model, the positive and negative sample ratios are more balanced, improving the training effect of the model. Then, each first training sample in the first training data set is predicted by using the first sub-model to obtain each first prediction value, and each second training sample in the second training data set is predicted by using the second sub-model to obtain each second prediction value. Next, a joint loss value is determined by using the first sample label of each first training sample in the first training data set, each first prediction value, and the second sample label of each second training sample in the second training data set. In this way, the trained recommendation model obtained by training the recommendation model to be trained by using the joint loss value, the joint loss value can comprehensively consider the loss values obtained by different sampling methods, so as to more comprehensively evaluate the performance of the recommendation model, improve the prediction accuracy of the recommendation model, and at the same time reduce the overfitting risk of the recommendation model and improve the generalization ability of the recommendation model, so that the recommendation model has a better recommendation effect in the real scenario. Description of the Drawings

[0036] Figure 1 is a schematic structural diagram of a model training system 100 provided by an embodiment of the present application;

[0037] Figure 2A is a schematic structural diagram of a server 400 provided by an embodiment of the present application;

[0038] Figure 2B is another schematic structural diagram of a server 400 provided by an embodiment of the present application;

[0039] Figure 3A is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0040] Figure 3B is a schematic flowchart of obtaining an initial training data set provided by an embodiment of the present application;

[0041] Figure 3C It is a schematic flowchart for determining the inverse sampling probability corresponding to each recommended information identifier provided by an embodiment of this application;

[0042] Figure 3D It is a schematic flowchart for determining the first predicted value provided by an embodiment of this application;

[0043] Figure 3E It is a schematic flowchart for determining the combined loss value provided by an embodiment of this application;

[0044] Figure 3F It is a schematic flowchart for determining the second loss value corresponding to the second sub-model provided by an embodiment of this application;

[0045] Figure 3G It is a schematic flowchart for determining the first weight provided by an embodiment of this application;

[0046] Figure 4 It is a schematic flowchart of the information recommendation method provided by an embodiment of this application;

[0047] Figure 5 It is a schematic flowchart of the training process of the recommendation model provided by an embodiment of this application;

[0048] Figure 6 It is a schematic diagram of the structure of the wide&&deep network;

[0049] Figure 7 It is a schematic diagram of the structure of the recommendation model. Specific embodiments

[0050] In order to make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0051] In the following descriptions, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0052] In the following descriptions, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here.

[0053] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0054] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0055] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0056] 1) Wide&Deep Learning for Recommender Systems (Wide&Deep network): A neural network model that combines wide and deep features, mainly consisting of two parts, namely the wide part and the deep part. Among them, the wide part uses a linear model to process categorical features, while the deep part uses a deep neural network to learn low-dimensional dense continuous feature representations. In this way, both the breadth of features can be retained and the deep relationships between features can be mined, thereby improving the recall rate and accuracy of the model.

[0057] 2) Distillation Loss: A commonly used method in model compression, which "distills" the knowledge of a large and complex model into a small and simple model, thereby reducing the model parameters and computational amount. In this process, distillation means transferring the information in the large model to the small model, so that the small model can learn similar feature representations to the large model. Specifically, the distillation loss value is measured by comparing the differences between the outputs (usually probability distributions) of the large model and the small model. By minimizing the distillation loss value, the small model can learn the key features in the large model, thereby achieving the purpose of model compression.

[0058] The embodiments of the present application provide a model training method, an information recommendation method, a device, an electronic device, a computer program product, and a computer-readable storage medium, which can obtain a recommendation network model with more accurate recommendation effects and a wider recommendation range.

[0059] The following describes the exemplary applications of the electronic device provided in the embodiments of the present application. The electronic device provided in the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smart phones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can also be implemented as a server. Below, the exemplary applications when the electronic device is implemented as a server will be described.

[0060] See Figure 1 , Figure 1 is a schematic architecture diagram of the information recommendation system 100 provided in the embodiments of the present application. To support an exemplary application, the terminal 200 is connected to the server 400 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. The information recommendation system 100 may also include a database 500 for storing the trained recommendation model. The database 500 can be independent of the server 400 or located inside the server 400. In Figure 1 an example where the database 500 is independent of the server 400 is described.

[0061] When pushing the recommended information to be released to the target object, first, the terminal 200 will send the recommended information to be released to the server 400. After receiving these messages, the server 400 obtains the recommended information to be released and the candidate object information of multiple candidate recommended objects, and at the same time loads the trained recommendation model. Subsequently, the server 400 starts to process the recommended information to be released and the candidate object information of multiple candidate recommended objects. First, the server 400 will determine the characteristics of the recommended information to be released and the characteristics of each candidate object. For this purpose, the server 400 uses the feature processing module in the recommendation model to process the characteristics of the recommended information to be released and the candidate objects, and obtains a second set of feature vectors to be predicted. Next, the server 400 uses the first sub-model in the recommendation model to predict these second sets of feature vectors, thereby obtaining the target object of the recommended information to be released. Finally, the server 400 sends the recommended information to be released back to the terminal device corresponding to the target object.

[0062] In the embodiment of the present application, when the server 400 trains the recommendation model, first, the first sub-model in the recommendation model to be trained is used to predict each first training sample in the first training dataset to obtain each first prediction value. Then, the second sub-model in the recommendation model to be trained is used to predict each second training sample in the second training dataset to obtain each second prediction value. Among them, the recommendation model to be trained includes a first sub-model and a second sub-model. The first training dataset is obtained by randomly sampling the initial training dataset, and the second training dataset is obtained by inverse sampling the initial training dataset based on the inverse sampling probability negatively correlated with the number of training samples corresponding to the recommendation information identifiers in the initial training dataset. Then, the joint loss value is determined by using the first sample labels of each first training sample in the first training dataset, each first prediction value, and the second sample labels of each second training sample in the second training dataset. Finally, the recommendation model to be trained is trained by using the joint loss value, and thus the trained recommendation model is obtained. By using the second training dataset obtained by inverse sampling, the training of the recommendation model can avoid over-relying on popular recommendation information. In addition, the loss value obtained by random sampling is used to evaluate the relevance between the popular recommendation information of the recommendation result and the interests of the target object, that is, to judge whether the popular items in the recommendation information conform to the user's interests. The loss value obtained by inverse sampling is used to evaluate the relevance between the unpopular recommendation information of the recommendation result and the user's interests, that is, to check whether the unpopular items in the recommendation information conform to the user's interests. The joint loss value considering these two aspects can make the recommendation model pay more attention to the personalized needs of the target object, thereby improving the accuracy level of the recommendation.

[0063] In some embodiments, the server 400 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 200 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiment of the present application.

[0064] See Figure 2A , Figure 2A is a schematic structural diagram of the server 400 provided by the embodiment of the present application. Figure 2AThe server 400 shown includes: at least one processor 410-1, a memory 450-1, at least one network interface 420-1, and a user interface 430-1. Each component in the server 400 is coupled together through a bus system 440-1. It can be understood that the bus system 440-1 is used to implement the connection and communication between these components. In addition to including a data bus, the bus system 440-1 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2A all kinds of buses are labeled as the bus system 440-1.

[0065] The processor 410-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0066] The user interface 430-1 includes one or more output devices 431-1 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430-1 also includes one or more input devices 432-1, including user interface components that assist user input, such as keyboards, mice, microphones, touch screen displays, cameras, other input buttons, and controls.

[0067] The memory 450-1 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 450-1 optionally includes one or more storage devices that are physically located far from the processor 410-1.

[0068] The memory 450-1 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (Random Access Memory, RAM). The memory 450-1 described in the embodiments of the present application is intended to include any suitable type of memory.

[0069] In some embodiments, the memory 450-1 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.

[0070] The operating system 451-1 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0071] The network communication module 452-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420-1. Exemplary network interfaces 420-1 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;

[0072] The presentation module 453-1 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431-1 associated with the user interface 430-1 (such as a display screen, a speaker, etc.);

[0073] The input processing module 454-1 is used to detect one or more user inputs or interactions from one of one or more input devices 432-1 and translate the detected inputs or interactions.

[0074] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2A Shown is a model training device 455 stored in the memory 450-1, which can be software in the form of programs and plugins, etc., including the following software modules: a first acquisition module 4551, a sampling module 4552, a first prediction module 4553, a first determination module 4554, and a training module 4555. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0075] See Figure 2B , Figure 2B is another structural schematic diagram of the server 400 provided by the embodiments of the present application. Figure 2B The shown server 400 includes: at least one processor 410-2, a memory 450-2, at least one network interface 420-2, and a user interface 430-2. Each component in the server 400 is coupled together through a bus system 440-2. It can be understood that the bus system 440-2 is used to realize the connection and communication between these components. The bus system 440-2 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2B all kinds of buses are labeled as the bus system 440-2.

[0076] The processor 410-2 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0077] The user interface 430-2 includes one or more output devices 431-2 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430-2 also includes one or more input devices 432-2, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.

[0078] The memory 450-2 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450-2 optionally includes one or more storage devices that are physically located away from the processor 410-2.

[0079] The memory 450-2 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450-2 described in the embodiments of the present application is intended to include any suitable type of memory.

[0080] In some embodiments, the memory 450-2 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0081] The operating system 451-2 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0082] The network communication module 452-2 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420-2. Exemplary network interfaces 420-2 include: Bluetooth, Wi-Fi (Wireless Compatibility Certification), and Universal Serial Bus (USB), etc.;

[0083] A presentation module 453-2 for enabling the presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 431-2 associated with the user interface 430-2 (e.g., a display screen, a speaker, etc.);

[0084] An input processing module 454-2 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432-2.

[0085] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2B Shown is an information recommendation device 456 stored in the memory 450-2, which may be software in the form of a program and a plug-in, etc., including the following software modules: a second acquisition module 4561, a second determination module 4562, a processing module 4563, a second prediction module 4564, and a sending module 4565. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0086] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the model training method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0087] To better understand the model training method and information recommendation method provided by the embodiments of the present application, first, artificial intelligence, each branch of artificial intelligence, and the application fields involved in the model training method and information recommendation method provided by the embodiments of the present application will be described.

[0088] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0089] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The solution provided in the embodiments of this application mainly relates to machine learning technology, and the following is an explanation of this technology.

[0090] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0091] Adaptive computing: It refers to automatically adjusting the computational amount and precision of the model according to different input data to achieve the purpose of improving the computational efficiency of the model while maintaining the model's precision. Adaptive computing can flexibly adjust the computational amount and precision of the model on different input data, thereby better balancing the computational efficiency and precision of the model.

[0092] Model parallel computing: It refers to distributing the computational tasks of the model to multiple computing devices (such as CPUs, GPUs, TPUs, etc.) for simultaneous computing, thereby accelerating the training and inference of the model. Model parallel computing can effectively utilize computing resources and improve the computational efficiency and training speed of the model.

[0093] A pre-training model (PTM), also known as a foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a vast amount of unlabeled data, and by leveraging the function approximation ability of the large-parameter DNN, the PTM extracts common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, pre-training models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. Pre-training models are important tools for outputting artificial intelligence-generated content and can also serve as a general interface connecting multiple specific task models.

[0094] Distributed training means splitting and sharing the workload of training a model among multiple microprocessors. Large models have a large number of parameters and large amounts of training data, exceeding the capacity of a single machine. Therefore, distributed parallelism is required to speed up the training. Parallel mechanisms include data parallelism (DP), model parallelism (MP), pipeline parallelism (PP), and hybrid parallelism (HP). The architecture design includes constructs based on parameter servers, reduction, MPI, etc.

[0095] Model compression and quantization: refers to using compression and quantization techniques to help reduce the model size and accelerate model inference, thereby reducing the costs of model storage and computing. Model compression usually includes pruning, low-rank decomposition, knowledge distillation, etc. Model quantization refers to converting the floating-point parameters in the model into fixed-point or integer parameters to reduce the model size and accelerate model inference.

[0096] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0097] Next, the model training method provided by the embodiments of the present application will be described. As mentioned above, the electronic device implementing the model training method of the embodiments of the present application can be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter. Refer to Figure 3A , Figure 3A which is a schematic flowchart of the model training method provided by the embodiments of the present application, and will be described in combination with the steps shown in Figure 3A . Figure 3A The subject of the step is the server

[0098] In step 101, a recommendation model to be trained is obtained, and an initial training data set is obtained.

[0099] In some embodiments, the recommendation model to be trained can be a neural network model, such as a deep learning model, a recurrent neural network model, etc. The recommendation model to be trained includes a first sub-model and a second sub-model. Among them, the first sub-model can be understood as a main model, and the second sub-model can be understood as an auxiliary model. The first sub-model and the second sub-model can be neural network models or models dedicated to recommendation systems, such as Wide&DeepModel, Deep Factorization Machine Model (DeepFM model), Neural Collaborative Filtering Model (NCF model), etc. The first sub-model is used to predict the first training samples randomly sampled from the initial training data set, and the second sub-model is used to predict the second training samples obtained by inverse sampling from the initial training data set. In order to obtain a trained recommendation model, it is necessary to train the first sub-model and the second sub-model simultaneously. After obtaining the trained recommendation model, the first sub-model is used for prediction.

[0100] In some embodiments, refer to Figure 3B , Figure 3B which is a schematic flowchart of obtaining the initial training data set provided by the embodiments of the present application, Figure 3A The "obtaining the initial training data set" in step 101 shown in can be implemented through steps 1011 to 1016 as shown in Figure 3B . The following will be specifically described in combination with Figure 3B .

[0101] Step 1011: Obtain multiple pieces of recommendation information and the object information of the recommended objects corresponding to each piece of recommendation information.

[0102] In some embodiments, by interacting with a recommendation information system, multiple pieces of recommendation information are obtained from the recommendation information system. These pieces of recommendation information may include content such as the title, description, and picture of the recommendation information. At the same time, the object information of the recommended object corresponding to each piece of recommendation information is also obtained from the recommendation information system, such as the object identifier, gender, age, etc.

[0103] Step 1012: Extract features from each piece of recommendation information and the object information corresponding to each piece of recommendation information to obtain respective recommendation feature data and object feature data corresponding to the respective recommendation feature data.

[0104] In some embodiments, by extracting features from each piece of recommendation information and its corresponding object information, respective recommendation feature data and corresponding object feature data can be obtained. First, extract recommendation information features, including recommendation information identifier, recommendation information category, the manufacturer corresponding to the recommendation information, and information such as the price of the product and product attributes corresponding to the recommendation information. These features can describe different attributes and association relationships of the recommendation information. At the same time, object features also need to be extracted as part of the object information. Among them, object attribute features are constructed based on basic information such as the age and gender of the object, and these features can reflect the personal background and interest preferences of the object. In addition, object behavior features need to be extracted, such as article reading records, recommendation information click records, and recommendation information conversion records within the past 14 days and other behavior information. These features can reveal the activity track and consumption intention of the object.

[0105] Step 1013: Determine each piece of recommendation feature data and the object feature data corresponding to each piece of recommendation feature data, and determine the cross-feature data corresponding to each piece of recommendation feature data.

[0106] In some embodiments, the recommendation feature data may include multiple recommendation feature fields. For example, the recommendation information identifier and the recommendation information category are each a recommendation feature field. Similarly, the object feature data includes multiple object feature fields. Exemplarily, it may include object feature fields such as the object identifier, gender, age, etc. In the process of determining the cross-feature data corresponding to the recommendation feature data, one or more recommendation feature fields in the recommendation feature data and one or more object feature fields in the object feature data can be combined to obtain the cross-feature data. For example, concatenate the gender field of the object and the recommendation information identifier field to obtain cross-feature data, indicating the gender feature of the object for the recommendation information. Such cross-feature data can provide more dimensions and information for more accurately describing the relationship between the object and the recommendation information. Other types of cross-feature data can also be constructed, such as the cross-feature data of the object's geographical location and the type of recommendation information, the cross-feature data of the object's age and price, etc.

[0107] Step 1014: Construct each initial training sample with each recommended feature data, the object feature data corresponding to each recommended feature data, and the cross-feature data corresponding to each recommended feature data.

[0108] In some embodiments, integrating each recommended feature data, the object feature data corresponding to each recommended feature data, and the cross-feature data corresponding to each recommended feature data obtained in Step 1012 and Step 1013 can construct multiple initial training samples, providing a basis for subsequent model training. The initial training samples contain rich information, which can help the recommendation system more accurately predict the preferences or behaviors of the object.

[0109] Step 1015: Determine the sample labels of each initial training sample.

[0110] In some embodiments, infer based on the click behavior of the recommended information. For example, when the object browses a web page, the object may see multiple recommended information, some of which are clicked, while other recommended information is not clicked. Determine whether the recommended information is clicked by determining the click behavior of the object. For example, when the object clicks on a certain recommended information, take the sample label of the initial training sample corresponding to the recommended information as the positive sample label and give it a label value of 1, which means that the recommended information is the recommended information that the object is interested in and clicks. On the contrary, when the object does not click on a certain recommended information but the recommended information is still exposed, take the sample label of the initial training sample of the recommended information as the negative sample label and give it a label value of 0, which means that the recommended information is the recommended information that the object does not choose to click. For example, when the object clicks on recommended information A, consider the sample label of the initial training sample of recommended information A as the positive sample label and assign a label value of 1. If the object does not click on recommended information B but recommended information B is still exposed, then consider the sample label of the initial training sample of recommended information B as the negative sample label and assign a label value of 0.

[0111] Step 1016: Construct an initial training data set based on each initial training sample and the sample labels of each initial training sample.

[0112] In some embodiments, pair each initial training sample with its corresponding initial training sample label to construct an initial training data set.

[0113] In Steps 1011 to 1016, by interacting with the recommendation information system and parsing the object behavior data, multiple recommended information and the object information of the recommended objects corresponding to each recommended information can be accurately obtained. When training the recommendation model, using the cross-feature data corresponding to the recommended feature data can improve the richness of the training data, thereby improving the performance of the trained recommendation model.

[0114] In some embodiments, after obtaining the initial training dataset, it is also necessary to determine the inverse sampling probability corresponding to each recommendation information identifier. Refer to Figure 3C , Figure 3C which is a schematic flowchart for determining the inverse sampling probability corresponding to each recommendation information identifier provided by an embodiment of this application, and can be implemented through steps 001 to 005 as shown in Figure 3C . The following will specifically describe with reference to Figure 3C .

[0115] Step 001: Obtain the recommendation information identifier corresponding to each initial training sample.

[0116] In some embodiments, by parsing and extracting the recommendation information features in the initial training dataset, the recommendation information identifier corresponding to each initial training sample is obtained therefrom.

[0117] Step 002: Determine the number of initial training samples corresponding to each recommendation information identifier, and determine the maximum number of training samples from the number of initial training samples corresponding to each recommendation information identifier.

[0118] In some embodiments, after obtaining the recommendation information identifier corresponding to the initial training sample, the number of initial training samples corresponding to each recommendation information identifier is counted in the initial training dataset, and the number of initial training samples corresponding to each recommendation information is compared to obtain the maximum number of training samples in the initial training samples corresponding to each recommendation information. For example, the number of initial training samples corresponding to recommendation information A is 50, the number of initial training samples corresponding to recommendation information B is 100, the number of initial training samples corresponding to recommendation information C is 20, and the number of samples corresponding to recommendation information D is 50. By comparing the number of initial training samples of each recommendation information, the maximum number of training samples is the number of initial training samples of recommendation information B, which is 100.

[0119] Step 003: Determine the ratio of the maximum number of training samples to the number of initial training samples corresponding to each recommendation information identifier as the weighted weight corresponding to each recommendation information identifier.

[0120] In some embodiments, the weighted weight corresponding to each recommendation information identifier is equal to the maximum number of training samples divided by the number of initial training samples corresponding to each recommendation information identifier. The formula for the weighted weight corresponding to each recommendation information identifier is as follows:

[0121]

[0122] where N i is the number of initial training samples corresponding to each recommendation information identifier, N max is the maximum number of training samples, and N max = max(N i)。When the number of initial training samples corresponding to each recommended information identifier is larger, the weighted weight corresponding to each recommended information identifier is smaller. Correspondingly, when facing popular recommended information, the weighted weight corresponding to the popular recommended information identifier is smaller.

[0123] Step 004: Determine the sum of the weighted weights corresponding to each recommended information identifier as the total weight.

[0124] In some embodiments, the weighted weights corresponding to each recommended information identifier are accumulated to obtain the total weight corresponding to all recommended information identifiers. Exemplarily, the number of initial training samples corresponding to recommended information identifier A is 50, the number of initial training samples corresponding to recommended information identifier B is 100, the number of initial training samples corresponding to recommended information identifier C is 20, and the number of initial training samples corresponding to recommended information identifier D is 50. Then Nmax is 100, w A is 2, w B is 1, w C is 5, w D is 2. Therefore, the total weight is 2 + 1 + 5 + 2 which equals 10.

[0125] Step 005: Determine the ratio of the weighted weight corresponding to each recommended information identifier to the total weight as the inverse sampling probability corresponding to each recommended information identifier.

[0126] In some embodiments, the inverse sampling probability corresponding to each recommended information identifier is equal to the weighted weight corresponding to each recommended information identifier divided by the total weight. The inverse sampling probability corresponding to each recommended information identifier is as shown in Formula 2:

[0127]

[0128] where w i is the weighted weight corresponding to each recommended information identifier, is the total weight, and P i is the inverse sampling probability corresponding to each recommended information identifier.

[0129] Continuing with the above example, w A is 2, w B is 1, w C is 5, w D is 2. Then according to Formula (2), it can be obtained that P A is 0.2, P B is 0.1, P C is 0.5, and P D is 0.2. Thus, it can be seen that the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommended information identifier in the initial training dataset.

[0130] As described in the above steps 1017 to 10111, in the embodiment of the present application, by performing inverse sampling on the initial training data set based on the inverse sampling probability, a second training data set can be obtained. This process ensures that the samples in the second training data set have a more balanced distribution and can better represent the entire data set. By using the inverse sampling probability, the samples can be weighted according to the importance of each sample in the initial training data set, so that the samples with fewer occurrences receive more attention and emphasis in model training, thereby better capturing the features of these rare samples or categories, improving the prediction accuracy for them, enabling the recommendation system to better handle the cold start situation, and providing more accurate recommendation results for the recommended objects.

[0131] Continue to refer to Figure 3A , and continue to describe in connection with step 101.

[0132] In step 102, the initial training data set is randomly sampled to obtain a first training data set, and the initial training data set is inversely sampled based on the inverse sampling probability to obtain a second training data set.

[0133] In some embodiments, when randomly sampling the initial training data set, a part of the training samples are randomly selected from the entire initial training data set according to a preset sampling ratio as the first training samples. This process is completely random, and the probability of each training sample being selected is equal, thus ensuring the randomness and representativeness of the sampling result.

[0134] After obtaining the first training data set, the initial training data set is further processed based on the inverse sampling probability to obtain second training samples, and the second training sample set is used as the second training data set. The inverse sampling probability means that the probability of each second training sample being selected is inversely proportional to its frequency of occurrence in the initial training data set. That is, the probability of the initial training sample with a higher frequency of occurrence being selected is lower, while the probability of the initial training sample with a lower frequency of occurrence being selected is higher. Continuing with the above example, for inverse sampling, since P A is 0.2, P B is 0.1, P C is 0.5, P D is 0.2, then the number of second training samples of recommended information A in the second training data set is 10, the number of second training samples of recommended information B is 10, the number of second training samples of recommended information C is 10, and the number of second training samples of recommended information D is 10.

[0135] In step 103, each first training sample in the first training data set is predicted using the first sub-model to obtain respective first prediction values, and each second training sample in the second training data set is predicted using the second sub-model to obtain respective second prediction values.

[0136] In some embodiments, each first training sample in the first training dataset and each second training sample in the second training dataset are respectively input into the first sub-model and the second sub-model. The first sub-model and the second sub-model are used to respectively predict the first training sample and the second training sample, and the first prediction value and the second prediction value are respectively determined. Wherein, the first prediction value and the second prediction value can represent the probability of an object clicking on the recommended information, and are usually represented by a continuous score value, such as a probability value ranging from 0 to 1, which is not limited here.

[0137] In some embodiments, referring to Figure 3D , Figure 3D is a schematic flowchart of the process for determining the first prediction value provided by the embodiments of the present application, which can be implemented through steps 1031 to 1034 as shown in Figure 3D . The following will be specifically described in conjunction with Figure 3D .

[0138] Step 1031: For each first training sample, determine the first recommendation feature vector corresponding to the first recommendation feature data, the first object feature vector corresponding to the first object feature data, and the first cross-feature vector corresponding to the first cross-feature data in the first training sample.

[0139] In some embodiments, since the first training sample includes recommendation feature data, object feature data, and cross-feature data, and the recommendation feature data, object feature data, and cross-feature data respectively include at least one feature field. If a feature field is a non-continuous feature, for example, gender field, recommendation information identification field, object identification field, etc. are all non-continuous features, then a non-continuous feature can be mapped to a 64-dimensional feature vector through an embedding layer; if a feature field is a continuous feature field, such as commodity price field, sales volume field, etc. are continuous features, then the feature field can be discretized by using a bucketing method to obtain a discrete feature, and then the discrete feature is mapped to a 64-dimensional feature vector. In this way, after obtaining the feature vectors corresponding to each recommendation feature field in the first recommendation feature data, the feature vectors corresponding to each recommendation feature field are concatenated to obtain the first recommendation feature vector. Similarly, the feature vectors corresponding to each object feature field in the first object feature data are concatenated to obtain the first object feature vector, and the feature vectors corresponding to the recommendation feature field and the object feature field in the first cross-feature data are concatenated to obtain the first cross-feature vector.

[0140] Step 1032: Perform a concatenation process on the first recommendation feature vector, the first object feature vector, and the first cross-feature vector to obtain the first input feature vector.

[0141] In some embodiments, the obtained first recommended feature vector and the first object feature vector are concatenated, and the first cross - feature vector is concatenated to obtain the first input feature vector. For example, if the recommended information identifier of the recommended feature in recommended information A is recommended information 1, and the object identifier of the object feature in object B is object 2, then the identifier information of the cross - feature of recommended information A and object B is <recommended information 1, object 2>. If the first recommended feature vector corresponding to recommended information 1 is C, the first object feature vector corresponding to object B is D, and the first cross - feature vector corresponding to <recommended information 1, object 2> is E, then the corresponding first input feature vector of the three is [C, D, E].

[0142] Step 1033: Integrate the features of the first input feature vector to obtain the first feature vector to be predicted.

[0143] In some embodiments, the wide&&deep network can be used to integrate the features of the first input feature vector. First, load the pre - trained wide&&deep model. The wide&&deep model consists of two parts: a wide model and a deep model. These two models are used to process different types of features respectively. Next, pass the first input feature vector to the wide model. The wide model mainly focuses on features with high frequency and important influence. It encodes these features in the first input feature vector and performs feature integration through a simple linear model. Then, pass the first input feature vector to the deep model. The deep model mainly focuses on those very sparse but potentially complex interaction features. The deep model extracts and transforms features through multiple neural network layers. Each neural network layer learns different levels of feature representations, from low - level raw features to higher - level abstract features. Finally, integrate the features obtained from the wide model and the deep model to obtain the first feature vector to be predicted. This first feature vector to be predicted combines the important features captured by the wide model and the high - level abstract features learned by the deep model, aiming to make full use of the advantages of both models, considering both the correlation between specific features and the potential complex interaction relationships.

[0144] In addition to using the wide&&deep network, other networks can also be used for feature integration. For example, a convolutional neural network (CNN) can be used to capture the spatial correlation in the input features, especially suitable for image data; a recurrent neural network (RNN) can be used to process sequence data, such as natural language text or time - series data; in addition, an attention mechanism can be used to enhance the attention degree to different input features to improve the expression ability of the model.

[0145] Step 1034: Use the first sub-model to predict the first feature vector to be predicted, and obtain a first prediction value.

[0146] In some embodiments, the first feature vector to be predicted is input into the first sub-model, and the first sub-model is composed of a forward network. The forward network is an algorithm based on the neural network structure. It can transmit the input data to the hidden layer and the output layer, and finally generate a prediction value. In the forward network, the first feature vector to be predicted will propagate along the path of the neural network. Each layer will perform specific operations, such as weighted summation, non-linear activation functions, etc. These operations will gradually extract the recommended information features and object characteristics in the input data and transmit them to the next layer. When the first feature vector to be predicted passes through all the hidden layers, it will finally reach the output layer. The output layer will perform the final calculation based on the input data according to the design of the network structure. The final output result is the first prediction value, which may be a single number, probability, or a probability distribution.

[0147] In steps 1031 to 1034, for each first training sample, a first prediction feature vector is constructed based on the first recommendation feature vector corresponding to the first recommendation feature data, the first object feature vector corresponding to the first object feature data, and the first cross-feature vector corresponding to the first cross-feature data included therein. That is, the first prediction feature vector includes the recommendation feature vector, the object feature vector, and the cross-feature vector. Thus, using the first sub-model to predict the first feature vector to be predicted can obtain a more accurate first prediction value, thereby providing an accurate data basis for determining the loss value for model training subsequently, and thus improving the model training efficiency.

[0148] It should be noted that the implementation process of using the second sub-model to predict each second training sample in the second training dataset to obtain each second prediction value is similar to steps 1031 to 1034, and the implementation process of steps 1031 to 1034 can be referred to when implementing.

[0149] Continue to refer to Figure 3A below, and continue to describe step 103.

[0150] In step 104, a joint loss value is determined using the first sample labels of each first training sample in the first training dataset, each first prediction value, and the second sample labels of each second training sample in the second training dataset.

[0151] In some embodiments, refer to Figure 3E , Figure 3E is a schematic flowchart of the process for determining the joint loss value provided by the embodiments of the present application. Figure 3AThe shown step 104 can be implemented by steps 1041 to 1044 as shown in Figure 3E and will be specifically described below in conjunction with Figure 3E .

[0152] Step 1041: Determine the first loss value corresponding to the first sub-model by using the first sample labels and the respective first predicted values of the respective first training samples in the first training dataset.

[0153] In some embodiments, the first sample labels and the respective first predicted values of the respective first training samples can be used as input values of a preset loss function, so as to determine the first loss value corresponding to the first sub-model. Exemplarily, the preset loss function can be a cross-entropy loss function, or can also be other types of loss functions. In the embodiments of the present application, taking the preset loss function as the cross-entropy function as an example for illustration, the calculation formula of the cross-entropy loss function for determining the first loss value corresponding to the first sub-model is as follows:

[0154]

[0155] where P(x i ) is the first sample label, and Q(x i ) is the first predicted value. By calculating the cross-entropy loss value of each first training sample, the performance of the first sub-model on the entire first training dataset can be evaluated. The smaller the loss value, the higher the degree of coincidence between the prediction result of the model and the true label, and the lower the error between the predicted value and the sample label. On the contrary, it indicates a poor prediction effect.

[0156] Step 1042: Determine the second loss value corresponding to the second sub-model by using the respective first predicted values of the respective first training samples in the first training dataset and the second sample labels of the respective second training samples in the second training dataset.

[0157] In some embodiments, referring to Figure 3F , Figure 3F is a schematic flowchart of the process for determining the second loss value corresponding to the second sub-model provided by the embodiments of the present application, Figure 3E the shown step 1042 can be implemented by steps 421 to 425 as shown in Figure 3F and will be specifically described below in conjunction with Figure 3F .

[0158] Step 421: Determine the main loss value corresponding to the second sub-model by using the second sample labels and the respective second predicted values of the respective second training samples in the second training dataset.

[0159] In some embodiments, the method for determining the main loss value corresponding to the second sub-model is the same as the method for determining the first loss value in step 1041. The calculation formula of the cross-entropy loss function for determining the main loss value corresponding to the second sub-model is as follows:

[0160]

[0161] where P(y j ) is the second sample label, and Q(y j ) is the second predicted value.

[0162] Step 422: Determine whether the first loss value is less than the main loss value.

[0163] In some embodiments, when the first loss value is less than the main loss value, step 423 is entered; when the first loss value is greater than or equal to the main loss value, step 425 is entered.

[0164] Step 423: Determine the distillation loss value using the first predicted value and the second predicted value.

[0165] In some embodiments, when the first loss value is less than the main loss value, the cross-entropy loss is calculated using each first predicted value in the first training dataset and each second predicted value in the second training dataset to determine the distillation loss value. The calculation formula of the cross-entropy loss function for determining the distillation loss value is as follows:

[0166]

[0167] where P(x i ) is the first predicted value, and Q(y j ) is the second predicted value.

[0168] Step 424: Perform weighted summation on the main loss value and the distillation loss value to obtain the second loss value corresponding to the second sub-model.

[0169] In some embodiments, after obtaining the main loss value and the distillation loss value, the distillation loss value multiplied by the main loss value and the hyperparameter β is subjected to weighted summation to obtain the second loss value corresponding to the second sub-model. The calculation formula of the second loss value corresponding to the second sub-model is as follows:

[0170]

[0171] where Loss2 is the second loss value, is the main loss value, is the distillation loss value, β is the hyperparameter, and β is obtained through multiple training and debugging.

[0172] Step 425: Determine the main loss value as the second loss value corresponding to the second sub-model.

[0173] In some embodiments, when the first loss value is greater than or equal to the main loss value, it indicates that knowledge distillation is not required, and the main loss value is directly determined as the second loss value corresponding to the second sub-model.

[0174] Step 1043: Determine the first weight corresponding to the first loss value and the second weight corresponding to the second loss.

[0175] In some embodiments, referring to Figure 3G , Figure 3G is a schematic flowchart of the process for determining the first weight provided in the embodiments of the present application. Figure 3E The shown step 1043 can be implemented through steps 431 to 433 as shown in Figure 3G . The following will specifically describe with reference to Figure 3G .

[0176] Step 431: Obtain the current training round number and the total number of training rounds.

[0177] In some embodiments, when starting model training, a fixed total number of training rounds is set, indicating how many rounds the entire training process will proceed. This value is determined according to the nature of the specific task and dataset and can be obtained through manual setting. Before starting training, a counter is initialized to record which round of training is currently being carried out. After starting training, the value of the counter will be gradually increased to represent the number of training rounds that have been completed. Then, the model parameters are gradually updated through iterative training, and each completed iteration represents one round of training. Before each round of training starts, the current training round number will be obtained. The total number of training rounds is set before the start of training and can be used as a fixed parameter.

[0178] Step 432: Determine the ratio of the current training round number to the total number of training rounds as the second weight.

[0179] In some embodiments, the ratio obtained by dividing the current training round number by the total number of training rounds is the second weight. The second weight represents the degree of influence of the training progress on model training. As the training progress advances, the second weight gradually approaches 1, indicating that the model training is mainly affected by the current training round number. For example, if the current is the 5th round of training and the total number of training rounds is 20 rounds, then the ratio obtained by dividing the current training round number by the total number of training rounds is 5 / 20 = 0.25, that is, the second weight is 0.25. If the current training round number is larger, then the second weight will also be larger; conversely, if the current training round number is smaller, then the second weight will also be smaller.

[0180] Step 433: Determine the first weight based on the second weight.

[0181] In some embodiments, the sum of the first weight and the second weight is 1. Assuming the first weight is w1, the current training round is T, and the total number of training rounds is Tmax, the calculation formula for the first weight is as follows:

[0182] w1 = 1 - w2 (7)

[0183] where w2 is the second weight, and that is to say, the second weight is positively correlated with the current training round. As the current training round increases, the second weight will gradually increase. And it can be seen from formula (7) that the first weight is negatively correlated with the current training round, that is, as the current training round increases, the first weight will gradually decrease.

[0184] Thus, in the early stage of model training, the first weight is relatively high, and the model pays more attention to the training results of the first sub-model. As the training progresses, the second weight gradually increases, and the model focuses more on the training results of the second sub-model. This can enable the model to more flexibly adapt to the training requirements at different stages during training and improve the performance and generalization ability of the model.

[0185] Step 1044: Use the first weight and the second weight to perform weighted summation on the first loss value and the second loss value to obtain a combined loss value.

[0186] In some embodiments, let the combined loss value be Loss, then the formula for calculating the combined loss value is as follows:

[0187] Loss = w * Loss1 + (1 - w) * Loss2 (8)

[0188] where Loss1 is the first loss value corresponding to the first sub-model, and Loss2 is the second loss value corresponding to the second sub-model.

[0189] Therefore, by dynamically allocating the combined loss value, the model pays more attention to the randomly sampled first training samples in the early stage of training and more attention to the long-tail samples obtained by inverse sampling, that is, the second training samples, in the later stage of training.

[0190] Since the ratio obtained by dividing the current training round by the total number of rounds is used as the second weight, the model pays more attention to the first training samples obtained by random sampling in the early stage. Because in the initial stage, the model needs to establish the ability to accurately predict data under normal circumstances. Therefore, giving a higher weight to the first sub-model obtained by random sampling can prompt the model to learn general rules faster. Then, the combined loss value makes the model pay more attention to the long-tail samples obtained by inverse sampling, that is, the second training samples, during the later training. Long-tail samples often have a lower occurrence frequency, but they contain some special cases or potential valuable information. By increasing the weight of long-tail samples, the model can be guided to pay more attention to these uncommon but significant second training samples, thereby enhancing the recommendation model's ability to handle personalized needs and special cases.

[0191] In the above steps 1041 to 1044, using the combined loss value can dynamically allocate weights according to the training stage, enabling the model to focus on general rules in the early stage and special cases in the later stage. Such a training strategy helps the recommendation model better adapt to various data scenarios and improve prediction accuracy and personalized recommendation effects.

[0192] Continue to refer to Figure 3A , and continue to explain in connection with step 104.

[0193] In step 105, the combined loss value is used to train the recommendation model to be trained, and a trained recommendation model is obtained.

[0194] In some embodiments, the calculated combined loss value is used to train the recommendation model to be trained. By minimizing the combined loss value, the parameters of the model are adjusted so that the model can better fit the training data and improve the accuracy of recommendations. By iterating multiple training rounds, the model parameters are continuously optimized so that the model can more accurately predict the probability that the recommended information will be clicked by the recommended object. When the training is completed, a trained recommendation model will be obtained.

[0195] In the embodiments of the present application, the initial training dataset is randomly sampled and inversely sampled to obtain a first training dataset and a second training dataset. The inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training dataset. This inverse sampling method can increase the occurrence probability of long-tail samples. The samples in the first training dataset are predicted using a first sub-model to obtain a first predicted value, and the samples in the second training dataset are predicted using a second sub-model to obtain a second predicted value. To improve the recommendation effect of long-tail advertisements, a dynamic distillation loss value is used to train the second sub-model. At the same time, a joint loss value is determined and dynamically allocated based on the sample labels in the first training dataset, the first predicted value, and the sample labels in the second training dataset. In this way, more attention is paid to the first training samples obtained by random sampling during the early stage of model training, and more attention is paid to the second training samples obtained by inverse sampling during the later stage of training. Finally, the joint loss value is used to train the recommendation model to be trained, thereby obtaining a trained recommendation model. The trained recommendation model has undergone a training process using the joint loss value, can pay more attention to different types of samples, and has high prediction accuracy and recommendation effect.

[0196] Based on the foregoing embodiments, an information recommendation method is provided in the embodiments of the present application. Refer to Figure 4 , Figure 4 which is a flowchart of the information recommendation method provided in the embodiments of the present application and will be described in conjunction with Figure 4 the steps shown. Figure 4 The main body of the steps is the server.

[0197] Step 201: Obtain the to-be-published recommendation information, the candidate object information of multiple candidate recommendation objects corresponding to the to-be-published recommendation information, and the trained recommendation model.

[0198] In some embodiments, the to-be-published recommendation information may include information such as recommended content, recommended target, and recommended time. The candidate object information may include personal profiles, historical behaviors, interest preferences, etc. of the objects. To obtain the to-be-published recommendation information and the candidate object information, a connection can be established with the database of the recommendation system, and the to-be-published recommendation information can be obtained from the database through a query operation.

[0199] For a trained recommendation model obtained by the model training method provided in the embodiments of the present application, by performing inverse sampling on the initial training data set, the obtained second training data set can enhance the model's learning ability for long-tail samples. To further improve the recommendation model, the recommendation model is divided into a first sub-model and a second sub-model. By performing random sampling on the initial training data set, a first training data set is obtained for training the first sub-model. Then, by performing inverse sampling on the initial training data set, a second training data set is obtained for training the second sub-model. The recommendation model to be trained is trained using the joint loss value, so that the trained recommendation model has better generalization ability and recommendation effect. Through the way of joint training, the recommendation model can better utilize the advantages of the two sub-models, thus comprehensively considering the information brought by random sampling and inverse sampling, improving the performance of the recommendation model and the accuracy of recommendations. Therefore, the training method of the embodiments of the present application can enhance the learning ability of the recommendation model for minority-class samples, and further improve the generalization ability and recommendation effect of the recommendation model.

[0200] Step 202: Determine the to-be-published recommendation features of the to-be-published recommendation information and the candidate object features of each candidate object information.

[0201] In some embodiments, after obtaining the to-be-published recommendation information and the candidate object information, feature extraction is performed on the to-be-published recommendation information and the candidate object information to determine the to-be-published recommendation features of the to-be-published recommendation information and the candidate object features of the candidate object information. First, relevant features of the recommendation information are extracted, including recommendation information identifiers, recommendation information content, product types, keywords, etc. Then, information such as the personal profiles, historical behaviors, and hobbies of each candidate object is analyzed, and features related to the recommended object are extracted therefrom. These features can be the geographical location, age, gender, purchase preferences, etc. of the object.

[0202] Step 203: Use the feature processing module in the trained recommendation model to process the to-be-published recommendation features and the candidate object features of each candidate object information to obtain a second to-be-predicted feature vector.

[0203] In some embodiments, a trained recommendation model is loaded, and a feature processing module in the trained recommendation model is used to process the candidate object features of the to-be-released recommendation features and the candidate object information. The feature processing module discretizes, standardizes, or normalizes continuous features to obtain a 64-dimensional continuous feature vector. It can also perform data field mapping or embedding encoding on identification type features to convert them into 64-dimensional identification type feature vectors. Finally, the feature processing module outputs a second to-be-predicted feature vector, which may include the to-be-released recommendation feature vector, the candidate object feature vector, and the cross feature vector after being transformed by the feature processing module. The second to-be-predicted feature vector contains more informative feature representations and can better describe the relationship between the to-be-released recommendation information and the candidate object.

[0204] Step 204: Use the first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector to obtain the target object of the to-be-released recommendation information.

[0205] In some embodiments, a trained recommendation model can be obtained by jointly training the first sub-model and the second sub-model. When normally predicting recommendation information, since the sampled recommendation information is randomly selected, only the first sub-model in the recommendation model needs to be used for prediction. This is because the first sub-model has been fully trained and can accurately predict the representation of the target object to determine whether to send the recommendation information to the object. Therefore, the prediction task can be completed solely relying on the first sub-model without using the second sub-model. When the second to-be-predicted feature vector is received, it is passed to the first sub-model for prediction. The first sub-model will evaluate the degree of preference of the target object for the to-be-released recommendation information based on the relationships learned during the training process. In this way, a prediction result, that is, the target object of the to-be-released recommendation information, can be obtained.

[0206] Step 205: Send the to-be-released recommendation information to the terminal corresponding to the target object.

[0207] In some embodiments, the to-be-released recommendation information is predicted by the trained recommendation model for the second to-be-predicted feature vector, and the to-be-released recommendation information is sent to the terminal corresponding to the target object through network transmission or other means. When the recommendation information reaches the terminal of the target object, the terminal device will decode and process the received recommendation information. According to the format of the recommendation information, the terminal device can select a suitable application or browser to display the recommended content.

[0208] In the above manner, in the embodiments of the present application, by performing random sampling and inverse sampling on the initial data set, the inverse sampling can increase the occurrence probability of long-tail samples, thereby better covering the diversity requirements in information recommendation. Using the first sub-model to predict the first training sample, a first prediction value can be obtained, and then the performance of the first model can be evaluated. Using the second sub-model to predict the second training sample to obtain a second prediction value, so as to measure the performance of the second model. Calculating the loss value of the first prediction value can reflect the optimization effect of the first model on the randomly sampled samples. By calculating the second prediction value and combining it with the first prediction value, the main loss value and the distillation loss value of the second model can be obtained, further improving the recommendation effect of the long-tail recommendation information. Combining the main loss value and the distillation loss value of the second model can obtain the overall loss value of the second model, making it pay more attention to long-tail samples. At the same time, combining the loss value of the first model and the loss value of the second model can obtain a combined loss value for jointly training the two sub-models. Dynamically allocating the loss value can make the recommendation model pay more attention to the normal samples obtained by random sampling in the early stage and more attention to the long-tail samples obtained by inverse sampling in the later stage, so as to achieve key optimization of the recommendation model at different stages. Finally, when using the trained recommendation model, only the first training sample obtained by randomly sampling the initial data set using the first sub-model is used for prediction to ensure that appropriate recommendation information is sent to the object.

[0209] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0210] It can be understood that in the embodiments of the present application, data related to user information and the like are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0211] In a recommendation system, the categories of recommended information usually show a long-tail distribution, that is, a small number of recommended information categories account for most of the exposures and consumptions, while most recommended information is in the long-tail position. However, the current model training methods often focus too much on popular recommended information, resulting in poor recommendation effects for long-tail recommended information. This is because in the data set, most samples belong to the categories of popular recommended information, and the learning of the model mainly focuses on these popular recommended information. This situation makes the model's understanding and recommendation ability for long-tail recommended information relatively insufficient. However, there is a major drawback in the existing technical solutions, that is, it will make the model overly biased towards learning and recommending popular recommended information, and the recommendation effect for long-tail recommended information is not good. This is a challenge for the recommendation information system because long-tail recommended information also needs to be accurately recommended to the object and obtain more exposures and clicks.

[0212] To solve this problem, the embodiments of the present application propose a new training method for a recommendation model, aiming to significantly improve the recommendation effect of long-tail recommendation information while not affecting the recommendation effect of popular recommendation information. As Figure 5 shown, Figure 5 is a schematic diagram of the training process of the recommendation model provided by the embodiments of the present application.

[0213] Step 301: Initially select the recommendation information in the initial training samples and the object information of the recommended objects corresponding to the recommendation information.

[0214] In some embodiments, by interacting with the recommendation information system, multiple pieces of recommendation information are obtained from the recommendation information system. These pieces of recommendation information may include content such as the title, description, and picture of the recommendation information. At the same time, the object information of the recommended objects corresponding to each piece of recommendation information, such as object identifiers, genders, ages, etc., is also obtained from the recommendation information system. To select the recommendation information as the initial training samples, it is determined whether the object clicks on the recommendation information based on the object behavior data analysis. When the object clicks on a certain piece of recommendation information, the sample label of the initial training sample corresponding to this piece of recommendation information is used as the positive sample label, and it is given a label value of 1, indicating that this piece of recommendation information is the recommendation information that the object is interested in and clicks on. On the contrary, when the object does not click on a certain piece of recommendation information but still exposes this piece of recommendation information, the sample label of the initial training sample of this piece of recommendation information is used as the negative sample label, and it is given a label value of 0, indicating that this piece of recommendation information is the recommendation information that the object does not choose to click on.

[0215] Step 302: Extract features from the recommendation information to obtain recommendation feature data.

[0216] In some embodiments, by extracting features from each piece of recommendation information, recommendation feature data can be obtained. In the recommendation of recommendation information, extracting the features of the recommendation information includes information such as recommendation information identifiers, recommendation information categories, recommendation information groups, and the products corresponding to the recommendation information. These features can describe different attributes and association relationships of the recommendation information.

[0217] Step 303: Extract features from the object information corresponding to the recommendation information to obtain object feature data.

[0218] In some embodiments, the object feature data may include object attribute features and may also include object behavior features. Among them, the object attribute features are constructed based on basic information such as the age and gender of the object, and these features can reflect the personal background and interest preferences of the object. In addition, it is also necessary to extract object behavior features, such as behavior information such as article reading records, recommendation information click records, and recommendation information conversion records within the past 14 days. These features can reveal the activity trajectories and consumption intentions of the object.

[0219] Step 304: Determine cross-feature data.

[0220] In some embodiments, during the process of determining the cross-feature data corresponding to the recommended feature data, one or more recommended feature fields in the recommended feature data are combined with one or more object feature fields in the object feature data to obtain the cross-feature data. For example, the gender of the object and the recommended information identifier are concatenated to form a cross-feature data, representing the gender feature of the object in the recommended information.

[0221] Step 305: Construct an initial training sample.

[0222] In some embodiments, integrating the recommended feature data, the object feature data corresponding to the recommended feature data, and the cross-feature data corresponding to the recommended feature data can construct multiple initial training samples, providing a basis for subsequent model training.

[0223] Step 306: Randomly sample the initial training dataset to obtain the first training dataset.

[0224] In some embodiments, when randomly sampling the initial training dataset, a part of the training samples are randomly selected from the entire initial training dataset according to a preset sampling ratio as the first training samples. This process is completely random, and the probability of each training sample being selected is equal, thus ensuring that the sampling result has randomness and representativeness. The random sampling probability can be set manually or determined according to the training requirements.

[0225] Step 307: Inverse sample the initial training dataset to obtain the second training dataset.

[0226] In some embodiments, to obtain the second training dataset, the inverse sampling probability needs to be calculated first. The inverse sampling probability refers to the probability that each second training sample is selected being inversely proportional to its frequency of occurrence in the initial training dataset. The inverse sampling probability corresponding to the recommended information identifier is as follows:

[0227]

[0228] where w i is the weighted weight corresponding to each recommended information identifier, is the total weight, and P i is the inverse sampling probability corresponding to each recommended information identifier.

[0229] The formula for the weighted weight corresponding to the recommended information identifier is as follows:

[0230]

[0231] where Ni The initial number of training samples corresponding to each recommendation information identifier, N max The maximum number of training samples, N max = max(N i ).

[0232] Step 308: Train the recommendation model to be trained using the first training dataset and the second training dataset.

[0233] In some embodiments, first, the first input feature vectors of the respective first training samples included in the first training dataset are determined separately. Since the first training samples include recommendation feature data, object feature data, and cross feature data, and the recommendation feature data, object feature data, and cross feature data each include at least one feature field, if a feature field is a non - continuous feature, for example, the gender field, the recommendation information identifier field, the object identifier field, etc. are all non - continuous features, then a non - continuous feature can be mapped to a 64 - dimensional feature vector through an embedding layer; if a feature field is a continuous feature field, for example, the commodity price field, the sales volume field, etc. are continuous features, then the feature field can be discretized by means of bucketing to obtain a discrete feature, and then the discrete feature is mapped to a 64 - dimensional feature vector. Thus, through the above steps, the first recommendation feature vector of the recommendation feature data, the first object feature vector corresponding to the object feature data, and the first cross feature vector corresponding to the cross feature data included in the first training sample can be determined. Then, the first recommendation feature vector, the first object feature vector, and the first cross feature vector are concatenated to obtain the first input feature vector.

[0234] Next, the first input feature vector is used with a wide && deep network for feature integration to obtain a first feature vector to be predicted. In implementation, load the pre - trained wide && deep model, Figure 6 which is a schematic diagram of the structure of the wide && deep network, as Figure 6 shown. The pre - trained wide && deep model includes a wide model 601 and a deep model 602. The recommendation feature data, object feature data, and cross feature data are passed to the wide model 601, and the recommendation feature data, object feature data, and cross feature data are passed to the deep model 602. The features obtained from the wide model and the deep model are integrated to obtain a first feature vector to be predicted.

[0235] Then, through steps similar to those for obtaining the first feature vector to be predicted, the second input feature vector corresponding to the second training sample in the second training dataset is subjected to feature integration using a wide && deep network to obtain a second feature vector to be predicted. The specific processes of obtaining the second input feature vector and the second feature vector to be predicted are similar to those of obtaining the first input feature vector and the first feature vector to be predicted, and will not be described in detail here.

[0236] Figure 7 is a schematic structural diagram of a recommendation model, as Figure 7 shown. The recommendation model to be trained includes a main model 701 (corresponding to the first sub-model in other embodiments) and an auxiliary model 702 (corresponding to the second sub-model in other embodiments). Finally, the first feature vector to be predicted is input into the main model 701 (the first sub-model) for prediction to obtain a first predicted value. The second feature vector to be predicted is input into the auxiliary model 702 (the second sub-model) for prediction to obtain a second predicted value. The second feature vector to be predicted is obtained by concatenating the obtained second recommendation feature vector, the second object feature vector, and the second cross feature vector.

[0237] An embodiment of this application designs a dynamic distillation network to distill the prediction result of the first sub-model to the second sub-model. Specifically as follows:

[0238] First, determine the first loss value corresponding to the first sub-model. The first loss value corresponding to the first sub-model is determined through a cross-entropy loss function. During this period, the cross-entropy loss is calculated using the first sample label of each first training sample in the first training dataset and each first predicted value. Among them, the first sample labels are divided into positive samples and negative samples. The label corresponding to a positive sample is 1; the label corresponding to a negative sample is 0. The cross-entropy loss function is used to calculate the first predicted value and the first sample label. The calculation formula of the cross-entropy loss function for determining the first loss value corresponding to the first sub-model is as follows:

[0239] Loss1 = cross_entropy(logit1, y1), (11)

[0240] where y1 is the first sample label and logit1 is the first predicted value. In this formula, the first sample label and the first predicted value are input to calculate the cross-entropy loss value.

[0241] Then, determine the main loss value corresponding to the second sub-model. Exemplarily, the main loss value can be determined by the following formula (12):

[0242]

[0243] Among them, y2 is the second sample label and logit2 is the second predicted value. Substitute the labels and predicted values of the positive samples in the second training dataset into the positions of the second sample label and the second predicted value, and then substitute the labels and predicted values of the negative samples into the positions of the second sample label and the second predicted value. Calculate the cross-entropy loss values of the two parts respectively, and add them up to obtain the final loss value.

[0244] Next, determine the distillation loss value corresponding to the second sub-model. Exemplarily, the distillation loss value can be determined by the following formula (13):

[0245]

[0246] Among them, logit1 is the first predicted value and logit2 is the second predicted value. That is, distillation is performed if the first loss value is less than the main loss value, otherwise distillation is not performed. When the first loss value is less than the main loss value, use the respective first predicted values in the first training dataset and the respective second predicted values in the second training dataset to calculate the cross-entropy loss and determine the distillation loss value.

[0247] After determining the main loss value and the distillation loss value of the auxiliary model, the second loss value corresponding to the auxiliary model can be determined using the following formula (14):

[0248]

[0249] Among them, Loss2 is the second loss value, is the main loss value, is the distillation loss value, and β is a hyperparameter, which is adjusted through multiple trainings.

[0250] Finally, determine the joint loss value.

[0251] Let the joint loss value be Loss, then the formula for calculating the joint loss value is as follows:

[0252] Loss = w * Loss1+(1 - w) * Loss2 (15)

[0253] Among them, Loss1 is the first loss value corresponding to the first sub-model, Loss2 is the second loss value corresponding to the second sub-model, and w is the first weight. The calculation formula for the first weight is:

[0254]

[0255] Among them, T is the current training round, and Tmax is the total number of training rounds. The sum of the first weight and the second weight is 1. The first weight is negatively correlated with the current training round, that is, as the current training round increases, the first weight will gradually decrease. The second weight is positively correlated with the current training round, and as the current training round increases, the second weight will gradually increase. That is to say, during the training of the recommendation model, it will first focus on learning the randomly sampled samples, and then as the current training round increases, it will gradually increase the inversely sampled samples, that is, the long-tail samples, and gradually use the inversely sampled samples to train the recommendation model.

[0256] Step 309: Train to obtain a trained recommendation model.

[0257] In some embodiments, the recommendation model to be trained is trained using the calculated joint loss value. By minimizing the joint loss value, the hyperparameters of the model are adjusted so that the model can better fit the training data and improve the accuracy of recommendations. By iterating through multiple training rounds, the model parameters are continuously optimized so that the model can more accurately predict the interests of the object. When the training is completed, a trained recommendation model will be obtained.

[0258] After obtaining the trained recommendation model, the trained recommendation model is used to predict the recommendation information. Only the first training samples obtained by randomly sampling the initial data set using the first sub-model in the trained recommendation model are used for prediction to obtain the click probability of the recommended object for the to-be-published recommendation information.

[0259] In some embodiments, a trained recommendation model can be obtained by jointly training the first sub-model and the second sub-model. During normal prediction of recommendation information, since the sampled recommendation information is randomly selected, only the first sub-model in the recommendation model needs to be used for prediction. This is because the first sub-model has been fully trained and can accurately predict the representation of the target object, thereby determining whether to send the recommendation information to the object. Therefore, the prediction task can be completed solely relying on the first sub-model without using the second sub-model. When the second to-be-predicted feature vector is received, it is passed to the first sub-model for prediction. The first sub-model will evaluate the preference degree of the target object for the to-be-published recommendation information according to the relationships learned during the training process. In this way, a prediction result, that is, the target object of the to-be-published recommendation information, can be obtained.

[0260] In some embodiments, the trained recommendation model is used to predict the second feature vector to be predicted, obtaining the recommended information to be published and the click probability corresponding to the recommended information to be published. The recommended information to be published with a high recommendation probability is sent to the terminal corresponding to the target object through means such as network transmission. When the recommended information reaches the terminal of the target object, the terminal device decodes and processes the received recommended information. According to the format of the recommended information, the terminal device can select a suitable application or browser to display the recommended content.

[0261] In summary, the embodiments of the present application adopt the methods of random sampling and inverse sampling, which improve the appearance probability of long-tail samples in information recommendation, thus better meeting the diversity requirements. The first sub-model is used to predict the training samples to evaluate the performance of the first model and calculate its optimization effect on the randomly sampled samples. At the same time, the second sub-model is used to predict another set of training samples, and combined with the prediction results of the first model, the main loss value and the distillation loss value of the second model are obtained, further improving the effect of long-tail recommended information. By dynamically allocating the loss value, the recommendation model can focus on optimizing normal samples and long-tail samples at different stages. Finally, when using the trained recommendation model, only the training samples obtained by randomly sampling the initial data set by the first sub-model are used for prediction to ensure that appropriate recommended information is published to the object. These method improvements bring a 1% increase in the cold start success rate of the recommended information of the recommendation model and a 0.1% improvement in the model training performance.

[0262] Next, the implementation of the model training device 455 provided by the embodiments of the present application as a software module will be further described. In some embodiments, as Figure 2A shown, the software module in the model training device 455 stored in the memory 450-1 may include:

[0263] The first acquisition module 4551 is used to acquire the recommendation model to be trained and acquire the initial training data set. The recommendation model to be trained includes a first sub-model and a second sub-model;

[0264] The sampling module 4552 is used to randomly sample the initial training data set to obtain a first training data set, and perform inverse sampling on the initial training data set based on the inverse sampling probability to obtain a second training data set. The inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommended information identifier in the initial training data set;

[0265] The first prediction module 4553 is used to use the first sub-model to predict each first training sample in the first training data set to obtain each first prediction value, and use the second sub-model to predict each second training sample in the second training data set to obtain each second prediction value;

[0266] The first determination module 4554 is configured to determine a combined loss value by using the first sample labels of the respective first training samples in the first training dataset, the respective first predicted values, and the second sample labels of the respective second training samples in the second training dataset;

[0267] The training module 4555 is configured to train the recommendation model to be trained by using the combined loss value to obtain a trained recommendation model.

[0268] In some embodiments, the first acquisition module is further configured to acquire a plurality of recommendation information and object information of the recommended objects corresponding to the respective recommendation information; perform feature extraction on the respective recommendation information and the object information corresponding to the respective recommendation information to obtain respective recommendation feature data and object feature data corresponding to the respective recommendation feature data; determine cross-feature data corresponding to the respective recommendation feature data from the respective recommendation feature data and the object feature data corresponding to the respective recommendation feature data; construct respective initial training samples from the respective recommendation feature data, the object feature data corresponding to the respective recommendation feature data, and the cross-feature data corresponding to the respective recommendation feature data; determine sample labels of the respective initial training samples; and construct an initial training dataset based on the respective initial training samples and the sample labels of the respective initial training samples.

[0269] In some embodiments, the first acquisition module is further configured to acquire recommendation information identifiers corresponding to the respective initial training samples; determine the number of initial training samples corresponding to the respective recommendation information identifiers, and determine the maximum number of training samples from the number of initial training samples corresponding to the respective recommendation information identifiers; determine the ratio of the maximum number of training samples to the number of initial training samples corresponding to the respective recommendation information identifiers as the weighted weight corresponding to the respective recommendation information identifiers; determine the sum of the weighted weights corresponding to the respective recommendation information identifiers as the total weight; and determine the ratio of the weighted weight corresponding to the respective recommendation information identifier to the total weight as the inverse sampling probability corresponding to the respective recommendation information identifier.

[0270] In some embodiments, the first prediction module is further configured to, for each first training sample, determine a first recommendation feature vector corresponding to the first recommendation feature data, a first object feature vector corresponding to the first object feature data, and a first cross feature vector corresponding to the first cross feature data in the first training sample; perform a splicing process on the first recommendation feature vector, the first object feature vector, and the first cross feature vector to obtain a first input feature vector; perform feature integration on the input feature vector to obtain a first feature vector to be predicted; correspondingly, using the first sub-model to predict each first training sample in the first training dataset to obtain each first prediction value, including: using the first sub-model to predict the first feature vector to be predicted to obtain a first prediction value.

[0271] In some embodiments, the first determination module is further configured to use the first sample label of each first training sample in the first training dataset and each first prediction value to determine a first loss value corresponding to the first sub-model; use each first prediction value of each first training sample in the first training dataset and the second sample label of each second training sample in the second training dataset to determine a second loss value corresponding to the second sub-model; determine a first weight corresponding to the first loss value and a second weight corresponding to the second loss, where the first weight is negatively correlated with the current training round number, and the second weight is positively correlated with the current training round number; use the first weight and the second weight to perform a weighted sum of the first loss value and the second loss value to obtain the joint loss value.

[0272] In some embodiments, the first determination module is further configured to use the second sample label of each second training sample in the second training dataset and each second prediction value to determine a main loss value corresponding to the second sub-model; when the first loss value is less than the main loss value, use the first prediction value and the second prediction value to determine a distillation loss value; perform a weighted sum of the main loss value and the distillation loss value to obtain a second loss value corresponding to the second sub-model.

[0273] In some embodiments, the first determination module is further configured to, when the first loss value is greater than or equal to the main loss value, determine the main loss value as the second loss value corresponding to the second sub-model.

[0274] In some embodiments, the first determination module is further configured to obtain the current training round number and the total number of training rounds; determine the ratio of the current training round number to the total number of training rounds as the second weight; based on the second weight, determine the first weight, where the sum of the first weight and the second weight is 1.

[0275] Next, the exemplary structure of the information recommendation device 455-2 provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 2B shown, the software modules stored in the information recommendation device 456 in the memory 450-2 may include:

[0276] A second acquisition module 4561, configured to acquire the to-be-published recommendation information, candidate object information of a plurality of candidate recommendation objects corresponding to the to-be-published recommendation information, and a trained recommendation model;

[0277] A second determination module 4562, configured to determine the to-be-published recommendation features of the to-be-published recommendation information and the candidate object features of each candidate object information;

[0278] A processing module 4563, configured to use the feature processing module in the trained recommendation model to process the to-be-published recommendation features and the candidate object features of each candidate object information, so as to obtain a second to-be-predicted feature vector;

[0279] A second prediction module 4564, configured to use the first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector, so as to obtain the target object of the to-be-published recommendation information;

[0280] A sending module 4565, configured to send the to-be-published recommendation information to the terminal corresponding to the target object.

[0281] The embodiments of the present application provide a computer program product, which includes a computer program or computer-executable instructions. The computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the model training method or information recommendation method described above in the embodiments of the present application.

[0282] The embodiments of the present application provide a computer-readable storage medium storing computer-executable instructions, where computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the model training method or information recommendation method provided in the embodiments of the present application. For example, as Figures 3A to 4 the method.

[0283] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0284] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0285] As an example, the computer-executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).

[0286] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0287] In summary, the embodiments of the present application adopt the methods of random sampling and inverse sampling, which improve the occurrence probability of long-tail samples in information recommendation, thus better meeting the diversity requirements. The performance of the first model is evaluated by predicting the training samples through the first sub-model, and the optimization effect on the randomly sampled samples is calculated. At the same time, the second sub-model predicts another set of training samples, and combined with the prediction results of the first model, the main loss value and the distillation loss value of the second model are obtained, further improving the effect of long-tail recommendation information. By dynamically allocating the loss value, the recommendation model can focus on optimizing normal samples and long-tail samples at different stages. Finally, when using the trained recommendation model, only the training samples randomly sampled from the initial data set by the first sub-model are used for prediction to ensure that appropriate recommendation information is published to the object.

[0288] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A model training method, characterized in that, The method includes: Obtaining a recommendation model to be trained and obtaining an initial training data set, where the recommendation model to be trained includes a first sub-model and a second sub-model; Performing random sampling on the initial training data set to obtain a first training data set, and performing inverse sampling on the initial training data set based on an inverse sampling probability, where the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training data set; Using the first sub-model to predict each first training sample in the first training data set to obtain respective first prediction values, and using the second sub-model to predict each second training sample in the second training data set to obtain respective second prediction values; Determining a joint loss value using the first sample labels of each first training sample in the first training data set, the respective first prediction values, and the second sample labels of each second training sample in the second training data set; Training the recommendation model to be trained using the joint loss value to obtain a trained recommendation model.

2. The method according to claim 1, wherein The obtaining of the initial training data set includes: Obtaining a plurality of recommendation information and object information of the recommended objects corresponding to each recommendation information; Performing feature extraction on each recommendation information and the object information corresponding to each recommendation information to obtain respective recommendation feature data and object feature data corresponding to the respective recommendation feature data; Determining the respective recommendation feature data and the object feature data corresponding to the respective recommendation feature data, and determining cross feature data corresponding to the respective recommendation feature data; Constructing respective initial training samples from the respective recommendation feature data, the object feature data corresponding to the respective recommendation feature data, and the cross feature data corresponding to the respective recommendation feature data; Determining the sample labels of each initial training sample; Constructing an initial training data set based on each initial training sample and the sample labels of each initial training sample.

3. The method according to claim 2, characterized in that, The method further includes: Obtaining the recommendation information identifier corresponding to each initial training sample; Determining the number of initial training samples corresponding to each recommendation information identifier and determining the maximum number of training samples from the number of initial training samples corresponding to each recommendation information identifier; Determining the ratio of the maximum number of training samples to the number of initial training samples corresponding to each recommendation information identifier as the weighted weight corresponding to each recommendation information identifier; Determining the sum of the weighted weights corresponding to each recommendation information identifier as the total weight; Determining the ratio of the weighted weight corresponding to each recommendation information identifier to the total weight as the inverse sampling probability corresponding to each recommendation information identifier.

4. The method according to claim 1, wherein The recommendation model further includes a feature processing module, the first training sample includes first recommendation feature data, first object feature data, and first cross feature data, and the method further includes: For each first training sample, determining a first recommendation feature vector corresponding to the first recommendation feature data in the first training sample, a first object feature vector corresponding to the first object feature data, and a first cross feature vector corresponding to the first cross feature data; Concatenate the first recommended feature vector, the first object feature vector, and the first cross feature vector to obtain a first input feature vector; Integrate the features of the input feature vector to obtain a first feature vector to be predicted; Correspondingly, the step of using the first sub-model to predict each first training sample in the first training dataset to obtain each first prediction value includes: Use the first sub-model to predict the first feature vector to be predicted to obtain a first prediction value.

5. The method according to claim 1, wherein The step of determining the joint loss value using the first sample labels of each first training sample in the first training dataset, each first prediction value, and the second sample labels of each second training sample in the second training dataset includes: Use the first sample labels of each first training sample in the first training dataset and each first prediction value to determine the first loss value corresponding to the first sub-model; Use each first prediction value of each first training sample in the first training dataset and the second sample labels of each second training sample in the second training dataset to determine the second loss value corresponding to the second sub-model; Determine the first weight corresponding to the first loss value and the second weight corresponding to the second loss. The first weight is negatively correlated with the current training round, and the second weight is positively correlated with the current training round; Use the first weight and the second weight to perform weighted summation on the first loss value and the second loss value to obtain the joint loss value.

6. The method according to claim 5, wherein The step of using each first prediction value of each first training sample in the first training dataset and the second sample labels of each second training sample in the second training dataset to determine the second loss value corresponding to the second sub-model includes: Use the second sample labels of each second training sample in the second training dataset and each second prediction value to determine the main loss value corresponding to the second sub-model; When the first loss value is less than the main loss value, use the first prediction value and the second prediction value to determine the distillation loss value; Perform weighted summation on the main loss value and the distillation loss value to obtain the second loss value corresponding to the second sub-model.

7. The method according to claim 6, wherein The step of using each first prediction value of each first training sample in the first training dataset and the second sample labels of each second training sample in the second training dataset to determine the second loss value corresponding to the second sub-model includes: When the first loss value is greater than or equal to the main loss value, determine the main loss value as the second loss value corresponding to the second sub-model.

8. The method according to claim 5, characterized in that, The step of determining the first weight corresponding to the first loss value and the second weight corresponding to the second loss includes: Obtain the current training round and the total number of training rounds; Determine the ratio of the current training round to the total number of training rounds as the second weight; Based on the second weight, determine the first weight, and the sum of the first weight and the second weight is 1.

9. An information recommendation method, characterized in that, The method includes: Obtain the recommended information to be released, the candidate object information of multiple candidate recommended objects corresponding to the recommended information to be released, and the trained recommendation model; Determine the to-be-released recommendation features of the to-be-released recommendation information and the candidate object features of each candidate object information; Use the feature processing module in the trained recommendation model to process the to-be-released recommendation features and the candidate object features of each candidate object information to obtain a second to-be-predicted feature vector; Use the first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector to obtain the target object of the to-be-released recommendation information; Send the to-be-released recommendation information to the terminal corresponding to the target object.

10. A model training device, characterized in that, The device includes: A first acquisition module, configured to acquire a recommendation model to be trained and acquire an initial training data set, where the recommendation model to be trained includes a first sub-model and a second sub-model; A sampling module, configured to randomly sample the initial training data set to obtain a first training data set, and perform inverse sampling on the initial training data set based on an inverse sampling probability, where the inverse sampling probability is negatively correlated with the number of training samples corresponding to the recommendation information identifier in the initial training data set; A first prediction module, configured to use the first sub-model to predict each first training sample in the first training data set to obtain each first prediction value, and use the second sub-model to predict each second training sample in the second training data set to obtain each second prediction value; A first determination module, configured to determine a joint loss value by using the first sample labels of each first training sample in the first training data set, the respective first prediction values, and the second sample labels of each second training sample in the second training data set; A training module, configured to train the recommendation model to be trained by using the joint loss value to obtain a trained recommendation model.

11. An information recommendation device, characterized in that, The device includes: A second acquisition module, configured to acquire the recommended information to be released, the candidate object information of multiple candidate recommended objects corresponding to the recommended information to be released, and the trained recommendation model; A second determination module, configured to determine the to-be-released recommendation features of the to-be-released recommendation information and the candidate object features of each candidate object information; A processing module, configured to use the feature processing module in the trained recommendation model to process the to-be-released recommendation features and the candidate object features of each candidate object information to obtain a second to-be-predicted feature vector; A second prediction module, configured to use the first sub-model in the trained recommendation model to predict the second to-be-predicted feature vector to obtain the target object of the to-be-released recommendation information; A sending module, configured to send the to-be-released recommendation information to the terminal corresponding to the target object.

12. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the model training method according to any one of claims 1 to 8 or the information recommendation method according to claim 9 when executing the computer-executable instructions stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by a processor, the model training method described in any one of claims 1 to 8 is implemented, or when executed by a processor, the information recommendation method described in claim 9 is implemented.

14. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by a processor, the model training method described in any one of claims 1 to 8 is implemented, or when executed by a processor, the information recommendation method described in claim 9 is implemented.