Information recommendation model training method, recommendation method, device, equipment and medium
By adjusting the physical distance in the training samples of the information recommendation model, the model's distance sensitivity was improved, solving the problem of poor recommendation performance caused by the target object being too far away from the information delivery entity, and achieving a higher shallow conversion rate and a better user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2023-05-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing information recommendation models lack distance sensitivity, resulting in poor recommendation performance when the target object is too far away from the information delivery entity, which affects the shallow conversion rate and user experience.
By obtaining the original training samples of the information recommendation model, a perturbation factor is randomly generated to adjust the physical distance, and a target training sample is generated. The model is then trained by combining the original and target training samples to improve the model's distance sensitivity.
It improves the shallow conversion rate of information recommendation models, saves training time and sample acquisition costs, and enhances the user experience for target users.
Smart Images

Figure CN118964718B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to information recommendation technology, and more particularly to information recommendation model training methods, information recommendation methods, devices, electronic equipment, computer program products, and storage media. Background Technology
[0002] Artificial Intelligence (AI) is a comprehensive technology within computer science that studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. AI technology is a multidisciplinary field, encompassing a wide range of areas, including natural language processing and machine learning / deep learning. It is believed that with technological advancements, AI will be applied in more fields and play an increasingly important role.
[0003] In traditional technologies, various information recommendation systems typically use general data recommendation models to ensure recommendation speed when recommending relevant information to target objects. However, general data recommendation models are trained on general data from multiple domains, resulting in overly conventional recommendation results that lack industry-specific information recommendations. Conversely, while models specifically designed for a particular industry may provide recommendations that meet the target object's needs, the shallow conversion process based on the recommended information often fails due to the distance between the target object and the entity that delivered the information. This leads to poor recommendation performance. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an information recommendation model training method, apparatus, electronic device, and storage medium. The technical solution of the embodiments of the present invention is implemented as follows: This invention provides a method for training an information recommendation model, comprising: Obtain the original training samples for the information recommendation model; The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample label carried by the original training samples is: the shallow conversion rate of the recommendation information for the target object; the shallow conversion rate is used to indicate the probability that the target object reaches the location of the delivery entity. Randomly generate a perturbation factor corresponding to the physical distance; The physical distance is randomly adjusted according to the disturbance factor to obtain the target distance; The target training sample is obtained by replacing the physical distance in the original training sample with the target distance; The information recommendation model is trained by combining the original training samples and the target training samples; The information recommendation model is used to recommend the information to be recommended to the target object based on the physical distance between the target object and the entity to which the information to be recommended belongs.
[0005] This invention also provides an information recommendation method, including: Retrieve the information to be recommended from the data source of the recommendation information; By using an information recommendation model to predict different types of information to be recommended, the shallow conversion rate of different types of information to be recommended can be determined. The recommendation order of the recommended information is adjusted according to the shallow conversion rate of different information to be recommended, and the information is recommended according to the recommendation order.
[0006] This invention also provides an information recommendation model training device, the device comprising: The information transmission module is used to acquire the original training samples of the information recommendation model; The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample label carried by the original training samples is: the shallow conversion rate of the recommendation information for the target object; the shallow conversion rate is used to indicate the probability that the target object reaches the location of the delivery entity. The information processing module is used to randomly generate a perturbation factor corresponding to the physical distance; The information processing module is used to randomly adjust the physical distance according to the disturbance factor to obtain the target distance; The information processing module is used to replace the physical distance in the original training sample with the target distance to obtain the target training sample; The information processing module is used to train the information recommendation model by combining the original training samples and the target training samples; The information recommendation model is used to recommend the information to be recommended to the target object based on the physical distance between the target object and the entity to which the information to be recommended belongs.
[0007] In some embodiments, the information transmission module is used to obtain the identifier of the recommendation information; The information transmission module is used to determine the shallow conversion action of the recommendation information for the target object based on the identifier of the recommendation information.
[0008] In some embodiments, the information transmission module is used to determine the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, based on the identifier of the recommendation information; The information transmission module is used to determine the positive samples in the original training samples based on the sample labels, the shallow conversion actions of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The information transmission module is used to determine the negative sample in the original training sample based on the sample label, the shallow conversion action of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The information transmission module is used to combine the positive sample and the negative sample to obtain the original training sample.
[0009] In some embodiments, the information processing module is configured to randomly generate a disturbance factor corresponding to the physical distance; The information processing module is used to randomly adjust the physical distance according to the disturbance factor to obtain the target distance, or; The information processing module is used to obtain fixed adjustment rules corresponding to the physical distance; The information processing module is used to adjust the physical distance according to the fixed adjustment rules to obtain the target distance; The information processing module is used to replace the physical distance in the original training sample with the target distance to obtain the target training sample.
[0010] In some embodiments, the information processing module is used to train the information recommendation model using the original training samples to determine the first model parameters of the information recommendation model; The information processing module is used to train the information recommendation model using the target training samples, adjust the first model parameters, and obtain the second model parameters of the information recommendation model. The information processing module is used to train the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model.
[0011] In some embodiments, the information processing module is used to determine the main loss function of the information recommendation model; The information processing module is used to train the information recommendation model using the original training samples and the main loss function; The information processing module is used to determine the first model parameters of the information recommendation model when the main loss function reaches the corresponding convergence condition.
[0012] In some embodiments, the information processing module is used to determine the contrastive learning loss function of the information recommendation model; The information processing module is used to train the information recommendation model using the target training samples, the first model parameters, and the contrastive learning loss function. The information processing module is used to adjust the first model parameters to obtain the second model parameters of the information recommendation model when the contrastive learning loss function reaches the corresponding convergence condition.
[0013] In some embodiments, the information processing module is configured to calculate the total loss function of the information recommendation model based on the main loss function and the contrastive learning loss function; The information processing module is used to train the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model. The information processing module is used to adjust the second model parameters to obtain the model parameters of the information recommendation model when the total loss function reaches the corresponding convergence condition.
[0014] This invention also provides an information recommendation device, the device comprising: The data transmission module is used to obtain the information to be recommended from the recommendation information data source; The data processing module is used to predict different information to be recommended using an information recommendation model, and to determine the shallow conversion rate of different information to be recommended. The data processing module is used to adjust the recommendation order of the recommended information according to the shallow conversion rate of different information to be recommended, and to recommend information according to the recommendation order.
[0015] This invention also provides an electronic device, the electronic device comprising: Memory, used to store executable instructions; When a processor executes executable instructions stored in the memory, it implements the aforementioned information recommendation model training method or the aforementioned information recommendation method.
[0016] This invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the aforementioned information recommendation model training method or the aforementioned information recommendation method.
[0017] The embodiments of the present invention have the following beneficial effects: This invention obtains the original training samples of an information recommendation model; randomly generates a perturbation factor corresponding to physical distance; randomly adjusts the physical distance according to the perturbation factor to obtain a target distance; replaces the physical distance in the original training samples with the target distance to obtain a target training sample; and trains the information recommendation model by combining the original training sample and the target training sample. This improves the distance sensitivity of the information recommendation model, resulting in a higher shallow conversion rate for information recommendations. Furthermore, adjusting the physical distance in the original training samples directly yields the target training sample, eliminating the need for data feedback from the entity delivering the recommendation information, thus saving training time and the cost of acquiring training samples, and improving the user experience for the target audience. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating a usage scenario of the information recommendation model training method provided in this embodiment of the invention; Figure 2 An optional flowchart illustrating the training method of the information recommendation model provided in the embodiments of this application; Figure 3 This is a schematic diagram of the training logic of the information recommendation model in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the process of training the information recommendation model in an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the information recommendation process using an information recommendation model in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating an optional information recommendation in an embodiment of the present invention; Figure 7 A schematic diagram of an optional process for training an information recommendation model provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the training logic of the information recommendation model in an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating the training process of the information recommendation model in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of an electronic device 100 for performing the training method of the information recommendation model provided in this application, as provided in an embodiment of this application. Figure 11 This is a schematic diagram of the structure of an electronic device 1000 provided in this application embodiment for performing the information recommendation method provided in this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0021] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0022] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention will be explained, and the nouns and terms involved in the embodiments of the present invention shall be interpreted as follows.
[0023] 1) Shallow conversion generally refers to the conversion behavior that occurs within the advertising domain and targets the advertiser's advertising goals.
[0024] 2) Augmented samples generally refer to samples that are artificially modified from existing samples and are similar to or dissimilar to the original samples.
[0025] 3) Model training: Multi-class classification learning on the image dataset. This model can be built using deep learning frameworks such as TensorFlow and Torch, employing multiple layers of neural networks like CNNs to form a multi-class classification model. The model's input is the training samples in tuples, and its output is the multi-class probability. Algorithms such as softmax are used to ultimately output the predicted shallow conversion rate. During training, the model approximates the correct trend using objective functions such as cross-entropy.
[0026] 4) Neural Network (NN): Artificial Neural Network (ANN), also known as neural network or neural network-like network, is a mathematical or computational model in the fields of machine learning and cognitive science that imitates the structure and function of biological neural networks (the central nervous system of animals, especially the brain) and is used to estimate or approximate functions.
[0027] 5) Multi-objective recall: This involves considering multiple objectives within a single recall model. In recommendation systems, it's often necessary to optimize multiple business objectives simultaneously to generate greater business revenue. For example, in e-commerce, the goal is to simultaneously optimize click-through rate and conversion rate, giving the platform more targeted benefits; in news feeds, the aim is to increase click-through rate while also encouraging engagement such as following, liking, and commenting, creating a better community atmosphere and thus improving retention.
[0028] 6) Recommendation Accuracy: The recommended information has a certain effect over a period of time, and this effect is measured by the target audience's interest in the recommended information. Accuracy plays an important role in target audience retention, click-through rate, and CTR on the mobile device.
[0029] 7) softmax: A very commonly used and important function in machine learning, especially in multi-class classification scenarios. It maps some inputs to real numbers between 0 and 1, and normalization ensures that the sum is 1.
[0030] 8) Recommended information refers to various forms of information available on the Internet, such as advertisements, video files, recommended information, and news information displayed on clients or smart devices.
[0031] 9) BPR Loss (Bayesian Personalized Ranking Loss): This is a loss function used to train ranking models. The purpose of this loss function is to help the model learn how to arrange data in a specific order by calculating the distance between two data points.
[0032] 10) Contrastive Learning: Contrastive learning is a type of self-supervised learning, meaning it does not rely on labeled data. In related technologies, contrastive learning seems to be in a state of "no clear definition, but with guiding principles." Its guiding principle is: by automatically constructing similar and dissimilar instances, it aims to learn a representation learning model. Through this model, similar instances are made closer in the projection space, while dissimilar instances are made farther apart in the projection space.
[0033] During their research, the inventors discovered that in related technologies, when recommending information through information recommendation models, for local life information, the target audience's in-store consumption is directly affected by the distance between the target audience and the entity delivering the information. Specifically, in the local life information industry (such as restaurant information, beauty and hair salon information, entertainment information, and offline education information), the business objectives of the information delivery entity all require the target audience to ultimately make a purchase in-store. Therefore, many information delivery entities aim to get the target audience to make a purchase in-store. Within the information domain, the target audience can generate shallow conversion actions and deep conversion actions based on the recommended information. Shallow conversion actions correspond to shallow conversion rates, and deep conversion actions correspond to deep conversion rates. Shallow conversion actions include downloading, activating, and registering, while deep conversion actions include paying, day 2 retention, day 3 retention, and day 7 retention.
[0034] In this scenario, the distance between the information delivery entity and the target audience becomes a crucial factor in whether the target audience ultimately visits the store. The inventors discovered that current local life post-link optimization algorithms have the following shortcomings: 1) Some information delivery entities report that target audiences, after clicking on the information, do not ultimately visit the store, affecting the final effectiveness of the information delivery and consequently impacting the information delivery entity's budget, thus hindering the increase of the click-through rate of the information delivery system. 2) Because the information recommendation model does not constrain distance sensitivity, target audiences who click on local life information may be unable to visit the store due to its distance or have a very low willingness to do so, affecting the user experience.
[0035] To address the aforementioned shortcomings, this application provides a method for training information recommendation models that makes them more distance-sensitive.
[0036] Figure 1 This is a schematic diagram illustrating a usage scenario of the training method for the information recommendation model provided in this embodiment of the invention. (See attached diagram.) Figure 1Electronic device 100 or electronic device 1000 may include two different types of terminals or electronic devices, namely terminal 10-1 and terminal 10-2, and the target object can be selected according to usage requirements. Terminal 10-1 and terminal 10-2 are equipped with clients or mini-programs capable of displaying recommendation information. The terminals connect to server 200 via network 300, which can be a wide area network, a local area network, or a combination of both, using a wireless link to achieve data transmission. The recommendation information includes, but is not limited to, videos, images, GIF animations, and advertising information. The types of recommendation information obtained by the terminals (including terminal 10-1 and terminal 10-2) from the corresponding server 200 via network 300 can be the same or different. For example, the terminals (including terminal 10-1 and terminal 10-2) can obtain video advertisements or image advertisements from the same industry from the corresponding server 200 via network 300. The specific types are not limited in this application. Server 200 can store different recommendation information, among which the recommendation information as advertisements can be content in different dynamic formats, such as gif, mp4, mov, etc.
[0037] During the process of obtaining and displaying information recommended by the information recommendation model from the server 200 via the network 300, the target audience can perform different operations on the recommended information presented in the playback window through the terminals (terminals 10-1 and / or 10-2). Terminals 10-1 and / or 10-2 can record data on the different usage processes of the target audience. For example, when the recommended information is a video advertisement, the target audience can share and / or like the exposed video advertisement while watching the information, or click on the link provided by the video advertisement to jump to the product purchase page. When the recommended information is a local life advertisement, during the exposure of the advertisement through the terminals (terminals 10-1 and / or 10-2), the target audience can forward and / or comment on the advertisement, or decide whether to make a purchase in-store by checking the physical distance between the local life advertisement's target entity and the target audience.
[0038] In some embodiments of the present invention, the trained information recommendation model can also recommend financial information to meet the financial needs of the target audience. For example, it can recommend financial product information or financial trading venue information to the target audience in the financial industry, so as to meet the target audience in the financial industry to conduct financial activities through virtual or physical resources, or to go to the financial trading venue for offline financial transactions.
[0039] As an example, when server 200 determines what information to recommend to the target audience via terminal 10-1 or 10-2, it needs to adjust the recommended information in a timely manner. For example, based on the shallow conversion rate calculated by the information recommendation model, recommended information with a shallow conversion rate lower than a preset threshold is replaced to adapt to the recommendation needs of target audiences at different distances.
[0040] Taking local lifestyle restaurant advertisement recommendations as an example, the information recommendation model provided by this invention can calculate the shallow conversion rate of different restaurant advertisements stored in server 200. The shallow conversion rate is used to indicate the probability that the target object will reach the location of the restaurant advertisement placement entity. By using the shallow conversion rates of different restaurant advertisements, the recommendation order of the recommended restaurant advertisements can be adjusted. The restaurant advertisement with the highest shallow conversion rate is recommended first, and finally the recommended restaurant advertisement is presented on the UI (User Interface) of terminal 10-1 or 10-2.
[0041] In some embodiments of the present invention, the shallow conversion rate of the information to be recommended obtained by the information recommendation model can also be called by other applications. For example, the shallow conversion rate of the video advertisement placed by the same placement entity to which the recommendation information belongs can be transferred to the image advertisement recommendation process, or to the video advertisement recommendation process in the mini program or the video advertisement recommendation process in the webpage. This not only saves the data processing time of the information recommendation model, but also saves the waiting time for the target object to obtain the recommendation information.
[0042] As an example, server 200 is used to deploy a corresponding information recommendation model to implement the training method and information recommendation method of the information recommendation model provided by this invention. Specifically, during the training of the information recommendation model, the original training samples of the information recommendation model are first obtained; then, a perturbation factor corresponding to the physical distance is randomly generated; the physical distance is randomly adjusted according to the perturbation factor to obtain the target distance; the physical distance in the original training samples is replaced by the target distance to obtain the target training sample; finally, the information recommendation model is trained by combining the original training sample and the target training sample.
[0043] The training method for the information recommendation model provided in this application is based on artificial intelligence (AI). AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions.
[0044] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0045] In the embodiments of this application, the main artificial intelligence software technologies involved include the aforementioned speech processing technologies and machine learning. For example, it may involve Automatic Speech Recognition (ASR) technology in speech technology, including speech signal preprocessing, speech signal frequency analyzing, speech signal feature extraction, speech signal feature matching / recognition, and speech training.
[0046] For example, this could involve machine learning (ML), a multidisciplinary field encompassing probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning typically includes techniques such as deep learning, which includes artificial neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep neural networks (DNNs).
[0047] It is understood that the training method and information recommendation method of the information recommendation model provided in this application can be applied to intelligent devices. Intelligent devices can be any device with information display function, such as intelligent terminals, smart home devices (such as smart speakers, smart washing machines, etc.), smart wearable devices (such as smartwatches), in-vehicle intelligent central control systems (which display recommended information to target objects through applets that perform different tasks), or AI intelligent medical devices (which recommend the nearest medical institutions through the information recommendation model), etc.
[0048] See Figure 2 , Figure 2 This is an optional flowchart illustrating the training method for the information recommendation model provided in this application embodiment. The embodiment of the present invention can be implemented using cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. It can also be understood as a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on cloud computing business models. The backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites; therefore, cloud technology needs cloud computing as its support.
[0049] It's important to note that cloud computing is a computing model that distributes computing tasks across a resource pool comprised of numerous computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" are infinitely scalable, readily available, and can be used on demand, expanded at any time, and paid for based on usage. As the foundational providers of cloud computing capabilities, they establish cloud resource pool platforms, often referred to as cloud platforms or Infrastructure as a Service (IaaS). These platforms deploy various types of virtual resources within the resource pool for external customers to choose from. The cloud resource pool primarily includes: computing devices (which can be virtualized machines containing operating systems), storage devices, and network devices. Different operational permissions are assigned to the target objects processing data through the cloud server.
[0050] Cloud storage is a new concept that extends and develops from cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to aggregate a large number of various types of storage devices (also called storage nodes) across a network through application software or application interfaces to work together and provide data storage and business access functions. In related technologies, the storage method of a storage system is as follows: Logical volumes are created. When creating a logical volume, physical storage space is allocated to each logical volume. This physical storage space may consist of the disks of one or several storage devices. Clients store data on a logical volume, which means storing the data on the file system. The file system divides the data into many parts, each part being an object. Each object contains not only data but also additional information such as a data identifier (ID, IDentity). The file system writes each object to the physical storage space of the logical volume and records the storage location information of each object. Therefore, when a client requests data access, the file system can allow the client to access the data based on the storage location information of each object. The process of a storage system allocating physical storage space to a logical volume involves the following steps: Based on a capacity estimate of the objects stored on the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of Redundant Array of Independent Disks (RAID), the physical storage space is pre-divided into strips. A logical volume can be understood as a strip, thus allocating physical storage space to the logical volume. When applied to cloud products, the front end of the cloud product can be a Web UI component, used to receive operations from the target object on the target interface to achieve the target function.
[0051] Understandably, Figure 2 The steps shown can be implemented by a terminal or server running a training device with an information recommendation model, or by a terminal and server working together, or by a cloud server or a cloud server cluster working together. Here, we will take the server implementation as an example, which includes the following steps: Step 201: Obtain the original training samples for the information recommendation model.
[0052] The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample labels carried by the original training samples; the shallow conversion rate of the recommendation information for the target object; and the shallow conversion rate, which indicates the probability that the target object will reach the location of the delivery entity.
[0053] In some embodiments of the present invention, different types of recommendation information correspond to different shallow conversion actions. Therefore, it is necessary to obtain the identifier of the recommendation information. Based on the identifier of the recommendation information, the shallow conversion action of the recommendation information for the target object is determined. For example, when the recommendation information is a local life advertisement, the business objectives of the local life advertisement placement entity all require the target object to eventually go to the store to make a purchase. The shallow conversion actions in this process include at least one of the following: clicking on the exposed advertisement, following the placement entity, activating the placement entity's form, and the target object going to the store. In addition, the deep conversion actions include at least one of the following: going to the store to make a purchase, placing an order using the in-store mini-program, and purchasing the service products provided by the local life advertisement.
[0054] It should be understood that exposure refers to the act of a visitor seeing the recommended information. Click refers to the act of a visitor clicking on the recommended information after seeing it. Activation refers to actions such as downloading or registering a membership after clicking on the recommended information. It is evident that since exposure, activation, and download are all instantaneous actions, the level of user interaction with the recommended information is relatively low. In contrast, in-store consumption, using in-store navigation information, and purchasing services and products offered by local lifestyle advertisements are actions that occur over a longer period, indicating a higher level of user interaction with the recommended information. In some embodiments of the present invention, the original training samples of the information recommendation model can be obtained in the following ways: Based on the identifier of the recommendation information, determine the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; based on the sample labels, the shallow conversion actions of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information, determine the positive samples in the original training samples; based on the sample labels, the shallow conversion actions of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information, determine the negative samples in the original training samples; combine the positive and negative samples to obtain the original training samples. The physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs can be represented as S. ab Where a is the coordinates of the target object and b is the coordinates of the entity to which the recommendation information belongs (e.g., the location of a store, the location of a service location, or the specified location of an advertisement).
[0055] Since the information recommendation model provided in this application supports contrastive learning during training and does not rely on labeled data, it is necessary to obtain positive and negative samples when acquiring training samples. The Bayesian personalized ranking loss function is used to maximize the score difference between positive and negative samples. Therefore, it is necessary to obtain positive and negative samples for training the information recommendation model. Taking the shallow conversion rate of the recommendation information for the target object as a binary classification probability as an example, target objects that received the recommendation information in the past 14 days and underwent shallow conversion but ultimately did not visit the store can be extracted as negative samples and marked as 0. In the back-link data returned by the target object, target objects that visited the store after receiving the recommendation information are treated as positive samples and marked as 1, with information identifier (id) and physical distance (distance to the nearest store of the information delivery entity) added. The final original training samples can be represented as: Sample label (0 / 1), shallow conversion action of target object (target object visits store), recommendation information identifier, physical distance (distance d from nearest store).
[0056] In some embodiments of the present invention, in order to improve the accuracy of the shallow conversion rate calculated by the information recommendation model, the original training samples can be adjusted. For example, the type, content, and source of the recommendation information can be added; or object information such as the target object's age, basic information, and hobby information can be added. This application does not impose specific limitations in this regard.
[0057] Step 202: Randomly generate a perturbation factor corresponding to the physical distance, and randomly adjust the physical distance according to the perturbation factor to obtain the target distance.
[0058] Step 203: Replace the physical distance in the original training samples with the target distance to obtain the target training samples.
[0059] Since the physical distance between the target object and the entity to which the corresponding recommendation information belongs directly affects the shallow conversion rate of the target object, for example, if the distance between the store location and the target object is too large, the target object will give up on shallow conversion after receiving the recommendation information, let alone deep conversion. Therefore, when training the information recommendation model in this application, target training samples need to be used to increase the distance sensitivity of the information recommendation model. In some embodiments of the present invention, there are two optional ways to adjust the physical distance in the original training samples: 1) randomly adjust the original training samples; 2) fix the original training samples according to a preset specific rule, which will be described below.
[0060] In some embodiments of the present invention, the original training samples are randomly adjusted, which can be achieved by: randomly generating a perturbation factor corresponding to the physical distance; randomly adjusting the physical distance according to the perturbation factor to obtain the target distance; and replacing the physical distance in the original training samples with the target distance to obtain the target training samples. Specifically, the nearest store distance d, which is the physical distance in the original training samples, is randomly perturbed, and a perturbation factor factor is randomly generated within the selectable value range [0.5, 1.5]. Then, the physical distance in the generated target training samples is factor*d. Assuming the generated perturbation factors are 0.8 and 1.2 respectively (if factor>1, the expected probability of arrival at the store output by the recommendation model should be lower than the probability of arrival at the store in the original samples; if factor<1, the expected probability of arrival at the store output by the recommendation model should be higher than the probability of arrival at the store in the original samples), the following two target training samples are generated: Target training sample 1: Sample label (0 / 1), shallow conversion action of the target object (target object visits the store), information identifier, target distance (distance to the nearest store is 0.8d).
[0061] Target training sample 2: Sample label (0 / 1), shallow conversion action of the target object (target object visits the store), information identifier, target distance (distance to the nearest store is 1.2d).
[0062] In some embodiments of the present invention, when the original training samples are fixedly adjusted according to a preset specific rule, a distance enhancement parameter can be preset. A selectable distance enhancement parameter is 1.5. The physical distance is adjusted according to the fixed adjustment rule to obtain the target distance. By replacing the physical distance in the original training samples with the target distance, the following target training samples are generated: Target training sample 3: Sample label (0 / 1), shallow conversion action of the target object (target object visits the store), information identifier, target distance (distance to the nearest store is 1.5d).
[0063] Through the processing in step 203, the target training samples can be obtained directly without the need for the recommendation information delivery entity to send back data, thus saving the training time and cost of obtaining training samples for the multimedia information recommendation model.
[0064] Step 204: Train the information recommendation model by combining the original training samples and the target training samples.
[0065] Among them, the information recommendation model is used to recommend information to the target object based on the physical distance between the target object and the entity to which the information to be recommended belongs.
[0066] refer to Figure 3 , Figure 3 This is a schematic diagram of the training logic of the information recommendation model in an embodiment of the present invention, such as... Figure 3 The model shown first maps the initial training samples obtained in step 201 to 64-dimensional vectors by using an embedding layer. All vectors are then concatenated into a long vector of length m*64, where m is the number of feature fields. After processing by the information recommendation model, the predicted shallow conversion rate is output. Simultaneously, the target training samples generated in step 202 are used to train the information recommendation model, outputting the shallow conversion rate corresponding to the enhanced samples. This process uses the BPR loss function as the loss for contrastive learning. The total loss function = w * main loss function + (1-w) * contrastive learning loss function, where w is a hyperparameter between (0, 1).
[0067] The following is combined with Figure 3 The training logic shown illustrates the training process of the information recommendation model provided in this application. (Refer to...) Figure 4 , Figure 4 This is a schematic diagram illustrating the training process of the information recommendation model in an embodiment of the present invention. It can be understood that... Figure 4 The steps shown can be implemented by a terminal or server running a training device with an information recommendation model, or by a terminal and server working together, or by a cloud server or a cloud server cluster working together. Here, we will take the server implementation as an example, which includes the following steps: Step 401: Train the information recommendation model using the original training samples to determine the first model parameters of the information recommendation model.
[0068] In step 401, the main loss function of the information recommendation model is first determined; then, the information recommendation model is trained using the original training samples and the main loss function; finally, when the main loss function reaches the corresponding convergence condition, the first model parameters of the information recommendation model are determined. The main loss function of the information recommendation model can use the cross-entropy loss function. For example, taking the shallow conversion rate of the recommendation information to the target object as a multi-class probability as an example, the main loss function can adopt the category cross-entropy loss function as shown in formula (1): (1); Where n is the index of the training sample, N is the total number of training samples, and c is the class index of the multi-class probability. This indicates whether the nth training sample belongs to the cth category, with 0 representing no and 1 representing yes. It represents the probability of the c-th category obtained after training the n-th set of information recommendation inputs.
[0069] By executing step 401, the information recommendation model can process the original training samples to obtain the corresponding shallow conversion rate. However, since the training samples lack perturbation of physical distance, the information recommendation model is not sensitive to distance information at this time and cannot recommend information to the target object based on the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs.
[0070] Step 402: Train the information recommendation model using the target training samples, adjust the first model parameters, and obtain the second model parameters of the information recommendation model.
[0071] In some embodiments of the present invention, the information recommendation model is trained using target training samples, which can be achieved in the following ways: Determine the contrastive learning loss function for the information recommendation model; train the information recommendation model using the target training samples, the first model parameters, and the contrastive learning loss function; when the contrastive learning loss function reaches the corresponding convergence condition, adjust the first model parameters to obtain the second model parameters of the information recommendation model. Specifically, the first model parameters of the information recommendation model can be shared during the training process using the target training samples. At this point, the contrastive learning loss function of the information recommendation model is first determined; the contrastive learning loss function can use Bayesian Personalized Ranking, and the function reference is (Formula 2): (Formula 2) Where si represents the positive sample in the target training sample, and sj represents the negative sample in the target training sample.
[0072] The information recommendation model is trained using target training samples, first model parameters, and a contrastive learning loss function. When the contrastive learning loss function reaches the corresponding convergence condition, the first model parameters are adjusted to obtain the second model parameters of the information recommendation model. At this point, the information recommendation model trained with the target training samples has acquired distance sensitivity.
[0073] Step 403: Train the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model.
[0074] Through training in steps 401 and 402, the information recommendation model has acquired distance sensitivity. To ensure that the model can recommend information that meets both the nearest distance requirement and the target object's usage needs, the loss functions used in training steps 401 and 402 can be combined to further train the model. Specifically, this includes the following steps: First, calculate the total loss function of the information recommendation model based on the main loss function and the contrastive learning loss function; Total loss function = w * main loss function + (1-w) * contrastive learning loss function, where w is a hyperparameter between (0, 1). Refer to Formula 3: (Formula 3) The information recommendation model is trained based on the second model parameters, the target training samples, and the total loss function of the information recommendation model. When the total loss function reaches the corresponding convergence condition, the second model parameters are adjusted to obtain the model parameters of the information recommendation model. Thus, through steps 401-403, the training of the information recommendation model using initial training samples and target training samples can be completed. The trained information recommendation model can not only accurately calculate the shallow conversion rate of different recommended information for target objects, but also, due to the use of target training samples during training, the information recommendation model has better distance sensitivity, and can recommend recommended information delivery entities that are physically closer to the target object, thereby guiding the target object to perform the corresponding shallow conversion action.
[0075] Once the information recommendation model is trained, it can be deployed on a server to recommend information to different target audiences. (Refer to...) Figure 5 , Figure 5 This is a schematic diagram illustrating the information recommendation process using an information recommendation model in an embodiment of the present invention. It can be understood that... Figure 5The steps shown can be implemented by a terminal or server running a training device with an information recommendation model, or by a terminal and server working together, or by a cloud server or a cloud server cluster working together. Here, we will take the server implementation as an example, which includes the following steps: Step 501: Obtain the information to be recommended from the recommendation information data source.
[0076] Step 502: Predict different information to be recommended using an information recommendation model to determine the shallow conversion rate of different information to be recommended.
[0077] Step 503: Adjust the recommendation order of the information to be recommended according to the shallow conversion rate of different information to be recommended, and recommend information according to the recommendation order.
[0078] For example, consider two pieces of information to be recommended: a local lifestyle video ad A (the physical distance between the target audience and the entity to which video ad A belongs is 100m) and ad B (the physical distance between the target audience and the entity to which video ad A belongs is 200m). If the information recommendation model provided in this application determines that the shallow conversion rate of video ad A is 0.8 and that of video ad B is 0.5, it indicates that video ad A has a higher shallow conversion rate, meaning the target audience is more likely to reach the location of the entity to which video ad A belongs. Therefore, the recommendation order is adjusted, prioritizing video ad A and allocating more playback traffic to ad A to increase its exposure and provide a better viewing experience for users, thereby increasing the trigger rate of ad A.
[0079] In some embodiments of the present invention, when performing step 503, the recommendation order of the recommended information is adjusted according to the shallow conversion rate of different recommended information, and when recommending information through the recommendation order, the recommendation order of the information can be adjusted according to the adjusted recommendation order.
[0080] In some embodiments of the present invention, see Figure 6 , Figure 6 This is a schematic diagram of an optional information recommendation in an embodiment of the present invention, wherein the category of the information to be recommended can be determined; in response to the category of the information to be recommended, a matching information data source is triggered. For example, when it is determined that the category of the information to be recommended is local life advertising information, the trained information recommendation model is triggered. However, the information recommendation model provided in this application is not applicable to the recommendation of video information, short video recommendations, and online shopping product recommendations.
[0081] like Figure 6As shown, taking local lifestyle ads as an example, video ads from different entities in the same industry can be played sequentially in different ad slot video ad playback windows (for example, ad slots 1, 2, and 3 play video ads A, B, and C from the catering industry, respectively). Alternatively, when all different ad slot video ad playback areas on the display interface are occupied by the same entity, the same video ad can be displayed in a loop, that is, ad slots 1, 2, and 3 will cycle through video ad A from the same entity. Figure 6 The advertised ad slots 1, 2, and 3 shown are... Figure 6 The advertisement is played sequentially from right to left to maximize its exposure.
[0082] Using the information recommendation model provided in this application, it is determined that for the same target audience, video ad A has a shallow conversion rate of 0.5, video ad B has a shallow conversion rate of 0.8, and video ad C has a shallow conversion rate of 1. Therefore, the recommendation order of video ads is adjusted according to the shallow conversion rates of different video ads to obtain better video ad recommendation results. Figure 6 For example, when adjusting the recommendation order using the information recommendation model provided in this application, it is possible to... Figure 6 When recommending different ads to users viewing ads in ad slots 1, 2, and 3, the recommendation order of the ad videos will be adjusted to ad video C, ad video B, and ad video A. That is, ad slot 1 displays ad C, ad slot 2 displays ad B, and ad slot 3 displays ad A. This ensures that users see video ads from entities closer to them, while increasing the shallow conversion rate of ads to achieve better ad placement results.
[0083] In some embodiments of the present invention, such as Figure 6 As shown, when all different video ad playback areas in the display interface are contracted by different placement entities, the same video ad can be displayed in a loop. For example, if ad slots 1, 2, and 3 have 100, 100, and 70 exposures respectively, and the information recommendation model provided in this application determines that for the same target audience, video ad A has a shallow conversion rate of 0.5, video ad B has a shallow conversion rate of 1, and video ad C has a shallow conversion rate of 1, then since the shallow conversion rates of video ad B and video ad C are the same or equal, the following two playback orders are possible: 1) For ad slots 1 and 2 with the same number of exposures, ad slot 1 can play video ad B, ad slot 2 can play video ad C, and ad slot 3 can play video ad A.
[0084] 2) Video ad B is played in ad slot 2, video ad C is played in ad slot 1, and video ad A is played in ad slot 3.
[0085] when Figure 6 When all the different video ad playback areas in the display interface shown are contracted by the same placement entity, the placement entities of video ad B and video ad C will bid to determine the playback content of ad slot 1, ad slot 2 and ad slot 3.
[0086] To better illustrate the information recommendation model training method provided in this application, the following uses an advertisement as an example to explain the working process of the information recommendation model training method provided in this application. The inventors have found the following defects in related technologies when performing advertisement recommendations: When recommending ads using information recommendation models, for local lifestyle ads, user in-store consumption is directly affected by the distance between the user and the advertiser's store. Specifically, in the local lifestyle advertising industry (such as restaurant ads, beauty salon ads, entertainment ads, and offline education ads), the advertiser's business goal requires users to ultimately make a purchase in the store. Therefore, many advertisers' advertising targets users to make in-store purchases. The conversion behavior of users making in-store purchases caused by the advertiser's ads within the advertising domain can be called shallow conversion.
[0087] In this scenario, the distance between the advertiser's store and the user becomes a crucial factor in whether the user ultimately visits the store. The inventors discovered that current local lifestyle post-link optimization advertising algorithms have the following shortcomings: 1) Some advertisers reported that users did not ultimately visit the store after clicking on the ad, affecting the final effect of the ad campaign and consequently impacting the advertiser's advertising budget, thus hindering the increase of the ad campaign's click-through rate. 2) Because the information recommendation model does not constrain distance sensitivity, users who click on local lifestyle ads may be unable to visit the store due to its distance or have a very low willingness to do so, negatively impacting the user experience.
[0088] To address the aforementioned shortcomings, this application provides a method for training information recommendation models, which makes the models more distance-sensitive. See [link to relevant documentation] Figure 7 , Figure 7 This is an optional flowchart illustrating the information recommendation model training method provided in this embodiment of the invention. It can be understood that... Figure 7 The steps shown can be performed by various electronic devices running the information recommendation model training device, such as a dedicated terminal, server, or server cluster with an information recommendation model training device. The dedicated terminal with the information recommendation model training device can be the electronic device with the information recommendation model training device shown in the previous embodiments. The following focuses on... Figure 7 The steps shown are explained.
[0089] Step 701: The information recommendation model training device acquires historical data of the target object, wherein the historical data is used to record the shallow conversion behavior of the target object based on the advertising information.
[0090] In some embodiments of the present invention, historical data of the target object can be obtained in the following ways: The process involves identifying the identifiers of advertising information; determining the number of times a target audience successfully made a shallow conversion using the advertising information based on these identifiers; and determining the number of times a target audience failed to make a shallow conversion using the advertising information based on these identifiers. Historical data can be the sum of data from various industry advertising recommendations, such as product recommendations, ad recommendations, e-commerce ad recommendations, and offline education ad recommendations, encompassing the sum of all basic data across multiple advertising recommendation environments. When acquiring historical data, effective extraction can be achieved from raw logs of user usage data, such as extracting the user's device ID (user account), advertising type, ad viewing duration, ad recommendation environment, advertising identifier, the number of times a target audience successfully made a shallow conversion using the advertising information (e.g., making a purchase based on a restaurant ad), and the number of times a target audience failed to make a shallow conversion using the advertising information (e.g., receiving a restaurant ad but not making a purchase).
[0091] It is understood that in the specific implementation of this application, user-related data such as historical data and industry historical data in the media information recommendation environment are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0092] Step 702: The information recommendation model training device determines the original training samples that match the information recommendation model based on the historical data of the target object.
[0093] In some embodiments of the present invention, determining the original training samples that match the information recommendation model based on the historical data of the target object can be achieved by: determining the original distance parameters in the original training samples based on the identifier of the advertising information; Based on the number of times the target audience successfully made a shallow conversion using advertising information, the original distance parameter, and the ad information identifier, positive examples are determined in the original training sample. Based on the number of times the target audience failed to make a shallow conversion using advertising information, the original distance parameter, and the ad information identifier, negative examples are determined in the original training sample. The positive and negative examples are then combined to obtain the original training sample. For example, when recommending local lifestyle ads using an information recommendation model, users who received a recommended ad in the past 14 days and made a shallow conversion but ultimately did not visit the store can be extracted as negative samples and marked as 0. In the user's feedback data, users who visited the store after receiving the recommended ad can be used as positive samples and marked as 1. The ad identifier (id) and the distance parameter (distance to the nearest store of the advertiser) are added. The final original training sample can be represented as: Sample label (0 / 1), user shallow conversion action (user visits the store), advertising identifier, distance parameter (distance d from the nearest store).
[0094] Step 703: The information recommendation model training device randomly perturbs the original distance parameters in the original training samples to obtain the result.
[0095] In some embodiments of the present invention, the original distance parameters in the original training samples are randomly perturbed to obtain distance-enhanced training samples, which can be achieved in the following ways: A perturbation factor corresponding to the original distance parameter is randomly generated. Based on the perturbation factor, the original distance parameter is randomly perturbed to obtain the target distance parameter. The original distance parameter in the original training sample is replaced with the target distance parameter to obtain the distance-enhanced training sample. Specifically, the nearest store distance *d*, which is used as a distance parameter in the original training sample, is randomly perturbed, and a perturbation factor *factor* is randomly generated within the range [0.5, 1.5]. Therefore, the distance of the distance parameter *d* in the generated distance-enhanced training sample is *factor*d. Assuming the generated *factor* is 0.8, the following distance-enhanced training sample is generated: Sample label (0 / 1), user shallow conversion action (user visits the store), advertising identifier, distance parameter (distance to the nearest store is 0.8d).
[0096] Step 704: The information recommendation model training device trains the information recommendation model using the original training samples and distance-enhanced training samples to obtain the model parameters of the information recommendation model.
[0097] refer to Figure 8 , Figure 8 This is a schematic diagram of the training logic of the information recommendation model in an embodiment of the present invention, such as... Figure 8As shown, for the same user 'a' and the same advertisement identifier 'b', the information recommendation model first maps the data fields of the identifier type into 64-dimensional vectors through an embedding layer on the initial training samples generated in step 702, and concatenates the vectors corresponding to each identifier into a long vector of length m*64, where m is the number of feature fields. After processing by the information recommendation model, it outputs the predicted store arrival probability. At the same time, the distance enhancement training samples (distance factor*d) generated in step 703 are also used to train the information recommendation model, outputting the store arrival probability of the enhanced samples (if factor>1, it is expected that the output store arrival probability should be lower than the store arrival probability of the original samples; if factor<1, it is expected that the output store arrival probability should be higher than the store arrival probability of the original samples). The BPR loss function is used as the loss for contrastive learning. The total loss function = w*main loss function + (1-w)*contrastive learning loss function, where w is a hyperparameter between (0,1).
[0098] Figure 9 This is a schematic diagram of the training process of the information recommendation model in this embodiment of the invention, which specifically includes the following steps: Step 901: The information recommendation model training device trains the information recommendation model using the original training samples to determine the first model parameters of the information recommendation model.
[0099] In some embodiments of the present invention, the first model parameters of the information recommendation model can be determined in the following ways: Determine the main loss function of the information recommendation model; train the information recommendation model using the original training samples and the main loss function; when the main loss function reaches the corresponding convergence condition, determine the first model parameters of the information recommendation model.
[0100] The main loss function used is the cross-entropy loss function.
[0101] For example, the weighted cross-entropy loss function can be used (Formula 4): (Formula 4) in This represents the input of the k-th training sample. This represents the estimated probability of the k-th vector sample (e.g., 0 or 1 for binary classification). This represents the actual probability of the k-th vector sample. This represents the weight of the training samples (for local lifestyle ads, the weights of positive and negative training samples can be the same).
[0102] Step 902: The information recommendation model training device trains the information recommendation model using distance-enhanced training samples, adjusts the first model parameters, and obtains the second model parameters of the information recommendation model.
[0103] In some embodiments of the present invention, adjusting the first model parameters to obtain the second model parameters of the information recommendation model can be achieved in the following ways: Determine the contrastive learning loss function for the information recommendation model; train the information recommendation model using distance-enhanced training samples, the first model parameters, and the contrastive learning loss function; when the contrastive learning loss function reaches the corresponding convergence condition, adjust the first model parameters to obtain the second model parameters for the information recommendation model.
[0104] The contrastive learning loss function can be Bayesian Personalized Ranking, and the function reference is (Formula 5): (Formula 5) Where si represents the positive samples in the distance augmentation training samples, and sj represents the negative samples in the distance augmentation training samples. For the sigmoid function, This is a penalty term used to prevent overfitting in information recommendation models.
[0105] Step 903: The information recommendation model training device trains the information recommendation model based on the second model parameters, distance augmentation training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model.
[0106] In some embodiments of the present invention, training the information recommendation model to obtain the model parameters can be achieved in the following ways: The total loss function of the information recommendation model is calculated based on the main loss function and the contrastive learning loss function. The information recommendation model is then trained based on the second model parameters, distance-enhanced training samples, and the total loss function of the information recommendation model. When the total loss function reaches the corresponding convergence condition, the second model parameters are adjusted to obtain the model parameters of the information recommendation model.
[0107] Wherein, the total loss function = w * main loss function + (1-w) * contrastive learning loss function, where w is a hyperparameter between (0, 1). See (Formula 6): (Formula 6) The trained information recommendation model can recommend local lifestyle ads placed by advertisers. For example, for ads placed by advertisers whose scope is local lifestyle, the trained information recommendation model can predict the probability of in-store conversion (shallow conversion rate) generated by the ad.
[0108] For example, consider two video ads, A and B, related to the food and beverage industry. If the recommendation model provided in this application determines that the probability of ad A reaching a customer's store is 0.5 and that of ad B is 1, then ad B has a higher priority than ad A. This indicates that the user is likely more interested in video ad B, and the store associated with ad B is closer to the user. Therefore, based on the probability of reaching a customer's store, ad B is recommended to the user first, and more playback bandwidth is allocated to ad B to increase its exposure and provide a better viewing experience for the user. This, in turn, increases the likelihood of shallow conversions generated by ad B and improves the success rate of shallow conversions.
[0109] Therefore, this application obtains historical data of the target object, which records the shallow conversion behavior of the target object based on advertising information; determines original training samples that match the information recommendation model based on the historical data of the target object; randomly perturbs the original distance parameters in the original training samples to obtain distance-enhanced training samples; and trains the information recommendation model using the original training samples and the distance-enhanced training samples to obtain the model parameters of the information recommendation model. This improves the distance sensitivity of the information recommendation model, resulting in a higher success rate of shallow conversions for advertising information recommendations based on the information recommendation model, which can be improved by at least 3%. Simultaneously, randomly perturbing the original distance parameters in the original training samples directly yields distance-enhanced training samples, eliminating the need for data feedback from the advertising information publisher, saving training time and data acquisition costs for the information recommendation model, and improving the user experience.
[0110] See Figure 10 , Figure 10 This is a schematic diagram of the structure of an electronic device 100 for executing the training method of the information recommendation model provided in this application, according to an embodiment of this application. Figure 10 The illustrated electronic device 100 includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components of the electronic device 100 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 10 The general labeled all buses as Bus System 440.
[0111] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0112] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0113] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0114] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0115] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0116] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with user interface 430 (e.g., a display screen, a speaker, etc.). The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0117] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 10 A training device 455 for an information recommendation model stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an information transmission module 4551 and an information processing module 4552. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0118] The following section further explains the functions of each software module in the training device for the information recommendation model. The information transmission module 4551 is used to acquire the original training samples of the information recommendation model; The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample labels carried by the original training samples are: the shallow conversion rate of the recommendation information for the target object; and the shallow conversion rate, which is used to indicate the probability that the target object will reach the location of the delivery entity. Information processing module 4552 is used to randomly generate a disturbance factor corresponding to the physical distance; The information processing module 4552 is used to randomly adjust the physical distance according to the disturbance factor to obtain the target distance; Information processing module 4552 is used to replace the physical distance in the original training sample with the target distance to obtain the target training sample; The information processing module 4552 is used to train the information recommendation model by combining the original training samples and the target training samples; Among them, the information recommendation model is used to recommend information to the target object based on the physical distance between the target object and the entity to which the information to be recommended belongs.
[0119] In some embodiments, the information transmission module 4551 is used to obtain the identifier of the recommendation information; The information transmission module 4551 is used to determine the shallow conversion action of the recommendation information for the target object based on the identifier of the recommendation information.
[0120] In some embodiments, the information transmission module 4551 is used to determine the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, based on the identifier of the recommendation information. The information transmission module 4551 is used to determine the positive samples in the original training samples based on the sample labels, the shallow conversion actions of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The information transmission module 4551 is used to determine the negative samples in the original training samples based on the sample labels, the shallow conversion actions of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The information transmission module 4551 is used to combine positive and negative samples to obtain the original training samples.
[0121] In some embodiments, the information processing module 4552 is used to randomly generate a disturbance factor corresponding to the physical distance; The information processing module 4552 is used to randomly adjust the physical distance according to the disturbance factor to obtain the target distance, or; Information processing module 4552 is used to obtain fixed adjustment rules corresponding to physical distance; The information processing module 4552 is used to adjust the physical distance according to a fixed adjustment rule to obtain the target distance; The information processing module 4552 is used to replace the physical distance in the original training samples with the target distance to obtain the target training samples.
[0122] In some embodiments, the information processing module 4552 is used to train the information recommendation model using the original training samples and determine the first model parameters of the information recommendation model; The information processing module 4552 is used to train the information recommendation model using target training samples, adjust the first model parameters, and obtain the second model parameters of the information recommendation model. The information processing module 4552 is used to train the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model, so as to obtain the model parameters of the information recommendation model.
[0123] In some embodiments, the information processing module 4552 is used to determine the main loss function of the information recommendation model; The information processing module 4552 is used to train the information recommendation model using the original training samples and the main loss function; The information processing module 4552 is used to determine the first model parameters of the information recommendation model when the main loss function reaches the corresponding convergence condition.
[0124] In some embodiments, the information processing module 4552 is used to determine the contrastive learning loss function of the information recommendation model; The information processing module 4552 is used to train the information recommendation model using the target training samples, the first model parameters, and the contrastive learning loss function. The information processing module 4552 is used to adjust the first model parameters to obtain the second model parameters of the information recommendation model when the contrastive learning loss function reaches the corresponding convergence condition.
[0125] In some embodiments, the information processing module 4552 is used to calculate the total loss function of the information recommendation model based on the main loss function and the contrastive learning loss function; The information processing module 4552 is used to train the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model. The information processing module 4552 is used to adjust the parameters of the second model when the total loss function reaches the corresponding convergence condition, so as to obtain the model parameters of the information recommendation model.
[0126] See Figure 11 , Figure 11 This is a schematic diagram of the structure of an electronic device 1000 for performing the information recommendation method provided in this application, according to an embodiment of this application. Figure 11 The illustrated electronic device 1000 includes at least one processor 1410, a memory 1450, at least one network interface 1420, and a user interface 1430. The various components in the electronic device 1000 are coupled together via a bus system 1440. It is understood that the bus system 1440 is used to implement communication between these components. In addition to a data bus, the bus system 1440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 11 The general labeled all buses as Bus System 1440.
[0127] The processor 1410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0128] User interface 1430 includes one or more output devices 1431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 1430 also includes one or more input devices 1432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0129] The memory 1450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 1450 may optionally include one or more storage devices physically located away from the processor 1410.
[0130] The memory 1450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 1450 described in this application embodiment is intended to include any suitable type of memory.
[0131] In some embodiments, memory 1450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0132] Operating system 1451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 1452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 1420, exemplary network interfaces 1420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 1453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 1431 (e.g., a display screen, a speaker, etc.) associated with user interface 1430. The input processing module 14514 is used to detect and translate one or more user inputs or interactions from one or more input devices 1432.
[0133] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 11 An information recommendation device 1455 stored in memory 1450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an information transmission module 14551 and an information processing module 14552. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0134] The functions of each software module in the information recommendation device will be explained below. Data transmission module 14551 is used to obtain the information to be recommended from the recommendation information data source; Data processing module 14552 is used to predict different information to be recommended through information recommendation model and determine the shallow conversion rate of different information to be recommended. Data processing module 14552 is used to adjust the recommendation order of information based on the shallow conversion rate of different information to be recommended, and to recommend information based on the recommendation order. The information recommendation model is based on... Figure 3 The method for training the information recommendation model shown is obtained.
[0135] The above are merely embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A training method for an information recommendation model, characterized in that, The method includes: Obtain the original training samples for the information recommendation model; The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample label carried by the original training samples is the shallow conversion rate of the recommendation information for the target object; the shallow conversion rate is used to indicate the probability that the target object will generate a shallow conversion action in response to the recommendation information. A perturbation factor corresponding to the physical distance is randomly generated; the perturbation factor is within a preset value range. The physical distance is randomly adjusted according to the disturbance factor to obtain the target distance; The target training sample is obtained by replacing the physical distance in the original training sample with the target distance; The information recommendation model is trained using the original training samples to determine the first model parameters of the information recommendation model; The information recommendation model is trained using the target training samples, and the first model parameters are adjusted to obtain the second model parameters of the information recommendation model. The information recommendation model is trained based on the second model parameters, the target training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model. Wherein, if the perturbation factor is greater than 1, it is expected that the probability of arriving at the store output by the information recommendation model is lower than the probability of arriving at the store in the original training sample; if the perturbation factor is less than 1, it is expected that the probability of arriving at the store output by the information recommendation model is higher than the probability of arriving at the store in the original training sample; the information recommendation model is used to recommend the information to be recommended to the target object based on the physical distance between the target object and the delivery entity to which the information to be recommended belongs.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the identifier of the recommendation information; Based on the identifier of the recommendation information, determine the shallow conversion action of the recommendation information for the target object.
3. The method according to claim 2, characterized in that, The original training samples of the information acquisition recommendation model include: Based on the identifier of the recommendation information, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs is determined; For a target object that arrives at the location of the delivery entity after receiving the recommendation information, positive samples in the original training samples are determined based on the sample label, the shallow conversion action of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. For a target object that undergoes a shallow conversion after receiving the recommendation information but ultimately fails to reach the location of the delivery entity, negative sample samples in the original training samples are determined based on the sample label, the shallow conversion action of the recommendation information on the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The positive and negative samples are combined to obtain the original training samples.
4. The method according to claim 1, characterized in that, The method further includes: Obtain the fixed adjustment rules corresponding to the physical distance; The physical distance is adjusted according to the fixed adjustment rules to obtain the target distance.
5. The method according to claim 1, characterized in that, The step of training the information recommendation model using the original training samples to determine the first model parameters of the information recommendation model includes: Determine the main loss function of the information recommendation model; The information recommendation model is trained using the original training samples and the main loss function. When the main loss function reaches the corresponding convergence condition, the first model parameters of the information recommendation model are determined.
6. The method according to claim 5, characterized in that, The step of training the information recommendation model using the target training samples and adjusting the first model parameters to obtain the second model parameters of the information recommendation model includes: Determine the contrastive learning loss function for the information recommendation model; The information recommendation model is trained using the target training samples, the first model parameters, and the contrastive learning loss function. When the contrastive learning loss function reaches the corresponding convergence condition, the first model parameters are adjusted to obtain the second model parameters of the information recommendation model.
7. The method according to claim 6, characterized in that, The step of training the information recommendation model based on the second model parameters, the target training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model includes: Calculate the total loss function of the information recommendation model based on the main loss function and the contrastive learning loss function; The information recommendation model is trained based on the second model parameters, the target training samples, and the total loss function of the information recommendation model. When the total loss function reaches the corresponding convergence condition, the second model parameters are adjusted to obtain the model parameters of the information recommendation model.
8. An information recommendation method, characterized in that, The method includes: Retrieve the information to be recommended from the data source of the recommendation information; The shallow conversion rate of different recommended information is determined by predicting different recommended information using an information recommendation model. The recommendation order of the information to be recommended is adjusted according to the shallow conversion rate of different information to be recommended, and information is recommended through the recommendation order, wherein the information recommendation model is trained based on any one of claims 1-7.
9. An information recommendation model training device, characterized in that, The device includes: The information transmission module is used to acquire the original training samples of the information recommendation model; The original training samples include: the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs; the sample label carried by the original training samples is the shallow conversion rate of the recommendation information for the target object; the shallow conversion rate is used to indicate the probability that the target object will generate a shallow conversion action in response to the recommendation information. An information processing module is used to randomly generate a disturbance factor corresponding to the physical distance; the disturbance factor is within a preset value range. The information processing module is also used to randomly adjust the physical distance according to the disturbance factor to obtain the target distance; The information processing module is also used to replace the physical distance in the original training sample with the target distance to obtain the target training sample; The information processing module is further configured to train the information recommendation model using the original training samples and determine the first model parameters of the information recommendation model. The information recommendation model is trained using the target training samples, and the first model parameters are adjusted to obtain the second model parameters of the information recommendation model. The information recommendation model is trained based on the second model parameters, the target training samples, and the total loss function of the information recommendation model to obtain the model parameters of the information recommendation model. Wherein, if the perturbation factor is greater than 1, it is expected that the probability of arriving at the store output by the information recommendation model is lower than the probability of arriving at the store in the original training sample; if the perturbation factor is less than 1, it is expected that the probability of arriving at the store output by the information recommendation model is higher than the probability of arriving at the store in the original training sample; the information recommendation model is used to recommend the information to be recommended to the target object based on the physical distance between the target object and the delivery entity to which the information to be recommended belongs.
10. The apparatus according to claim 9, characterized in that, The information transmission module is also used to obtain the identifier of the recommendation information; Based on the identifier of the recommendation information, determine the shallow conversion action of the recommendation information for the target object.
11. The apparatus according to claim 10, characterized in that, The information transmission module is also used to determine the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, based on the identifier of the recommendation information; For a target object that arrives at the location of the delivery entity after receiving the recommendation information, positive samples in the original training samples are determined based on the sample label, the shallow conversion action of the recommendation information for the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. For a target object that undergoes a shallow conversion after receiving the recommendation information but ultimately fails to reach the location of the delivery entity, negative sample samples in the original training samples are determined based on the sample label, the shallow conversion action of the recommendation information on the target object, the physical distance between the target object and the delivery entity to which the corresponding recommendation information belongs, and the identifier of the recommendation information. The positive and negative samples are combined to obtain the original training samples.
12. The apparatus according to claim 9, characterized in that, The information processing module is also used to obtain fixed adjustment rules corresponding to the physical distance; The physical distance is adjusted according to the fixed adjustment rules to obtain the target distance.
13. The apparatus according to claim 9, characterized in that, The information processing module is also used to determine the main loss function of the information recommendation model; The information recommendation model is trained using the original training samples and the main loss function. When the main loss function reaches the corresponding convergence condition, the first model parameters of the information recommendation model are determined.
14. The apparatus according to claim 13, characterized in that, The information processing module is also used to determine the contrastive learning loss function of the information recommendation model; The information recommendation model is trained using the target training samples, the first model parameters, and the contrastive learning loss function. When the contrastive learning loss function reaches the corresponding convergence condition, the first model parameters are adjusted to obtain the second model parameters of the information recommendation model.
15. The apparatus according to claim 14, characterized in that, The information processing module is further configured to calculate the total loss function of the information recommendation model based on the main loss function and the contrastive learning loss function; The information recommendation model is trained based on the second model parameters, the target training samples, and the total loss function of the information recommendation model. When the total loss function reaches the corresponding convergence condition, the second model parameters are adjusted to obtain the model parameters of the information recommendation model.
16. An information recommendation device, characterized in that, The device includes: The data transmission module is used to obtain the information to be recommended from the recommendation information data source; The data processing module is used to predict different information to be recommended using an information recommendation model, and to determine the shallow conversion rate of different information to be recommended. The data processing module is further configured to adjust the recommendation order of the information to be recommended according to the shallow conversion rate of different information to be recommended, and to recommend information through the recommendation order, wherein the information recommendation model is trained based on any one of claims 1-7.
17. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the information recommendation model training method according to any one of claims 1 to 7, or the information recommendation method according to claim 8.
18. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; When a processor executes executable instructions stored in the memory, it implements the information recommendation model training method according to any one of claims 1 to 7, or the information recommendation method according to claim 8.
19. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the information recommendation model training method according to any one of claims 1 to 7, or the information recommendation method according to claim 8.