Information processing method and device, equipment and storage medium
By using machine learning models to generate predicted user feature representations in the content recommendation system, the recommendation problems in new content, new users and sparse interactive data scenarios are solved, effective exposure and feedback of new content is achieved, and the overall performance of the recommendation system is improved.
Patent Information
- Application Number
- CN202510389158.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The content recommendation system is difficult to effectively recommend new content, new users or sparse interactive data scenarios, resulting in insufficient exposure of new content and lack of feedback, forming a vicious cycle.
By in response to the user sequence associated with the recommendation object, the input feature sequence of the machine learning model is determined based on the object feature representation of the recommendation object and the user feature representation in the user sequence, the predicted user feature representation is generated, the recommendation request initiates the user's user feature representation, and the recommendation object is provided to the user according to the similarity requirements.
It realizes effective user expansion of new content in cold start scenarios, improves the exposure and feedback of new content by the recommendation system, enhances the enthusiasm of content creators, and improves the fairness, interactivity and diversity of the recommendation system.
Smart Images

Figure CN120234458A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to an information processing method, apparatus, device, and storage medium. Background Art
[0002] Content recommendation is a personalized information distribution technology. The cold start problem of content recommendation is one of the challenges in introducing new content (such as articles, products, videos, etc.), new users, or processing sparse interaction data in a recommendation system. For example, when new content is first added to a content recommendation system, due to the lack of sufficient user interaction data (such as clicks, views, likes, etc.), it is difficult for the content recommendation system to effectively recommend the new content to target users. This results in the new content having difficulty obtaining user feedback due to insufficient exposure, and the lack of feedback further compresses its exposure opportunities, thus forming a vicious cycle. Summary of the Invention
[0003] In a first aspect of the present disclosure, an information processing method is provided. The method includes: in response to the existence of a user sequence associated with a recommendation object, determining an input feature sequence for a machine learning model based on at least one object feature representation of the recommendation object and at least one first user feature representation corresponding to at least one user in the user sequence, where the user sequence is determined based on users who have interacted with the recommendation object; generating, based on the input feature sequence, a predicted user feature representation for the recommendation object using the machine learning model, where the predicted user feature representation is used to indicate users who will interact with the recommendation object in the future; in response to receiving a recommendation request, extracting a second user feature representation corresponding to the initiating user of the recommendation request; and in response to the similarity between the second user feature representation and the predicted user feature representation meeting a similarity requirement, providing the recommendation object to the initiating user of the recommendation request.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for information processing. The apparatus includes: an input feature sequence determination module configured to, in response to the presence of a user sequence associated with a recommendation object, determine an input feature sequence for a machine learning model based on at least one object feature representation of the recommendation object and at least one first user feature representation corresponding to at least one user in the user sequence, where the user sequence is determined based on users who have interacted with the recommendation object; a predicted user feature representation generation module configured to generate, based on the input feature sequence and using the machine learning model, a predicted user feature representation for the recommendation object, where the predicted user feature representation is used to indicate a user who will interact with the recommendation object at a future time; a user feature representation extraction module configured to, in response to receiving a recommendation request, extract a second user feature representation corresponding to the initiating user of the recommendation request; and a recommendation object providing module configured to, in response to the similarity between the second user feature representation and the predicted user feature representation meeting the similarity requirement, provide the recommendation object to the initiating user of the recommendation request.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory, where the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. Computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, there is provided a computer program product. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0008] Embodiments of the present disclosure can be applied to video-oriented generative recommendations in cold start scenarios. It should be understood that the content described in this part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A flowchart showing an example process of an information processing method according to some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram showing an example of generating a predicted user feature representation according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic structural block diagram of a device for information processing according to some embodiments of the present disclosure; and
[0014] Figure 5 A block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. Detailed implementation manners
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0016] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0020] As briefly described above, in the content recommendation scenario, the cold start problem is one of the challenges that the recommendation system faces when introducing new content, registering new users, or processing sparse interaction data. For example, in the video recommendation scenario, solving the cold start problem directly determines whether high-quality new content can be distributed more, providing key feedback to creators, and ultimately enabling creators to actively create content. Traditional recommendation systems are plagued by the cold start problem due to their over-reliance on historical interaction data. For example, due to the lack of exposure and feedback of new content, the fairness, interactivity, and diversity of the recommendation system are damaged.
[0021] Taking the video recommendation scenario as an example, high-quality new videos uploaded by creators may be buried due to insufficient initial traffic. Even if a new video has the potential to become popular content, the recommendation algorithm will find it difficult to accurately assess its content value in the absence of user interaction data, resulting in distribution efficiency far lower than that of mature content. This "cold start problem" may not only fail to promote the enthusiasm of content creators, but also reduce the distribution efficiency and effectiveness of content on video platforms.
[0022] It is worth emphasizing that the cold start problem has multi-dimensional complexity. For example, it is difficult to build an interest model for new users due to the lack of behavioral history, and it is difficult to evaluate the quality of new content due to the lack of interaction data. In sparse interaction scenarios (such as long-tail content areas), the model fails due to insufficient data density. These three types of problems are intertwined and together constitute the main obstacles to the evolution of recommendation systems towards higher accuracy and wider coverage. How to break through this bottleneck has become a technical problem that needs to be solved urgently.
[0023] In the research, the inventors found that lookalike algorithms provide a possible solution to the cold start problem of new content. Different from traditional recommendation algorithms that view from the perspective of users, lookalike algorithms construct a user expansion mechanism for the recommended object by viewing from the perspective of the recommended object (also known as the recommended content). Specifically, for each recommended object, the lookalike algorithm can extract the common features of the seed users of the recommended object (such as users who have interacted with the recommended object). Then, based on these common features, the lookalike algorithm searches for users with similar features (also known as lookalike users) in a larger user group, and then provides the recommended object to these lookalike users to expand the users of the recommended object. However, in the actual application process, limited by the number and quality of seed users in the cold start scenario, traditional lookalike algorithms are difficult to extract effective common features, and thus difficult to accurately find lookalike users.
[0024] In view of this, embodiments of the present disclosure provide an information processing solution. According to this solution, first, in response to the existence of a user sequence associated with the recommended object, an input feature sequence for a machine learning model is determined based on at least one object feature representation of the recommended object and at least one first user feature representation corresponding to at least one user in the user sequence, where the user sequence is determined based on users who have interacted with the recommended object. Then, based on the input feature sequence, a predicted user feature representation for the recommended object is generated using the machine learning model, where the predicted user feature representation is used to indicate users who will interact with the recommended object in the future. Next, in response to receiving a recommendation request, a second user feature representation corresponding to the initiating user of the recommendation request is extracted. Subsequently, in response to the similarity between the second user feature representation and the predicted user feature representation meeting the similarity requirement, the recommended object is provided to the initiating user of the recommendation request.
[0025] It can be more clearly understood through the description below that, unlike the similar user expansion algorithm, the scheme of the present disclosure proposes a next user prediction algorithm, which makes it better applicable to the cold start problem. Specifically, the scheme of the present disclosure no longer simply extracts common features from the seed users of the recommended object, but is based on the user feature representation (such as the first user feature representation) of the seed users of the recommended object (such as new content) (such as the user who interacts with the recommended object) and the object feature representation of the recommended object, with the help of a machine learning model (such as a generative model based on the Transformer architecture) to determine the user feature representation (i.e., predicted user feature representation) that the future potential users of the recommended object (such as users who may interact with the recommended object in the future, such users can also be called the next users predicted for the recommended object) may have. Compared with the traditional scheme of extracting common features from the seed users of the recommended object, the scheme of the present disclosure uses a machine learning model to mine the potential associations and complex patterns between the first user feature representations and between the first user feature representations and the object features, thereby generating a predicted user feature representation with higher accuracy. In this way, even if there are only a small number of seed users or the quality of the seed users is poor, the solution disclosed in the present invention can generate a predicted user feature representation with high accuracy, so that similar users can still be accurately determined from the user group (such as the initiator of the recommendation request).
[0026] In this way, the solution of the present disclosure solves the problem of poor applicability of similar user expansion algorithms in cold start problems, thereby making user expansion for new content possible and improving the recommendation system's ability to cope with cold start problems.
[0027] Various example implementations of the solution will be described in detail below in conjunction with the accompanying drawings.
[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In the environment 100, a user 110 may be a provider of a recommendation object 151. For example, in a video recommendation scenario, the recommendation object 151 may be a video, and the user 110 may be the author of the video. The user 110 may create and manage the recommendation object 151 through an associated terminal device 120.
[0029] In environment 100, users 130-1 to 130-N can be users who have interacted with the recommended object 151. For example, in a video recommendation scenario, users 130-1 to 130-N can be viewers of a video, where N is a positive integer. Users 130-1 to 130-N can watch the video and interact with the video through their respective associated terminal devices 140-1 to 140-N. Such interactions include, but are not limited to, liking, commenting, and / or forwarding, etc. For ease of discussion, hereinafter, users 130-1 to 130-N will also be collectively or individually referred to as user 130, and terminal devices 140-1 to 140-N will also be collectively or individually referred to as terminal device 140.
[0030] In environment 100, terminal device 120 and terminal device 140 can communicate with content recommendation system 150 through communication means such as a network. Content recommendation system 150 can be an application, a website, a web page, and other accessible platforms. In a video recommendation scenario, terminal device 120 and terminal device 140 can be installed with an application for accessing content recommendation system 150, or terminal device 120 and terminal device 140 can access content recommendation system 150 in any suitable manner to watch or manage videos in content recommendation system 150.
[0031] In environment 100, user 160 can be a user who has not interacted with the recommended object 151. User 160 can send a recommendation request to content recommendation system 150 through the associated terminal device 170. Content recommendation system 150 can be configured to select whether to provide the recommended object 151 to user 160 based on the corresponding policy.
[0032] It should be noted that the above uses the video recommendation scenario as an example to illustrate the recommended object 151, but this does not constitute a limitation to the present disclosure. For example, in some applications, the recommended object 151 can also be live content, products, services, or advertisements, etc., which can be specifically determined according to actual needs.
[0033] In environment 100, terminal devices 120, 140, and 170 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal devices 120, 140, and 170 can also support any type of interface for users (such as "wearable" circuits, etc.).
[0034] In environment 100, the content recommendation system 150 can be deployed in any type of server device. The server device can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server device can for example include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on.
[0035] It should be understood that the structures and functions of the various elements in environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.
[0036] Figure 2 A flowchart of an example process 200 of an information processing method according to some embodiments of the present disclosure is shown. Process 200 can be implemented at the content recommendation system 150.
[0037] Refer to Figure 2 , at block 210, in response to the presence of a user sequence associated with the recommendation object, the content recommendation system 150 determines an input feature sequence for a machine learning model based on at least one object feature representation of the recommendation object and at least one first user feature representation corresponding to at least one user in the user sequence. The user sequence is determined based on users who have performed interactions with the recommendation object.
[0038] At block 220, the content recommendation system 150 generates a predicted user feature representation for the recommendation object using the machine learning model 320 based on the input feature sequence 310. The predicted user feature representation is used to indicate a user who will perform an interaction with the recommendation object at a future time.
[0039] As an example, the recommendation object can be an entity that the content recommendation system 150 wants to recommend to a user, and it can be various types of recommendable content. For example, the recommendation object can be any one of a video, live content, a product, a service, and / or an advertisement, which can be specifically determined according to actual needs, and the embodiments of the present disclosure do not limit this.
[0040] In some embodiments, a user sequence may be an ordered set determined based on users who have interacted with a recommendation object. In some embodiments, the interaction performed by a user with a recommendation object may refer to a positive interaction between the user and the recommendation object. For example, for a video, a positive interaction may be operations such as liking, favoriting, and / or commenting performed by the user after browsing the video. For a video, assuming that 100 users have watched and liked it successively, then, the content recommendation system 150 may arrange these 100 users based on a corresponding sorting strategy (such as a time-based sorting strategy) to form a user sequence.
[0041] It should be noted that the description of positive interaction here is only an exemplary description, and according to the different recommendation objects, the positive interaction may also be different. For example, for a product or service, the positive interaction may also include, but is not limited to, operations such as purchasing.
[0042] In some embodiments, the content recommendation system 150 selects a predetermined number of candidate users whose interaction execution time meets the time requirement from the candidate users who have interacted with the recommendation object to be users in the user sequence.
[0043] As an example, candidate users may be a group of users who have interacted with the recommendation object. For example, assuming that 100 users have watched and liked the recommendation object successively, then, the content recommendation system 150 may determine these 100 users as candidate users.
[0044] As an example, the time requirement is a limiting condition for screening candidate users. For example, for candidate users who have interacted with the recommendation object, the time requirement may indicate that the content recommendation system 150 selects the candidate users with the latest interaction execution time from these candidate users as users in the user sequence. In addition, the predetermined number may be used for further screening of the selected candidate users. For example, the predetermined number may indicate the upper limit of the number of users in the user sequence. In some embodiments, if the number of candidate users who meet the time requirement is greater than or equal to the predetermined number, the content recommendation system 150 may use the predetermined number of candidate users who meet the time requirement as users in the user sequence. If the number of candidate users who meet the time requirement is less than the predetermined number, the content recommendation system 150 may use all candidate users who meet the time requirement as users in the user sequence.
[0045] For example, assume that the predetermined number is 50. If the content recommendation system 150 determines the 100 users described above as candidate users, then the content recommendation system 150 can select the top 50 users with the latest like time from these 100 users to construct a user sequence. Assume that the content recommendation system 150 only determines 30 candidate users, then the content recommendation system 150 can use these 30 users to construct a user sequence.
[0046] In this way, the content recommendation system 150 can effectively handle massive data (including but not limited to a massive number of recommended objects or a massive number of interactive users of recommended objects), thereby reducing computing power consumption. In addition, for the content recommendation system 150 that needs to process a large number of requests in a short time, the above screening method can also improve the system response speed while reducing the storage resource occupancy, and improve the versatility of the content recommendation system 150.
[0047] In some embodiments, for the candidate users selected to construct the user sequence, the content recommendation system 150 can sort these candidate users based on the chronological order of their interactions with the recommended object.
[0048] It should be noted that the above descriptions of time requirements and predetermined numbers are only exemplary descriptions, which do not constitute a limitation on the embodiments of the present disclosure. According to actual needs, the time requirements can also indicate other content, and the predetermined number is not limited to 50.
[0049] Next, the description of block 210 and block 220 will continue. As an example, the object feature representation can correspond one-to-one with the object features of the recommended object. The object feature representation can be a vector or a numerical value obtained by encoding the corresponding object features by the content recommendation system 150. Taking the recommended object as a video as an example, the object features of the recommended object can indicate the identity (ID) of the recommended object, the application identifier (AID) of the recommended object, and / or the category of the recommended object, etc.
[0050] Figure 3 FIG. shows a schematic diagram of an example 300 for generating a predicted user feature representation according to some embodiments of the present disclosure.
[0051] Refer to Figure 3, the recommended object has multiple object feature representations 301. The multiple object feature representations 301 may include an object feature representation 301-1 corresponding to the ID of the recommended object, an object feature representation 301-2 corresponding to the AID of the recommended object, and so on. For the convenience of discussion, hereinafter, the object feature representation 301-1 and the object feature representation 301-2 will also be referred to individually or collectively as the object feature representation 301. The content recommendation system 150 can encode the object features of the recommended object into an object feature representation 301 that can be processed by a computer through embedding technology, feature engineering, or any suitable algorithm, so as to be processed in a machine learning model.
[0052] As an example, the first user feature representation can be one-to-one corresponding to the users in the user sequence. The first user feature representation can be a vector, a numerical value, etc. obtained after the content recommendation system 150 encodes the user features of the corresponding users. In some embodiments, the user features may indicate the basic information of the user and / or the behavior information of the user, etc. In addition, the user features may also indicate the interaction information (such as likes or collections, etc.) performed by the user with the recommended object. This information can be obtained after obtaining the authorization of the user. Similar to the object features, the content recommendation system 150 can encode these user features into a user feature representation that can be processed by a computer through embedding technology, feature engineering, or any suitable algorithm, so that the machine learning model can perform an association process on these user feature representations and the object feature representation 301.
[0053] Continue to refer to Figure 3 , as an example, assume that the content recommendation system 150 selects the top m users with the latest like time from the 100 users described above to construct a user sequence, where m < 100 and m is a positive integer. Then, the m first user feature representations corresponding to the m users include: the first user feature representation 302-1 corresponding to the first user, the first user feature representation 302-2 corresponding to the second user,..., and the first user feature representation 302-m corresponding to the mth user. For the convenience of discussion, hereinafter, the first user feature representation 302-1, the first user feature representation 302-2 to the first user feature representation 302-m will also be referred to individually or collectively as the first user feature representation 302.
[0054] As an example, the machine learning model 320 can be any suitable generative machine learning model 320. For example, the machine learning model 320 can be a generative model based on the Transformer architecture. In addition, the machine learning model 320 can also be a generative model based on other architectures, and the embodiments of the present disclosure will not be listed one by one here.
[0055] As an example, the input feature sequence 310 can be an ordered feature set for the machine learning model 320 constructed at least based on the object feature representation 301 and the first user feature representation 302. The input feature sequence 310 will be used as the input of the machine learning model 320, and the machine learning model 320 will generate the predicted user feature representation 304 based on this input.
[0056] As an example, the predicted user feature representation 304 can be represented in the form of a vector or a numerical value. The predicted user feature representation 304 is an estimation of the user features of a user (also referred to as a potential user) who may interact with the recommended object in the future by the machine learning model 320. In some embodiments, the predicted user feature representation 304 is used to describe the user features that the potential user of the recommended object is most likely to have. In some embodiments, the machine learning model 320 can generate the predicted user feature representation 304 through an autoregressive generation operation such as for the user feature representation.
[0057] As an example, the autoregressive generation operation is a data generation method based on a sequence. The autoregressive generation operation assumes that each element (also referred to as a token) in the sequence depends on the previously generated elements. When generating the predicted user feature representation 304, the machine learning model 320 can gradually predict the next possible user feature representation in the output feature sequence 330 at least based on the object feature representation 301 and the first user feature representation 302 in the input feature sequence 310 until the predicted user feature representation 304 is generated.
[0058] It should be noted that the above generation of the predicted user feature representation 304 is only an exemplary illustration. According to actual needs, the machine learning model can also generate the predicted user feature representation 304 through other prediction operations, and the embodiments of the present disclosure will not list them one by one here.
[0059] In some embodiments, the content recommendation system 150 sorts at least one first user feature representation 302 based on the execution time of the interaction to obtain a first user feature sequence. Then, the content recommendation system 150 determines a prefix feature sequence for the first user feature sequence based on at least one object feature representation 301. Subsequently, the content recommendation system 150 determines the input feature sequence at least based on the concatenation of the prefix feature sequence and the first user feature sequence.
[0060] Continue to refer to Figure 3, the content recommendation system 150 can arrange the first user feature representations 302-1, 302-2 to 302-m from far to near (compared to the current moment) according to the execution time of the interactions, so as to obtain the first user feature sequence 312. Such a first user feature sequence 312 can indicate the temporal correlation relationship among multiple users. For example, taking the recommended object as a video, the interaction (such as a comment) performed by the 2nd user on the recommended object may be caused by the interaction (such as a comment) first performed by the 1st user on the recommended object. In this way, the machine learning model 302 can learn the causal relationship among multiple users in the interactions with respect to the recommended object, thereby improving the accuracy of predicting the user feature representation 304.
[0061] Continue to refer to Figure 3 , the content recommendation system 150 can arrange the object feature representations 301-1, 301-2, etc. based on any suitable method, so as to obtain the prefix feature sequence 311. After obtaining the first user feature sequence 312 and the prefix feature sequence 311, the content recommendation system 150 can construct the input feature sequence 310 by appending the first user feature sequence 312 after the prefix feature sequence 311.
[0062] In this way, the content recommendation system 150 can embed multiple object feature representations 301 as prefix hints to make up for the limitations of the first user feature sequence 312, thereby achieving multi-modal feature fusion.
[0063] In some embodiments, the content recommendation system 150 determines the input feature sequence 310 based on the concatenation of the prefix feature sequence 311, the first user feature sequence 312, and the classification token 313 associated with the machine learning model 320. The classification token 313 is appended after the prefix feature sequence 311 and the first user feature sequence 312. Then, the content recommendation system 150 provides the input feature sequence 310 to the machine learning model 320 as an input. The machine learning model 320 performs an autoregressive generation operation for the user feature representation based on the prefix feature sequence 311 and the first user feature sequence 312 to generate a second user feature sequence 331 corresponding to the first user feature sequence 312 and a third user feature representation 332 appended after the second user feature sequence 331. Subsequently, the content recommendation system 150 generates the predicted user feature representation 304 based on the classification token 313, the prefix feature sequence 311, the first user feature sequence 312, and the third user feature representation 332.
[0064] Continue to refer to Figure 3, the content recommendation system 150 can sequentially concatenate the prefix feature sequence 311, the first user feature sequence 312, and the classification token 313 to construct the input feature sequence 310. As described above, the machine learning model 320 can generate the predicted user feature representation 304 by predicting the next possible user feature representation based on the input feature sequence 310.
[0065] For example, in the process of generating the predicted user feature representation 304, the machine learning model 320 can first predict the possible fourth user feature representation 303-1 in the second user feature sequence 331 based on the prefix feature sequence 311. Then, the machine learning model 320 predicts the possible fourth user feature representation 303-2 in the second user feature sequence 331 based on the prefix feature sequence 311 and the first user feature representation 302-1. Next, the machine learning model 320 predicts the possible fourth user feature representation 303-3 in the second user feature sequence 331 based on the prefix feature sequence 311, the first user feature representation 302-1, and the second user feature representation 302-2. And so on, the machine learning model 320 predicts the possible fourth user feature representation 303-m in the second user feature sequence 331 based on the prefix feature sequence 311, the first user feature representation 302-1,..., the first user feature representation 303-m-1. After that, the machine learning model 320 will further predict the possible third user feature representation 332 attached after the second user feature sequence 331 based on the prefix feature sequence 311, the first user feature representation 302-1,..., the first user feature representation 303-m (i.e., the first user feature sequence 312). Such a third user feature representation 332 can also be referred to as the next user feature representation of the fourth user feature representations 303-1 to 303-m, that is, the possible fourth user feature representation 303-m+1.
[0066] In some embodiments, the fourth user feature representations 303-1 to 303-m and the third user feature representation 332 have the same feature dimension as the first user feature representation 302-1. For the convenience of discussion, hereinafter, the fourth user feature representations 303-1 to 303-m will also be individually or collectively referred to as the fourth user feature representation 303.
[0067] The inventors found in their research that if the predicted user feature representation 304 is directly generated based on the prefix feature sequence 311 and the first user feature sequence 312, such a predicted user feature representation 304 will be similar to the third user feature representation 332 and have the same feature dimension as the first user feature representation 302-1. However, the user feature representation (such as the user feature representation of user 160) contained in a real recommendation request (such as a recommendation request from user 160) often has a more complex feature dimension. To align the predicted user feature representation 304 and the user feature representation contained in the real recommendation request in terms of the feature dimension (which can also be referred to as the feature domain), embodiments of the present disclosure introduce a classification token 313 into the input feature sequence 310. Through the classification token 313, the machine learning model 320 can incorporate more information during the generation process of the predicted user feature representation 304, thereby obtaining a more complex predicted user feature representation 304, and further aligning it with the user feature representation contained in the real recommendation request in terms of the feature dimension.
[0068] In some embodiments, the classification token 313 is determined during the training process of the machine learning model 320. Such a classification token 313 can also be referred to as a learnable classification token. In this way, the content recommendation system 150 can encode prior knowledge for aligning the feature dimension into the machine learning model 320, so as to be able to dynamically guide the machine learning model 320 to generate the predicted user feature representation 304 in an appropriate manner.
[0069] In some embodiments, the content recommendation system 150 determines a global feature representation related to at least one of the prefix feature sequence 311 and the first user feature sequence 312 based on the classification token 313. Subsequently, the content recommendation system 150 generates the predicted user feature representation 304 based on the prefix feature sequence 311, the first user feature sequence 312, the third user feature representation, and the global feature representation.
[0070] Through the global feature representation, the machine learning model 320 can consider, during the generation process of the predicted user feature representation 304, such as: the association relationships between multiple first user feature representations 302, the association relationships between multiple first user feature representations 302 and multiple object feature representations 301, and the association relationships between multiple object feature representations 301 themselves, etc. As an example, the association relationships here include but are not limited to causal relationships, etc.
[0071] In some embodiments, in response to the absence of a user sequence associated with the recommended object, the content recommendation system 150 determines the input feature sequence 310 for the machine learning model 320 based on at least one object feature representation 301.
[0072] As an example, in response to the non-existence of a user sequence associated with the recommended object, the content recommendation system 150 determines an input feature sequence 310 for the machine learning model 320 based on one less object feature representation 301 and a classification token 313 appended after the one less object feature representation 301.
[0073] As described above, the machine learning model 320 can generate a predicted user feature representation 304 by predicting the next possible user feature representation based on the input feature sequence 310. For example, in the process of generating the predicted user feature representation 304, the machine learning model 320 can first predict a possible fourth user feature representation 303-1 in the second user feature sequence 331 based on the prefix feature sequence 311. Then, the machine learning model 320 generates the predicted user feature representation 304 based on the prefix feature sequence 311, the fourth user feature representation 303-1, and the classification token 313. The specific generation process here can be referred to the above, so it will not be elaborated.
[0074] In this way, for the recommended object, even if there is no user who has interacted with it, the content recommendation system 150 can use the machine learning model 320 to generate the predicted user feature representation 304, so as to accurately find potential users of the recommended object. Thus, the embodiments of the present disclosure can further expand the applicable scope of the content recommendation system 150 for the cold start problem.
[0075] In some embodiments, the machine learning model 320 can be a machine learning model with a causal attention mechanism. The causal attention mechanism ensures that in the autoregressive generation operation of the machine learning model 320, at each position in the output feature sequence 330 (such as the fourth user feature representation 303, the third user feature representation 332, or the predicted user feature representation 333), it only depends on the current position and previous information, rather than future information, thereby maintaining the causality of the data by introducing a causal mask 323.
[0076] As an example, continuing to refer to Figure 3 , for each position of the output feature sequence 330, when calculating the attention weights, the causal mask 323 configures a mask 3231 at the position corresponding to the future information, so that the weight at the position corresponding to the future information is zero. In this way, when generating each position of the output feature sequence 330, the machine learning model 320 can only depend on the current position and previous information.
[0077] As an example, continuing to refer to Figure 3, for the fourth feature representation 303-1, the causal mask 323 configures 3231 at the positions of each first user feature representation 302 in the first user feature sequence 312 and at the position of the classification token 313. For the fourth feature representation 303-2, the causal mask 323 configures the mask 3231 at the positions of all first user feature representations 302 in the first user feature sequence 312 except the first user feature representation 302-1 and at the position of the classification token 313. For the fourth feature representation 303-3, the causal mask 323 configures the mask 3231 at the positions of all first user feature representations 302 in the first user feature sequence 312 except the first user feature representation 302-1 and the first user feature representation 302-2 and at the position of the classification token 313. And so on.
[0078] In some embodiments, the causal attention mechanism determines query features, key features, and value features based on the input feature sequence 310, and then generates an output feature sequence 330 based on the query features, key features, and value features. As an example, continue to refer to Figure 3 , the content recommendation system 15 can provide the position-encoded input feature sequence 310 to the machine learning model 320. After receiving the position-encoded input feature sequence 310, the machine learning model 320 calculates the key features and value features through the first neural network layer 321. As an example, the first neural network layer 321 can include a cascaded self-attention layer 3211, a residual connection and normalization layer 3212, a feed-forward neural network layer 3213, and a residual connection and normalization layer 3214. As an example, the self-attention layer 3211 can have a self-attention mechanism that can combine the causal attention mechanism and prefix information (such as the object feature representation 301).
[0079] It should be noted that although Figure 3 only shows one first neural network layer 321, however, according to actual needs, the number of the first neural network layers 321 can be 2 or more, and multiple first neural network layers 321 can be cascaded in sequence, and the embodiments of the present disclosure do not limit this.
[0080] As an example, continue to refer to Figure 3, the machine learning model 320 may provide the key features and value features output by the first neural network layer 321, together with the query features, to the second neural network layer 322. As an example, the second neural network layer 322 may include a cascaded self-attention layer 3221, a residual connection and normalization layer 3222, a feed-forward neural network layer 3223, and a residual connection and normalization layer 3224. As an example, the self-attention layer 3221 may have a self-attention mechanism that can combine a causal attention mechanism and prefix information (such as the object feature representation 301).
[0081] It should be noted that although Figure 3 only one first neural network layer 321 is shown, according to actual needs, the number of the second neural network layers 322 may be two or more, and the multiple second neural network layers 322 may be cascaded in sequence, and the embodiments of the present disclosure do not limit this. The machine learning model 320 may generate query features in any suitable manner, and the embodiments of the present disclosure do not limit this.
[0082] It should also be noted that the above structure of the machine learning model 320 is only an exemplary illustration. According to actual needs, the machine learning model 320 may also adopt other structures (such as other neural network layers, etc.).
[0083] In some embodiments, the encoding process of the machine learning model 320 for the input feature sequence 310 may be represented by formula (1):
[0084]
[0085] where p i ∈R d and u i ∈R d respectively represent the i-th object feature representation 301 and the i-th first user feature representation 302. CLS ∈ R d represents the classification token 313, which is attached to the end of the input feature sequence 310, represents the output of the encoder Encoder of the i-th token of type k (k = p, representing the object feature representation 301, k = u, representing the first user feature representation 302).
[0086] In some embodiments, the decoding process of the machine learning model 320 may be represented by formula (2).
[0087]
[0088] where q ∈ R (n+2)×d represents a learnable query feature, i ∈ [1, n] represents the i-th fourth feature representation 303, represents the third user feature representation 332, represents the predicted user feature representation 333.
[0089] It should be noted that the above formulas and parameters for the encoding process and the decoding process are only for illustrative purposes. According to actual needs, the encoding process and the decoding process can also be represented by other formulas and parameters.
[0090] The training process of the machine learning model 320 will be described below. It should be noted that to distinguish the inference process from the training process, the inputs and outputs involved in the inference process (such as the recommended object, the object feature representation 301, the predicted user feature representation 333, etc.) will all be expressed as corresponding "samples", that is, in a way similar to sample recommended object, sample object feature representation, predicted sample user feature representation, etc.
[0091] During the training process, since the true potential users of the sample recommended object can be known (that is, the next possible interaction user predicted by the machine learning model 320 based on the users who have interacted with the sample recommended object, and thus can also be called the next user), therefore, the most likely user to interact with the recommended sample object can be modeled by maximizing the likelihood function, and this process can be represented by formula (3):
[0092] argmaxP(u, f u )|model({(u1),..., (u j ),..., (u n ), (i, f i )); (3) where P(·) represents the probability distribution function, which indicates the probability that user u interacts with the sample recommended object i under the given conditions. model(·) represents the machine learning model 320, {(u1), ……, (u j ), ……, (u x )} represents the sample user sequence, f i represents the sample object feature representation of the sample recommended object i, and f u represents the sample user feature representation of user u.
[0093] In some embodiments, during the training process of the machine learning model 320, the loss function used for the training process is determined based on at least one of the following: contrast loss, cross-entropy loss, and / or auxiliary loss.
[0094] The Contrastive Loss is used to make the distance between similar sample pairs in the feature space as small as possible, while making the distance between dissimilar sample pairs as large as possible. The Cross-Entropy Loss is used to measure the difference between two probability distributions, where one probability distribution is predicted by the machine learning model 320 and the other is the true probability distribution. The goal of the Cross-Entropy Loss is to make the predicted probability distribution of the machine learning model 320 as close as possible to the true probability distribution. The Auxiliary Loss is an additional loss term added on the basis of the main loss functions (such as the Contrastive Loss and the Cross-Entropy Loss) and is used to assist in the training of the machine learning model 320. The Auxiliary Loss can help the machine learning model 320 better learn the features or structure of the data, thereby improving the generalization ability and convergence speed of the model. By combining the three loss functions to guide the training of the machine learning model 320, the generation ability of the machine learning model 320 and its robustness in the cold start problem can be enhanced.
[0095] In some embodiments, the loss function L for the training process genreative can be represented by Equation (4):
[0096]
[0097] where L contrasitve represents the Contrastive Loss, λ1 represents the weight of the Contrastive Loss L contrasitve , L CE represents the Cross-Entropy Loss, λ2 represents the weight of the Cross-Entropy Loss L CE , L auxiliary represents the Auxiliary Loss, and λ3 represents the weight of the Auxiliary Loss L auxiliary .
[0098] In some embodiments, the Contrastive Loss is configured to increase the similarity between the predicted sample user feature representation generated by the machine learning model 320 and the true user feature representation and decrease the similarity between the predicted sample user feature representation and the non-true user feature representation during the training process. The predicted sample user feature representation is generated by the machine learning model 320 based on training samples related to the true user feature representation.
[0099] As an example, the predicted sample user feature representation may refer to the user feature representation that the machine learning model 320 may have for the potential user (as described above, this potential user may also be referred to as the next user) predicted for the sample recommended object during the training process. The true user feature representation may be the user feature representation of the true next user of the sample recommended object. The non-true user feature representation may be the user feature representation of a user randomly sampled from the executed interaction users of the sample recommended object. The training samples related to the true user feature representation may include, but are not limited to, the sample user feature representations of the executed interaction users of the sample recommended object, the sample object feature representations of the sample recommended object, and the sample classification tokens, etc.
[0100] For the unexhausted content of the predicted sample user feature representation, reference can be made to the description of the predicted user feature representation 333 in the machine learning model inference process above. For the unexhausted content of the training samples, reference can be made to the description of the object feature representation 301, the first user feature representation 302, and the classification token 313 in the machine learning model inference process above, so it will not be elaborated here.
[0101] In some embodiments, the contrastive loss L contrasitve can be represented by formula (5):
[0102]
[0103] where indicates that after the sample recommended object is provided to the user corresponding to the i-th sample user feature representation u i this user has executed an interaction with the sample recommended object, which can also be referred to as the true label, represents the predicted sample user feature representation generated by the machine learning model. In the case of u i represents the true user feature representation, u j represents the non-true user feature representation, f(·) represents the similarity function, and τ is the temperature parameter.
[0104] In some embodiments, the cross-entropy loss is configured to increase the matching degree between the predicted sample user feature representation generated by the machine learning model 320 and the true label of the predicted sample user feature representation during the training process. The true label indicates whether the user has executed an interaction with the sample recommended object after the sample recommended object is provided to the user.
[0105] Since the user feature representations of users who were exposed but did not interact with the sample recommendation objects (i.e., the true value label indicates that the user did not interact with the sample recommendation object after it was provided to the user) are also collected, it is impossible to predict which user will interact next for these user feature representations, so the contrast loss L cannot be used for these user feature representations. contrasitve However, these samples can pass through the recommendation funnel (including retrieval, pre-ranking, ranking, and re-ranking) and reach exposure, which indicates that they have relatively higher quality and information content compared with non-true value user feature representations. CE This information can be put to better use.
[0106] In some embodiments, the cross entropy loss L CE It can be expressed by formula (6):
[0107]
[0108] where σ(·) is the sigmoid function, Indicates that the sample recommendation object is provided to the i-th sample user feature representation u i After the corresponding user, the user did not perform any interaction on the sample recommended object.
[0109] In some embodiments, the auxiliary loss is configured to increase the matching degree between the prefix sample user feature sequence generated by the machine learning model 320 and the training samples related to the true value user feature representation during the training process. The prefix sample user feature sequence is used by the machine learning model 320 to generate the predicted sample user feature representation.
[0110] The prefix sample user feature sequence can be a sample user feature representation that is generated by the machine learning model 320 through an autoregressive generation operation and is located before the predicted sample user feature representation. The specific content of the prefix sample user feature sequence can be found in the description of the fourth sample feature representation 303 and the third sample feature representation 332 in the machine learning model reasoning process, so it will not be repeated here. Through the auxiliary loss, it helps to enhance the learning of the prefix sample user feature sequence by the machine learning model 320.
[0111] In some embodiments, the auxiliary loss L auxiliary It can be expressed by formula (7):
[0112]
[0113] Where sg(·) means stop moving Propagate gradients to avoid machine learning model 320 crash.
[0114] In some embodiments, the machine learning model 320 is updated through an online learning process, and the predicted user feature representation 304 for the recommended object is updated using the updated machine learning model 320 .
[0115] Online learning is a machine learning paradigm. Unlike traditional batch learning, it does not require the collection of a large amount of training data at one time, but can update the model immediately when new data arrives. In online learning, the machine learning model 320 can continuously receive new sample data and adjust its own parameters based on the new data to gradually optimize the performance of the machine learning model 320. In the embodiments of the present disclosure, the online learning process of the machine learning model 320 can refer to the description of model training in the previous text, so it will not be repeated here.
[0116] Once the predicted user feature representation 304 of the recommended object is determined, the content recommendation system 150 can store the predicted user feature representation 304 and wait for recommendation requests from other users. Figure 2 , in block 230 , in response to receiving the recommendation request, the content recommendation system 150 extracts a second user feature representation corresponding to the initiator of the recommendation request.
[0117] As an example, a recommendation request may be initiated by a variety of channels. For example, in a mobile application, when a user opens the application homepage, enters a specific section (such as a video classification section), or clicks a "get recommendations" button, the application may send a recommendation request to the content recommendation system 150. On the web side, a user browsing a specific page or performing certain operations (such as searching for keywords) may also trigger a recommendation request.
[0118] As an example, the second user feature representation can be a vector or a numerical value obtained by encoding the user features of the user who initiated the recommendation request. In addition to the user features that can be indicated by the first user feature representation, the second user feature representation can also indicate more complex user features. As an example, the content recommendation system 150 can encode these user features into a second user feature representation that matches the predicted user feature representation 333 through embedding technology, feature engineering, or any appropriate algorithm, so that the recommendation system 150 can associate the predicted user feature representation 333 with the second user feature representation. As an example, the layered navigation small world map
[0119] In block 240 , in response to the similarity between the second user feature representation and the predicted user feature representation 304 satisfying the similarity requirement, the content recommendation system 150 provides a recommendation object to the initiating user of the recommendation request.
[0120] As an example, after extracting the second user feature representation, the content recommendation system 150 may determine the similarity between the second feature representation and the predicted user feature representations 304 of each recommended object stored in advance based on any appropriate search algorithm. Then, the content recommendation system 150 finds the predicted user feature representation 304 similar to the second user feature representation from these pre-stored predicted user feature representations 304 based on whether the similarity meets the similarity requirement (for example, higher than a predetermined similarity threshold). Then, the content recommendation system 150 provides the recommended object corresponding to the found predicted user feature representation 304 to the initiating user of the recommendation request.
[0121] As an example, the search algorithm includes but is not limited to the Hierarchical Navigable Small World (HNSW) algorithm. HNSW is an efficient Approximate Nearest Neighbor Search (ANN) algorithm, which is particularly suitable for similarity retrieval of large-scale, high-dimensional data sets. It is based on the concept of Small World Networks and achieves fast and efficient search by building a multi-level graph structure and calculating the dot product.
[0122] According to the various embodiments described above, it can be clearly understood that the embodiments of the present disclosure define user sequence users with positive interactions (the features of these interactions, for example, have the same sparsity) as seed users, and construct the first user feature sequence 312 through the real-time fine-grained first user feature representation 302. In this way, the first user feature sequence 312 introduces real-time dynamic interaction information, thereby enhancing the weaker user feature representation in the cold start problem. In addition, the embodiments of the present disclosure make the first user feature sequence 312 have temporal characteristics, thereby helping the machine learning model 320 to better learn the causal relationship between multiple first user feature representations 302 and improve the accuracy of the generation of the predicted user feature representation 304. In order to overcome the limitation that the machine learning model 320 only relies on the first user feature sequence 312, the embodiments of the present disclosure also integrate the object feature representation 301 of the recommended object as a prefix prompt embedding to enhance the first user feature sequence 312, thereby assisting the generation of the predicted user feature representation 304. In addition, the embodiments of the present disclosure also introduce learnable classification tokens, so as to bridge the gap between the predicted user feature representation 304 and the user feature representation in the real recommendation request in the feature domain.
[0123] Furthermore, the embodiments of the present disclosure utilize a machine learning model based on the Transformer architecture, and design three dedicated loss functions to model the generation process of the predicted user feature representation 304 (also referred to as the user feature representation of the next user). The contrast loss function can convert discrete entity predictions into entity feature representation learning. Through the cross-entropy loss function, the embodiments of the present disclosure can utilize training samples that have exposure but no interaction. The auxiliary loss function can enhance the learning of user feature representation by the machine learning model. These designs enable the generated predicted user feature representation 304 to seamlessly connect to the HNSW search algorithm.
[0124] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for information processing according to some embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in a content recommendation system 150. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0125] Reference Figure 4 , the device 400 includes an input feature sequence determination module 410, a predicted user feature representation generation module 420, a user feature representation extraction module 430 and a recommended object providing module 440. The input feature sequence determination module 410 is configured to determine an input feature sequence for a machine learning model based on at least one object feature representation of the recommended object and at least one first user feature representation corresponding to at least one user in the user sequence in response to the existence of a user sequence associated with the recommended object, wherein the user sequence is determined based on the user who interacts with the recommended object. The predicted user feature representation generation module 420 is configured to generate a predicted user feature representation for the recommended object based on the input feature sequence and using the machine learning model, the predicted user feature representation being used to indicate the user who interacts with the recommended object at a future time. The user feature representation extraction module 430 is configured to extract a second user feature representation corresponding to the initiating user of the recommendation request in response to receiving a recommendation request. The recommended object providing module 440 is configured to provide a recommended object to the initiating user of the recommendation request in response to the similarity between the second user feature representation and the predicted user feature representation satisfying the similarity requirement.
[0126] In some embodiments, the input feature sequence determination module 410 is further configured to: sort at least one first user feature representation based on the execution time of the interaction to obtain a first user feature sequence; determine a prefix feature sequence for the first user feature sequence based on at least one object feature representation; and determine the input feature sequence based at least on the cascade of the prefix feature sequence and the first user feature sequence.
[0127] In some embodiments, the input feature sequence determination module 410 is further configured to: determine an input feature sequence based on a concatenation of a prefix feature sequence, a first user feature sequence, and a classification token associated with a machine learning model, where the classification token is appended after the prefix feature sequence and the first user feature sequence. And the predicted user feature representation generation module 420 is further configured to: perform an autoregressive generation operation on the user feature representation based on the prefix feature sequence and the first user feature sequence by using the machine learning model, to generate a second user feature sequence corresponding to the first user feature sequence and a third user feature representation appended after the second user feature sequence, and generate a predicted user feature representation based on the classification token, the prefix feature sequence, the first user feature sequence, and the third user feature representation.
[0128] In some embodiments, the predicted user feature representation generation module 420 is further configured to: determine a global feature representation related to at least one of the prefix feature sequence and the first user feature sequence based on the classification token; and generate a predicted user feature representation based on the prefix feature sequence, the first user feature sequence, the third user feature representation, and the global feature representation.
[0129] In some embodiments, the classification token is determined during the training process of the machine learning model.
[0130] In some embodiments, the machine learning model is updated through an online learning process, and the predicted user feature representation for the recommended object is updated by using the updated machine learning model.
[0131] In some embodiments, the input feature sequence determination module 410 is further configured to: in response to the absence of a user sequence associated with the recommended object, determine an input feature sequence for the machine learning model based on at least one object feature representation.
[0132] In some embodiments, the apparatus 400 further includes a user sequence determination module. The user sequence determination module is configured to: select a predetermined number of candidate users whose interaction execution time meets the time requirement from candidate users who perform an interaction on the recommended object, as the users in the user sequence.
[0133] In some embodiments, during the training process of the machine learning model, the loss function for the training process is determined based on at least one of the following: contrastive loss, cross-entropy loss, and / or auxiliary loss.
[0134] In some embodiments, the contrastive loss is configured to increase the similarity between the predicted sample user feature representation generated by the machine learning model and the ground-truth user feature representation and decrease the similarity between the predicted sample user feature representation and the non-ground-truth user feature representation during the training process, where the predicted sample user feature representation is generated by the machine learning model based on the training samples related to the ground-truth user feature representation. The cross-entropy loss is configured to increase the matching degree between the predicted sample user feature representation generated by the machine learning model and the ground-truth label of the predicted sample user feature representation during the training process, where the ground-truth label indicates whether the user performs an interaction with the sample recommendation object after the sample recommendation object is provided to the user. The auxiliary loss is configured to increase the matching degree between the prefix sample user feature sequence generated by the machine learning model and the training samples related to the ground-truth user feature representation during the training process, where the prefix sample user feature sequence is used by the machine learning model to generate the predicted sample user feature representation.
[0135] Figure 5 FIG. shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. The electronic device 500 may be used, for example, to implement a content recommendation system 150 as shown in Figure 1 or a device 400 as shown in Figure 4 . It should be understood that Figure 5 the electronic device 500 shown is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein.
[0136] Referring to Figure 5 , the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.
[0137] An electronic device 500 generally includes multiple computer storage media. Such media can be any available media accessible to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be capable of storing information and / or data and can be accessed within the electronic device 500.
[0138] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules that are configured to execute various methods or actions of the various embodiments of the present disclosure.
[0139] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.
[0140] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed via the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0141] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0142] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0143] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0144] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0146] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The determination of the terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. An information processing method, comprising: In response to the existence of a user sequence associated with the recommended object, determining an input feature sequence for a machine learning model based on at least one object feature representation of the recommended object and at least one first user feature representation corresponding to at least one user in the user sequence, wherein the user sequence is determined based on users who interact with the recommended object; Based on the input feature sequence, using the machine learning model, generating a predicted user feature representation for the recommended object, wherein the predicted user feature representation is used to indicate a user who will interact with the recommended object in the future; In response to receiving a recommendation request, extracting a second user feature representation corresponding to a user initiating the recommendation request; as well as In response to the similarity between the second user feature representation and the predicted user feature representation satisfying a similarity requirement, providing the recommended object to the initiator of the recommendation request.
2. The method of claim 1, wherein determining an input feature sequence for a machine learning model comprises: sorting the at least one first user feature representation based on the execution time of the interaction to obtain a first user feature sequence; determining a prefix feature sequence for the first user feature sequence based on the at least one object feature representation; as well as The input feature sequence is determined based at least on a concatenation of the prefix feature sequence and the first user feature sequence.
3. The method according to claim 2, wherein determining the input feature sequence based at least on a concatenation of the prefix feature sequence and the first user feature sequence comprises: determining the input feature sequence based on a concatenation of the prefix feature sequence, the first user feature sequence, and a classification token associated with the machine learning model, wherein the classification token is appended to the prefix feature sequence and the first user feature sequence; and Generating the predicted user feature representation for the recommended object includes: using the machine learning model, performing an autoregressive generation operation on a user feature representation based on the prefix feature sequence and the first user feature sequence to generate a second user feature sequence corresponding to the first user feature sequence and a third user feature representation attached to the second user feature sequence, and The predicted user feature representation is generated based on the classification token, the prefix feature sequence, the first user feature sequence, and the third user feature representation.
4. The method according to claim 3, wherein generating the predicted user feature representation based on the classification token, the prefix feature sequence, the first user feature sequence and the third user feature representation comprises: determining, based on the classification token, a global feature representation associated with at least one of the prefix feature sequence and the first user feature sequence; as well as The predicted user feature representation is generated based on the prefix feature sequence, the first user feature sequence, the third user feature representation, and the global feature representation.
5. The method of claim 3, wherein the classification tokens are determined during the training process of the machine learning model.
6. The method according to claim 1, wherein the machine learning model is updated through an online learning process, and the predicted user feature representation for the recommended object is updated using the updated machine learning model.
7. The method according to claim 1, further comprising: In response to the absence of the user sequence associated with the recommended object, determining the input feature sequence for the machine learning model based on the at least one object feature representation.
8. The method according to claim 1, wherein the user sequence is determined by: A predetermined number of candidate users whose execution time of the interaction meets the time requirement are selected from the candidate users who perform the interaction on the recommended object as users in the user sequence.
9. The method of claim 1, wherein in the training process of the machine learning model, a loss function used for the training process is determined based on at least one of the following: Contrastive loss, Cross entropy loss, or Auxiliary loss.
10. The method according to claim 9, wherein: The contrast loss is configured to increase the similarity between the predicted sample user feature representation generated by the machine learning model and the true value user feature representation and reduce the similarity between the predicted sample user feature representation and the non-true value user feature representation during the training process, wherein the predicted sample user feature representation is generated by the machine learning model based on training samples related to the true value user feature representation; The cross entropy loss is configured to increase the degree of match between the predicted sample user feature representation generated by the machine learning model and the true value label of the predicted sample user feature representation during the training process, wherein the true value label indicates whether the user performs an interaction with the sample recommendation object after the sample recommendation object is provided to the user; as well as The auxiliary loss is configured to increase the matching degree between the prefix sample user feature sequence generated by the machine learning model and the training samples related to the true value user feature representation during the training process, wherein the prefix sample user feature sequence is used by the machine learning model to generate the predicted sample user feature representation.
11. An apparatus for information processing, comprising: an input feature sequence determination module, configured to determine, in response to the existence of a user sequence associated with a recommended object, an input feature sequence for a machine learning model based on at least one object feature representation of the recommended object and at least one first user feature representation corresponding to at least one user in the user sequence, wherein the user sequence is determined based on users who interact with the recommended object; a predicted user feature representation generating module, configured to generate a predicted user feature representation for the recommended object based on the input feature sequence and using the machine learning model, wherein the predicted user feature representation is used to indicate a user who will interact with the recommended object at a future time; A user feature representation extraction module is configured to extract a second user feature representation corresponding to a user who initiates the recommendation request in response to receiving the recommendation request; as well as The recommended object providing module is configured to provide the recommended object to the initiator of the recommendation request in response to the similarity between the second user feature representation and the predicted user feature representation satisfying the similarity requirement.
12. An electronic device comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.
13. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Object recommendation method and device based on artificial intelligence and electronic equipment
CN111125420A
Information recommendation method and device based on artificial intelligence, electronic equipment and storage medium
CN115080836A
Recommendation method and device, electronic equipment and storage medium
CN115082141A
Video data processing method and device, computer equipment and storage medium
CN117171389A
Large language model reasoning optimization method based on cascade and speculative decoding strategy
CN119047579A