Information processing method, system, storage medium and server
By training the feature extraction network in the information recommendation system, and using the comparison learning method of multiple sample pairs, the problem of low model accuracy caused by sparse user click records is solved, and the accuracy of the recommendation system and user interaction rate are improved.
Patent Information
- Application Number
- CN202210753559.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-06-28
AI Technical Summary
In the large-scale recommendation system, the existing information recommendation model has sparse records of user clicks/not clicks on items, resulting in insufficient training and low accuracy.
By obtaining positive sample pairs that are consistent with attribute features and inconsistent negative sample pairs, positive sample pairs that are similar based on semantic features and dissimilar negative sample pairs, positive sample pairs that are co-occurring based on session co-occurring, and negative sample pairs that are not co-occurring, the feature extraction network is trained using a comparison learning method, considering the attributes and semantic features of the sample object itself and the implicit correlation of the session.
It improves the accuracy of feature extraction network, improves the accuracy of information recommendation and user interaction rate, such as increasing the number of clicks per person, reading time and accuracy of recommended results.
Smart Images

Figure CN115130580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology based on artificial intelligence, and in particular to an information processing method, system, storage medium and server. Background Art
[0002] In the recommendation system, the recommendation background generally combines the user's information and the characteristics of each object stored in the system (such as news, videos, etc.) to determine multiple objects to be recommended, and sends the recommended objects to the user terminal for display, realizing personalized recommendation information service.
[0003] An existing information recommendation method mainly uses an artificial intelligence-based information recommendation model to calculate the characteristics of users and each object to determine the objects to be recommended. The information recommendation model needs to be trained based on a large amount of sample data. However, in the real world, large-scale recommendation systems often have millions of users and items (i.e., objects). The interaction records of users clicking / not clicking on items are particularly sparse compared to the overall possible interactions. Therefore, when using these interaction records as sample data to train the information recommendation model, insufficient training may occur, resulting in low accuracy of the information recommendation model. Summary of the Invention
[0004] The embodiments of the present invention provide an information processing method, system, storage medium and server, which improve the accuracy of a trained feature extraction network.
[0005] An embodiment of the present invention provides an information processing method, including:
[0006] Determine the initial network for feature extraction;
[0007] Obtain a first training sample pair, the first training sample pair including at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects;
[0008] Extracting features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object;
[0009] According to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network, the feature extraction initial network is adjusted to train the feature extraction network.
[0010] Another embodiment of the present invention provides an information processing system, including:
[0011] A network determination unit, used to determine an initial network for feature extraction;
[0012] A training sample unit is configured to obtain a first training sample pair, wherein the first training sample pair includes at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects;
[0013] a feature extraction unit, configured to extract features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object;
[0014] The adjustment training unit is used to adjust the feature extraction initial network according to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network to train the feature extraction network.
[0015] Another aspect of the embodiment of the present invention further provides a computer-readable storage medium, which stores a plurality of computer programs. The computer programs are suitable for being loaded by a processor and executing the information processing method as described in one aspect of the embodiment of the present invention.
[0016] Another aspect of the present invention provides a server, including a processor and a memory;
[0017] The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and execute the information processing method as described in one aspect of an embodiment of the present invention; the processor is used to implement each computer program in the multiple computer programs.
[0018] It can be seen that in the method of this embodiment, during the process of training the feature extraction network, a first training sample pair is obtained, such as a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features, a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features, a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence, thereby adjusting the feature extraction initial network based on the relationship information between the feature information of the two sample objects in any sample pair obtained by the feature extraction initial network, and realizing the training of the feature extraction network through comparative learning. In this process, not only the attribute features and semantic features of the sample objects themselves are taken into account, but also the implicit correlation between the sample objects based on the session, so that the accuracy of the trained feature extraction network is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 is a schematic diagram of an information processing method provided by an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of an information processing method provided by an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of an information processing method provided by an application embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of a distributed system to which an information processing method in another application embodiment of the present invention is applied;
[0024] Figure 5 is a schematic diagram of a block structure in another application embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of the logical structure of an information processing system provided by an embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of the logical structure of a server provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the invention described herein can, for example, be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or apparatus.
[0029] The embodiment of the present invention provides an information processing method, such as Figure 1 As shown, the feature extraction network can be trained mainly through the following steps. The feature map extraction network is used to extract the feature information of any object, including:
[0030] Determine the initial network for feature extraction;
[0031] Obtain a first training sample pair, the first training sample pair including at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects;
[0032] Extracting features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object;
[0033] According to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network, the feature extraction initial network is adjusted to train the feature extraction network.
[0034] The feature extraction network described above is a machine learning model based on artificial intelligence (AI). Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0035] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0036] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0037] In this way, in this process, not only the attribute characteristics and semantic characteristics of the sample objects themselves are taken into account, but also the implicit correlation between sample objects based on conversations, so that the accuracy of the trained feature extraction network is improved.
[0038] The embodiment of the present invention provides an information processing method, which is mainly executed by an information recommendation system. The flow chart is as follows: Figure 2 Shown, including:
[0039] Step 101: Determine an initial network for feature extraction.
[0040] It is understood that when determining the initial feature extraction network, the information recommendation system will determine the initial values of the parameters in the multi-layer structure and each layer of the network. The parameters of the initial feature extraction network refer to the fixed parameters used in the calculation process of each layer of the network, which do not need to be constantly assigned values. These parameters include parameter scale, number of network layers, user vector length, and other parameters.
[0041] Specifically, the feature extraction network can be a deep neural network (DNN), a convolutional neural network (CNN), etc.
[0042] Step 102: Obtain a first training sample pair, where the first training sample pair includes at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects.
[0043] In this embodiment, when training the feature extraction network, the contrastive learning training method is mainly used to train the feature extraction network. Here, contrastive learning is a self-supervised learning method that is widely used and verified in the fields of natural language processing and image processing. It mainly compares the similarities between the two sample objects in each sample pair (including positive sample pairs and negative sample pairs) to achieve the training of the feature extraction network. Therefore, in order to train the feature extraction network, it is necessary to include positive sample pairs and negative sample pairs in the first training sample pair obtained, which may include but are not limited to the following positive sample pairs and negative sample pairs:
[0044] (1) The first positive sample pair based on consistent attribute features and the first negative sample pair based on inconsistent attribute features
[0045] Here, attribute features refer to the features of the attributes of the sample objects, such as the identification, label, classification and creation time of the sample objects. Consistent attribute features refer to the same attribute features. In this case, the first positive sample pair may include two identical sample objects, while the first negative sample pair includes two different sample objects.
[0046] (2) The second positive sample pair based on similar semantic features and the second negative sample pair based on dissimilar semantic features
[0047] Specifically, when obtaining the second positive sample pair and the second negative sample pair, the similarity between the semantic features of any sample object and the semantic features of other sample objects can be calculated. If the similarity is greater than a preset value, the sample object and the other sample objects are combined into a second positive sample pair; if the similarity is not greater than the preset value, the sample object and the other sample objects are combined into a second negative sample pair. The semantic feature here refers to information that can describe the semantics contained in the sample object.
[0048] (3) The third positive sample pair based on session co-occurrence and the third negative sample pair based on session non-co-occurrence
[0049] Specifically, when obtaining the third positive sample pair and the third negative sample pair, multiple sessions can be determined, each session including multiple sample objects operated by the user in a continuous time period, and the co-occurrence information of any sample object and other sample objects in multiple sessions is counted, and the co-occurrence information includes the number of co-occurrences; any sample object whose number of co-occurrences is greater than a preset number is combined with other sample objects to form a third positive sample pair, and any sample object whose number of co-occurrences is not greater than the preset number is combined with other sample objects to form a third negative sample pair.
[0050] Here, a session refers to a user operating an object within a continuous period of time. In a session, a user may continuously operate (for example, click) multiple objects through the user terminal. There is some kind of sequential behavioral correlation between these multiple objects beyond semantics, which is an implicit correlation.
[0051] In the process of obtaining the third positive sample pair and the third negative sample pair, if the multiple sessions determined are K1, K2, ..., Kn, and each session includes multiple sample objects k1, k2, ..., km, then the relationship between the multiple sample objects in each session is a co-occurrence relationship. If two sample objects appear in p sessions at the same time, the number of co-occurrences of the two sample objects is p, and the third positive sample pair and the third negative sample pair are determined by the number of co-occurrences.
[0052] For example, a user clicks on an article related to [Beijing Travel], then clicks on an article related to [Beijing Roast Duck], then clicks on an article related to [The Great Wall], and finally clicks on an article related to [Qin Dynasty] within a continuous time period. The articles [Beijing Travel], [Beijing Roast Duck], [The Great Wall], and [Qin Dynasty] have a certain degree of correlation based on the sequential click behavior and can form a session, but their attribute characteristics and semantic correlation are not obvious.
[0053] Step 103: extract features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object.
[0054] Specifically, when obtaining feature information of each sample object in the first positive sample pair and the second negative sample pair, the attributes of each sample object included therein can be input into the feature extraction initial network, thereby obtaining the attribute features of each sample object.
[0055] When obtaining feature information of each sample object in the second positive sample pair and the second negative sample pair, the semantic representation information of each sample object included therein may be input into the feature extraction initial network, thereby obtaining the semantic features of each sample object.
[0056] When obtaining the feature information of each sample object in the above-mentioned third positive sample pair and third negative sample pair, various information of each sample object therein (including attributes and semantic representation information, etc.) can be input into the feature extraction initial network to obtain the comprehensive feature information of each sample object.
[0057] Step 104 , adjusting the feature extraction initial network according to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network to train the feature extraction network.
[0058] Specifically, the information processing system will first calculate the overall loss function related to the feature extraction initial network based on the results obtained by the feature extraction initial network in the above step 103. The overall loss function is mainly a loss function for comparative learning, which is used to indicate the similarity between the feature information of the two sample objects in each positive sample pair extracted by the feature extraction initial network, and the comparison with the similarity between the feature information of the two sample objects in the negative sample pair.
[0059] The training process of the feature extraction network is to minimize the value of the above-mentioned overall loss function. The training process is to continuously optimize the parameter values of the parameters in the feature extraction initial network determined in step 101 through a series of mathematical optimization methods such as backpropagation derivation and gradient descent, and to minimize the calculated value of the above-mentioned overall loss function. Specifically, when the function value of the calculated overall loss function is large, such as greater than a preset value, it is necessary to change the parameter value, such as reducing the weight value of a certain neuron connection, so that the function value of the overall loss function calculated according to the adjusted parameter value is reduced.
[0060] Specifically, since the first training sample pairs involve the positive sample pairs and negative sample pairs in the above-mentioned manner, when calculating the overall loss function of the feature extraction initial network, the following loss functions may be included but are not limited to:
[0061] (1) Calculating a first loss function, where the first loss function includes the ratio of the correlation between the attribute features of one sample object and the augmented attribute features of the other sample object in the first positive sample pair to the correlation between the attribute features of the two sample objects in the first negative sample pair.
[0062] (2) Calculating a second loss function, where the second loss function includes a ratio of a correlation between semantic features of two sample objects in the second positive sample pair to a correlation between semantic features of two samples in the second negative sample pair.
[0063] (3) Calculating a third loss function, where the third loss function includes a ratio of a correlation between feature information of two sample objects in the third positive sample pair and a correlation between semantic features of two samples in the third negative sample pair.
[0064] When calculating the overall loss function of the initial feature extraction network, the overall loss function related to the initial feature extraction network is calculated based on the first loss function, the second loss function, and the third loss function. The parameter values of the parameters in the initial feature extraction network are then adjusted based on the overall loss function. For example, the weighted sum of the first loss function, the second loss function, and the third loss function is used as the overall loss function.
[0065] It should be noted that the above steps 103 to 104 are an adjustment of the parameter values in the feature extraction initial network based on the feature information of each sample object detected by the feature extraction initial network. In actual applications, it is necessary to continuously loop through the above steps 103 to 104 until the adjustment of the parameter values meets certain stopping conditions.
[0066] Therefore, after executing steps 101 to 104 of the above embodiment, the information processing system needs to determine whether the current adjustment of the parameter value meets the preset stop condition. If so, the process ends and the initial feature extraction network after the parameter value adjustment in step 104 is used as the trained feature extraction network. If not, the process returns to executing steps 103 to 104 for the initial feature extraction network after the parameter value adjustment. The preset stop condition includes, but is not limited to, any one of the following conditions: the difference between the currently adjusted parameter value and the last adjusted parameter value is less than a threshold, i.e., the adjusted parameter value reaches convergence; and the number of parameter value adjustments is equal to a preset number, etc.
[0067] Furthermore, in addition to the first training sample pairs described above, the information processing system may also obtain second training sample pairs, each of which includes a fourth positive sample pair having an operational relationship and a fourth negative sample pair not having an operational relationship. Each sample pair in the second training sample includes a sample user and a sample object. The operational relationship means that the sample user has performed an operation on the sample object, such as a click or comment.
[0068] In this way, in the process of obtaining the feature information of each sample object in executing the above step 103, when obtaining the feature information of each sample object and sample user in the fourth positive sample pair and the fourth negative sample pair, various information of each sample object included therein (including attributes and semantic representation information, etc.) can be input into the feature extraction initial network, thereby obtaining the comprehensive feature information of each sample object, and various information of the sample user (including static attributes and dynamic operation information, etc.) can be input into the feature extraction initial network, thereby obtaining the comprehensive feature information of each sample user.
[0069] Furthermore, during the execution of step 104, a fourth loss function needs to be calculated. The fourth loss function includes the ratio of the correlation between the feature information of the sample user and the sample object in the fourth positive sample pair to the correlation between the feature information of the sample user and the sample object in the fourth negative sample pair. In this case, when calculating the overall loss function, the overall loss function associated with the initial feature extraction network is calculated based on the first loss function, the second loss function, the third loss function, and the fourth loss function. For example, the weighted sum of the first loss function, the second loss function, the third loss function, and the fourth loss function is used as the overall loss function.
[0070] It can be seen that in the method of this embodiment, during the process of training the feature extraction network, a first training sample pair is obtained, such as a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features, a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features, a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence, thereby adjusting the feature extraction initial network based on the relationship information between the feature information of the two sample objects in any sample pair obtained by the feature extraction initial network, and realizing the training of the feature extraction network through comparative learning. In this process, not only the attribute features and semantic features of the sample objects themselves are taken into account, but also the implicit correlation between the sample objects based on the session, so that the accuracy of the trained feature extraction network is improved.
[0071] The following is a specific application example to illustrate the information processing method of the present invention. In this embodiment, the trained feature extraction network is mainly applied to the application of information recommendation. In this way, the trained feature extraction network can be used to extract the first feature information of the user to be recommended and the second feature information of each object to be recommended in the system, thereby determining whether to recommend the object to be recommended to the user to be recommended based on the correlation between the first feature information and the second feature information.
[0072] Specifically, this feature extraction network can be applied to the recall module of an information recommendation system. This module can quickly retrieve hundreds of items from the recommended item pool and output them to the downstream ranking module, ultimately producing recommendation results. Items here refer to the various objects involved in the information recommendation system, such as news, videos, text, and audio.
[0073] like Figure 3 As shown in the figure, when the feature extraction network is trained offline, it can be achieved through the following steps, including:
[0074] Step 201: Determine the structure of the initial feature extraction network, which may be a DNN or a CNN, and determine the initial values of the parameters in the initial feature extraction network.
[0075] Step 202: Obtain a first training sample pair, which may specifically include a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects.
[0076] When obtaining the first positive sample pair and the first negative sample pair, each item d in the information recommendation system can be i , item d i Can be used with this item i Itself constitutes the first positive sample pair, item d i Can be combined with other items j Form the first negative-positive sample pair.
[0077] When obtaining the second positive sample pair and the second negative sample pair, for each item d in the information recommendation system, i , obtain the semantic vector of each item, and select the item d according to the distance between the semantic vectors of the items (such as the cosine distance between the semantic vectors of the titles of two items). i The nearest k items are obtained with item d i Positive example set P of items with similar semantic features i t, and other items can form the negative example set of items Thus, item d i Can be combined with item examples The items in the second positive sample pair are composed of items d i Can be combined with item negative example set The items in compose the second negative-positive sample pairs.
[0078] When obtaining the third positive sample pair and the third negative sample pair, multiple sessions recorded in the information recommendation system may be used, each session including items d operated by the user through the user terminal in a continuous time period. i , count the number of co-occurrences of any two items in multiple sessions. The greater the number of co-occurrences, the greater the session-based sequence behavior similarity between the two items. i , we can select a group of items with the largest number of co-occurrences with the item to form item d i Item positive example set based on session co-occurrence Other items can form the negative example set of items Thus, item d i Can be combined with item examples The items in the third positive sample pair are d i Can be combined with item negative example set The items in compose the third negative-positive pair.
[0079] Step 203 : Obtain second training sample pairs, including a fourth positive sample pair having an operation relationship and a fourth negative sample pair having no operation relationship. Each sample pair in the second training sample includes a sample user and a sample object.
[0080] Step 204 : Acquire feature information of each sample object and feature information of each sample object respectively through the above feature extraction initial network.
[0081] Step 205 , calculate the overall loss function associated with the feature extraction initial network.
[0082] Specifically, in this embodiment, the overall loss function may include but is not limited to the following loss functions:
[0083] (1) The first loss function L fea
[0084] Specifically, for any of the above first positive samples, the item d i , obtain item d through the above feature extraction initial network i The attribute characteristics f i d, and use a certain method to augment the attribute feature to obtain the augmented attribute feature For example, using the element-level dropout method, some elements in the attribute features are randomly set to 0 with a certain probability, as shown in the following formula 1-1:
[0085]
[0086] Then determine a similarity function g f () can be expressed by the following formula 1-2, where MLP is a multi-layer perceptron neural network:
[0087]
[0088] Then the first loss function can be expressed by the following formula 1-3:
[0089]
[0090] Among them, d i ∈B means that the item belongs to the item set of the training batch B, and d j ∈N Bi Indicates item d i ∈B, but That is, except for item d i All other items in the item set of batch B except , τ represents the hyper parameter of temperature in contrastive learning.
[0091] During the feature extraction network training process, the initial feature extraction network needs to be optimized based on the first loss function so that the distance between the attribute features of an item and its augmented attribute features is close, while the distance between the attribute features of other items is farther.
[0092] (2) Second loss function L sem
[0093] Specifically, the second loss function can be expressed by the following formula 2:
[0094]
[0095] Among them, d i and d k are the semantic features of each item in the second positive sample pair obtained by the initial network for feature extraction, d i and d j Semantic features of each item in the second negative sample pair obtained by the initial network for the above feature extraction,
[0096] In the process of training the feature extraction network, it is necessary to optimize the feature extraction initial network based on the second loss function so that the positive example set P of the item i t The feature information of any two items in the set is more similar, and the negative example set of items The distance between the feature information of any two items in is greater.
[0097] (3) The third loss function L sess
[0098] Specifically, the third loss function can be expressed by the following formula 3:
[0099]
[0100] In the process of training the feature extraction network, it is necessary to optimize the feature extraction initial network based on the third loss function so that an item and the item positive example set P i s The feature information of the items in the set is closer, and the negative example set of items The distance between the feature information of items is greater.
[0101] (4) The fourth loss function L S
[0102] Specifically, the fourth loss function can be expressed by the following formula 4:
[0103]
[0104] Among them, exp() is the exponential function, u i ,d k The fourth positive sample pair S obtained by the initial network for feature extraction + The feature information corresponding to the user and item, u i ,d j The fourth negative sample pair S obtained by the initial network for feature extraction - The feature information corresponding to users and items respectively.
[0105] In the process of training the feature extraction network, it is necessary to optimize the feature extraction initial network based on the fourth loss function so that the distance between the feature information of an item and the items with which it has an operational relationship is closer, and the distance between the feature information of an item and the items with which it has no operational relationship is farther.
[0106] It should be noted that the first loss function, the second loss function and the third loss function are calculated for the first training sample pair obtained in the above step 202, and the fourth loss function is calculated for the second training sample pair obtained in the above step 203.
[0107] Furthermore, after obtaining the first loss function, the second loss function, the third loss function, and the fourth loss function, the overall loss function related to the feature extraction initial network can be calculated, which can be specifically expressed by the following formula 5, where λ1, λ2, and λ3 are the weight values of the corresponding loss functions:
[0108] L=L s +λ1L fea +λ2L sem +λ3L sess (5)
[0109] Step 206: Adjust the parameter values of the parameters in the feature extraction initial network determined above according to the overall loss function calculated above.
[0110] In step 207, it is determined whether the adjustment of the parameter values of the parameters in the feature extraction initial network meets the preset conditions. If so, the feature extraction initial network after the parameter values are adjusted in step 206 is used as the trained feature extraction network; if not, the process returns to execute the above step 202.
[0111] It can be seen that in the process of training the feature extraction network in this embodiment, the attribute characteristics, semantic characteristics and feature information of the sample objects based on session co-occurrence are integrated to train the feature extraction network, which can make up for the problem caused by sample sparsity in the information recommendation system and improve the accuracy of the feature extraction network to a certain extent.
[0112] In actual application, information recommendation models (including FM, DeepFM, AutoInt, and RFM) were trained using existing methods on the video dataset Video-636M and the news dataset News-76M. A feature extraction network was also trained using the method of this embodiment and applied to the information recommendation system. Performance measurement parameters were calculated, such as hit rate@x, which represents the accuracy of information among the top x recommended pieces of information. As shown in Table 1 below, it can be seen that the feature extraction network trained using the embodiment of the present invention achieved significant improvements:
[0113] Table 1
[0114]
[0115] Furthermore, the feature extraction network trained according to the method of this embodiment is applied to the information recommendation system, and statistics are collected on users accessing the information recommendation system through user terminals to obtain parameters such as average number of clicks per person (ACC), average reading time per person (ADT), average number of clicks per person for cold start users (ACC-c), average reading time per person for cold start users (ADT-c), average number of likes per person (ALR), and average number of shares per person (AFR). As shown in Table 2 below, the proportion of information recommended by users through user terminal operations in the information recommendation system has increased to a certain extent:
[0116] Table 2
[0117]
[0118] It should also be noted that the above-mentioned feature extraction network is mainly trained using the data set in the information recommendation system, so that the trained feature extraction network is applied to the information recommendation system. In other embodiments, the feature extraction network can also be trained using data sets in other application systems, and the trained feature extraction network can be applied to other application systems. This will not be elaborated here.
[0119] The following is another specific application example to illustrate the information processing method of the present invention. The information processing system in the embodiment of the present invention is mainly a distributed system 100, which may include a client 300 and multiple nodes 200 (any form of computing device in the access network, such as a server, a user terminal), and the client 300 and the node 200 are connected through network communication.
[0120] Taking the distributed system as the blockchain system as an example, see Figure 4 This is a schematic diagram of an optional architecture for a distributed system 100 provided in an embodiment of the present invention, applied to a blockchain system. The system consists of multiple nodes 200 (any type of computing device connected to a network, such as a server or user terminal) and clients 300. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP). In a distributed system, any machine, such as a server or terminal, can join and become a node. Nodes include hardware, middleware, operating system, and application layers.
[0121] See also Figure 4 The functions of each node in the blockchain system shown include:
[0122] 1) Routing: A basic function of a node, used to support communication between nodes.
[0123] In addition to the routing function, nodes can also have the following functions:
[0124] 2) Applications, deployed in the blockchain, implement specific services based on actual business needs, record data related to the implementation of functions to form record data, carry digital signatures in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system for other nodes to add the record data to a temporary block when they successfully verify the source and integrity of the record data.
[0125] For example, the services implemented by an application include: code that implements information processing functions, which mainly include:
[0126] Determine an initial network for feature extraction; obtain a first training sample pair, the first training sample pair including at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects; perform feature extraction on each sample object in the first training sample pair through the initial network for feature extraction to obtain feature information of each sample object; adjust the initial network for feature extraction based on the relationship information between the feature information of two sample objects in any sample pair obtained by the initial network for feature extraction to train the feature extraction network.
[0127] 3) Blockchain, including a series of blocks that are connected to each other in the order of their generation. Once a new block is added to the blockchain, it will not be removed. The block records the record data submitted by the nodes in the blockchain system.
[0128] See also Figure 5 This is an optional schematic diagram of the block structure provided by an embodiment of the present invention. Each block includes the hash value of the transaction records stored in the block (the hash value of the current block) and the hash value of the previous block. The blocks are connected by hash values to form a blockchain. In addition, the block may also include information such as the timestamp when the block was generated. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains relevant information used to verify the validity of the information (anti-counterfeiting) and generate the next block.
[0129] The embodiment of the present invention also provides an information processing system, the structural diagram of which is shown as follows: Figure 6 Specifically, it may include:
[0130] The network determination unit 10 is used to determine the initial network for feature extraction.
[0131] The training sample unit 11 is used to obtain a first training sample pair, which includes at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects.
[0132] Specifically, if the first training sample pair includes a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features, the training sample unit 11 is specifically used to calculate the similarity between the semantic features of any sample object and the semantic features of other sample objects; if the similarity is greater than a preset value, the any sample object and the other sample objects are combined into a second positive sample pair; if the similarity is not greater than the preset value, the any sample object and the other sample objects are combined into a second negative sample pair.
[0133] If the first training sample pair includes a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence, the training sample unit 11 is specifically used to determine multiple sessions, each session including multiple sample objects operated by the user in a continuous time period; count the co-occurrence information of any sample object and other sample objects in the multiple sessions, the co-occurrence information including the number of co-occurrences; any sample object whose number of co-occurrences is greater than a preset number is combined with other sample objects to form a third positive sample pair, and any sample object whose number of co-occurrences is not greater than the preset number is combined with other sample objects to form a third negative sample pair.
[0134] Furthermore, the training sample unit 11 is also used to obtain second training sample pairs, which include: a fourth positive sample pair with an operation relationship and a fourth negative sample pair without an operation relationship, and each sample pair in the second training sample includes a sample user and a sample object.
[0135] The feature extraction unit 12 is configured to extract features from each sample object in the first training sample pair obtained by the training sample unit 11 using the feature extraction initial network determined by the network determination unit 10 to obtain feature information of each sample object.
[0136] The feature extraction unit 12 is specifically used to calculate a first loss function, which includes the ratio of the correlation between the attribute features of one sample object and the augmented attribute features of another sample object in the first positive sample pair to the correlation between the attribute features of the two sample objects in the first negative sample pair; calculate a second loss function, which includes the ratio of the correlation between the semantic features of the two sample objects in the second positive sample pair to the correlation between the semantic features of the two samples in the second negative sample pair; calculate a third loss function, which includes the ratio of the correlation between the feature information of the two sample objects in the third positive sample pair to the correlation between the semantic features of the two samples in the third negative sample pair; calculate an overall loss function related to the feature extraction initial network based on the first loss function, the second loss function and the third loss function; and adjust the parameter values of the parameters in the feature extraction initial network based on the overall loss function.
[0137] Furthermore, the feature extraction unit 12 is also used to calculate a fourth loss function, wherein the fourth loss function includes the ratio of the correlation between the feature information of the sample user and the sample object in the fourth positive sample pair and the correlation between the feature information of the sample user and the sample object in the fourth negative sample pair; when calculating the overall loss function related to the feature extraction initial network according to the first loss function, the second loss function and the third loss function, it is specifically used to calculate the overall loss function related to the feature extraction initial network according to the first loss function, the second loss function, the third loss function and the fourth loss function.
[0138] The adjustment training unit 13 is used to adjust the feature extraction initial network according to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network in the feature extraction unit 12 to train the feature extraction network.
[0139] The adjustment training unit 13 is further configured to stop adjusting the parameter value when the number of times the parameter value is adjusted equals a preset number, or when the difference between the currently adjusted parameter value and the last adjusted parameter value is less than a threshold.
[0140] It can be seen that in the process of training the feature extraction network of the system of this embodiment, the training sample unit 11 will obtain a first training sample pair, such as a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features, a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features, a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence, thereby adjusting the relationship information between the feature information of the two sample objects in any sample pair obtained by the training unit 13 based on the initial feature extraction network, adjusting the initial feature extraction network, and realizing the training of the feature extraction network through comparative learning. In this process, not only the attribute features and semantic features of the sample objects themselves are taken into account, but also the implicit correlation between the sample objects based on the session, so that the accuracy of the trained feature extraction network is improved.
[0141] The embodiment of the present invention further provides a server, the structural diagram of which is shown in FIG. Figure 7 As shown, the server may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 20 (e.g., one or more processors) and memory 21, and one or more storage media 22 (e.g., one or more mass storage devices) storing application programs 221 or data 222. Memory 21 and storage medium 22 may be temporary storage or persistent storage. The program stored in storage medium 22 may include one or more modules (not shown), each module may include a series of instruction operations on the server. Furthermore, the CPU 20 may be configured to communicate with the storage medium 22 to execute a series of instruction operations in the storage medium 22 on the server.
[0142] Specifically, the application 221 stored in the storage medium 22 includes an information processing application, and the application may include the network determination unit 10, training sample unit 11, feature extraction unit 12, and adjustment training unit 13 in the above-mentioned information processing system, which are not described in detail here. Furthermore, the central processing unit 20 can be configured to communicate with the storage medium 22 and execute a series of operations corresponding to the information processing application stored in the storage medium 22 on the server.
[0143] The server may also include one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input and output interfaces 25, and / or one or more operating systems 223, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0144] The steps performed by the information processing system in the above method embodiment can be based on the Figure 7 The structure of the server is shown.
[0145] Furthermore, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of computer programs, wherein the computer programs are suitable for being loaded by a processor and executing the information processing method executed by the above-mentioned information processing system.
[0146] Another aspect of the present invention provides a server, including a processor and a memory;
[0147] The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and executed by the information processing method executed by the above-mentioned information processing system; the processor is used to implement each computer program in the multiple computer programs.
[0148] In addition, according to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the information processing methods provided in the various optional implementations described above.
[0149] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0150] The above is a detailed introduction to an information processing method, system, storage medium and server provided in an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. An information processing method, characterized in that: include: Determine the initial network for feature extraction; Obtain a first training sample pair, the first training sample pair including at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects, the attribute refers to an attribute of the sample object, the session refers to a sample object operated by a user within a continuous time period, and the sample object is an article to be browsed; Extracting features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object; Adjusting the feature extraction initial network based on the relationship between the feature information of two sample objects in any sample pair obtained by the feature extraction initial network to train the feature extraction network; wherein the method includes: calculating a first loss function, the first loss function including the ratio of the correlation between the attribute features of one sample object in the first positive sample pair and the augmented attribute features of the other sample object to the correlation between the attribute features of the two sample objects in the first negative sample pair; the first loss function is related to the attribute features of the sample objects themselves; Calculating a second loss function, where the second loss function includes a ratio of a correlation between semantic features of two sample objects in the second positive sample pair to a correlation between semantic features of two samples in the second negative sample pair; the second loss function is related to the semantic features of the sample objects themselves; Calculating a third loss function, the third loss function including a ratio of a correlation between feature information of two sample objects in the third positive sample pair to a correlation between semantic features of two samples in the third negative sample pair; the third loss function being related to an implicit correlation between the sample objects based on the conversation; Calculating an overall loss function associated with the feature extraction initial network based on the first loss function, the second loss function, and the third loss function; The parameter values of the parameters in the feature extraction initial network are adjusted according to the overall loss function.
2. The method according to claim 1, wherein The first training sample pair includes a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features. The obtaining of the first training sample pair specifically includes: Calculate the similarity between the semantic features of any sample object and the semantic features of other sample objects; If the similarity is greater than a preset value, any sample object and other sample objects are combined into a second positive sample pair; if the similarity is not greater than the preset value, any sample object and other sample objects are combined into a second negative sample pair.
3. The method according to claim 1, wherein The first training sample pair includes a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence, and obtaining the first training sample pair specifically includes: Determine a plurality of sessions, each session including a plurality of sample objects operated by a user within a continuous time period; Counting co-occurrence information between any sample object and other sample objects in the multiple sessions, wherein the co-occurrence information includes the number of co-occurrences; Any sample object whose co-occurrence times are greater than a preset number is combined with other sample objects to form a third positive sample pair, and any sample object whose co-occurrence times are not greater than the preset number is combined with other sample objects to form a third negative sample pair.
4. The method according to any one of claims 1 to 3, wherein The adjusting the feature extraction initial network according to the relationship information between the feature information of two sample objects in any sample pair obtained by the feature extraction module specifically includes: Calculating a first loss function, where the first loss function includes a ratio of a correlation between an attribute feature of one sample object and an augmented attribute feature of another sample object in the first positive sample pair to a correlation between the attribute features of the two sample objects in the first negative sample pair; Calculating a second loss function, where the second loss function includes a ratio of a correlation between semantic features of two sample objects in the second positive sample pair to a correlation between semantic features of two samples in the second negative sample pair; Calculating a third loss function, the third loss function including a ratio of a correlation between feature information of two sample objects in the third positive sample pair and a correlation between semantic features of two samples in the third negative sample pair; Calculating an overall loss function associated with the feature extraction initial network based on the first loss function, the second loss function, and the third loss function; The parameter values of the parameters in the feature extraction initial network are adjusted according to the overall loss function.
5. The method according to claim 4, wherein The method further comprises: A second training sample pair is obtained, where the second training sample pair includes: a fourth positive sample pair having an operation relationship and a fourth negative sample pair having no operation relationship, and each sample pair in the second training sample includes a sample user and a sample object.
6. The method according to claim 4, wherein When the number of times the parameter value is adjusted is equal to a preset number, or when the difference between the currently adjusted parameter value and the last adjusted parameter value is less than a threshold, the adjustment of the parameter value is stopped.
7. An information processing system, characterized in that: include: A network determination unit, used to determine an initial network for feature extraction; A training sample unit is configured to obtain a first training sample pair, wherein the first training sample pair includes at least one of the following sample pairs: a first positive sample pair based on consistent attribute features and a first negative sample pair based on inconsistent attribute features; a second positive sample pair based on similar semantic features and a second negative sample pair based on dissimilar semantic features; a third positive sample pair based on session co-occurrence and a third negative sample pair based on session non-co-occurrence; wherein each sample pair includes two sample objects, the attribute refers to an attribute of the sample object, the session refers to a sample object operated by a user within a continuous time period, and the sample object is an article to be browsed; a feature extraction unit, configured to extract features from each sample object in the first training sample pair using the feature extraction initial network to obtain feature information of each sample object; An adjustment training unit is used to adjust the feature extraction initial network to train the feature extraction network based on the relationship information between the feature information of the two sample objects in any sample pair obtained by the feature extraction initial network; the adjustment training unit is specifically used to: calculate a first loss function, the first loss function includes the correlation between the attribute features of one sample object in the first positive sample pair and the augmented attribute features of the other sample object, and the ratio of the correlation between the attribute features of the two sample objects in the first negative sample pair; the first loss function is related to the attribute features of the sample object itself; calculate a second loss function, the second loss function includes the semantic features of the two sample objects in the second positive sample pair The method comprises the following steps: calculating a correlation between the feature information of the two sample objects in the third positive sample pair and a ratio of the correlation between the semantic features of the two samples in the second negative sample pair; calculating a correlation between the feature information of the two sample objects in the third positive sample pair and a ratio of the correlation between the semantic features of the two samples in the third negative sample pair; calculating a correlation between the feature information of the two sample objects in the third positive sample pair and a ratio of the correlation between the semantic features of the two samples in the third negative sample pair; calculating a correlation between the feature objects and the implicit correlation between the sample objects based on the session; calculating an overall loss function related to the feature extraction initial network according to the first loss function, the second loss function and the third loss function; and adjusting the parameter values of the parameters in the feature extraction initial network according to the overall loss function.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of computer programs, and the computer programs are suitable for being loaded by a processor and executing the information processing method according to any one of claims 1 to 6.
9. A server, characterized in that: including processor and memory; The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and executed by the information processing method according to any one of claims 1 to 6; the processor is used to implement each computer program in the multiple computer programs.
Citation Information
Patent Citations
Processing method and system for search services
CN107665220A
Text processing method, device and equipment and storage medium
CN112084789A
Short text topic distribution reasoning method and system, computer equipment and storage medium
CN112183108A
Image processing model training method, image classification method and device
CN114299363A