A model denoising method and device, electronic equipment and storage medium
By denoising noisy labels in the information recommendation system and using loss functions to optimize model training, the problem of noisy labels affecting model accuracy is solved, achieving higher accuracy in information recommendation and sharing.
Patent Information
- Application Number
- CN202210097794.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-01-27
AI Technical Summary
In information recommendation systems, the existence of noisy labels leads to low model accuracy, which affects the accuracy of information recommendation.
By obtaining a set of samples carrying noise labels and inputting them into the trained target model to perform noise label denoising, the loss function is used to optimize model training, the noise labels are directly removed, and the model accuracy is improved.
The accuracy of the model is improved, thereby improving the accuracy of information recommendation and information sharing.
Smart Images

Figure CN114418123B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the fields of deep learning, information recommendation, etc. BACKGROUND
[0002] With the development of mobile Internet, terminal devices (such as mobile phones, tablets, etc.) can implement information sharing, information recommendation and other information processing operations relying on mobile Internet. Artificial intelligence technology can be used to better and faster implement these information processing operations.
[0003] Taking an information recommendation system as an example, a target model (i.e., a trained model obtained by modeling using artificial intelligence technology) can be deployed in the software / hardware design of the information recommendation system to improve the processing accuracy of the software / hardware through the target model. However, there is noise interference in the modeling process, which leads to low model precision and thus affects the accuracy of information recommendation. SUMMARY
[0004] The present disclosure provides a model denoising method, an information recommendation method, an information sharing method, an apparatus, an electronic device, and a storage medium.
[0005] According to an aspect of the present disclosure, a model denoising method is provided, comprising:
[0006] obtaining a sample set carrying noise labels;
[0007] inputting the sample set carrying noise labels into a trained target model to perform denoising processing of the noise labels, and obtaining a denoised model output result.
[0008] According to another aspect of the present disclosure, an information recommendation method is provided, comprising:
[0009] obtaining interest preferences for a plurality of information contents based on information recommendation operations;
[0010] performing feature extraction on the interest preferences for the plurality of information contents to obtain a plurality of first features;
[0011] inputting the plurality of first features into the trained target model of the above-mentioned model denoising method to perform denoising processing of noise labels corresponding to the plurality of first features, and obtaining a denoised model output result;
[0012] performing information recommendation processing according to the denoised model output result to obtain an information recommendation result.
[0013] According to another aspect of the present disclosure, an information sharing method is provided, comprising:
[0014] obtaining a plurality of operation contents based on information sharing operations;
[0015] perform feature extraction on the plurality of operation contents to obtain a plurality of second features;
[0016] input the plurality of second features into the trained target model to perform noise label denoising processing corresponding to the plurality of second features, to obtain a denoised model output result;
[0017] perform information sharing processing according to the denoised model output result to obtain an information sharing result.
[0018] According to another aspect of the present disclosure, a model denoising device is provided, comprising:
[0019] an acquisition unit configured to acquire a sample set carrying noise labels;
[0020] a denoising unit configured to input the sample set carrying noise labels into a trained target model to perform noise label denoising processing, to obtain a denoised model output result.
[0021] According to another aspect of the present disclosure, an information recommendation device is provided, comprising:
[0022] a preference acquisition unit configured to acquire interest preferences for a plurality of information contents based on information recommendation operations;
[0023] a first feature extraction unit configured to perform feature extraction on the interest preferences for the plurality of information contents to obtain a plurality of first features;
[0024] a first label denoising unit configured to input the plurality of first features into the trained target model to perform noise label denoising processing corresponding to the plurality of first features, to obtain a denoised model output result;
[0025] an information recommendation unit configured to perform information recommendation processing according to the denoised model output result to obtain an information recommendation result.
[0026] According to another aspect of the present disclosure, an information sharing device is provided, comprising:
[0027] an operation content acquisition unit configured to acquire a plurality of operation contents based on information sharing operations;
[0028] a second feature extraction unit configured to perform feature extraction on the plurality of operation contents to obtain a plurality of second features;
[0029] A second label denoising unit is configured to input the plurality of second features into the trained target model described in the above-mentioned model denoising method, perform denoising processing on the noise labels corresponding to the plurality of second features, and obtain a denoised model output result;
[0030] The information sharing unit is used to perform information sharing processing according to the denoised model output result to obtain an information sharing result.
[0031] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0032] at least one processor; and
[0033] a memory communicatively connected to the at least one processor; wherein,
[0034] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by any embodiment of the present disclosure.
[0035] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to enable the computer to execute the method provided by any one of the embodiments of the present disclosure.
[0036] According to another aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which implement the method provided in any embodiment of the present disclosure when executed by a processor.
[0037] By adopting the present disclosure, a sample set carrying noise labels can be obtained, the sample set carrying noise labels can be input into a trained target model, and the noise labels can be denoised to obtain a model output result after denoising, thereby improving the model accuracy. Taking the trained target model as an example, the improvement of model accuracy can also improve the accuracy of information recommendation by deploying the trained target model in an information recommendation system.
[0038] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0040] Figure 1 is a schematic diagram of a distributed cluster processing scenario according to an embodiment of the present disclosure;
[0041] Figure 2is a flow chart of a model denoising method according to an embodiment of the present disclosure;
[0042] Figure 3 A flowchart of an information recommendation method according to an embodiment of the present disclosure;
[0043] Figure 4 A flowchart of an information sharing method according to an embodiment of the present disclosure;
[0044] Figure 5 is a schematic diagram of the structure of a model noise reduction device according to an embodiment of the present disclosure;
[0045] Figure 6 is a schematic diagram of the structure of an information recommendation device according to an embodiment of the present disclosure;
[0046] Figure 7 is a schematic diagram of the structure of an information sharing device according to an embodiment of the present disclosure;
[0047] Figure 8 It is a block diagram of an electronic device used to implement the model noise reduction method, information recommendation method and information sharing method of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0048] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0049] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.
[0050] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0051] According to an embodiment of the present disclosure, Figure 1 This is a schematic diagram of a distributed cluster processing scenario according to an embodiment of the present disclosure. The distributed cluster system is an example of a cluster system. It exemplarily describes how the distributed cluster system can be used for model denoising. The present disclosure is not limited to model denoising on a single machine or multiple machines. The use of distributed processing can further improve the accuracy of model denoising. Figure 1 As shown, the distributed cluster system includes multiple nodes (such as server cluster 101, server 102, server cluster 103, server 104, server 105, and server 105 can also be connected to electronic devices, such as mobile phone 1051 and desktop computer 1052), and one or more model denoising tasks can be performed together between multiple nodes, and between multiple nodes and connected electronic devices. Optionally, multiple nodes in the distributed cluster system can adopt a data-parallel model training method, and multiple nodes can perform model denoising training tasks based on the same training method to better train the model; if multiple nodes in the distributed cluster system adopt a model-parallel model training method, multiple nodes can perform model denoising training tasks based on different training methods to better train the model. Optionally, after each round of model training is completed, data exchange (such as data synchronization) can be performed between multiple nodes.
[0052] According to an embodiment of the present disclosure, a model denoising method is provided. Figure 2 This is a flow chart of the model denoising method according to an embodiment of the present disclosure. The method can be applied to a model denoising device. For example, the device can be deployed in a terminal or server or other processing device in a single machine, multi-machine or cluster system to implement model denoising and other processing. The terminal can be a user equipment (UE, User Equipment), a mobile device, a personal digital assistant (PDA, Personal Digital Assistant), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementations, the method can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 2 As shown, this method is applied to Figure 1 Any node or electronic device (mobile phone or desktop computer, etc.) in the cluster system shown includes:
[0053] S201: Obtain a sample set carrying noise labels.
[0054] S202: Input the sample set carrying the noise label into the trained target model, perform noise reduction processing on the noise label, and obtain the model output result after noise reduction.
[0055] In an example of S201-S202, there is noise in the sample label (a label with noise is called a noise label), and the sample set includes sample data and sample labels. According to different application scenarios, the sample data and sample labels in the sample set are different. Taking the information recommendation scenario as an example, the sample data is used to characterize the residence status after the first operation is performed on the first information. For example, for web browsing, after the user clicks on a resource on the web page (such as link information), the sample data can be the residence status of the information content displayed on the current information page (the residence status can be a statistical calculation of how long the user stays on the current information page); the sample label is used to characterize the residence time after the first operation is performed on the first information. For example, the sample label can be the residence time of the user on the current information page. The longer the residence time, the more interested the user is in the information content or similar information content in the current information page, and the user preference is obtained. Therefore, by statistically analyzing the user preference, it is possible to better recommend information to the user and improve the accuracy of information recommendation. Considering the existence of noise labels in the sample labels, since noise labels cannot well count user preferences and are interference factors for information recommendation, it is necessary to remove the noise labels. Noise reduction can be achieved through modeling. The sample set carrying noise labels is input into the trained target model, and the noise labels are denoised to obtain the model output results after denoising.
[0056] By using the embodiments of the present disclosure, a sample set carrying noise labels can be obtained, and the sample set carrying noise labels can be input into a trained target model, and the noise labels can be denoised to obtain a model output result after denoising. Thus, the model accuracy can be improved. Taking the trained target model as an example, the deployment of the trained target model in an information recommendation system can improve the accuracy of information recommendation.
[0057] In one embodiment, a sample set carrying noise labels is input into a trained target model, noise labels are denoised, and a model output result after denoising is obtained. This includes: when the target model is a time-lapse regression model, the sample set carrying noise labels is input into the time-lapse regression model, noise labels are denoised, and the model output result is a target value that does not include noise. Using this embodiment, taking the information recommendation scenario as an example, a time-lapse regression model can be used. That is, by using the time-lapse regression model and based on the sample data input into the time-lapse regression model and the sample labels carrying noise labels (such as the time a user stays on the current information page), better information recommendation results can be obtained. Furthermore, noise labels are directly denoised by the time-lapse regression model, and the model output result is a target value that does not include noise. The model accuracy after denoising is higher. Taking information recommendation as an example, deploying the time-lapse regression model in an information recommendation system can also achieve more accurate information recommendation results. It should be noted that the target model is not limited to the above-mentioned time-lapse regression model according to different application scenarios, but can also be other models for continuous value estimation.
[0058] In one embodiment, the method further includes: according to a sample set carrying a noise label (such as a sample label in the sample set is denoted as y i , the noise labels are included in the y i The target model is trained by back propagation of the loss function to obtain a trained target model. In this embodiment, the target model can be trained based on the sample set carrying noise labels and the predicted value for maximizing the probability of noise labels (such as denoted as p i ) to obtain the loss function, and use the loss function to train the target model to obtain a trained target model. The model output result of the trained target model is the target value excluding noise. The model accuracy is higher after denoising. Taking information recommendation as an example, the trained target model is deployed in the information recommendation system, making the information recommendation result more accurate.
[0059] In one embodiment, a loss function is obtained based on a sample set carrying noise labels and a predicted value for maximizing the probability of the noise labels, including: extracting sample labels from the sample set carrying noise labels, obtaining a first difference between the sample labels and the predicted values, and obtaining a loss function based on the first difference and the variance of the sample labels. In this embodiment, the loss function is obtained by calculating the first method of the loss function, for example, based on the first difference between the sample labels and the predicted values and the variance of the sample labels. After that, a derivative is performed in the back propagation of the loss function to train the target model, thereby obtaining a trained target model. Since the model output of the trained target model is a target value that does not include noise, the model accuracy is higher after denoising. Taking information recommendation as an example, the trained target model is deployed in the information recommendation system, making the information recommendation results more accurate.
[0060] In one embodiment, a loss function is derived based on a set of samples carrying noisy labels and a predicted value that maximizes the probability of the noisy labels. This includes extracting sample labels from the set of samples carrying noisy labels, obtaining a first difference between the sample labels and the predicted value, and obtaining a loss function based on the first difference, the variance of the sample labels, and a hyperparameter used for smoothing. This embodiment optimizes the calculation of the loss function by using a second method, for example, based on the first difference between the sample labels and the predicted value, the variance of the sample labels, and a hyperparameter used for smoothing, compared to the first method of calculating the loss function described above, thereby avoiding the occurrence of zero derivative results in subsequent derivation. Subsequently, a derivative is performed in the backpropagation of the loss function to train a target model, thereby obtaining a trained target model. Because the trained target model outputs a target value that excludes noise, the model accuracy is higher after denoising. For example, deploying this trained target model in an information recommendation system can result in more accurate recommendation results, taking information recommendation as an example.
[0061] According to an embodiment of the present disclosure, a method for recommending information is provided. Figure 3 is a flow chart of the information recommendation method according to an embodiment of the present disclosure, such as Figure 3 Shown, including:
[0062] S301: Obtain interest preferences for multiple information contents based on an information recommendation operation.
[0063] S302: Extract features of interest preferences for multiple information contents to obtain multiple first features.
[0064] S303: Input the multiple first features into the trained target model, perform denoising on the noise labels corresponding to the multiple first features, and obtain a denoised model output result.
[0065] S304: Execute information recommendation processing according to the denoised model output result to obtain an information recommendation result.
[0066] In an example of S301-S304, the trained target model obtained based on the loss function of the above-mentioned embodiment of the present disclosure can be applied to obtain interest preferences for multiple information contents based on the information recommendation operation, and then feature extraction is performed on the interest preferences for multiple information contents to obtain multiple first features. The multiple first features are input into the trained target model, and noise reduction processing is performed on the noise labels corresponding to the multiple first features to obtain the model output result after noise reduction, so as to perform information recommendation processing according to the model output result after noise reduction, so as to obtain more accurate information recommendation results.
[0067] According to an embodiment of the present disclosure, a method for sharing information is provided. Figure 4 is a flow chart of the information recommendation method according to an embodiment of the present disclosure, such as Figure 4 Shown, including:
[0068] S401: Acquire multiple operation contents based on the information sharing operation.
[0069] S402: Extract features from the multiple operation contents to obtain multiple second features.
[0070] S403: Input the plurality of second features into the trained target model, perform denoising on the noise labels corresponding to the plurality of second features, and obtain a denoised model output result.
[0071] S404: Execute information sharing processing according to the denoised model output result to obtain an information sharing result.
[0072] In an example of S401-S404, the trained target model obtained based on the loss function of the above-mentioned embodiment of the present disclosure can be applied. After obtaining multiple operation contents based on the information sharing operation, feature extraction is performed on the multiple operation contents to obtain multiple second features. The multiple second features are input into the trained target model, and noise reduction processing is performed on the noise labels corresponding to the multiple second features to obtain the model output result after noise reduction. Then, information sharing processing is performed according to the model output result after noise reduction, so as to obtain a more accurate information sharing result.
[0073] The following is an example of the model denoising method provided by the above-mentioned embodiment of the present disclosure.
[0074] Taking information recommendation as an example, an information recommendation system can predict a user's interests and preferences based on their actions on products (apps, websites, etc.), ultimately providing personalized information recommendations. Throughout the recommendation process, data may be generated from various sources, including the user, the target object, the user's actions, and the user's context. With the rapid development of information and the vast amount of information available, it is necessary to avoid information overload (i.e., prevent users from receiving excessive amounts of invalid information, so that they can directly receive the recommended information they need). Information recommendation systems can be used to filter this vast amount of information, ultimately ensuring that users receive the recommended information they need. By combining artificial intelligence technology, information recommendation systems can be deployed to predict users' ratings or preferences for a target object (such as information content or an item) based on their actions on products (apps, websites, etc.), thereby enabling information recommendation.
[0075] In the field of recommendation systems, estimating user dwell time is a crucial task. For example, after a user clicks on a resource on a webpage, such as a link, and browses the displayed information on the current information page, the estimated dwell time can be calculated as the user's dwell time on the current information page. A longer dwell time indicates a greater interest in the information on the current information page or similar information. Based on this, information recommendations can be made to the user, improving recommendation accuracy. Estimating user dwell time primarily involves converting the dwell time bucketing problem into a multi-classification problem or directly using a model to fit the dwell time estimate. However, during the modeling process, the sample labels (e.g., those used to represent the dwell time estimate) inherently contain noise. In other words, the presence of such noisy labels in the sample labels leads to low model accuracy and inefficient processing. Related art modeling techniques either ignore the impact of noisy labels and directly use sample labels containing noisy labels for modeling, or identify the noisy labels and then manually correct or discard them before modeling. This application example is not only universal and applicable to modeling fields such as continuous value regression, including duration estimation, but also considers the impact of noise labels in advance and directly removes noise labels during the modeling process, thereby improving model accuracy and processing efficiency.
[0076] This application example uses a duration regression model as an example. This model is also applicable to other continuous-value regression tasks beyond duration estimation (such as a completion rate model used to estimate the completion rate of video viewing by users), and is not limited to the information recommendation task targeted by this duration regression model. This duration regression model can be deployed on a terminal or server-side information recommendation system to improve model accuracy after noise reduction, ultimately increasing the accuracy of the information recommendation system.
[0077] A sample set can be composed of several features, where the features are sample data (which can be called sample features), and the sample labels are used to describe the features. Machine learning can be simply divided into supervised learning (i.e., sample data has corresponding sample labels), unsupervised learning (i.e., sample data does not have corresponding sample labels), and semi-supervised learning (i.e., some sample data in the sample data has corresponding sample labels). The duration regression model of this application example can use supervised learning (the sample set includes sample data and sample labels corresponding to the sample data. The goal of supervised learning is to train the model through sample data and sample labels corresponding to the sample data), that is, the sample data has corresponding sample labels y i For example, in an information recommendation system, dwell time estimation is used to describe how long a user stays on a resource after clicking on it. The sample label is dwell time. Sample data can be defined based on different business scenarios, such as user interests, user attributes, resource characteristics, or combined user and resource characteristics.
[0078] Assume that the sample label corresponding to sample data i is y i , considering the existence of noise labels in the sample labels, make y i The distribution of adopts the normal distribution shown in the following formula (1):
[0079]
[0080] In formula (1), Refers to the normal distribution, u i Refers to y i The mean of a normal distribution, Refers to y i It should be pointed out that the model actually needs to fit u i , instead of fitting y directly i , thus directly modeling the sample noise.
[0081] In order to make the model's predicted value p i More accurate, directly models the sample noise, and makes y i In order to maximize the probability of the noise label appearing in the noise (based on the maximum likelihood estimation method to achieve this probability maximization), it is necessary to select the predicted value that maximizes the probability of the noise label appearing. Considering the need to maximize the probability, the following formula (2) is used to calculate the probability maximization.
[0082]
[0083] In formula (2), Π i is the multiplication symbol, σ i is the standard deviation, σ i 2 is the variance, yi is the sample label, p i Is the predicted value, take the negative logarithm of the result of the operation obtained by formula (2), and convert the result of the continuous multiplication into a continuous addition operation, which is more convenient. Considering that the optimization of the loss function is generally to consider the minimization of the gradient descent algorithm, taking the negative logarithm can also convert the maximization problem into a minimization problem. After discarding the constant term, the loss function can be obtained by using the following formula (3). Formula (4) is used to implement the loss function for the predicted value p i Perform the derivation process to obtain the loss function for the predicted value p i The gradient of the loss function back propagated by partial derivative is:
[0084]
[0085]
[0086] In formulas (3)-(4), L is the loss function, y i is the sample label, p i is the predicted value, σ i 2 is the variance of the sample labels, Is the loss function for the predicted value p i Find the gradient of the loss function back propagated by the partial derivative.
[0087] The loss function of each sample label will also use the inverse of its corresponding sample variance to adjust the weight. The sample variance can reflect the noise of the sample. The larger the variance, the greater the noise, and the smaller the loss gradient contribution. The loss function obtained by the above formula (3) and the loss function obtained by the further formula (4) are back-propagated, which can make the smaller the variance, the smaller the noise, and the greater the loss gradient contribution. Considering that the use of formula (4) may cause the sample variance to be zero, further optimization of the loss function and the corresponding derivative processing can be achieved by using the following formulas (5) and (6):
[0088]
[0089]
[0090] In formulas (5)-(6), L is the loss function, y i is the sample label, p i is the predicted value, σ i 2 is the variance of the sample labels, σ 2 is the hyperparameter used for smoothing, Is the loss function for the predicted value p i Find the gradient of the loss function back propagated by the partial derivative.
[0091] For σ i 2 There are two calculation methods. The first method is to calculate σ based on the time distribution of the same resource under different users. i 2 The second calculation method is to divide the duration distribution of the same resource under different users into buckets, and only use the sample labels near the bucket where the current sample label is located to calculate σ i 2 The second calculation method is more accurate than the first one. As for the duration bucketing, considering that duration is a continuous value, and the sample labels faced by the classification problem are discrete values, such as classification 1, classification 2, classification 3, etc., the duration bucketing method can better perform accurate σ for discrete values. i 2 Duration bucketing is a classification process. First, the duration is bucketed and then divided into corresponding buckets according to different types. There are many ways to bucket duration, such as equal-interval bucketing (sampling sample labels with a duration between 0 and 10 seconds into the first bucket, and sample labels with a duration between 11 and 20 seconds into the second bucket, etc.), ensuring that the number of sample labels in each bucket is similar.
[0092] In this application example, when there are noisy labels in the sample labels, there is no need to first identify the noisy labels through the model and then manually process them (such as discarding them) and other tedious processes. Instead, the impact of the noisy labels is taken into account during modeling. Through systematic modeling, the optimized loss function obtained by the calculation formula of the above loss function is used to obtain a trained target model. This target model then implements one-stop noise reduction processing and directly removes the noisy labels. In other words, through systematic modeling, the noisy labels carrying noise in the sample labels are identified, and the noise reduction processing is performed directly. Therefore, the target model can effectively remove the impact of noise and improve model accuracy.
[0093] According to an embodiment of the present disclosure, a model denoising device is provided. Figure 5 Schematic diagram of the structure of the model noise reduction device according to the embodiment of the present disclosure. Figure 5 As shown, the model denoising device 500 includes: an acquisition unit 501, used to obtain a sample set carrying noise labels; a denoising unit 502, used to input the sample set carrying noise labels into a trained target model, perform denoising processing on the noise labels, and obtain a denoised model output result.
[0094] In one embodiment, the sample set carrying the noise label includes: sample data, and a sample label containing the noise label; wherein the sample data is used to characterize the residence status after the first operation is performed on the first information; and the sample label is used to characterize the residence time after the first operation is performed on the first information.
[0095] In one embodiment, the noise reduction unit is used to: when the target model is a time-length regression model, input the sample set carrying the noise label into the time-length regression model, perform noise reduction processing on the noise label, and the obtained model output result is a target value that does not include noise.
[0096] In one embodiment, a training unit is further included, which is used to: obtain a loss function based on the sample set carrying the noise label and the prediction value used to maximize the probability of the noise label; train the target model through back propagation of the loss function to obtain the trained target model.
[0097] In one embodiment, the training unit is used to: extract sample labels from the sample set carrying noise labels to obtain a first difference between the sample labels and the predicted values; and obtain the loss function based on the first difference and the variance of the sample labels.
[0098] In one embodiment, the training unit is used to: extract sample labels from the sample set carrying noise labels to obtain a first difference between the sample labels and the predicted values; and obtain the loss function based on the first difference, the variance of the sample labels and the hyperparameters used for smoothing.
[0099] According to an embodiment of the present disclosure, a model denoising device is provided. Figure 6 is a schematic diagram of the structure of the information recommendation device according to an embodiment of the present disclosure. Figure 6 As shown, the information recommendation device 600 includes: a preference acquisition unit 601, which is used to obtain interest preferences for multiple information contents based on the information recommendation operation; a first feature extraction unit 602, which is used to extract features of the interest preferences for multiple information contents to obtain multiple first features; a first label denoising unit 603, which is used to input the multiple first features into the trained target model, perform denoising processing on the noise labels corresponding to the multiple first features, and obtain a denoised model output result; an information recommendation unit 604, which is used to perform information recommendation processing according to the denoised model output result to obtain an information recommendation result.
[0100] According to an embodiment of the present disclosure, a model denoising device is provided. Figure 7 is a schematic diagram of the structure of the information sharing device according to an embodiment of the present disclosure, such as Figure 7 As shown, the information sharing device 700 includes: an operation content acquisition unit 701, which is used to obtain multiple operation contents obtained based on the information sharing operation; a second feature extraction unit 702, which is used to perform feature extraction on the multiple operation contents to obtain multiple second features; a second label denoising unit 703, which is used to input the multiple second features into the trained target model, perform denoising processing on the noise labels corresponding to the multiple second features, and obtain the denoised model output result; an information sharing unit 704, which is used to perform information sharing processing according to the denoised model output result to obtain an information sharing result.
[0101] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0102] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0103] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0104] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0105] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0106] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the model denoising method. For example, in some embodiments, the model denoising method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the model denoising method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the model denoising method by any other appropriate means (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0113] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.
[0114] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a denoising model, comprising: Acquire a sample set carrying a noise label; the sample set includes sample data and a sample label including the noise label; The sample data is used to characterize the residence status after the first operation is performed on the first information; the sample label is used to characterize the residence time after the first operation is performed on the first information; Extracting a sample label from the sample set carrying the noise label, and obtaining a first difference between the sample label and a predicted value for maximizing the probability of the noise label; Obtaining a loss function according to the first difference and the variance of the sample label; The target model is trained by the loss function to obtain a trained target model that can be subjected to noise reduction processing; Wherein, assuming that the sample labels containing noisy labels obey the normal distribution, the loss function is expressed as: L represents the loss function, represents the sample label of the i-th sample in the sample set, represents the predicted value of the i-th sample in the sample set, Represents the variance of the sample labels in the sample set where the i-th sample is located; where the sample labels Normal distribution: , Represents the mean of the normal distribution, which is the target that the target model needs to fit.
2. The method according to claim 1, wherein The target model is a duration regression model.
3. A method for training a denoising model, comprising: Acquire a sample set carrying a noise label; the sample set includes sample data and a sample label including the noise label; The sample data is used to characterize the residence status after the first operation is performed on the first information; the sample label is used to characterize the residence time after the first operation is performed on the first information; Extracting a sample label from the sample set carrying the noise label, and obtaining a first difference between the sample label and a predicted value for maximizing the probability of the noise label; Obtaining a loss function according to the first difference, the variance of the sample label, and a hyperparameter for smoothing; The target model is trained by the loss function to obtain a trained target model that can be subjected to noise reduction processing; Wherein, assuming that the sample labels containing noisy labels obey the normal distribution, the loss function is expressed as: L represents the loss function, represents the sample label of the i-th sample in the sample set, represents the predicted value of the i-th sample in the sample set, Represents the variance of the sample labels in the sample set where the i-th sample is located; represents the hyperparameter used for smoothing; where the sample label Normal distribution: , Represents the mean of the normal distribution, which is the target that the target model needs to fit.
4. The method according to claim 3, wherein: The target model is a duration regression model.
5. An information recommendation method, comprising: Obtaining interest preferences for multiple information contents based on information recommendation operations; Extracting features of the interest preferences for the plurality of information contents to obtain a plurality of first features; Inputting the plurality of first features into a trained target model, performing denoising processing on noise labels corresponding to the plurality of first features, and obtaining a denoised model output result; The target model is obtained based on the training method of the denoising model according to any one of claims 1 to 4; An information recommendation process is performed based on the denoised model output result to obtain an information recommendation result.
6. An information sharing method comprising: Obtain multiple operation contents obtained based on information sharing operations; Performing feature extraction on the multiple operation contents to obtain multiple second features; Inputting the plurality of second features into the trained target model, performing denoising processing on the noise labels corresponding to the plurality of second features, and obtaining a denoised model output result; The target model is obtained by the training method of the denoising model according to any one of claims 1 to 4; An information sharing process is performed according to the denoised model output result to obtain an information sharing result.
7. A training device for a noise reduction model, comprising: An acquisition unit, configured to acquire a sample set carrying a noise label; the sample set includes sample data and a sample label including the noise label; The sample data is used to characterize the residence status after the first operation is performed on the first information; the sample label is used to characterize the residence time after the first operation is performed on the first information; A training unit is configured to extract sample labels from the sample set carrying noise labels, obtain a first difference between the sample labels and a predicted value representing the maximum probability of the noise labels; obtain a loss function based on the first difference and the variance of the sample labels; and train a target model using the loss function to obtain a trained target model capable of performing noise reduction processing; Wherein, assuming that the sample labels containing noisy labels obey the normal distribution, the loss function is expressed as: L represents the loss function, represents the sample label of the i-th sample in the sample set, represents the predicted value of the i-th sample in the sample set, Represents the variance of the sample labels in the sample set where the i-th sample is located; where the sample labels Normal distribution: , Represents the mean of the normal distribution, which is the target that the target model needs to fit.
8. The device according to claim 7, wherein The target model is a duration regression model.
9. A training device for a noise reduction model, comprising: An acquisition unit, configured to acquire a sample set carrying a noise label; the sample set includes sample data and a sample label including the noise label; The sample data is used to represent the residence status after the first operation is performed on the first information; the sample label is used to represent the residence time after the first operation is performed on the first information; A training unit is configured to extract sample labels from the sample set carrying noise labels, obtain a first difference between the sample labels and a predicted value that maximizes the probability of the noise labels; obtain a loss function based on the first difference, the variance of the sample labels, and a hyperparameter for smoothing; and train a target model using the loss function to obtain a trained target model that can perform noise reduction processing; Wherein, assuming that the sample labels containing noisy labels obey the normal distribution, the loss function is expressed as: L represents the loss function, represents the sample label of the i-th sample in the sample set, represents the predicted value of the i-th sample in the sample set, Represents the variance of the sample labels in the sample set where the i-th sample is located; represents the hyperparameter used for smoothing; where the sample label Normal distribution: , Represents the mean of the normal distribution, which is the target that the target model needs to fit.
10. The device according to claim 9, wherein The target model is a duration regression model.
11. An information recommendation device, comprising: A preference acquisition unit, configured to acquire interest preferences for multiple information contents based on the information recommendation operation; A first feature extraction unit is configured to extract features of the interest preferences for the plurality of information contents to obtain a plurality of first features; A first label denoising unit is configured to input the plurality of first features into a trained target model, perform denoising on noise labels corresponding to the plurality of first features, and obtain a denoised model output result; The target model is obtained by the training method of the denoising model according to any one of claims 1 to 4; The information recommendation unit is used to perform information recommendation processing according to the denoised model output result to obtain an information recommendation result.
12. An information sharing device comprising: An operation content acquisition unit, configured to acquire multiple operation contents obtained based on the information sharing operation; A second feature extraction unit is used to extract features from the multiple operation contents to obtain multiple second features; A second label denoising unit is configured to input the plurality of second features into a trained target model, perform denoising on noise labels corresponding to the plurality of second features, and obtain a denoised model output result; The target model is obtained by the training method of the denoising model according to any one of claims 1 to 4; The information sharing unit is used to perform information sharing processing according to the denoised model output result to obtain an information sharing result.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Face recognition model training method for anti-noise data
CN110879985A
Training method of noise reduction model and related device
CN112598597A