Satisfaction recognition method and device, electronic equipment and readable storage medium
By constructing a sample training index set and pseudo-labeling a sample pseudo-label index set, and iteratively training the satisfaction recognition model, the problem of low accuracy in user satisfaction recognition in existing technologies is solved, and higher accuracy in user satisfaction recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-04-07
AI Technical Summary
The accuracy of existing model-based user satisfaction identification is low, mainly due to the limited amount of sample data and the lack of objectivity caused by reliance on human intervention.
By acquiring historical user information to construct a sample training indicator set, using the satisfaction recognition model to be trained to calculate the satisfaction attribution probability value and similarity of sample users, pseudo-labeling the sample pseudo-label indicator set, iteratively training the satisfaction recognition model, expanding the sample training indicator set, and realizing the objectivity of automatically expanding sample data.
It improves the accuracy of user satisfaction identification, overcomes the problem of lack of objectivity in identification results caused by human intervention under limited sample data, and constructs a satisfaction identification model with higher fitting accuracy.
Smart Images

Figure CN116738321B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, in particular to a satisfaction recognition method and device, electronic equipment and readable storage medium. BACKGROUND
[0002] With the continuous development of technology, for operators, user satisfaction recognition based on a model has become a mainstream recognition method, so building a recognition model with high recognition accuracy has become one of the important research directions of operators.
[0003] At present, there are usually two ways to build a recognition model: one is to assign weights to different evaluation indicators in the model by summarizing objective data and combining subjective experience, and then score the satisfaction of different users based on the weighted evaluation indicator system, and finally output the list of dissatisfied users through the set threshold; the second is to process the existing sample data set by using oversampling or undersampling, and then select a related algorithm to build a recognition model after balancing the positive and negative samples to form a modeling data set, and manually optimize the sample set of the recognition model, that is, due to the limitation of sample data, the sample data set for building the recognition model needs to rely on manual participation, which leads to the lack of objectivity of the satisfaction recognition result output by the recognition model, so the recognition accuracy of the current user satisfaction recognition based on a model is low. SUMMARY
[0004] The main purpose of the present application is to provide a satisfaction recognition method, device, electronic equipment and readable storage medium, which aims to solve the technical problem of low recognition accuracy of user satisfaction recognition based on a model in the prior art.
[0005] To achieve the above purpose, the present application provides a satisfaction recognition method, which comprises:
[0006] obtaining a sample training indicator set constructed by historical user information, wherein the sample training indicator set comprises a sample label indicator set carrying a satisfaction label and a real sample indicator set not carrying the satisfaction label, and the sample label indicator set comprises a positive sample label indicator set and a negative sample label indicator set;
[0007] inputting the real sample indicator set into a to-be-trained satisfaction recognition model, performing satisfaction recognition on a sample user corresponding to the real sample indicator set through the to-be-trained satisfaction recognition model, and obtaining a satisfaction attribution probability value of the sample user;
[0008] Calculate the positive sample group similarity and negative sample group similarity of the sample users respectively, wherein the positive sample group similarity is used to characterize the average similarity between the sample users and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample users and the negative sample group corresponding to the negative sample label index set;
[0009] Based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity, the real sample indicator set is labeled with pseudo-labels to obtain the sample pseudo-label indicator set;
[0010] Based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, the satisfaction recognition model to be trained is iteratively trained to obtain the satisfaction recognition model.
[0011] The satisfaction recognition model is used to identify the satisfaction level of the indicator to be identified, and the satisfaction recognition result is obtained.
[0012] To achieve the above objectives, this application also provides a satisfaction recognition device, the satisfaction recognition device comprising:
[0013] The acquisition module is used to acquire a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying a satisfaction label and a real sample indicator set without the satisfaction label, and the sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set.
[0014] The first identification module is used to input the real sample index set into the satisfaction identification model to be trained, and to identify the satisfaction of the sample users corresponding to the real sample index set through the satisfaction identification model to be trained, so as to obtain the satisfaction belonging probability value of the sample user.
[0015] The calculation module is used to calculate the positive sample group similarity and negative sample group similarity of the sample user, wherein the positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set.
[0016] The annotation module is used to annotate the real sample indicator set with pseudo-labels based on the satisfaction attribution probability value, the positive sample group similarity and the negative sample group similarity, so as to obtain the sample pseudo-label indicator set;
[0017] The training module is used to iteratively train the satisfaction recognition model to be trained based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, so as to obtain the satisfaction recognition model.
[0018] The second identification module is used to identify the satisfaction level of the indicator to be identified through the satisfaction identification model, and obtain the satisfaction identification result.
[0019] This application also provides an electronic device, the electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the satisfaction recognition method as described above.
[0020] This application also provides a computer-readable storage medium storing a program implementing a satisfaction recognition method, wherein when the program is executed by a processor, it implements the steps of the satisfaction recognition method as described above.
[0021] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the satisfaction recognition method described above.
[0022] This application provides a satisfaction recognition method, apparatus, electronic device, and readable storage medium. Specifically, it involves acquiring a sample training indicator set constructed from historical user information. This sample training indicator set includes a sample label indicator set carrying satisfaction labels and a real sample indicator set without the satisfaction labels. The sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set. The real sample indicator set is input into a satisfaction recognition model to be trained. The model then identifies the satisfaction of the sample users corresponding to the real sample indicator set, obtaining the satisfaction belonging probability value of each sample user. Finally, the positive sample group similarity and negative sample group similarity of the sample users are calculated. Body similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set. Based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity, the real sample index set is labeled with pseudo-labels to obtain a sample pseudo-label index set. Based on the target training index set formed by merging the sample pseudo-label index set and the sample label index set, the satisfaction recognition model to be trained is iteratively trained to obtain a satisfaction recognition model. The satisfaction recognition model is used to identify the satisfaction of the index to be identified to obtain the satisfaction recognition result.
[0023] In this application, when performing satisfaction identification, a sample training indicator set constructed from historical user information is first obtained. Since the sample training indicator set contains a set of real sample indicators without satisfaction labels, a set of positive sample label indicators with satisfaction labels, and a set of negative sample label indicators with satisfaction labels, the satisfaction identification model awaiting training first outputs the satisfaction attribution probability value of the sample users corresponding to the real sample indicator set. Then, the positive sample group similarity of the sample users is calculated based on the positive sample label indicator set, and the negative sample group similarity of the sample users is calculated based on the negative sample label indicator set. Then, the sample pseudo-label indicator set is obtained by pseudo-labeling using the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity. Finally, the target training indicator set, formed by merging the sample pseudo-label indicator set and the sample label indicator set, is used to iteratively train the satisfaction identification model to be trained, thus obtaining the satisfaction identification model. Since the sample pseudo-label indicator set is obtained based on the real sample indicator set without satisfaction labels, the purpose of automatically expanding the sample training indicator set used for iterative training of the satisfaction identification model to be trained can be achieved.
[0024] Since a pseudo-label index set can be actively obtained by expanding the sample training index set, the limitation of sample data volume is eliminated by training the satisfaction recognition model to be trained through the target training index set. At the same time, since the sample pseudo-label index set carries real information about user satisfaction, the objectivity of the sample data expansion process is ensured. Finally, through the expanded target training index set, a satisfaction recognition model with higher fitting accuracy can be trained. Since the satisfaction recognition model is used to identify user satisfaction, the user satisfaction can be accurately identified by constructing the satisfaction recognition model.
[0025] Based on this, this application obtains a sample pseudo-label index set by pseudo-labeling the real sample dataset in the sample training index set. Since the sample pseudo-label index set is determined by the satisfaction attribution probability value, the similarity of positive sample groups, and the similarity of negative sample groups, the sample indicators in the sample pseudo-label index set possess objectivity. Furthermore, the target training index set, formed by merging the sample pseudo-label index set and the sample training index set, is used for iterative training of the satisfaction recognition model to be trained, thus obtaining the trained satisfaction recognition model. This achieves the goal of constructing a more reliable satisfaction recognition model by actively expanding the objective sample training index set, rather than relying on manual participation in expanding the sample data when the data volume is limited. In other words, it overcomes the technical defect that the limited sample data volume requires manual intervention when constructing the sample dataset for the recognition model, leading to a lack of objectivity in the satisfaction recognition results output by the recognition model. Therefore, it improves the recognition accuracy of user satisfaction based on the model. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating the satisfaction recognition method provided in Embodiment 1 of this application;
[0029] Figure 2 A schematic diagram illustrating the process of constructing a sample training index set for the satisfaction recognition method provided in Embodiment 1 of this application;
[0030] Figure 3This is a flowchart illustrating the satisfaction recognition method provided in Embodiment 2 of this application;
[0031] Figure 4 A flowchart of the model training process for the satisfaction recognition model to be trained in the satisfaction recognition method provided in Embodiment 1 of this application;
[0032] Figure 5 This is a schematic diagram of the satisfaction recognition device provided in Embodiment 3 of this application;
[0033] Figure 6 This is a schematic diagram of the structure of the electronic device provided in Embodiment 4 of this application.
[0034] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0035] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Example 1
[0037] User satisfaction has always been a crucial indicator for operators to evaluate their operational performance. However, because user satisfaction data is collected through outbound calls or surveys, the sample size is typically small. This highly imbalanced data makes building a user satisfaction identification model a critical challenge for operators. Currently, two main approaches are used: employing AHP (The Analytic Hierarchy Process) and its improved algorithms to construct a comprehensive weighted evaluation model, or using conventional logistic regression, support vector machines, or random forests, and training the model using oversampling or undersampling. However, both of these approaches have limitations. While the comprehensive weighted evaluation model can be used to build user satisfaction, user stability, or other application models, it relies heavily on a combination of subjective judgment and objective data in its data indicator construction. This means that the AHP indicator judgment matrices constructed by different experts vary significantly, leading to differences in the final indicator weights, making manual intervention highly necessary. This can reduce the accuracy of the recognition model to some extent. Furthermore, due to the extreme imbalance of small sample training data, other recognition model construction methods often employ crude oversampling techniques when building customer satisfaction recognition models. This involves repeatedly sampling minority class samples to generate new modeling datasets for model construction. During this process, minority class sample data is easily duplicated and superimposed to form new datasets for model construction. This leads to the repeated learning of minority class features during model training, making the recognition model prone to overfitting. In summary, limited by factors such as human subjectivity and the size of the sample dataset, the current recognition model has low accuracy in recognizing user satisfaction. Therefore, there is an urgent need for a method to improve the accuracy of model-based user satisfaction recognition.
[0038] This application provides a satisfaction recognition method. In the first embodiment of the satisfaction recognition method of this application, referring to... Figure 1 The satisfaction recognition method includes:
[0039] Step S10: Obtain a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying a satisfaction label and a real sample indicator set not carrying the satisfaction label, and the sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set.
[0040] Step S20: Input the real sample index set into the satisfaction recognition model to be trained, and use the satisfaction recognition model to identify the satisfaction of the sample users corresponding to the real sample index set to obtain the satisfaction belonging probability value of the sample users.
[0041] Step S30: Calculate the positive sample group similarity and negative sample group similarity of the sample user, respectively. The positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set.
[0042] Step S40: Based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity, the real sample indicator set is labeled with pseudo-labels to obtain the sample pseudo-label indicator set;
[0043] Step S50: Based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, iteratively train the satisfaction recognition model to be trained to obtain the satisfaction recognition model.
[0044] Step S60: Use the satisfaction recognition model to identify the indicators to be identified and obtain the satisfaction recognition result.
[0045] In this embodiment, it should be noted that, although Figure 1The logical order is shown, but in some cases, the steps shown or described may be performed in a different order than that shown here. The satisfaction recognition method is applied to a satisfaction recognition device, which refers to an electronic device with the processing power required for model training, specifically a computer or mainframe computer. The satisfaction recognition model to be trained is a base model for recognizing user satisfaction, awaiting training. Historical user information refers to information generated by users in historical time periods that is related to the modeling samples. Specifically, it may include user information actively collected by the satisfaction recognition device and user network quality satisfaction scores collected from the user side by operators through information collection strategies such as outbound calls or SMS. Specifically, user information may include user call duration, user call quality, and user data usage. Using historical user information, sample training metrics required for training the satisfaction recognition model to be trained can be constructed. User satisfaction can be fed back from different evaluation dimensions. For example, in one feasible approach, the specific evaluation dimensions corresponding to the sample training indicators may include customer attributes, cross-network defection, traffic usage, and 5G network quality. Among these, 5G network quality is an important evaluation dimension affecting user satisfaction. The sample training indicators included under the 5G network quality dimension may include: the average number of times each video stutters on Internet TV in the past month, the percentage of Internet TV stutter duration, the average playback response time of Internet TV (in minutes), the Internet TV playback success rate, the average game ping latency in the past month (in minutes), the number of broadband network quality complaints in the past six months, the number of network failures in the past three months, the webpage opening success rate, the average webpage access response rate, whether the community broadband access method supports the broadband package speed, the average monthly actual network speed (in MB / s), and the number of alarms from users with poor Internet TV quality.
[0046] Additionally, it should be noted that the sample training indicator set is used to represent the set of indicator sets composed of sample training indicators required for training the satisfaction recognition model to be trained. These sample training indicators are specifically constructed from historical user information. The sample training indicator set includes a sample label indicator set carrying satisfaction labels and a real sample indicator set without satisfaction labels. The sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set. The sample label indicator set refers to the set of sample training indicators carrying satisfaction labels, while the real sample indicator set refers to the set of sample training indicators without satisfaction labels. Satisfaction labels are used to identify user satisfaction and can be "satisfied" or "dissatisfied," etc., specifically including positive and negative sample labels, which are different categories. For example, in one feasible approach, "satisfied" can be used as a negative sample label and "dissatisfied" as a positive sample label. Then, user information for t-3, t-2, and t-1, as well as user network quality satisfaction scores for t-2 and t-1 collected by the operator through outbound calls or SMS strategies, are collected. When constructing sample training indicators using user information from t-3, a satisfaction label for user satisfaction is established using the user network quality satisfaction score from t-2. Similarly, when constructing sample training indicators using user information from t-2, a satisfaction label for user satisfaction is established using the user quality satisfaction score from t-1. The misalignment of satisfaction labels and sample training indicators serves as an early warning mechanism, allowing the satisfaction label of the current month to predict the satisfaction of the previous month's sample training indicators. Through this construction process, a sample training indicator set can ultimately be constructed.
[0047] Additionally, it should be noted that the satisfaction recognition model to be trained can be constructed based on conventional model building methods. For example, in one feasible approach, the sample label indicator set and the real sample indicator set are first processed using data cleaning and feature engineering techniques to improve the quality of the initial training indicators for different evaluation dimensions. Then, the sample label indicator set is split into modeling indicator sets and test indicator sets according to a certain ratio. Next, the population is initialized, the fitness function is set, and the number of iterations or termination conditions of the model are determined. In the semi-supervised learning framework, a cost-sensitive support vector machine is used as the fully supervised classifier construction algorithm. An improved genetic algorithm is applied to optimize the model parameters of the fully supervised model classifier to find the optimal combination of fully supervised model parameters. One of the core components of the semi-supervised learning model is the underlying fully supervised model classifier; therefore, obtaining the optimal underlying fully supervised model is crucial to the overall performance of the semi-supervised learning model. Specifically, the population initialization steps can be as follows: using a cost-sensitive support vector machine as the basis function, where the cost-sensitive support vector machine has two core parameters C and gam. ma, where C is the penalty coefficient, i.e., the tolerance for error. A higher penalty coefficient means it is less tolerant of errors, thus easily leading to overfitting; a lower penalty coefficient means it is easily leading to underfitting. gamma implicitly determines the distribution of sample data after mapping to the new feature space. A larger gamma indicates fewer support vectors, and a smaller gamma indicates more support vectors. The number of support vectors will affect the training and prediction speed. Therefore, initializing the population means randomly initializing the population using the cost-sensitive support vector machine parameters C and gamma as chromosomes; then setting the fitness function. The fitness function evaluates the "survival ability" of individuals in a population. For single-objective optimization, the objective function can be directly used as the fitness function. In this embodiment, the average cross-validation accuracy of the fully supervised cost-sensitive support vector machine base model is used as the fitness value, that is, the average cross-validation calculation function of the fully supervised learning model is used as the fitness function. Furthermore, a termination iteration condition is set, that is, the basis for stopping the evolution of the genetic algorithm is set. For example, it can be "the maximum number of iterations is 100 or the average cross-validation accuracy of the satisfaction recognition model to be trained changes less than a threshold 1*e" for 10 consecutive iterations. -6"As a termination condition, the output model is used as the satisfaction recognition model. That is, based on the processed target training index set, the cost-sensitive support vector machine algorithm is selected to complete the initial generation of the genetic algorithm population training (based on the initialized genetic algorithm population and the processed positive and negative sample index set, the base model training of each individual in the population is completed). The fitness value of each individual in the initial population is calculated based on the fitness function. Through a specific individual selection algorithm, high fitness value individuals in the population are selected, mutated, and crossovered to form a new population. If the current iteration still does not meet the set termination condition, the training continues with the new population until the termination condition is met, and then the optimal solution is output."
[0048] Additionally, it should be noted that the satisfaction attribution probability value is used to characterize the probability value of a sample user belonging to the satisfaction label, the positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set. Specifically, it can be calculated based on the conventional similarity calculation formula, and this application embodiment does not limit it in this way.
[0049] Additionally, it should be noted that the satisfaction recognition model is used to identify user satisfaction. The indicators to be identified are those awaiting identification, specifically corresponding to the target object. The target object represents the user whose satisfaction needs to be identified. The satisfaction recognition model outputs a result such as "satisfied," "dissatisfied," or a specific satisfaction rating. The sample pseudo-label indicator set represents the set of pseudo-labeled training indicators from the real sample indicator set. These pseudo-labeled training indicators correspond to the positive or negative sample label indicators in the positive sample label indicator set. The negative sample label metrics have extremely high similarity and do not require manual parameter annotation. The target training metric set is a set of sample training metrics used to train the satisfaction recognition model. For example, in one feasible approach, the target training metric set includes customer attribute dimension metrics: user age, user star rating, user online duration, user broadband usage duration, whether the user is a secondary card user, whether the user is a key person in the group, the month-on-month change in the number of service interruptions in the past three months, and whether the user uses a number from another network; cross-network defection metrics: the number of inbound or outbound customer service calls from other networks in the past three months; and call service metrics: the month-on-month change in the number of mutual call transfers in the past three months and the month-on-month change in the number of calls in the past three months. The following metrics were used in the past three months: average monthly percentage of inter-network calls, month-on-month comparison of the number of calls to our customer service, and average monthly voice call overage charges; data traffic metrics: month-on-month comparison of data traffic usage, month-on-month comparison of data traffic usage saturation, and average monthly data traffic usage saturation; consumption metrics: number of recharges, recharge amount, month-end balance, monthly ARPU (Average Revenue Per User), monthly voice revenue, and monthly data traffic revenue; and broadband attribute metrics: number of months until broadband expiration, number of days without using our broadband network this month, month-on-month comparison of broadband internet access days in the past three months, and whether or not broadband was used in the past three months. Broadband renewal, whether broadband was renewed in the past three months, month-on-month comparison of call charges in the past three months, and broadband type; network quality satisfaction indicators: average number of times each video stuttered on Internet TV in the past month, percentage of Internet TV stuttering time, average playback response time of Internet TV, Internet TV playback success rate, average game latency in the past month, number of broadband network quality complaints in the past six months, webpage opening success rate, average webpage access response rate, whether the community broadband access method supports the broadband package speed, the difference between the broadband package and the community's supported speed, average monthly actual network speed, number of days each month with bandwidth less than half, and frequency of community fault interception.
[0050] As an example, steps S10 to S60 include: constructing a sample label indicator set and a real sample indicator set based on historical user information; combining the sample label indicator set and the real sample indicator set into a sample training indicator set; inputting the real sample indicator set into a satisfaction recognition model to be trained; using the satisfaction recognition model to identify the satisfaction of the sample users constituting the real sample indicator set to obtain the satisfaction attribution probability value of the sample users, wherein the sample users can be one or more; and calculating the similarity between the sample users and the positive sample groups corresponding to the positive sample label indicator set based on a preset similarity calculation formula. The average similarity and the average similarity between the sample users and the negative sample groups corresponding to the negative sample label index set are calculated. Based on the satisfaction probability value, the similarity of the positive sample group, and the similarity of the negative sample group, the real sample index set is labeled with pseudo-labels to obtain a sample pseudo-label index set. The sample pseudo-label index set and the sample label index set are merged into a target training index set. The satisfaction recognition model to be trained is iteratively trained based on the target training index set until the target loss function converges, thus obtaining the satisfaction recognition model. The satisfaction recognition model is used to identify the satisfaction of the indicators to be identified, resulting in the satisfaction recognition result. Since the sample pseudo-label index set is obtained by labeling the real sample index set in the sample training index set with pseudo-labels, the purpose of expanding the sample training indicators used to train the satisfaction recognition model to be trained is achieved without human intervention. Finally, the expanded target training index set is used to iteratively train the satisfaction recognition model to be trained, resulting in a satisfaction recognition model with high fitting accuracy. This improves the accuracy of user satisfaction recognition based on the satisfaction recognition model, thus enhancing the accuracy of user satisfaction recognition based on the model.
[0051] In one feasible approach, the specific steps for obtaining a sample pseudo-label index set by labeling the real sample index set with pseudo-labels based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity are as follows: Calculate a first difference between the satisfaction attribution probability value and the positive sample group similarity, and a second difference between the satisfaction attribution probability value and the negative sample group similarity. If the first difference is detected to be greater than the second difference, then the sample training index pseudo-label corresponding to the sample user in the real sample index set is labeled as a positive sample label index. If the first difference is detected to be less than the second difference, then the sample training index pseudo-label corresponding to the sample user in the real sample index set is labeled as a negative sample label index.
[0052] The real sample indicator set includes at least one real sample indicator, the positive sample label indicator set includes at least one positive sample label indicator, and the negative sample label indicator set includes at least one negative sample label indicator. The steps of calculating the positive sample group similarity and the negative sample group similarity of the sample users respectively include:
[0053] Step A10: Obtain the first index value of each of the real sample indicators, the second index value of each of the positive sample label indicators, and the third index value of each of the negative sample label indicators, respectively.
[0054] Step A20: Calculate the positive sample group similarity of the sample users based on each of the first indicator values, each of the second indicator values, and the positive sample label indicator quantity corresponding to the positive sample label indicator set.
[0055] Step A30: Calculate the negative sample group similarity of the sample users based on each of the first indicator values, each of the third indicator values, and the negative sample label indicator quantity corresponding to the negative sample label indicator set.
[0056] In this embodiment, it should be noted that conventional group similarity calculation methods do not consider the relationship between sample users and other sample users in the real sample indicator set. That is, they do not consider the weight of different indicator dimensions in the real sample indicator set, which leads to low accuracy of the obtained positive sample group similarity or negative sample group similarity. Therefore, this embodiment introduces a first indicator value, a second indicator value, and a third indicator value to calculate the positive sample group similarity and the negative sample group similarity, respectively. The first indicator value is used to characterize the indicator value of the sample user on different sample training indicators in the real sample indicator set. The second indicator value is used to characterize the indicator value of the positive sample user in the positive sample group on different positive sample label indicators in the positive sample label indicator set. The third indicator value is used to characterize the indicator value of the negative sample user in the negative sample group on different negative sample label indicators in the negative sample label indicator set.
[0057] As an example, steps A10 to A30 include: obtaining the first indicator value of each of the real sample indicators, the second indicator value of each of the positive sample label indicators, and the third indicator value of each of the negative sample label indicators; inputting the first indicator value and the second indicator value into a first preset similarity calculation formula to obtain a first similarity; and calculating the positive sample group similarity of the sample users based on the first similarity and the number of positive sample label indicators corresponding to the positive sample label indicator set, wherein the calculation formula corresponding to the positive sample group similarity is as follows:
[0058]
[0059]
[0060] Among them, S t For the positive sample group similarity, sma pi Let k1, k2, ..., k be the first similarity. n Let p1, p2, ..., p be the first index values of n sample users. t Let k be the second indicator value for t positive sample users, where t is the positive sample label indicator value, and k is the second indicator value for t positive sample users. i For the i-th sample user, p i For the i-th positive sample user, the first indicator value and the third indicator value are input into the second preset similarity calculation formula to obtain the second similarity. Based on the second similarity and the negative sample label indicator quantity corresponding to the negative sample label indicator set, the negative sample group similarity of the sample user is calculated. The calculation formula for the negative sample group similarity is as follows:
[0061]
[0062]
[0063] Among them, S u For the negative sample group similarity, sma ui Let k1, k2, ..., k be the second similarity. n Let u1, u2, ..., u be the first index values of n sample users. y Let y be the second indicator value of y negative sample users, t be the negative sample label indicator quantity, and k be the second indicator value of y negative sample users. i For the i-th sample user, u y Let i be the i-th negative sample user. Because the above algorithm takes into account the weights of different indicator dimensions in the real sample indicator set, it improves the accuracy of the similarity between positive and negative sample groups.
[0064] The step of obtaining the sample pseudo-label index set by labeling the real sample index set with pseudo-labels based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity includes:
[0065] Step B10: Calculate the positive probability of the sample user belonging to the positive sample indicator set based on the first weight of the satisfaction belonging probability value and the second weight of the positive sample group similarity; and calculate the negative probability of the sample user belonging to the negative sample indicator set based on the first weight and the third weight of the negative sample group similarity.
[0066] Step B20: Based on the positive probability and the negative probability, the sample users corresponding to the real sample indicator set are labeled with pseudo-labels to obtain the sample pseudo-label indicator set.
[0067] As an example, steps B10 to B20 include: inputting the first weight of the satisfaction attribution probability value and the second weight of the positive sample group similarity into a preset positive probability calculation formula to calculate the positive probability that the sample user belongs to the positive sample indicator set; and inputting the first weight of the satisfaction attribution probability value and the third weight of the negative sample group similarity into a preset negative probability calculation formula to calculate the negative probability that the sample user belongs to the negative sample indicator set, wherein the preset positive probability calculation formula is as follows:
[0068] Lable p =λ*S p +δ*pro
[0069] Where λ is the first weight of the satisfaction attribution probability value, and Label p The positive probability is denoted by δ, the second weight of the positive sample group similarity is δ, and pro is the category inference probability output by the model to be trained for the real sample index set. The preset negative probability is calculated as follows:
[0070] Lable u =λ*S u +δ*pro
[0071] Where λ is the first weight of the satisfaction attribution probability value, and Label u Let δ be the negative probability, δ be the second weight of the positive sample group similarity, and pro be the category inference probability output by the model to be trained for the real sample indicator set; based on the positive probability and the negative probability, the sample users corresponding to the real sample indicator set are labeled with pseudo-labels to obtain the sample pseudo-label indicator set.
[0072] Specifically, based on the positive and negative probabilities, the sample users corresponding to the real sample indicator set are labeled with pseudo-labels to obtain the sample pseudo-label indicator set. This pseudo-labeling can be achieved using a "threshold method." For example, in one feasible approach, both λ and δ can be set to 0.5, and the pseudo-labeling threshold can be 0.9. When the Label of the sample users in the real sample indicator set... p Or Label u When the value exceeds the pseudo-labeling threshold, the real sample metrics corresponding to the sample user are pseudo-labeled and merged into the sample label metric set to form the target training metric set. That is, if the sample user's Label p If the value is greater than 0.9, the real sample metric corresponding to the sample user will be labeled as a positive sample pseudo-label metric. If the sample user's Label... uIf the value is greater than 0.9, the real sample indicator corresponding to the sample user will be labeled as a negative sample pseudo-label indicator.
[0073] The step of obtaining the sample training dataset constructed from historical user information includes:
[0074] Step C10: Based on historical user information, construct at least one indicator to be included in the model, wherein one of the indicators to be included in the model corresponds to one of the real sample indicator, the positive sample label indicator, and the negative sample label indicator;
[0075] Step C20: Filter each of the indicators to be included in the model to obtain the filtered indicators to be included in the model.
[0076] Step C30: Construct a sample training index set based on the selected indices to be incorporated into the model.
[0077] In this embodiment, it should be noted that the indicators to be included in the model are used to characterize the sample training indicators waiting to be included in the training of the satisfaction recognition model for iterative training. Since the stability of different sample training indicators varies, after constructing the indicators to be included in the model through historical user information, it is necessary to screen some sample training indicators to improve model performance. For example, in one feasible approach, sample training indicators in the sample training indicators of three months in which the user satisfaction status is conflicting and no promotional order behavior has occurred are taken as sample training indicators to be screened out, thereby performing the corresponding elimination operation. Specifically, the indicators to be included in the model can be real sample indicators, positive sample label indicators, or negative sample label indicators. That is, the indicators can be screened for the real sample indicator set and the sample label indicator set respectively. When screening the sample training indicators, corresponding indicator screening rules can be established based on different sample training indicators.
[0078] As an example, steps C10 to C20 include: constructing at least one candidate indicator based on historical user information, wherein each candidate indicator corresponds to one of the real sample indicator, the positive sample label indicator, and the negative sample label indicator; filtering each candidate indicator according to a preset indicator filtering rule to obtain filtered candidate indicators; and using at least one filtered candidate indicator as a sample training indicator set. By filtering the candidate indicators constructed based on historical user information one by one through the preset indicator filtering rule, multiple filtered candidate indicators are obtained. These multiple filtered candidate indicators are then used together as a sample training indicator set, thereby filtering out noisy indicators generated during the process of constructing sample training indicators from historical user information. This improves the quality of the sample training indicator set and lays the foundation for improving the performance of the satisfaction recognition model.
[0079] In one feasible approach, refer to Figure 2 , Figure 2 To illustrate the process of constructing a sample training indicator set, the user label corresponds to the satisfaction label in the embodiment of this application, and user indicator feature 1, user indicator feature 3, ..., user indicator feature r are indicators to be included in the model for different sample periods.
[0080] In this embodiment of the application, when performing satisfaction identification, a sample training indicator set constructed from historical user information is first obtained. Since the sample training indicator set contains a set of real sample indicators without satisfaction labels, a set of positive sample label indicators with satisfaction labels, and a set of negative sample label indicators with satisfaction labels, the satisfaction identification model awaiting training first outputs the satisfaction attribution probability value of the sample users corresponding to the real sample indicator set. Then, the positive sample group similarity of the sample users is calculated based on the positive sample label indicator set, and the negative sample group similarity of the sample users is calculated based on the negative sample label indicator set. Then, the sample pseudo-label indicator set is obtained by pseudo-labeling using the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity. Finally, the target training indicator set, formed by merging the sample pseudo-label indicator set and the sample label indicator set, is used to iteratively train the satisfaction identification model to be trained, thus obtaining the satisfaction identification model. Since the sample pseudo-label indicator set is obtained based on the real sample indicator set without satisfaction labels, the purpose of automatically expanding the sample training indicator set used for iterative training of the satisfaction identification model to be trained can be achieved.
[0081] Since a pseudo-label index set can be actively obtained by expanding the sample training index set, the limitation of sample data volume is eliminated by training the satisfaction recognition model to be trained through the target training index set. At the same time, since the sample pseudo-label index set carries real information about user satisfaction, the objectivity of the sample data expansion process is ensured. Finally, through the expanded target training index set, a satisfaction recognition model with higher fitting accuracy can be trained. Since the satisfaction recognition model is used to identify user satisfaction, the user satisfaction can be accurately identified by constructing the satisfaction recognition model.
[0082] Based on this, this application obtains a sample pseudo-label index set by pseudo-labeling the real sample dataset in the sample training index set. Since the sample pseudo-label index set is determined by the satisfaction attribution probability value, the similarity of positive sample groups, and the similarity of negative sample groups, the sample indicators in the sample pseudo-label index set possess objectivity. Furthermore, the target training index set, formed by merging the sample pseudo-label index set and the sample training index set, is used for iterative training of the satisfaction recognition model to be trained, thus obtaining the trained satisfaction recognition model. This achieves the goal of constructing a more reliable satisfaction recognition model by actively expanding the objective sample training index set, rather than relying on manual participation in expanding the sample data when the data volume is limited. In other words, it overcomes the technical defect that the limited sample data volume requires manual intervention when constructing the sample dataset for the recognition model, leading to a lack of objectivity in the satisfaction recognition results output by the recognition model. Therefore, it improves the recognition accuracy of user satisfaction based on the model.
[0083] Example 2
[0084] Furthermore, referring to Figure 3 In another embodiment of this application, content that is the same as or similar to that in Embodiment 1 described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of filtering each of the indicators to be included in the model to obtain the filtered indicators includes:
[0085] Step D10: For any of the indicators to be included in the model, perform data drift detection on the indicator to be included in the model based on the indicator skewness value and the indicator kurtosis value.
[0086] Step D20: After data drift detection is performed on all the indicators to be included in the model, the indicators that do not have data drift are selected as the filtered indicators to be included in the model.
[0087] In this embodiment, it should be noted that setting separate index screening rules for each training index would undoubtedly increase the workload during the screening process. Furthermore, if some training indices are screened only qualitatively, the completeness of the index set cannot be guaranteed. Therefore, this embodiment introduces data drift detection technology to detect data drift in the indices to be included in the model, thereby enabling accurate and efficient screening. For example, in one feasible approach, the data drift detection target is all the indices to be included in the model. The sample label index set includes 54,208 training indices, with 13,552 positive and 40,656 negative labels. The real sample index set includes 6.28 million training indices. Data drift detection is then performed based on the skewness and kurtosis values of each training indices.
[0088] Additionally, it should be noted that the skewness value is used to measure the degree of asymmetry in the numerical distribution of the indicators to be included in the model. When the indicators become more symmetrical, their skewness value will approach zero; otherwise, the larger the absolute value, the more skewed the indicator distribution. The kurtosis value is used to characterize the peak value of the probability density distribution curve at the mean value, also known as the kurtosis coefficient. Intuitively, the kurtosis value reflects the sharpness of the peak of the probability density distribution curve. Indicators to be included in the model that cause data drift are eliminated. For example, in one feasible approach, assuming there are 46 indicators to be included in the model, data drift detection can eliminate two major indicators: "number of user alarms for poor Internet TV quality" and "number of network failures in the past three months," resulting in 44 filtered indicators to be included in the model.
[0089] As an example, steps D10 to D20 include: for any candidate indicator to be included in the model, calculating the indicator skewness value and the indicator kurtosis value of the candidate indicator; performing data indicator detection on the candidate indicator based on the indicator skewness value and the indicator kurtosis value; and after performing data drift detection on each candidate indicator, selecting the candidate indicator without data drift as the filtered candidate indicator.
[0090] The step of detecting data drift of the indicator to be included in the model based on its skewness and kurtosis values includes:
[0091] Step E10: Input the index skewness value of the index to be included in the model to a preset skewness mapping function to obtain the index distribution skewness value of the index to be included in the model within the sample period; and input the index kurtosis value of the index to be included in the model to a preset kurtosis mapping function to obtain the index distribution kurtosis value of the index to be included in the model within the sample period.
[0092] Step E20: By inputting the index distribution skewness value and the index distribution skewness value together into a preset data drift function, the data drift value is obtained;
[0093] Step E30: Based on the relationship between the data drift value and the preset data drift threshold, perform data drift detection on the indicator to be included in the model.
[0094] As an example, steps E10 to E30 include: inputting the index skewness value of the index to be included in the model to a preset skewness mapping function to obtain the index distribution skewness value of the index to be included in the model within the sample period; and inputting the index kurtosis value of the index to be included in the model to a preset kurtosis mapping function to obtain the index distribution kurtosis value of the index to be included in the model within the sample period, wherein the preset skewness mapping function is as follows:
[0095]
[0096] Where, λ(X) T γ is the skewness value of the index distribution of the index to be included in the model within the sample period T. T γ is the skewness value of the index distribution of the index to be included in the model in month T. T-1 Let X be the skewness value of the index distribution of the index to be included in the model in month T-1, and let X be the index to be included in the model, which can be obtained from the value sequence in the sample training index set. The preset kurtosis mapping function is as follows:
[0097]
[0098] Where, τ(X) T τ is the skewness of the index distribution of the index to be included in the model within the sample period T. T τ is the skewness value of the index distribution of the index to be included in the model in month T. T-1 The skewness value of the indicator distribution to be included in the model is the indicator distribution skewness value in month T-1; by inputting the indicator distribution skewness value and the indicator distribution skewness value together into a preset data drift function, the data drift value is obtained, wherein the preset data drift function is as follows:
[0099]
[0100] Where f(x) is the data drift value; based on the relationship between the data drift value and the preset data drift threshold, the data drift detection is performed on the index to be included in the model.
[0101] In one feasible approach, if f(x) > 0, then the index X to be included in the model has experienced data drift within the sample period T, thereby determining that the index to be included in the model is an index to be filtered out; if f(x) = 0, then the index X to be included in the model has not experienced data drift within the sample period T.
[0102] Prior to the step of detecting data drift of the target indicator based on its skewness and kurtosis values, the satisfaction recognition method further includes:
[0103] Step F10: Obtain the mean of the sample indicator set and the standard deviation of the corresponding sample indicator set for the indicator to be included in the model.
[0104] Step F20: Input the mean and standard deviation of the sample indicator set into a first preset mean function to obtain the indicator skewness value of the indicator to be included in the model, and input the mean and standard deviation of the sample indicator set into a second preset mean function to obtain the indicator kurtosis value of the indicator to be included in the model.
[0105] As an example, steps F10 to F20 include: obtaining the mean and standard deviation of the sample indicator set corresponding to the indicator to be included in the model; inputting the mean and standard deviation of the sample indicator set into a first preset mean function to obtain the indicator skewness value of the indicator to be included in the model; and inputting the mean and standard deviation of the sample indicator set into a second preset mean function to obtain the indicator kurtosis value of the indicator to be included in the model, wherein the first preset mean function is as follows:
[0106]
[0107] Where λ is the skewness value of the indicator, μ is the mean of the sample indicator set, that is, the arithmetic mean of all modeling samples in the sample training indicator set corresponding to the indicator to be included in the model, and σ is the standard deviation of the sample indicator set, that is, the standard deviation of all modeling samples in the sample training indicator set corresponding to the indicator to be included in the model. E represents the mean operation, which is the mathematical expression for calculating the mean of a sequence. Introducing the mean operation makes the expression of second-order and third-order central matrices more intuitive; that is, k2 is the second-order central matrix, and k3 is the third-order central matrix. The second preset mean function is as follows:
[0108]
[0109] Where τ is the kurtosis value of the indicator. By using the mean operation to obtain the skewness and kurtosis values of the indicators to be included in the model, the second-order and third-order centrality matrices in the calculation process can be displayed more intuitively, thus laying the foundation for accurate and efficient filtering of the indicators to be included in the model.
[0110] In one feasible approach, refer to Figure 4 , Figure 4 The flowchart shows the training process for the satisfaction recognition model. Data cleaning, feature engineering, and data drift detection are the indicator selection processes, and the best output is the satisfaction recognition model.
[0111] This application provides a method for displaying satisfaction recognition results. Specifically, the satisfaction recognition results are split into a first satisfaction recognition result corresponding to a first data anomaly task and a second satisfaction recognition result jointly corresponding to each of the second data anomaly tasks. A first display parameter is matched to the first satisfaction recognition result, and a second display parameter is matched to the second satisfaction recognition result, wherein the first display parameter is greater than the second display parameter. Based on the first and second display parameters, the first and second satisfaction recognition results are jointly displayed on a preset verification management interface. In displaying satisfaction recognition results, this application splits the satisfaction recognition results into first and second satisfaction recognition results corresponding to different data tasks, matches display parameters to both the first and second satisfaction recognition results, and finally displays the first and second satisfaction recognition results together on the preset verification management interface based on different display parameters. This achieves the goal of distinguishing and displaying different satisfaction recognition results on the preset verification management interface, thus laying the foundation for improving the satisfaction recognition experience for verification personnel.
[0112] Example 3
[0113] This application embodiment also provides a satisfaction recognition device, referring to... Figure 5 The satisfaction recognition device includes:
[0114] The acquisition module 101 is used to acquire a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying a satisfaction label and a real sample indicator set not carrying the satisfaction label, and the sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set.
[0115] The first identification module 102 is used to input the real sample index set into the satisfaction identification model to be trained, and to identify the satisfaction of the sample users corresponding to the real sample index set through the satisfaction identification model to be trained, so as to obtain the satisfaction belonging probability value of the sample users.
[0116] The calculation module 103 is used to calculate the positive sample group similarity and the negative sample group similarity of the sample user, wherein the positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set.
[0117] The annotation module 104 is used to annotate the real sample indicator set with pseudo-labels based on the satisfaction attribution probability value, the positive sample group similarity and the negative sample group similarity, to obtain the sample pseudo-label indicator set;
[0118] Training module 105 is used to iteratively train the satisfaction recognition model to be trained based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, so as to obtain the satisfaction recognition model.
[0119] The second identification module 106 is used to identify the satisfaction level of the indicator to be identified through the satisfaction identification model, and obtain the satisfaction identification result.
[0120] Optionally, the real sample indicator set includes at least one real sample indicator, the positive sample label indicator set includes at least one positive sample label indicator, and the negative sample label indicator set includes at least one negative sample label indicator. The calculation module 103 is further configured to:
[0121] The first index value of each real sample index, the second index value of each positive sample label index, and the third index value of each negative sample label index are obtained respectively.
[0122] The positive sample group similarity of the sample users is calculated based on each of the first indicator values, each of the second indicator values, and the positive sample label indicator quantity corresponding to the positive sample label indicator set.
[0123] The negative sample group similarity of the sample users is calculated based on each of the first indicator values, each of the third indicator values, and the negative sample label indicator quantity corresponding to the negative sample label indicator set.
[0124] Optionally, the annotation module 104 is further configured to:
[0125] Based on the first weight of the satisfaction attribution probability value and the second weight of the positive sample group similarity, the positive probability of the sample user belonging to the positive sample indicator set is calculated, and based on the first weight and the third weight of the negative sample group similarity, the negative probability of the sample user belonging to the negative sample indicator set is calculated.
[0126] Based on the positive and negative probabilities, sample users corresponding to the real sample indicator set are labeled with pseudo-labels to obtain the sample pseudo-label indicator set.
[0127] Optionally, the acquisition module 101 is further configured to:
[0128] Based on historical user information, at least one indicator to be included in the model is constructed, wherein the indicator to be included in the model corresponds to one of the real sample indicator, the positive sample label indicator, and the negative sample label indicator;
[0129] The selected indicators to be included in the model are then filtered to obtain the filtered indicators to be included in the model.
[0130] Based on the selected indicators to be incorporated into the model, a set of sample training indicators is constructed.
[0131] Optionally, the acquisition module 101 is further configured to:
[0132] For any of the indicators to be included in the model, data drift detection is performed on the indicator based on the indicator skewness value and the indicator kurtosis value of the indicator to be included in the model.
[0133] After data drift detection is performed on all the indicators to be included in the model, the indicators that do not have data drift are selected as the filtered indicators to be included in the model.
[0134] Optionally, the acquisition module 101 is further configured to:
[0135] The index skewness value of the index to be included in the model is input into a preset skewness mapping function to obtain the index distribution skewness value of the index to be included in the model within the sample period; and the index kurtosis value of the index to be included in the model is input into a preset kurtosis mapping function to obtain the index distribution kurtosis value of the index to be included in the model within the sample period.
[0136] The data drift value is obtained by inputting the index distribution skewness value and the index distribution skewness value into a preset data drift function;
[0137] Based on the relationship between the data drift value and the preset data drift threshold, data drift detection is performed on the indicator to be included in the model.
[0138] Optionally, the satisfaction recognition device is further used for:
[0139] Obtain the mean and standard deviation of the sample indicator set corresponding to the indicator to be included in the model;
[0140] The mean and standard deviation of the sample indicator set are input into a first preset mean function to obtain the indicator skewness value of the indicator to be included in the model, and the mean and standard deviation of the sample indicator set are input into a second preset mean function to obtain the indicator kurtosis value of the indicator to be included in the model.
[0141] The satisfaction recognition device provided by this invention employs the satisfaction recognition method in the above embodiments, solving the technical problem of low recognition accuracy in model-based user satisfaction recognition. Compared with the prior art, the beneficial effects of the satisfaction recognition device provided by this invention are the same as those of the satisfaction recognition method provided in the above embodiments, and other technical features in this satisfaction recognition device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0142] Example 4
[0143] This invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the satisfaction recognition method in Embodiment 1 above.
[0144] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0145] like Figure 6 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus.
[0146] Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0147] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.
[0148] The electronic device provided by this invention employs the satisfaction recognition method in the above embodiments, solving the technical problem of low recognition accuracy in model-based user satisfaction recognition. Compared with the prior art, the beneficial effects of the electronic device provided by this invention are the same as those of the satisfaction recognition method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0149] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0150] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0151] Example 5
[0152] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the satisfaction recognition method in the above embodiment.
[0153] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0154] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0155] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying satisfaction labels and a real sample indicator set not carrying the satisfaction labels, the sample label indicator set including a positive sample label indicator set and a negative sample label indicator set; input the real sample indicator set into a satisfaction recognition model to be trained, and use the satisfaction recognition model to identify the satisfaction of the sample users corresponding to the real sample indicator set, obtaining the satisfaction belonging probability value of the sample users; and calculate the positive sample group similarity and negative sample group similarity of the sample users, respectively. In this model, the positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label indicator set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label indicator set. Based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity, the real sample indicator set is labeled with pseudo-labels to obtain a sample pseudo-label indicator set. Based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, the satisfaction recognition model to be trained is iteratively trained to obtain a satisfaction recognition model. The satisfaction recognition model is used to identify the satisfaction of the indicator to be identified to obtain the satisfaction recognition result.
[0156] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0158] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0159] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described satisfaction recognition method, thus solving the technical problem of low recognition accuracy in model-based user satisfaction recognition. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this invention are the same as those of the satisfaction recognition method provided in the above-described embodiments, and will not be repeated here.
[0160] Example 6
[0161] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the satisfaction recognition method described above.
[0162] The computer program product provided in this application solves the technical problem of low accuracy in model-based user satisfaction identification. Compared with the prior art, the beneficial effects of the computer program product provided in this embodiment are the same as those of the satisfaction identification method provided in the above embodiments, and will not be repeated here.
[0163] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for identifying satisfaction levels, characterized in that, The satisfaction recognition method includes: Obtain a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying a satisfaction label and a real sample indicator set not carrying the satisfaction label, and the sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set. The real sample index set is input into the satisfaction recognition model to be trained, and the satisfaction of the sample users corresponding to the real sample index set is identified by the satisfaction recognition model to be trained, so as to obtain the satisfaction belonging probability value of the sample user. Calculate the positive sample group similarity and negative sample group similarity of the sample users respectively, wherein the positive sample group similarity is used to characterize the average similarity between the sample users and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample users and the negative sample group corresponding to the negative sample label index set; Based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity, the real sample indicator set is labeled with pseudo-labels to obtain the sample pseudo-label indicator set; Based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, the satisfaction recognition model to be trained is iteratively trained to obtain the satisfaction recognition model. The satisfaction recognition model is used to identify the satisfaction level of the indicator to be identified, and the satisfaction recognition result is obtained. The step of obtaining the sample pseudo-label index set by labeling the real sample index set with pseudo-labels based on the satisfaction attribution probability value, the positive sample group similarity, and the negative sample group similarity includes: Based on the first weight of the satisfaction attribution probability value and the second weight of the positive sample group similarity, the positive probability of the sample user belonging to the positive sample indicator set is calculated, and based on the first weight and the third weight of the negative sample group similarity, the negative probability of the sample user belonging to the negative sample indicator set is calculated. Based on the positive and negative probabilities, sample users corresponding to the real sample indicator set are labeled with pseudo-labels to obtain the sample pseudo-label indicator set.
2. The satisfaction recognition method as described in claim 1, characterized in that, The set of real sample metrics includes at least one real sample metric, the set of positive sample label metrics includes at least one positive sample label metric, and the set of negative sample label metrics includes at least one negative sample label metric. The steps of calculating the positive sample group similarity and negative sample group similarity of the sample users respectively include: The first index value of each real sample index, the second index value of each positive sample label index, and the third index value of each negative sample label index are obtained respectively. The positive sample group similarity of the sample users is calculated based on each of the first indicator values, each of the second indicator values, and the positive sample label indicator quantity corresponding to the positive sample label indicator set. The negative sample group similarity of the sample users is calculated based on each of the first indicator values, each of the third indicator values, and the negative sample label indicator quantity corresponding to the negative sample label indicator set.
3. The satisfaction recognition method as described in claim 1, characterized in that, The steps for obtaining the sample training dataset constructed from historical user information include: Based on historical user information, at least one indicator to be included in the model is constructed, wherein the indicator to be included in the model corresponds to one of the real sample indicator, the positive sample label indicator, and the negative sample label indicator; The selected indicators to be included in the model are then filtered to obtain the filtered indicators to be included in the model. Based on the selected indicators to be incorporated into the model, a set of sample training indicators is constructed.
4. The satisfaction recognition method as described in claim 3, characterized in that, The step of filtering each of the indicators to be included in the model to obtain the filtered indicators to be included in the model includes: For any of the indicators to be included in the model, data drift detection is performed on the indicator based on the indicator skewness value and the indicator kurtosis value of the indicator to be included in the model. After data drift detection is performed on all the indicators to be included in the model, the indicators that do not have data drift are selected as the filtered indicators to be included in the model.
5. The satisfaction recognition method as described in claim 4, characterized in that, The step of detecting data drift of the indicator to be included in the model based on the indicator skewness value and the indicator kurtosis value of the indicator to be included in the model includes: The index skewness value of the index to be included in the model is input into a preset skewness mapping function to obtain the index distribution skewness value of the index to be included in the model within the sample period; and the index kurtosis value of the index to be included in the model is input into a preset kurtosis mapping function to obtain the index distribution kurtosis value of the index to be included in the model within the sample period. The data drift value is obtained by inputting the index distribution skewness value and the index distribution kurtosis value into a preset data drift function; Based on the relationship between the data drift value and the preset data drift threshold, data drift detection is performed on the indicator to be included in the model.
6. The satisfaction recognition method as described in claim 4, characterized in that, Before the step of detecting data drift of the target indicator based on its skewness and kurtosis values, the satisfaction recognition method further includes: Obtain the mean and standard deviation of the sample indicator set corresponding to the indicator to be included in the model; The mean and standard deviation of the sample indicator set are input into a first preset mean function to obtain the indicator skewness value of the indicator to be included in the model, and the mean and standard deviation of the sample indicator set are input into a second preset mean function to obtain the indicator kurtosis value of the indicator to be included in the model.
7. A satisfaction recognition device, characterized in that, The satisfaction recognition device includes: The acquisition module is used to acquire a sample training indicator set constructed from historical user information, wherein the sample training indicator set includes a sample label indicator set carrying a satisfaction label and a real sample indicator set without the satisfaction label, and the sample label indicator set includes a positive sample label indicator set and a negative sample label indicator set. The first identification module is used to input the real sample index set into the satisfaction identification model to be trained, and to identify the satisfaction of the sample users corresponding to the real sample index set through the satisfaction identification model to be trained, so as to obtain the satisfaction belonging probability value of the sample user. The calculation module is used to calculate the positive sample group similarity and negative sample group similarity of the sample user, wherein the positive sample group similarity is used to characterize the average similarity between the sample user and the positive sample group corresponding to the positive sample label index set, and the negative sample group similarity is used to characterize the average similarity between the sample user and the negative sample group corresponding to the negative sample label index set. The annotation module is used to annotate the real sample indicator set with pseudo-labels based on the satisfaction belonging probability value, the positive sample group similarity, and the negative sample group similarity, to obtain a sample pseudo-label indicator set. The annotation module is also used to calculate the positive probability of a sample user belonging to a positive sample indicator set based on a first weight of the satisfaction belonging probability value and a second weight of the positive sample group similarity, and to calculate the negative probability of a sample user belonging to a negative sample indicator set based on the first weight and a third weight of the negative sample group similarity; and to annotate the sample users corresponding to the real sample indicator set with pseudo-labels based on the positive and negative probabilities, to obtain a sample pseudo-label indicator set. The training module is used to iteratively train the satisfaction recognition model to be trained based on the target training indicator set formed by merging the sample pseudo-label indicator set and the sample label indicator set, so as to obtain the satisfaction recognition model. The second identification module is used to identify the satisfaction level of the indicator to be identified through the satisfaction identification model, and obtain the satisfaction identification result.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the satisfaction recognition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing the satisfaction recognition method, which is executed by a processor to implement the steps of the satisfaction recognition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for expanding potential users
CN108875761A
Positive no-label learning method based on label propagation
CN113869445A