Method for quickly identifying user value
The user behavior data is calculated through the box graph and the user value is evaluated based on multi-dimensionality, which solves the problems of singularity and inefficiency of traditional evaluation methods, and achieves fast and accurate user value identification and enterprise decision support.
Patent Information
- Application Number
- CN202510434848.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-01
AI Technical Summary
The traditional user value evaluation method has a single dimension and low efficiency, making it difficult to deal with massive data and complex business scenarios, and cannot accurately measure users' willingness to pay.
The box graph calculation principle is used to process user behavior data. Based on the five dimensions of offline days, login period, average daily clicks, average daily login success times, and average daily browsing time, data intervals are divided through interquartile and limit value calculation methods, different scores are assigned, and the weights of each indicator are set to calculate the total user value score.
It realizes rapid and accurate identification of user values, provides reliable corporate decision-making basis, improves work efficiency, and enhances user retention and corporate competitiveness.
Smart Images

Figure CN120410583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to user value evaluation, and more particularly to a method for quickly identifying user value. Background Art
[0002] In today's digital age, the competition of Internet products is becoming increasingly fierce, and accurately identifying user value has become the key to the development of enterprises. With the continuous expansion of the user scale, user behavior data has grown explosively, and traditional user value evaluation methods are facing many difficulties. On the one hand, in the past, it mostly relied on a single dimension or a simple index system, which was difficult to comprehensively reflect the true value of users in complex business scenarios. For example, only focusing on the user's consumption amount while ignoring potential value factors such as user activity and loyalty. On the other hand, the difficulty of processing massive data is large, and traditional methods cannot efficiently process and analyze it, resulting in a lag in evaluation results and being unable to provide timely support for enterprise decision-making. Therefore, there is an urgent need for a method that can quickly, comprehensively, and accurately identify user value to adapt to market changes and enhance the competitiveness of enterprises.
[0003] The Chinese patent with the publication number CN115619434A proposes a user value evaluation method and system implementation method based on TF-IDF, including the following steps: obtaining user terminal behavior data; performing user grouping; arranging the behaviors in ascending order of occurrence time and generating behavior combinations in units of users; calculating the TF-IDF weight values of the time series combinations and calculating the conversion rates of the time series combinations; combining the results of step S104, using TF-IDF as the vertical axis and the conversion rate as the horizontal axis to draw a quadrant diagram; performing behavior value analysis according to the visualization presentation results to implement the user value evaluation method; weighting the conversion rate of the time series combination according to the TF-IDF value; identifying high-value user behavior combinations to implement the user value evaluation system. The present invention makes it convenient for business personnel to quickly identify high-value combined behaviors of users without preprocessing the original data by drawing a quadrant diagram by combining the TF-IDF weight values and conversion rates of behavior combinations. However, according to the prior art, it is found that in the operation of Internet products, traditional user value evaluation methods are single-dimensional and inefficient, difficult to cope with massive data and complex business scenarios, and unable to accurately measure the willingness of users to pay. Summary of the Invention
[0005] In order to better solve the above problems, the present invention provides a method for quickly identifying user value, and the method includes: According to business requirements, select five dimensions of offline days, login cycle, average daily click times, average daily successful login times, and average daily browsing duration to evaluate the willingness of users to pay; Use the box plot calculation principle to process the data of users in the five dimensions and determine the specific division criteria for each index level interval; Using the quartile and extreme value calculation method of the box plot, divide the data of each indicator into six data intervals: (-∞, lower limit], (lower limit, Q1], (Q1, Q2], (Q2, Q3], (Q3, upper limit], (upper limit, +∞), and assign different scores to different data intervals; Confirm the weights of each indicator; Calculate the total user value score according to the scores of each indicator and their corresponding weights; Set up a total score value comparison table, and identify the value level of each user at the current stage based on the comparison table.
[0006] As a preferred technical solution of the present invention, the imaging unit is further configured that: the data of the five dimensions for evaluating the user's willingness to pay the bill are obtained by collecting the user's behavior data during the use of the Internet product.
[0007] As a preferred technical solution of the present invention, the acquisition unit processes the user data by using the calculation principle of the box plot, including: first sorting the user data of each dimension in ascending order, and then calculating the lower quartile Q1, median Q2, upper quartile Q3 and interquartile range IQR of the sorted user data of each dimension, where IQR = Q3 - Q1, where Q1 is the quartile, Q2 is the median, Q3 is the upper quartile and IQR is the interquartile range.
[0008] As a preferred technical solution of the present invention, the lower limit value in the data interval is calculated by the formula Q1 - 1.5 * IQR, and the upper limit value in the data interval is calculated by the formula Q3 + 1.5 * IQR.
[0009] As a preferred technical solution of the present invention, the weights of each indicator are respectively: the number of offline days accounts for 25%, the login cycle accounts for 25%, the average daily click times account for 12%, the average daily successful login times account for 20%, and the average daily browsing duration accounts for 18%.
[0010] As a preferred technical solution of the present invention, the formula for calculating the total user value score is: the total user value score S = Σ(indicator score × corresponding weight), where Σ represents the sum of the products of the scores of each indicator and the corresponding weights.
[0011] As a preferred technical solution of the present invention, in the total score value comparison table, the higher the total score, the higher the corresponding value level, and the lower the total score, the lower the corresponding value level.
[0012] As a preferred technical solution of the present invention, before data collection, it is necessary to clearly set the scope and time span of data collection. The data scope includes all or part of the users of a specific Internet product, and the time span is determined as a continuous period according to business requirements. Before calculating the total user value score, the collected user data is cleaned and preprocessed, and a data screening algorithm is used to remove abnormal data and duplicate data. The abnormal data includes data that significantly exceeds the reasonable range and data that does not meet the data type requirements, and the duplicate data refers to data records with exactly the same data content.
[0013] The present invention also provides a system for quickly identifying user value. The system is used to implement the method for quickly identifying user value described in any one of the foregoing technical solutions. The system includes: A selection unit, configured to select five dimensions of offline days, login period, average daily click times, average daily successful login times, and average daily browsing duration according to business requirements to evaluate the willingness of users to place orders. A division unit, configured to process the data of users in the above five dimensions by using the calculation principle of box plots to determine the specific division criteria for each index level interval, and also configured to use the quartile and extreme value calculation methods of box plots to divide each index data into six data intervals: (-∞, lower limit], (lower limit, Q1], (Q1, Q2], (Q2, Q3], (Q3, upper limit], (upper limit, +∞), and assign different scores to different data intervals. A calculation unit, configured to confirm the weights of each index, and calculate the total user value score according to the scores of each index and their corresponding weights. A determination unit, configured to set a total score value comparison table, and identify the value level of each user at the current stage based on this comparison table.
[0014] The present invention also provides a storage medium storing program instructions, wherein when the program instructions run, the device where the storage medium is located is controlled to execute the method for quickly identifying user value described in any one of the foregoing technical solutions.
[0015] Compared with the prior art, the beneficial effects of the present invention are at least as follows: 1. Traditional user value judgment often relies on subjectivity and lacks theoretical support. Starting from business requirements, the present invention selects five key dimensions of offline days, login period, average daily click times, average daily successful login times, and average daily browsing duration to comprehensively reflect the interaction between users and products and potential consumption willingness. Using the calculation principle of box plots to process data can overcome the problem of large randomness of user behavior data and objectively and accurately divide the grade intervals. Combining with the scientifically set weights of each index to calculate the total score can accurately judge the value level of each user and provide a reliable basis for enterprise decision-making.
[0016] 2. In the era of big data, user behavior data is massive and complex. The boxplot calculation principle adopted by this invention is not affected by outliers, the model is stable, and it can quickly process large amounts of data. After data collection, operations such as sorting and quartile calculation can be completed quickly, and the data intervals of each indicator can be determined and assigned scores. Compared with traditional processing methods, this greatly shortens processing time, reduces enterprise data processing costs, and improves work efficiency.
[0017] 3. Based on accurate user value assessment results, companies can develop personalized retention strategies for users of varying value levels. For high-value users, exclusive services and personalized recommendations can be provided to enhance their loyalty; for low-value users, targeted activation campaigns can be designed to tap into their potential value. This differentiated strategy can effectively improve user retention and enable more efficient allocation of corporate resources.
[0018] 4. Accurately identifying user value helps companies gain a deeper understanding of user needs, enabling them to optimize their products and services. By continuously tapping into user potential, companies can maintain their advantage in the fiercely competitive market, continuously innovate their business, and achieve sustainable development. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A flowchart of a method for quickly identifying user value in the present invention; Figure 2 This is a structural diagram of a method system for quickly identifying user value in the present invention. DETAILED DESCRIPTION
[0020] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0021] like Figure 1 As shown, the present invention provides a method for quickly identifying user value, the method comprising: Based on business needs, we selected five dimensions: offline days, login cycle, average daily clicks, average daily successful logins, and average daily browsing time to assess users' willingness to pay. Use the boxplot calculation principle to process the user's data in five dimensions and determine the specific division standards for each indicator level interval; Using the quartile and limit value calculation method of the box plot, the indicator data are divided into six data intervals: (-∞, lower limit value], (lower limit value, Q1], (Q1, Q2], (Q2, Q3], (Q3, upper limit value], (upper limit value, +∞), and different scores are assigned to different data intervals; Confirm the weight of each indicator; Calculate the total user value score based on the scores of each indicator and their corresponding weights; Set up a total score value comparison table, and quickly identify the value level of each user at the current stage based on the comparison table.
[0022] Specifically, based on the correlation between user behavior and the willingness to pay, quantitatively evaluate the value of users in the business from multiple dimensions. Through comprehensive analysis of multi-dimensional data, more comprehensively and accurately measure the contribution and potential value of users to the business, and determine the selection of offline days, login cycle, average daily click times, average daily successful login times, and average daily browsing duration as evaluation dimensions according to business requirements. Collect user behavior data in these dimensions from data sources such as the log system and database of Internet products.
[0023] The data of the five dimensions used to evaluate the willingness of users to pay are obtained by collecting the behavior data of users during the use of Internet products.
[0024] Specifically, the behavior data is the real record of users during the use of Internet products, which can intuitively reflect the behavior patterns and preferences of users, provide a reliable basis for evaluating the willingness to pay. At the same time, use data collection tools such as Flume and Logstash to collect user behavior data in real time or regularly from different sources such as the front-end interaction logs and back-end service logs of the product, and store it in a data warehouse or distributed file system such as Hive and HDFS.
[0025] When using the box plot calculation principle to process user data, the specific steps are as follows: first, sort the user data of each dimension in ascending order, and then calculate the lower quartile Q1, median Q2, upper quartile Q3, and interquartile range IQR of the sorted user data of each dimension, where IQR = Q3 - Q1.
[0026] Specifically, the box plot calculation principle is based on the statistical characteristics of data, can effectively display the distribution of data, and provides an objective standard for dividing the data interval by determining the quartiles and interquartile range. Use the pandas library of Python or other data analysis tools to sort the user data of each dimension in ascending order. Call relevant functions to calculate the lower quartile Q1, median Q2, upper quartile Q3, and interquartile range IQR (IQR = Q3 - Q1). For example, in pandas, the describe() method can be used to obtain basic statistics, and then IQR can be obtained through custom calculation.
[0027] The lower limit value is calculated by the formula Q1 - 1.5 * IQR, and the upper limit value is calculated by the formula Q3 + 1.5 * IQR, where Q1 is the first quartile, Q2 is the median, Q3 is the third quartile, and IQR is the interquartile range.
[0028] Specifically, by calculating the lower limit value and the upper limit value through specific formulas, the normal range of data can be clearly defined, potential outliers can be identified, and the accuracy of data analysis can be ensured. According to the calculated Q1 and IQR, the formula Q1 - 1.5 * IQR is used to calculate the lower limit value, and the formula Q3 + 1.5 * IQR is used to calculate the upper limit value. In code implementation, mathematical operations are directly used for calculation, such as being completed through simple mathematical expressions in Python.
[0029] The weights of each indicator are as follows: the number of offline days accounts for 25%, the login cycle accounts for 25%, the average number of daily clicks accounts for 12%, the average number of daily successful logins accounts for 20%, and the average daily browsing duration accounts for 18%.
[0030] Specifically, assigning weights to each indicator takes into account the different importance degrees of different dimensions in the evaluation of user value. Through reasonable weight assignment, the comprehensive evaluation result can better conform to the actual business situation. At the same time, the weights of 25% for the number of offline days, 25% for the login cycle, 12% for the average number of daily clicks, 20% for the average number of daily successful logins, and 18% for the average daily browsing duration are determined. In subsequent calculations, these weights are used as fixed parameters to participate in the operations.
[0031] The formula for calculating the total user value score is: the total user value S = Σ(indicator score × corresponding weight), where Σ represents the sum of the products of the scores of each indicator and the corresponding weights.
[0032] Specifically, the weighted summation formula is used to calculate the total user value score, and the evaluation results of each dimension are integrated to obtain a quantitative value representing the overall user value. Traverse the scores of each indicator, multiply them by the corresponding weights, and then use the sum() function in Python or the accumulation function in other programming languages to sum the product results to obtain the total user value score S.
[0033] In the total score value comparison table, the higher the total score, the higher the corresponding value level, and the lower the total score, the lower the corresponding value level.
[0034] Specifically, set up a total score - value comparison table to divide the continuous total user value scores into different levels, which is convenient for intuitively classifying and managing users, formulating targeted strategies. Create a dictionary or table structure to store the corresponding relationship between the total score and the value level, such as {5: "extremely high", (4, 5): "high", (3, 4): "medium", (2, 3): "relatively low", (1, 2): "low", (0, 1): "extremely low"}. When evaluating the user value level, determine it by looking up this comparison table.
[0035] Before data collection, it is necessary to clearly set the scope and time span of data collection. The data scope includes all or part of the users of a specific Internet product, and the time span is determined as a continuous period according to business requirements.
[0036] Specifically, clarifying the data collection scope and time span can ensure that the collected data is targeted and timely, avoid interference from irrelevant data, improve data processing efficiency and the accuracy of analysis results. Before data collection, formulate a detailed data collection plan, clearly stipulate the data collection scope, such as new users, active users, etc. of a specific Internet product, and determine the time span, such as the last month, the last three months, etc. During the collection process, screen and collect data according to the set scope and time conditions.
[0037] Before calculating the total user value score, clean and pre - process the collected user data. Use a data screening algorithm to remove abnormal data and duplicate data. Abnormal data includes data that significantly exceeds the reasonable range and data that does not meet the data type requirements. Duplicate data refers to data records with exactly the same content.
[0038] Specifically, clean and pre - process the collected raw data to remove abnormal data and duplicate data, improve data quality, and ensure the reliability and accuracy of subsequent analysis results. Use a data screening algorithm, such as a rule - based filtering algorithm. According to the characteristics of abnormal data (such as significantly exceeding the reasonable range, data type error) and the definition of duplicate data (exactly the same data content), write code for data cleaning. In Python, the drop_duplicates() method of the pandas library can be used to remove duplicate data, and abnormal data can be screened out through conditional judgment and processed.
[0039] As Figure 2 shown, the present invention also provides a system for quickly identifying user value. The system is used to implement the method for quickly identifying user value in any one of the foregoing technical solutions. The system includes: A selection unit, which is used to select five dimensions of offline days, login period, average daily click-through rate, average daily successful login times, and average daily browsing duration according to business requirements to evaluate the user's willingness to place an order; A division unit, which is used to process the data of users in the above five dimensions by using the calculation principle of box plots to determine the specific division criteria for each index level interval, and is also used to divide the data of each index into six data intervals of (-∞, lower limit], (lower limit, Q1], (Q1, Q2], (Q2, Q3], (Q3, upper limit], and (upper limit, +∞) by using the quartile and extreme value calculation methods of box plots, and assign different scores to different data intervals; A calculation unit, which is used to confirm the weights of each index and calculate the total user value score according to the scores of each index and their corresponding weights; A determination unit, which is used to set a total score value comparison table and identify the value level of each user at the current stage based on this comparison table.
[0040] The present invention also provides a storage medium storing program instructions, which control the device where the storage medium is located to execute the method for quickly identifying user value according to any one of the foregoing technical solutions when the program instructions are running.
[0041] In summary, in specific use, the server-side log system undertakes the important task of recording the user's login and logout times. It accurately captures the moments of each user's login and logout through timestamp technology and stores these times in the database in a specific format. When calculating the offline days, the system will obtain the current time in real time and compare it with the user's last login time, and use the time difference calculation algorithm to accurately obtain the offline days. For example, if the user logs in at 10:00 on May 1 and has not logged in since then, when the system conducts data statistics at 10:00 on May 10, by calculating the time difference of "10:00 on May 10 - 10:00 on May 1", the offline days are obtained as 9 days. This recording and calculation method ensures the accuracy and real-time nature of the offline days data.
[0042] Secondly, it relies on the timestamps in the server-side logs. When a user logs in each time, the system not only records the login time but also associates this time with the previous login time. Through time series analysis algorithms, the system can automatically calculate the time intervals between adjacent logins. For the convenience of subsequent analysis and processing, the system classifies and stores these login cycle data according to the user ID. For example, if user A's login times within a week are 10:00 on Monday, 14:00 on Wednesday, and 16:00 on Friday, the system will calculate the login cycle from Monday to Wednesday as "14:00 on Wednesday - 10:00 on Monday" and the login cycle from Wednesday to Friday as "16:00 on Friday - 14:00 on Wednesday", and organize and store these data for subsequent statistics and analysis.
[0043] In addition, at the client side, data collection is achieved by embedding data collection code in key interactive elements of the application or web page. When a user performs a click operation, the click event triggers the collection code, and the code sends a record containing the click time, click location, etc. to the server. When the login is successful, the login success event triggers the code to record the login success time and related information. For the browsing duration, the system starts a timer when the user enters the page and stops the timer when the user leaves the page, and sends the recorded duration data to the server. The server side uses data aggregation algorithms to summarize and statistically analyze these data on a daily basis. For example, on a certain day, user B performed 100 click operations, logged in successfully 5 times, and the cumulative browsing duration was 120 minutes. At the end of the day, the server side aggregates and calculates these scattered records to obtain the user's daily average click count of 100 times, daily average login success count of 5 times, and daily average browsing duration of 120 minutes.
[0044] Subsequently, the data collected in each dimension is sorted separately, which is the basis for subsequent calculation of quartiles and determination of data intervals. The sorting algorithm uses an efficient sorting method, such as the quicksort algorithm. Taking the offline days data as an example, assuming that the offline days data of 100 users is collected, the quicksort algorithm will rearrange these data in ascending order, making the data in an ordered state for subsequent calculation operations.
[0045] Based on the ordered data, calculate the lower quartile Q1, median Q2, upper quartile Q3, and interquartile range IQR. For a set of data, Q1 represents the value at the 25% position of the data, Q2 represents the value at the 50% position (i.e., the median), and Q3 represents the value at the 75% position. When calculating the quartiles, determine the corresponding positions according to the number of data. For example, for 100 data, Q1 is at the 25th data position, Q2 is at the 50th data position, and Q3 is at the 75th data position. IQR is calculated by subtracting Q1 from Q3, which reflects the degree of dispersion of the data.
[0046] In addition, use the quartile and extreme value calculation method of the box plot to determine 6 data intervals for each indicator. The lower limit value is calculated by the formula Q1 - 1.5 * IQR, and the upper limit value is calculated by the formula Q3 + 1.5 * IQR. For example, for the data of the number of offline days, if Q1 is calculated to be 5 days, Q3 is 15 days, and IQR is 10 days, then the lower limit value is 5 - 1.5 * 10 = -10 days (in practical applications, the data cannot be negative, so the lower limit value is taken as 0 at this time), and the upper limit value is 15 + 1.5 * 10 = 30 days. Thus, the data interval of the number of offline days is determined as (-∞, 0], (0, 5], (5, 10], (10, 15], (15, 30], (30, +∞). This interval division method based on the box plot can effectively exclude the interference of outliers and more accurately reflect the distribution characteristics of the data.
[0047] Assigning different scores to different data intervals is a key step in converting data into intuitive user value evaluation indicators. The setting of scores is based on an in-depth analysis of the relationship between user behavior and user value. Generally, the higher the score, the higher the user value. For example, for the data interval of the number of offline days, the interval (-∞, 0] is assigned 1 point, which means that users in this interval have logged in recently and may maintain a certain degree of attention to the product. The interval (0, 5] is assigned 2 points, indicating that users have a relatively short offline time and relatively high activity. As the interval changes, the score gradually increases. The interval (30, +∞) is assigned 6 points, indicating that users in this interval have a very long offline time and may have a relatively low interest in the product. Data in other dimensions are divided into intervals and scores are assigned according to the same logic, so that the data in each dimension can be converted into comparable user value scores.
[0048] The weights of each indicator are set as follows: offline days 25%, login cycle 25%, average daily click-through times 12%, average daily successful login times 20%, and average daily browsing duration 18%. This is based on a comprehensive consideration of the importance of each dimension in the process of evaluating user value. The offline days and login cycle reflect the interaction frequency between users and the product, which is of great significance for judging user activity and loyalty. Therefore, relatively higher weights are assigned. The average daily click-through times and average daily successful login times reflect the operation behavior and participation degree of users within the product, with secondary weights. The average daily browsing duration reflects the degree of attention of users to the product content and also has certain reference value. Therefore, relatively lower weights are assigned. This way of weight distribution has been verified through multiple experiments and data analysis and can more accurately reflect the contribution degree of each dimension to user value.
[0049] The total user value score S is calculated using the formula "Total user value score S = Σ(indicator score * corresponding weight)". This formula is based on the principle of weighted average, and the scores of each dimension are weighted and calculated according to their corresponding weights to obtain a total score that comprehensively reflects user value. For example, assuming that the score of user C for offline days is 3 points, the score for the login cycle is 4 points, the average daily click-through times score is 2 points, the average daily successful login times score is 5 points, and the average daily browsing duration is 3 points, then the total value score S of this user = 3 * 25% + 4 * 25% + 2 * 12% + 5 * 20% + 3 * 18% = 3.53 points. Through this calculation method, the user behavior data of multiple dimensions can be integrated into a quantitative value indicator, which is convenient for unified evaluation and comparison of user value.
[0050] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as within the scope described in this specification.
[0051] The above embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
[0052] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for quickly identifying user value, characterized in that, The method includes: According to business requirements, select five dimensions: offline days, login period, average daily click times, average daily successful login times, and average daily browsing duration to evaluate the user's willingness to place an order; Use the box plot calculation principle to process the user data in the five dimensions and determine the specific division criteria for each index level interval; Use the quartile and extreme value calculation methods of the box plot to divide the data of each index into six data intervals and assign different scores to different data intervals; Confirm the weights of each index; Calculate the total user value score according to the score of each index and its corresponding weight; Set a total score value comparison table and identify the value level of each user at the current stage based on the comparison table.
2. The method for quickly identifying user value according to claim 1, wherein It includes: The data of the five dimensions used to evaluate the user's willingness to place an order are obtained by collecting the user's behavior data during the use of Internet products.
3. The method for quickly identifying user value according to claim 1, characterized in that, Using the box plot calculation principle to process user data includes: first sorting the user data of each dimension in ascending order, and then calculating the lower quartile Q1, median Q2, upper quartile Q3, and interquartile range IQR of the sorted user data of each dimension, where IQR = Q3 - Q1.
4. The method for quickly identifying user value according to claim 3, wherein It includes: The lower limit value in the data interval is calculated by the formula Q1 - 1.5 * IQR, and the upper limit value in the data interval is calculated by the formula Q3 + 1.5 * IQR, where Q1 is the quartile, Q2 is the median, Q3 is the upper quartile, and IQR is the interquartile range.
5. The method for quickly identifying user value according to claim 1, wherein It includes: The weights of each index are respectively: offline days account for 25%, login period accounts for 25%, average daily click times account for 12%, average daily successful login times account for 20%, and average daily browsing duration accounts for 18%.
6. The method for quickly identifying user value according to claim 1, wherein It includes: The formula for calculating the total user value score is: total user value S = Σ(indicator score × corresponding weight), where Σ represents the sum of the products of the scores of each indicator and the corresponding weights.
7. The method for quickly identifying user value according to claim 1, characterized in that It includes: In the total score value comparison table, the higher the total score, the higher the corresponding value level, and the lower the total score, the lower the corresponding value level.
8. The method for quickly identifying user value according to claim 1, characterized in that, It includes: Before data collection, it is necessary to clearly set the scope and time span of data collection. The data scope includes all or part of the users of a specific Internet product, and the time span is determined as a continuous period according to business requirements. Before calculating the total user value score, clean and preprocess the collected user data, and use a data screening algorithm to remove abnormal data and duplicate data. The abnormal data includes data that significantly exceeds the reasonable range and data that does not meet the data type requirements, and duplicate data refers to data records with exactly the same content.
9. A rapid user value recognition system, which is used to implement the rapid user value recognition method described in any one of claims 1 to 8, characterized in that, The system includes: A selection unit for selecting five dimensions: offline days, login period, average daily click times, average daily successful login times, and average daily browsing duration according to business requirements to evaluate the user's willingness to place an order; A division unit, which is used to process the user's data in the above five dimensions by using the calculation principle of box plots, determine the specific division criteria for each index level interval, and is also used to divide each index data into six data intervals by using the quartile and extreme value calculation methods of box plots, and assign different scores to different data intervals; A calculation unit, which is used to confirm the weight of each index and calculate the total user value score according to the score of each index and its corresponding weight; A determination unit, which is used to set a total score value comparison table and identify the value level of each user at the current stage based on this comparison table.
10. A storage medium, characterized in that, The storage medium stores program instructions, and when the program instructions run, it controls the device where the storage medium is located to execute the method for quickly identifying user value according to any one of claims 1 to 8.
Citation Information
Patent Citations
User value evaluation method and system implementation method based on TF-IDF
CN115619434A