Method and apparatus for determining a characteristic value range

Through the method of feature discretization and continuous, constraints and objective functions are established, and the optimal solution of the objective function is calculated, which solves the problems of inconsistent training effects and inconsistent operation management in the existing technology, and achieves high-quality business services and training effects.

CN114358985BActive Publication Date: 2025-06-27泰康保险集团股份有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111459466.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-06-27
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

When training relevant personnel in the existing technology, the training results of each training class are different, and the operation methods of each training class cannot be managed uniformly, resulting in poor service quality and business service effects.

Method used

By discretizing the features and calculating the index value of the features on each discrete value, establishing constraints and objective functions to continuously the features and calculating the optimal solution of the objective function, we obtain the best value range for the feature performance, maximize service quality and improve the service effect of business services.

Benefits of technology

It realizes the mining of valuable data from sample data, maximizes service quality, improves the service effect of business services, and determines the best value range for the mono- and multi-character characteristics, which improves the training effect and resource utilization rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358985B_ABST
    Figure CN114358985B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for determining a characteristic value range, relating to the field of computer technologies. A specific implementation manner of the method includes: discretizing the characteristic values of characteristics included in sample data to obtain the discrete values of the characteristics, and calculating the index values of the characteristics at different discrete values according to a set evaluation index and the class labels of the sample data; wherein, the sample data is data generated by a business service; establishing a constraint condition and an objective function for maximizing the service quality of the business service according to the discrete values and the index values; wherein, the constraint condition is obtained by integrating the index values of the evaluation index with the discrete values located in a set interval as the upper and lower limits of integration; and solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the characteristic. This implementation manner can maximize the service quality and improve the service effect of the business service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and apparatus for determining a characteristic value range. Background Art

[0002] In order to improve the professional levels of relevant personnel, such as salespersons and volunteers, many institutions will provide training services for training relevant personnel. In the prior art, when training relevant personnel, it is usually carried out in the form of offline training courses according to historical training experience. In this way, the training effects of each training course are different, and the operation modes of each training course cannot be uniformly managed. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method and apparatus for determining a characteristic value range. The method discretizes a characteristic in sample data, calculates an index value of the characteristic at each discrete value, and then based on the discrete value and the index value, establishes a constraint condition and an objective function to continuousize the characteristic and calculate an optimal solution of the objective function, so as to obtain a value range where the characteristic performs best, which can maximize the service quality and improve the service effect of business services.

[0004] To achieve the above object, according to one aspect of embodiments of the present invention, a method for determining a characteristic value range is provided.

[0005] A method for determining a characteristic value range according to an embodiment of the present invention includes: discretizing characteristic values of a characteristic included in sample data to obtain discrete values of the characteristic, and calculating index values of the characteristic at different discrete values according to a set evaluation index and a class label of the sample data; wherein the sample data is data generated by a business service; establishing a constraint condition and an objective function for maximizing the service quality of the business service according to the discrete value and the index value; wherein the constraint condition is obtained by integrating the index value of the evaluation index with the discrete value within a set interval as the upper and lower limits of integration; and solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the characteristic.

[0006] Optionally, the sample data includes a plurality of the characteristics and corresponding characteristic values; before the step of calculating the index values of the characteristic at different discrete values, the method further includes: performing a correlation analysis on the plurality of characteristics to determine whether there is a logical relationship between the plurality of characteristics; if there is a logical relationship between the plurality of characteristics, determining the plurality of characteristics as multivariate characteristics and determining the logical relationship between the plurality of characteristics; if there is no logical relationship between the plurality of characteristics, determining the characteristic as a univariate characteristic;

[0007] Calculating the index values of the feature at different discrete values includes: calculating the index values of the univariate feature at different discrete values, and the index values of the multivariate feature at different discrete value groups; wherein, the discrete value group is obtained by combining the discrete values of each feature in the multivariate feature.

[0008] Optionally, the evaluation indicators include the net excellent rate and the support rate. The net excellent rate represents the probability that the univariate feature reaches the set target condition at the current discrete value, or the multivariate feature reaches the set target condition at the current discrete value group; the support rate represents the proportion of the number of samples of the univariate feature at the current discrete value, or the multivariate feature at the current discrete value group, in the total number of samples.

[0009] The constraint conditions include the cumulative excellent rate and the cumulative support rate. The cumulative excellent rate is obtained by integrating the index values of the net excellent rate, and the cumulative support rate is obtained by integrating the index values of the support rate.

[0010] The method further includes: classifying the sample data in the sample data set according to the target condition to obtain the class label corresponding to the sample data.

[0011] Optionally, the objective function established for the univariate feature is:

[0012] f(x i ,x j )=cumsumNetRate(x i ,x j )×cumsumSuportRate(x i ,x j )

[0013] In the formula, f(x i ,x j ) represents a function with the discrete values x i and x j of the univariate feature as variables; cumsumNetRate(x i ,x j ) represents the cumulative excellent rate when the discrete values of the univariate feature are in the interval [x i ,x j ; cumsumSuporRate(x i ,x j ) represents the cumulative support rate when the discrete values of the univariate feature are in the interval [x i ,x j .

[0014] Optionally, solving the objective function under the constraint conditions through the dynamic programming algorithm to obtain the value range of the feature includes: determining a first initial value, a first step size, a first target interval, and a second target interval for iterating the unary feature; left boundary iteration: gradually reducing the first initial value to the left according to the first step size, calculating the value of the objective function in the first target interval, and determining to end the left boundary iteration and obtain the left boundary value when a set first stop condition is met during the iteration; right boundary iteration: gradually increasing the first initial value to the right according to the first step size, calculating the value of the objective function in the second target interval, and determining to end the right boundary iteration and obtain the right boundary value when a set second stop condition is met during the iteration; using the left boundary value as the minimum value of the unary feature and the right boundary value as the maximum value of the unary feature to obtain the value range of the unary feature.

[0015] Optionally, the objective function established for the multi - feature is:

[0016] f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)=

[0017] h(p r ,p s ,p t ,…)×cumsumNetRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)×

[0018] cumsumSuportRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)

[0019] In the formula, f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) represents the discrete value group (x rm ,x rn ,x sa ,xsb , x te , x tf , …) is a function of variables; h(p r , p s , p t , …) represents the logical relationship corresponding to the multivariate feature (r, s, t, …); cumsumNetRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative excellent rate when the discrete values of the respective features in the multivariate feature (r, s, t, …) are in the intervals ([x m , x n , [x a , x b , [x e , x f , …); cumsumSuportRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative support rate when the discrete values of the respective features in the multivariate feature (r, s, t, …) are in the intervals ([x m , x n , [x a , x b , [x e , x f , …).

[0020] Optionally, solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the feature includes: determining a second initial value, an iteration direction, a second step size, a third target interval, and a fourth target interval for iterative calculation of each feature in the multivariate feature; calculating the value range of each feature in different iteration directions according to the following steps, substituting the value ranges of each feature in different iteration directions into the objective function respectively, and using the value range that makes the objective function take the maximum value as the value range of the multivariate feature:

[0021] Left boundary iteration: According to the second step size of each feature, gradually reduce the second initial value of each feature to the left, calculate the value of the objective function in the third target interval of each feature, and during the iteration process, when it is determined that the set third stop condition is met, end the left boundary iteration to obtain the left boundary value of each feature; Right boundary iteration: According to the second step size of each feature, gradually increase the second initial value of each feature to the right, calculate the value of the objective function in the fourth target interval of each feature, and during the iteration process, when it is determined that the set fourth stop condition is met, end the right boundary iteration to obtain the right boundary value of each feature; and use the left boundary value of each feature as the minimum value of the corresponding feature, and the right boundary value of each feature as the maximum value of the corresponding feature to obtain the value interval of each feature.

[0022] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a device for determining a value interval of a feature.

[0023] A device for determining a value interval of a feature according to an embodiment of the present invention includes: a feature discretization module, configured to discretize the feature values of the features included in the sample data to obtain the discrete values of the features, and calculate the index values of the features at different discrete values according to the set evaluation index and the class label of the sample data; wherein, the sample data is data generated by business services; a feature continuous module, configured to establish a constraint condition and an objective function for maximizing the service quality of the business service according to the discrete values and the index values; wherein, the constraint condition is obtained by integrating the index values of the evaluation index with the discrete values within the set interval as the upper and lower limits of integration; an interval determination module, configured to solve the objective function under the constraint condition through a dynamic programming algorithm to obtain the value interval of the feature.

[0024] Optionally, the sample data includes a plurality of the features and corresponding feature values; the device further includes: a feature partitioning module, configured to perform a correlation analysis on the plurality of features to determine whether there is a logical relationship between the plurality of features; if there is a logical relationship between the plurality of features, determine that the plurality of features are multivariate features and determine the logical relationship between the plurality of features; if there is no logical relationship between the plurality of features, determine that the feature is a univariate feature;

[0025] The feature discretization module is further configured to calculate the index values of the univariate feature at different discrete values, and the index values of the multivariate feature at different discrete value groups; wherein, the discrete value group is obtained by combining the discrete values of each feature in the multivariate feature.

[0026] Optionally, the evaluation metrics include the net excellent rate and the support rate. The net excellent rate represents the probability that the unary feature reaches a set target condition at the current discrete value, or the multivariate feature reaches the set target condition at the current discrete value group. The support rate represents the proportion of the number of samples of the unary feature at the current discrete value, or the multivariate feature at the current discrete value group, in the total number of samples.

[0027] The constraint conditions include the cumulative excellent rate and the cumulative support rate. The cumulative excellent rate is obtained by integrating the index value of the net excellent rate, and the cumulative support rate is obtained by integrating the index value of the support rate.

[0028] The device further includes: a sample marking module, configured to classify the sample data of the sample data set according to the target condition to obtain the class label corresponding to the sample data.

[0029] Optionally, the objective function established for the unary feature is:

[0030] f(x i ,x j )=cumsumNetRate(x i ,x j )×cumsumSuportRate(x i ,x j )

[0031] In the formula, f(x i ,x j ) represents a function with the discrete values x i and x j of the unary feature as variables; cumsumNetRate(x i ,x j ) represents the cumulative excellent rate when the discrete value of the unary feature is in the interval [x i ,x j ; cumsumSuporRate(x i ,x j ) represents the cumulative support rate when the discrete value of the unary feature is in the interval [x i ,x j .

[0032] Optionally, the interval determination module is further configured to determine a first initial value, a first step size, a first target interval, and a second target interval for iterating the unary feature; Left boundary iteration: gradually reduce the first initial value to the left according to the first step size, calculate the value of the objective function in the first target interval, and during the iteration, when it is determined that a set first stop condition is satisfied, end the left boundary iteration to obtain a left boundary value; Right boundary iteration: gradually increase the first initial value to the right according to the first step size, calculate the value of the objective function in the second target interval, and during the iteration, when it is determined that a set second stop condition is satisfied, end the right boundary iteration to obtain a right boundary value; Use the left boundary value as the minimum value of the unary feature and the right boundary value as the maximum value of the unary feature to obtain the value interval of the unary feature.

[0033] Optionally, the objective function established for the multi-variate feature is:

[0034] f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)=

[0035] h(p r ,p s ,p t ,…)×cumsumNetRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)×

[0036] cumsumSuportRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)

[0037] In the formula, f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) represents the discrete value group (x rm ,x rn ,x sa ,x sb ,x te ,xtf , …) is a function of variables; h(p r , p s , p t , …) represents the logical relationship corresponding to the multivariate features (r, s, t, …); cumsumNetRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative excellent rate when the discrete values of the respective features in the multivariate features (r, s, t, …) are in the intervals ([x m , x n , [x a , x b , [x e , x f , …); cumsumSuportRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative support rate when the discrete values of the respective features in the multivariate features (r, s, t, …) are in the intervals ([x m , x n , [x a , x b , [x e , x f , …).

[0038] Optionally, the interval determination module is further configured to determine a second initial value, an iteration direction, a second step size, a third target interval, and a fourth target interval for each feature in the multivariate feature to be iterated; calculate the value intervals of each feature in different iteration directions according to the following steps, and substitute the value intervals of each feature in different iteration directions into the objective function, and use the value interval that makes the objective function take the maximum value as the value interval of the multivariate feature:

[0039] Left boundary iteration: According to the second step length of each feature, gradually reduce the second initial value of each feature to the left, calculate the value of the objective function in the third target interval of each feature, and during the iteration process, when it is determined that the set third stop condition is met, end the left boundary iteration to obtain the left boundary value of each feature; Right boundary iteration: According to the second step length of each feature, gradually increase the second initial value of each feature to the right, calculate the value of the objective function in the fourth target interval of each feature, and during the iteration process, when it is determined that the set fourth stop condition is met, end the right boundary iteration to obtain the right boundary value of each feature; and take the left boundary value of each feature as the minimum value of the corresponding feature, and the right boundary value of each feature as the maximum value of the corresponding feature to obtain the value range of each feature.

[0040] To achieve the above object, according to another aspect of the embodiments of the present invention, an electronic device is provided.

[0041] An electronic device according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement a method for determining a value range of a feature according to an embodiment of the present invention.

[0042] To achieve the above object, according to another aspect of the embodiments of the present invention, a computer-readable medium is provided.

[0043] A computer-readable medium according to an embodiment of the present invention has a computer program stored thereon, and when the program is executed by a processor, it implements a method for determining a value range of a feature according to an embodiment of the present invention.

[0044] One embodiment of the above invention has the following advantages or beneficial effects: By discretizing features and calculating the index values of features at each discrete value, and then based on the discrete values and index values, establishing constraint conditions and an objective function to continuousize the features and calculate the optimal solution of the objective function, the value range where the features perform best can be obtained, which can mine valuable data from sample data, maximize the service quality, and improve the service effect of business services.

[0045] Based on correlation analysis, features are divided into unary features and multi-variable features. Subsequently, objective functions are established for unary features and multi-variable features respectively, ensuring that during the process of continuousizing features, the value ranges where unary features and multi-variable features perform best can be determined respectively, and the credibility of the value ranges is high. Using the net excellent rate and support rate as evaluation indicators, and the cumulative excellent rate and cumulative support rate as constraint conditions, it is ensured that the value ranges of the subsequent obtained features are the best interval ranges of the features that meet the target conditions.

[0046] For unary features, during the process of feature continuousization, only the optimal interval of a single feature needs to be considered; for multi - feature cases, logical relationships need to be added during the feature continuousization process to determine the optimal interval of the combined feature, which has strong scalability. By using the dynamic programming algorithm to solve the objective function, an accurate feature value interval can be obtained. Subsequently, relevant personnel can be trained according to this value interval, improving the training effect and saving training resources.

[0047] The further effects of the above - mentioned non - conventional optional methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings are used to better understand the present invention and do not unduly limit the present invention. Among them:

[0049] Figure 1 is a schematic diagram of the main steps of the method for determining the feature value interval according to an embodiment of the present invention;

[0050] Figure 2 is a schematic diagram of the main process of the method for determining the feature value interval according to an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of the main process of solving the first objective function for the unary feature k according to an embodiment of the present invention;

[0052] Figure 4 is a schematic diagram of the main process of solving the second objective function for the multi - feature (r, s, t) according to an embodiment of the present invention;

[0053] Figure 5 is a schematic diagram of the main process of calculating the value intervals of each feature in one iteration direction according to an embodiment of the present invention;

[0054] Figure 6 is a schematic diagram of the main modules of the device for determining the feature value interval according to an embodiment of the present invention;

[0055] Figure 7 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;

[0056] Figure 8 is a schematic diagram of the structure of a computer device of an electronic device suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The exemplary embodiments of the present invention will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the descriptions of well-known functions and structures are omitted below.

[0058] Figure 1 is a schematic diagram of the main steps of the method for determining the characteristic value interval according to the embodiments of the present invention. As Figure 1 shown, the method for determining the characteristic value interval according to the embodiments of the present invention mainly includes the following steps:

[0059] Step S101: Discretize the characteristic values of the characteristics included in the sample data to obtain the discrete values of the characteristics, and calculate the index values of the characteristics at different discrete values according to the set evaluation index and the class labels of the sample data. Among them, the sample data is the data generated by the business service, including at least one characteristic and the corresponding characteristic value. The business service refers to various services provided by each institution, such as training services, insurance services, loan services, etc., and the characteristics of different business services are different.

[0060] Discretization means mapping the finite individuals in the infinite space to the finite space to improve the space-time efficiency of the algorithm. It can map it to positive integers on the premise of maintaining the size relationship of the original sequence. For example, if the characteristic is the number of participants in training and the characteristic values are [50, 100, 150, 200, 300], the above characteristic values can be mapped to [0, 1, 2, 3, 4] through discretization. In the embodiment, the discretization method can adopt equal-width discretization, equal-frequency discretization, clustering discretization, etc. The discretization method is a prior art and will not be elaborated here.

[0061] After discretizing the characteristic values of the characteristics, it is necessary to calculate the index values of the set evaluation index of the characteristics at different discrete values. Among them, the selection of the evaluation index is related to the specific business service. Taking the business service as a training service as an example, the evaluation index can be the net excellent rate, or the net excellent rate and the support rate. Among them, the net excellent rate represents the probability that the characteristic reaches the set target condition at each discrete value; the support rate represents the proportion of the sample quantity of the characteristic at each discrete value in the total sample quantity.

[0062] In order to calculate the above index values, it is necessary to know the class labels of the sample data. In one embodiment, the sample data of the sample data set can be classified according to the target condition to obtain the class labels corresponding to the sample data. After determining the class labels and discrete values, the index values of the characteristics at different discrete values can be calculated in combination with the class labels and the calculation expressions of the evaluation index.

[0063] Step S102: Based on the discrete value and the index value, establish a constraint condition and an objective function for maximizing the quality of service of the business service. Among them, the constraint condition is obtained by integrating the index value of the evaluation index with the discrete value within the set interval as the upper and lower limits of integration. This integral form of the constraint condition reflects the probability of the discrete value of the feature reaching the target condition when it is less than or equal to the interval range. When its slope is positive, it indicates that the probability of reaching the target condition in this interval range is relatively large.

[0064] The objective function needs to maximize the quality of service of the business service. The evaluation basis for the quality of service of different business services is usually different. Taking the business service as the training service as an example, the evaluation basis for the quality of service can be training benefits, exam scores, etc.

[0065] Step S103: Through the dynamic programming algorithm, solve the objective function under the constraint condition to obtain the value range of the feature. Dynamic programming applies the principle of optimization to solve the optimal solution of the problem under the target condition. In this embodiment, by adopting the dynamic programming algorithm, the optimal solution of the objective function under the above constraint condition is obtained, and this optimal solution is used as the value range of the feature, so as to obtain the value range where the feature performs best.

[0066] Figure 2 It is a schematic diagram of the main process for determining the value range of features according to the embodiment of the present invention. As Figure 2 shown, the method for determining the value range of features according to the embodiment of the present invention mainly includes the following steps:

[0067] Step S201: Classify the sample data in the sample data set according to the set target condition to obtain the class labels corresponding to the sample data. Among them, the sample data includes multiple features of the business service and the corresponding feature values, and the multiple features constitute a feature set. The target condition is the basis for classifying the sample data, which can be custom-set according to business requirements. The sample data that meets the target condition can be labeled as excellent samples, and the sample data that does not meet the target condition can be labeled as non-excellent samples. Optionally, numbers, letters, etc. can be used as class labels. For example, the number 1 is used to represent excellent samples, and the number 0 is used to represent non-excellent samples.

[0068] Taking the training service as an example, the sample data is the training data of multiple classes, and the feature set can include N features such as the location of the training, the level of the training, the number of training days, the number of participants, the number of trainees who completed the training, the number of employees who started work, the completion rate, the employment rate, the number of courses, and the class hours (N is a positive integer). The selection of features can be determined according to business requirements and is not limited here. After the classification process of this step, the labeled data shown in Table 1 can be obtained.

[0069] Table 1

[0070]

[0071] Step S202: Discretize the feature values of each feature in the sample data to obtain the discrete values corresponding to each feature. Discretize the feature values of each feature in the sample data separately to obtain the discrete values corresponding to each feature. Still taking the training service as an example, assuming there are 3 classes, feature 1 is the number of trainees who completed the training, and the feature values are {30, 20, 25}, then the discrete values can be {3, 1, 2}.

[0072] Step S203: Conduct a correlation analysis on each feature to determine whether there is a logical relationship between the features. If there is no logical relationship between the features, then execute Step S204; if there is a logical relationship between the features, then execute Step S205. Correlation analysis refers to analyzing two or more variable elements with correlation to measure the degree of correlation between the two variable elements.

[0073] This step is used to calculate the correlation between features. If the features are not correlated, it means the features are independent, then there is no logical relationship between the features; if the features are correlated, it means the features are not independent, then there is a logical relationship between the features. The calculation method of correlation analysis is a prior art and will not be elaborated here. Optionally, it is also possible to judge whether there is a logical relationship between features based on business experience. For example, based on business experience, it is known that there is a proportional relationship between the feature of the number of courses and the feature of class duration.

[0074] Step S204: Add the features without logical relationship as single features to the single feature set, and then calculate the index values of each single feature in the single feature set at different discrete values according to the set evaluation indicators and class labels, and execute Step S206. In the embodiment, the evaluation indicators include the net excellent rate and the support rate. Among them, the net excellent rate can be expressed by the following formula:

[0075] netRate = label_1 - label_0 Formula 1

[0076] In the above formula, netRate represents the net excellent rate, label_1 represents the excellent rate of the feature at the current discrete value, and label_0 represents the non-excellent rate of the feature at the current discrete value. The excellent rate and the non-excellent rate can be expressed by the following formulas:

[0077]

[0078] In the formula, y1 represents the number of excellent samples of the feature at the current discrete value, Total1 represents the total number of excellent samples, y0 represents the number of non-excellent samples of the feature at the current discrete value, and Total0 represents the total number of non-excellent samples.

[0079] The support rate can be expressed by the following formula:

[0080]

[0081] In the formula, supportRate represents the support rate, Total i represents the number of samples whose discrete value of the feature is the current discrete value, and Total represents the total number of samples (i.e., the total number of sample data in the sample dataset).

[0082] For a unary feature, the net excellent rate represents the value that the ratio of excellent samples (i.e., the excellent rate) to non-excellent samples (i.e., the non-excellent rate) of the unary feature at the current discrete value is higher. This evaluation index reflects the probability of the unary feature achieving excellence at each discrete value. If this evaluation index is greater than 0, it indicates that the probability of obtaining excellence at this discrete value is relatively large. The support rate represents the proportion of the number of samples of the unary feature at the current discrete value to the total number of samples.

[0083] According to Formula 1 - Formula 3, calculate the excellent rate, non-excellent rate, net excellent rate, and support rate of the unary feature k (1 ≤ k ≤ N) at each discrete value. The calculation results can be recorded in Table 2.

[0084] Table 2

[0085] Feature k Excellent rate Non-excellent rate Net excellent rate Support rate Discrete value 1 Discrete value 2 ……

[0086] Step S205: Add multiple features with logical relationships as multi - feature to the multi - feature set, determine the logical relationships between the features included in the multi - feature, calculate the index values of each multi - feature in the multi - feature set at different discrete value groups according to the set evaluation index and class label, and execute Step S208. For multiple features with logical relationships (i.e., multi - features), regression analysis or other analysis methods can be used to determine the logical relationships between these multiple features. The specific analysis method is prior art and will not be elaborated here.

[0087] For a multi - feature, the net excellent rate represents the probability that the multi - feature meets the target conditions at the current discrete value group, and the support rate represents the proportion of the number of samples of the multi - feature at the current discrete value group to the total number of samples. Among them, the discrete value group is obtained by combining the discrete values of each feature in the multi - feature.

[0088] Assume that a multi - feature is composed of feature r, feature s, and feature t. Then, this multi - feature can be expressed as (r, s, t). According to Formula 1 - Formula 3, calculate the excellent rate, non - excellent rate, net excellent rate, and support rate of the multi - feature (r, s, t) under each discrete value group, and the calculation results can be recorded in Table 3. Among them, the data in the second row of Table 3 represents the excellent rate, non - excellent rate, net excellent rate, and support rate calculated when the multi - feature (r, s, t) takes its respective discrete value 1 for feature r, feature s, and feature t (these three discrete values 1 form a discrete value group).

[0089] Table 3

[0090] Feature r Feature s Feature t Excellent rate Non-excellent rate Net excellent rate Support rate Discrete value 1 Discrete value 1 Discrete value 1 Discrete value 2 Discrete value 2 Discrete value 2 …… …… ……

[0091] Step S206: Based on the discrete value and the index value, establish the first constraint condition and the first objective function for each single - feature in the single - feature set. For a single - feature, its first constraint condition includes the cumulative excellent rate and the cumulative support rate. The cumulative excellent rate is obtained by integrating the index value of the net excellent rate of this single - feature, and the cumulative support rate is obtained by integrating the index value of the support rate of this single - feature. Specifically, it can be expressed by Formula 4:

[0092]

[0093] In the formula, cumsumNetRate(x i ,x j ) represents the cumulative excellent rate when the discrete value of the single - feature k is in the interval [x i ,x j ; netRate k represents the net excellent rate of the single - feature k; cumsumSuporRate(x i ,x j ) represents the cumulative support rate when the discrete value of the single - feature k is in the interval [x i ,x j ; suportRate k represents the support rate of the single - feature k.

[0094] The first objective function established for the single - feature needs to maximize the quality of the business service. For the training service, it is necessary to maximize the training benefit. In the embodiment, the first objective function is a function between the discrete value of each single - feature and the cumulative excellent rate and the cumulative support rate, and it can be specifically expressed by Formula 5:

[0095] f(x i ,x j )=cumsumNetRate(x i ,x j)×cumsumSuportRate(x i ,x j ) Formula 5

[0096] Wherein, f(x i ,x j ) represents a function with the discrete values x i and x j as variables.

[0097] Step S207: Using the dynamic programming algorithm, solve the first objective function under the first constraint condition to obtain the value range of each unary feature. This step is used to calculate the solution that makes the first objective function take the maximum value (i.e., max f(x i ,x j )) The specific implementation can be seen in the description of Figure 3 .

[0098] Step S208: According to the discrete values and index values, establish the second constraint condition and the second objective function for each multi - feature in the multi - feature set. For multi - features, its second constraint condition includes the cumulative excellent rate and the cumulative support rate. The cumulative excellent rate is obtained by integrating the index value of the net excellent rate of the multi - feature, and the cumulative support rate is obtained by integrating the index value of the support rate of the multi - feature.

[0099] The second objective function established for the multi - feature also needs to be able to maximize the service quality of the business service. In the embodiment, the second objective function is a function between the discrete value group of each multi - feature, its logical relationship, the cumulative excellent rate, and the cumulative support rate, which can be expressed by Formula 6:

[0100] f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) =

[0101] h(p r ,p s ,p t ,…)×cumsumNetRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)×

[0102] cumsumSuportRate(x rm ,x rn ,x sa ,x sb ,xte , x tf , …) Formula 6

[0103] In the formula, f(x rm , x rn , x sa , x sb , x te , x tf , …) represents a function with the discrete value group (x rm , x rn , x sa , x sb , x te , x tf , …) of multiple features (r, s, t, …) as variables; h(p r , p s , p t , …) represents the logical relationship between the features of multiple features (r, s, t, …); cumsumNetRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative excellent rate when the discrete values of the features in multiple features (r, s, t, …) are respectively in the intervals ([x m , x n , [x a , x b , [x e , x f , …); cumsumSuportRate(x rm , x rn , x sa , x sb , x te , x tf , …) represents the cumulative support rate when the discrete values of the features in multiple features (r, s, t, …) are respectively in the intervals ([x m , x n , [x a , x b , [x e , x f , …).

[0104] Step S209: Using the multiple dynamic programming algorithm, solve the second objective function under the second constraint condition to obtain the value range of each multiple feature. This step is used to calculate the maximum value of the second objective function (i.e., max f(x rm , x rn , x sa , x sb , xte , x tf , …)) The solution, for the specific implementation process, see the description regarding Figure 4 .

[0105] Figure 3 is the main process schematic diagram for solving the first objective function for the unary feature k according to an embodiment of the present invention. As Figure 3 shown, the process of solving the first objective function for the unary feature k in the embodiment of the present invention (i.e., step S207) mainly includes the following steps:

[0106] Step S301: Determine the first initial value, the first step size, the first target interval, and the second target interval for the iteration of the unary feature k. In the embodiment, the first initial value can be represented by k0, and the value can be the discrete value that maximizes the objective function of the unary feature k. The first step size is represented by Δk, and the value can be the ratio of the range of each discrete value of the unary feature k to the number of samples.

[0107] The first target interval is the target interval for the left boundary iteration, which can be represented by [k0 - n×Δk, k0], where n is the number of iterations for the left boundary iteration. The second target interval is the target interval for the right boundary iteration, which can be represented by [k0 + m×Δk, k0], where m is the number of iterations for the right boundary iteration.

[0108] Step S302: Gradually decrease the first initial value to the left according to the first step size, calculate the values of the objective function of the unary feature k in the first target interval, and during the iteration process of this step, when it is determined that the set first stop condition is satisfied, end the iteration to obtain the left boundary value. This step is used for the left boundary iteration, and the objective function of the unary feature k is the first objective function.

[0109] In this step, a threshold β1 needs to be preset in advance. The first stop condition can be set as a decreasing trend for β1 consecutive step sizes. Then, when the objective function is in the iteration process, if β1 consecutive step sizes are in a decreasing trend, end the left boundary iteration. At this time, the left boundary value is: k0 - (n + β1)×Δk.

[0110] Step S303: Gradually increase the first initial value to the right according to the first step size, calculate the values of the objective function of the unary feature k in the second target interval, and during the iteration process of this step, when it is determined that the set second stop condition is satisfied, end the iteration to obtain the right boundary value. This step is used for the right boundary iteration, and the objective function of the unary feature k is the first objective function.

[0111] In this step, a threshold β2 needs to be preset in advance. The second stop condition can be set as a decreasing trend for β2 consecutive step sizes. Then, when the objective function is in the iteration process, if β2 consecutive step sizes are in a decreasing trend, end the right boundary iteration. At this time, the right boundary value is: k0 + (m - β2)×Δk.

[0112] Step S304: Take the left boundary value as the minimum value of the unary feature k, and the right boundary value as the maximum value of the unary feature k, to obtain the value range of the unary feature k. After completing the left and right boundary iterative calculation process, the value range of the unary feature k can be determined as: [k0 - (n + β1) × Δk, k0 + (m - β2) × Δk]. In the above process, the threshold β1 and the threshold β2 can be taken according to the actual situation, and the two can be the same or different.

[0113] Figure 4 is a schematic diagram of the main process for solving the second objective function for the multi - feature (r, s, t) according to an embodiment of the present invention. As Figure 4 shown, the process of solving the second objective function for the multi - feature (r, s, t) in the embodiment of the present invention (i.e., step S209) mainly includes the following steps:

[0114] Step S401: Determine the second initial values, iteration directions, second step sizes, third target intervals, and fourth target intervals for each feature in the multi - feature (r, s, t) to be iterated. In the embodiment, the second initial values of the features r, s, and t in the multi - feature (r, s, t) are r0, s0, and t0 respectively, and the specific values can be the discrete value groups that make the objective function of the multi - feature (r, s, t) maximum.

[0115] Each feature of the multi - feature (r, s, t) has 2 iteration directions, left (decrease) and right (increase). There are 4 direction combinations for the three features r, s, and t, specifically: r, s, and t all go left; r and s go left, t goes right; s and t go left, r goes right; r and t go left, s goes right. In addition, the number of common direction combinations α for n features is:

[0116]

[0117] The second step sizes of the features r, s, and t in the multi - feature (r, s, t) are Δr, Δs, and Δt respectively. Δr, Δs, and Δt can be set as the ratio of the range of each discrete value group of the multi - feature (r, s, t) to the number of samples.

[0118] The third target intervals of the features r, s, and t in the multi - feature (r, s, t) are the target intervals for left - boundary iteration, which are [r0 - n r ×Δr, r0], [s0 - n s ×Δs, s0], and [t0 - n t ×Δt, t0] respectively. Wherein, n r 、n s and n tThe number of iterations for left boundary iteration for feature r, feature s, and feature t respectively.

[0119] The fourth target interval for feature r, feature s, and feature t in the multi - feature (r, s, t) is the target interval for right boundary iteration, which are respectively ([r0, r0 + m r ×Δr], [s0, s0 + m s ×Δs], and [t0, t0 + m t ×Δt]). Among them, m r 、m s and m t are the number of iterations for right boundary iteration for feature r, feature s, and feature t respectively.

[0120] Step S402: Calculate the value intervals of each feature in different iteration directions. The calculation methods of the value intervals of each feature in different iteration directions are the same. Figure 5 One of the iteration directions (i.e., feature r, feature s, and feature t all move left) is described.

[0121] Step S403: Substitute the value intervals of each feature in different iteration directions into the objective function respectively, and take the value interval that makes the objective function reach the maximum value as the value interval of the multi - feature (r, s, t). Substitute the 4 groups of value intervals of feature r, feature s, and feature t into the objective function of the multi - feature (r, s, t) (i.e., the second objective function), and take the group of value intervals that makes the objective function the largest as the optimal value interval of the multi - feature (r, s, t).

[0122] Figure 5 is the main flow diagram for calculating the value intervals of each feature in one iteration direction according to the embodiment of the present invention. As Figure 5 shown, the process of calculating the value intervals of each feature in one iteration direction according to the embodiment of the present invention mainly includes the following steps:

[0123] Step S501: According to the second step size of each feature, gradually reduce the second initial value of each feature to the left, calculate the value of the objective function of the multi - feature (r, s, t) in the third target interval of each feature, and during the iteration process of this step, when the set third stop condition is met, end the iteration to obtain the left boundary value of each feature.

[0124] This step is used for left boundary iteration. Specifically, according to the second step size Δr of feature r, gradually reduce the second initial value r0 of feature r; according to the second step size Δs of feature s, gradually reduce the second initial value s0 of feature s; according to the second step size Δt of feature t, gradually reduce the second initial value t0 of feature t; then calculate the value of the objective function in the third target interval (i.e., the third target interval of feature r [r0 - n r×Δr, r0], the third target interval of feature s is [s0 - n s ×Δs, s0], the third target interval of feature t is [t0 - n t ×Δt, t0]).

[0125] In this step, thresholds β r1 , β s1 and β t1 need to be set in advance for feature r, feature s, and feature t. The third stopping condition can be set such that feature r, feature s, and feature t are each continuously decreasing by β r1 , β s1 and β t1 step lengths. Then, during the iteration of the objective function, if feature r, feature s, and feature t are each continuously decreasing by β r1 , β s1 and β t1 step lengths, the left - boundary iteration ends. At this time, the left - boundary values are: for feature r: r0 - (n r - β r1 )×Δr, for feature s: s0 - (n s - β s1 )×Δs, for feature t: t0 - (n t - β t1 )×Δt.

[0126] Step S502: According to the second step - lengths of each feature, gradually increase the second initial values of each feature to the right, calculate the values of the objective function of the multi - feature (r, s, t) in the fourth target intervals of each feature, and during the iteration of this step, end the iteration when the set fourth stopping condition is met to obtain the right - boundary values of each feature.

[0127] This step is used for right - boundary iteration. Specifically, according to the second step - length Δr of feature r, gradually increase the second initial value r0 of feature r; according to the second step - length Δs of feature s, gradually increase the second initial value s0 of feature s; according to the second step - length Δt of feature t, gradually increase the second initial value t0 of feature t; then calculate the values of the objective function in the fourth target interval (i.e., the fourth target interval of feature r is [r0, r0 + m r ×Δr], the fourth target interval of feature s is [s0, s0 + m s ×Δs], the fourth target interval of feature t is [t0, t0 + m t ×Δt]).

[0128] In this step, thresholds β r2 , β s2 and β t2 need to be set in advance for feature r, feature s, and feature t. The third stopping condition can be set such that feature r, feature s, and feature t are each continuously decreasing by βr2 , β s2 and β t2 step lengths show a decreasing trend. Then, during the iteration of the objective function, if features r, s, and t are each continuously β r2 , β s2 and β t2 step lengths show a decreasing trend, end the right boundary iteration. At this time, the right boundary values are: Feature r: r0 + (m r -β r2 ) × Δr, Feature s: s0 + (m s -β s2 ) × Δs, Feature t: t0 + (m t -β t2 ) × Δt.

[0129] Step S503: Take the left boundary values of each feature as the minimum value of the corresponding feature, and the right boundary values of each feature as the maximum value of the corresponding feature to obtain the value range of each feature. After completing the left and right boundary iteration calculation process, the value ranges of features r, s, and t can be determined as follows:

[0130] Feature r: [r0 - (n r -β r1 ) × Δr, r0 + (m r -β r2 ) × Δr],

[0131] Feature s: [s0 - (n s -β s1 ) × Δs, s0 + (m s -β s2 ) × Δs],

[0132] Feature t: [t0 - (n t -β t1 ) × Δt, t0 + (m t -β t2 ) × Δt].

[0133] The method for determining the value range of features in the embodiments of the present invention first discretizes the features, calculates the net excellent rate and support rate at each discrete point, then continuousizes the features, and obtains the value range with the highest comprehensive score (i.e., the product of the cumulative excellent rate and the cumulative support rate) of the feature. And for features with logical relationships, logical relationships are added during the process of continuousizing the features to obtain the optimal value range under multiple features. The above method can obtain the optimal value range of all features under the target conditions.

[0134] In addition, during the calculation of the value range in this embodiment, the number of sample data with good performance in this value range is fully considered to provide data support for this value range, so the credibility is strong; moreover, for features with logical relationships, logical relationships are added during the process of feature continuousization, and the scalability is strong.

[0135] The method for determining the characteristic value range of the embodiment of the present invention can be applied to scenarios such as pre-job training for insurance, recruit training camps, gas stations, and take-off classes. The method for determining the characteristic value range of the embodiment of the present invention will be further described below in combination with the scenario of pre-job training.

[0136] This embodiment is used to mine the optimal value range of each feature from sample data, and combining these value ranges can maximize the training benefits of the training class. In this example, the sample data is the training data of multiple classes. The training characteristics include: holding location, holding level, number of training days, number of participants, number of trainees completing the training, number of employees taking up their posts, completion rate, employment rate, number of courses, and class duration, a total of 10 operation characteristics.

[0137] The target condition is: the ranking order according to the 3500P rate is greater than or equal to the set threshold. Among them, the 3500P rate refers to the ratio of the cumulative number of agents with a standard premium of more than 3500 yuan in the first 6 months to the number of employees taking up their posts in the class. The above threshold can be defined according to business requirements, for example, set to 30%. Based on the above target condition, classes with a 3500P rate ranking in the top 30% will be marked as excellent samples, and classes with a 3500P rate ranking in the bottom 70% will be marked as non-excellent samples.

[0138] Thus, 30,243 sample data with category labels as shown in Table 1 can be obtained. Then, the feature values of these 10 features are discretized to obtain the corresponding discrete values of each feature. And correlation analysis and regression analysis are performed on these 10 features, and it is concluded that the two features of holding location and holding level have no correlation with other features, which conforms to the calculation rules of the value range of unary features. There is a set of logical relationships among the number of participants, the number of trainees completing the training, the number of employees taking up their posts, the completion rate, and the employment rate; there is another set of logical relationships among the number of courses, the class duration, and the number of training days, all of which conform to the calculation rules of the value range of multi-feature features.

[0139] After determining the unary feature and multi-feature features, the excellent rate, non-excellent rate, net excellent rate, and support rate of the above 10 features are calculated using Formula 1-3 to obtain Table 2 and Table 3. Then, for the two features that conform to the calculation rules of the value range of unary features: holding location and holding level, the optimal value range is obtained by solving the dynamic programming problem in the manner of Step S206 - Step S207: holding location: hotel, holding level: branch company.

[0140] In addition, through business experience and correlation analysis, the logical relationships among the number of trainees, the number of trainees completing the training, the number of trainees taking up positions, the completion rate, and the employment rate are as follows: Completion rate = number of trainees / number of trainees completing the training, employment rate = number of trainees taking up positions / number of trainees completing the training, and the number of trainees ≥ the number of trainees completing the training ≥ the number of trainees taking up positions. Therefore, by adopting the method of steps S208 - S209, the optimal value ranges are obtained by solving a multi - variable dynamic programming problem: number of trainees: [11, 14], number of trainees completing the training: [7, 14], number of trainees taking up positions: [7, 14], completion rate: [0.55, 1], employment rate: 1.

[0141] Through business experience and correlation analysis, the logical relationships among the number of courses, the class duration, and the number of days of the training are as follows: Class duration = [0.6, 2] × number of courses, and class duration ≤ number of days of the training × 10. Therefore, by adopting the method of steps S208 - S209, the optimal value ranges are obtained by solving a multi - variable dynamic programming problem: number of courses: [16, 20], class duration: [30h, 40h], number of days of the training: 4 days.

[0142] It can be seen from this that in order to achieve better training benefits for pre - job training and enable salespersons to generally achieve good performance results in the first six months, the characteristic values of the training class should be: held in units of branch companies, with a 4 - day training in a hotel, teaching 16 - 20 courses, with a cumulative class duration of 30 - 40 hours, and the scale of the training class should be controlled in a small - class mode of 11 - 14 people, and the completion rate should be controlled within the range of 0.55 - 1.

[0143] Based on the performance of salespersons who have completed the training in the above - mentioned embodiments, it feeds back to the operation process of the training class. Through data analysis and model quantification, the optimal value range of the characteristics is determined to achieve intelligent class opening, provide high - quality learning and training services for salespersons, and help salespersons obtain the best training effect.

[0144] Figure 6 It is a schematic diagram of the main modules of the device for determining the value range of characteristics according to an embodiment of the present invention. As Figure 6 shown, the device 600 for determining the value range of characteristics according to an embodiment of the present invention mainly includes:

[0145] A feature discretization module 601, configured to discretize the feature values of the features included in the sample data to obtain the discrete values of the features, and calculate the index values of the features at different discrete values according to the set evaluation index and the class label of the sample data. The sample data is data generated by business services, including at least one feature and the corresponding feature value.

[0146] After discretizing the eigenvalues of the features, it is necessary to calculate the index values of the set evaluation indices of the features at different discrete values. To calculate the above index values, it is necessary to know the class labels of the sample data. In one embodiment, the sample data of the sample data set can be classified according to the target conditions to obtain the class labels corresponding to the sample data. After determining the class labels and discrete values, the index values of the features at different discrete values can be calculated by combining the class labels and the calculation expressions of the evaluation indices.

[0147] The feature continuous module 602 is used to establish a constraint condition and an objective function that maximizes the service quality of the service according to the discrete value and the index value. Among them, the constraint condition is obtained by integrating the index value of the evaluation index with the discrete value within the set interval as the upper and lower limits of the integral. This integral form of the constraint condition reflects the probability that the discrete value of the feature reaches the target condition when it is less than or equal to the interval range. When its slope is positive, it means that the probability of reaching the target condition in this interval range is relatively large.

[0148] The objective function needs to maximize the service quality of the service. The evaluation basis of the service quality of different services is usually different. Taking the service as a training service as an example, the evaluation basis of the service quality can be training benefits, test scores, etc.

[0149] The interval determination module 603 is used to solve the objective function under the constraint condition through a dynamic programming algorithm to obtain the value interval of the feature. Dynamic programming applies the principle of optimization to solve the optimal solution of the problem under the target conditions. In this embodiment, by using the dynamic programming algorithm, the optimal solution of the objective function that satisfies the above constraint conditions is obtained, and this optimal solution is used as the value interval of the feature, and the value interval with the best performance of the feature can be obtained.

[0150] In addition, the device 600 for determining the value interval of the feature in the embodiment of the present invention may further include: a feature division module and a sample marking module ( Figure 6 not shown in the figure). Among them, the feature division module is used to perform a correlation analysis on multiple features to determine whether there is a logical relationship between the multiple features; if there is a logical relationship between the multiple features, the multiple features are determined to be multivariate features, and the logical relationship between the multiple features is determined; if there is no logical relationship between the multiple features, the feature is determined to be a univariate feature. The sample marking module is used to classify the sample data of the sample data set according to the target conditions to obtain the class labels corresponding to the sample data.

[0151] As can be seen from the above description, by discretizing features, calculating the index values of features at each discrete value, and then establishing constraint conditions and objective functions based on the discrete values and index values to continuousize features and calculate the optimal solution of the objective function, the value range with the best feature performance can be obtained, valuable data can be mined from sample data, and the service quality can be maximized to improve the service effect of business services.

[0152] Figure 7 Fig. 700 shows an exemplary system architecture to which the method for determining the value range of features or the device for determining the value range of features according to the embodiments of the present invention can be applied.

[0153] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0154] Users can use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. The terminal devices 701, 702, 703 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0155] The server 705 may be a server providing various services, such as a background management server that processes execution requests sent by an administrator using the terminal devices 701, 702, 703. After receiving the execution request, the background management server can discretize the features of the sample data, then calculate the index values at each discrete value, and then continuousize the features and calculate the optimal solution (i.e., the value range) of the objective function, and feedback the calculation results (such as the value ranges of each feature) to the terminal devices.

[0156] It should be noted that the method for determining the value range of features provided by the embodiments of the present application is generally executed by the server 705. Correspondingly, the device for determining the value range of features is generally set in the server 705.

[0157] It should be understood that Figure 7 the numbers of the terminal devices, the network, and the server in

[0158] are merely illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers.

[0159] The electronic device of the present invention includes: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, enable the one or more processors to implement a method for determining a characteristic value range according to an embodiment of the present invention.

[0160] The computer-readable medium of the present invention stores a computer program thereon, and when the program is executed by a processor, it implements a method for determining a characteristic value range according to an embodiment of the present invention.

[0161] Reference is made below Figure 8 , which shows a schematic structural diagram of a computer system 800 suitable for implementing the electronic device according to an embodiment of the present invention. Figure 8 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0162] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0163] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as required, so that a computer program read from it can be installed into the storage section 808 as required.

[0164] In particular, according to the embodiments disclosed in the present invention, the process described in the above main step diagram can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the method shown in the main step diagram. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above functions defined in the system of the present invention are performed.

[0165] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0167] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a feature discretization module, a feature continuity module, and an interval determination module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the feature discretization module can also be described as "a module that discretizes the feature values of the features included in the sample data to obtain the discrete values of the features, and calculates the index values of the features at different discrete values according to the set evaluation index and the class label of the sample data".

[0168] As another aspect, the present invention also provides a computer-readable medium, which may be included in the devices described in the above embodiments; or may exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the device includes: discretizing the feature values of the features included in the sample data to obtain the discrete values of the features, and calculating the index values of the features at different discrete values according to the set evaluation index and the class label of the sample data; wherein the sample data is data generated by a business service; establishing a constraint condition and an objective function for maximizing the service quality of the business service according to the discrete values and the index values; wherein the constraint condition is obtained by integrating the index values of the evaluation index with the discrete values within a set interval as the upper and lower limits of the integral; and solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the feature.

[0169] According to the technical solution of the embodiment of the present invention, by discretizing features and calculating the index values of features at each discrete value, and then based on the discrete values and index values, establishing constraint conditions and objective functions to continuousize features and calculate the optimal solution of the objective function, the value range with the best feature performance can be obtained, valuable data can be mined from sample data, and the service quality can be maximized to improve the service effect of business services.

[0170] The above product can execute the method provided by the embodiment of the present invention and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.

[0171] The above specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for determining a characteristic value range, characterized in that Including: Discretize the feature values of the features included in the sample data to obtain the discrete values of the features, and calculate the index values of the features at different discrete values according to the set evaluation index and the class labels of the sample data; wherein, the sample data is data generated by business services. Establish a constraint condition and an objective function for maximizing the service quality of the business service according to the discrete values and the index values; wherein, the constraint condition is obtained by integrating the index values of the evaluation index with the discrete values within the set interval as the upper and lower limits of integration. Solve the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the feature. The sample data includes multiple such features and corresponding feature values; before the step of calculating the index values of the features at different discrete values, the method further includes: performing a correlation analysis on the multiple features to determine whether there is a logical relationship between the multiple features; if there is a logical relationship between the multiple features, determine the multiple features as multivariate features and determine the logical relationship between the multiple features; if there is no logical relationship between the multiple features, determine the feature as a univariate feature; establish an objective function for the corresponding univariate feature or multivariate feature.

2. The method according to claim 1, characterized in that, The calculating the index values of the features at different discrete values includes: Calculating the index values of the univariate feature at different discrete values and the index values of the multivariate feature at different discrete value groups; wherein, the discrete value groups are obtained by combining the discrete values of the features in the multivariate feature.

3. The method according to claim 2, wherein The evaluation index includes the net excellent rate and the support rate. The net excellent rate represents the probability that the univariate feature reaches the set target condition at the current discrete value, or the multivariate feature reaches the set target condition at the current discrete value group; the support rate represents the proportion of the sample quantity of the univariate feature at the current discrete value, or the multivariate feature at the current discrete value group, in the total sample quantity. The constraint condition includes the cumulative excellent rate and the cumulative support rate. The cumulative excellent rate is obtained by integrating the index values of the net excellent rate, and the cumulative support rate is obtained by integrating the index values of the support rate. The method further includes: classifying the sample data in the sample data set according to the target condition to obtain the class labels corresponding to the sample data.

4. The method according to claim 3, characterized in that, The objective function established for the univariate feature is: f(x i ,x j ) = cumsumNetRate(x i ,x j ) × cumsumSuportRate(x i ,x j ) where f(x i , x j ) represents a function with the discrete values x i and x j of the unary feature as variables; cumsumNetRate(x i , x j ) represents the cumulative excellent rate when the discrete values of the unary feature are in the interval [x i, x j ; cumsumSuporRate(x i , x j ) represents the cumulative support rate when the discrete values of the unary feature are in the interval [x i, x j .

5. The method according to claim 4, characterized in that Solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the value range of the feature includes: Determine the first initial value, the first step size, the first target interval and the second target interval for the iteration of the univariate feature. Left boundary iteration: Gradually reduce the first initial value to the left according to the first step size, calculate the value of the objective function in the first target interval, and during the iteration, end the left boundary iteration and obtain the left boundary value when the set first stop condition is met. Right boundary iteration: According to the first step size, gradually increase the first initial value to the right, calculate the value of the objective function in the second target interval, and during the iteration process, when the set second stop condition is satisfied, end the right boundary iteration to obtain the right boundary value; Take the left boundary value as the minimum value of the unary feature and the right boundary value as the maximum value of the unary feature to obtain the value range of the unary feature.

6. The method according to claim 3, characterized in that The objective function established for the multi-variable feature is: f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) = h(p r ,p s ,p t ,…)×cumsumNetRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…)× cumsumSuportRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) Where f(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) represents a function with discrete value groups (x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) of multiple features (r, s, t, …) as variables; h(p r ,p s ,p t ,…) represents the logical relationship corresponding to the multiple features (r, s, t, …); cumsumNetRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) represents the cumulative excellent rate when the discrete values of each feature in the multiple features (r, s, t, …) are respectively in the intervals ([x m, x n , [x a, x b , [x e, x f , …); cumsumSuportRate(x rm ,x rn ,x sa ,x sb ,x te ,x tf ,…) represents the cumulative support rate when the discrete values of each feature in the multiple features (r, s, t, …) are respectively in the intervals ([x m, x n , [x a, x b , [x e, x f , …).

7. The method according to claim 6, characterized in that, By using the dynamic programming algorithm, solve the objective function under the constraint conditions to obtain the value range of the feature, including: Determine the second initial values, iteration directions, second step sizes, third target intervals, and fourth target intervals for each feature in the multi-variable feature to perform iteration; Calculate the value ranges of each feature in different iteration directions according to the following steps, substitute the value ranges of each feature in different iteration directions into the objective function respectively, and take the value range that makes the objective function reach the maximum value as the value range of the multi-variable feature: Left boundary iteration: According to the second step size of each feature, gradually decrease the second initial value of each feature to the left, calculate the value of the objective function in the third target interval of each feature, and during the iteration process, when the set third stop condition is satisfied, end the left boundary iteration to obtain the left boundary value of each feature; Right boundary iteration: According to the second step size of each feature, gradually increase the second initial value of each feature to the right, calculate the value of the objective function in the fourth target interval of each feature, and during the iteration process, when the set fourth stop condition is satisfied, end the right boundary iteration to obtain the right boundary value of each feature; and Take the left boundary value of each feature as the minimum value of the corresponding feature and the right boundary value of each feature as the maximum value of the corresponding feature to obtain the value range of each feature.

8. An apparatus for determining a characteristic value range, characterized in that, Including: Feature discretization module, used to discretize the feature values of the features included in the sample data to obtain the discrete values of the features, and calculate the index values of the features at different discrete values according to the set evaluation index and the class label of the sample data; wherein, the sample data is the data generated by the business service; the sample data includes multiple such features and corresponding feature values; Feature continuous module, used to establish constraint conditions and an objective function that maximizes the service quality of the business service according to the discrete values and the index values; wherein, the constraint conditions are obtained by integrating the index values of the evaluation index with the discrete values located in the set interval as the upper and lower limits of the integral; Interval determination module, used to solve the objective function under the constraint conditions by using the dynamic programming algorithm to obtain the value range of the feature; A feature division module, configured to perform correlation analysis on a plurality of the features to determine whether there is a logical relationship between the plurality of the features; if there is a logical relationship between the plurality of the features, determine the plurality of the features as multivariate features and determine the logical relationship between the plurality of the features; if there is no logical relationship between the plurality of the features, determine the features as univariate features; and establish an objective function for the corresponding univariate features or multivariate features.

9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Path planning method through collaboration of multiple underwater robots

    CN109917817A

  • Non-linear planning model based production planning system, production planning method and computer-readable storage medium

    WO2021168783A1