Method and apparatus for determining a class of business data and evaluating business data

By selecting multi-dimensional features through feature engineering and significance indicators, and combining regression algorithms and machine learning models, the problem of unscientific and unfair classification and evaluation of business data in existing technologies has been solved, achieving more accurate and fair classification and evaluation of business data.

CN114219037BActive Publication Date: 2026-04-28SHENGDOUSHI SHANGHAI SCI & TECH DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENGDOUSHI SHANGHAI SCI & TECH DEV CO LTD
Filing Date
2021-12-17
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the methods for classifying and evaluating business data are monotonous and crude, failing to consider the multidimensional differences between individuals within a hierarchy, resulting in evaluations that are not scientific or fair.

Method used

By acquiring business data related to business scenarios, feature engineering and saliency indicators are used to select multi-dimensional features for initial and further classification to improve classification accuracy. Regression algorithms and machine learning models are used to determine salient features for more accurate business data classification and evaluation.

Benefits of technology

It enables more scientific and fair classification and evaluation of business data, taking into account the multi-dimensional differences in characteristics of individuals at different levels and within the same level, and provides more valuable classification and evaluation information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219037B_ABST
    Figure CN114219037B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for determining the category of business data and evaluating the business data based on the category. The method for determining the category of business data comprises obtaining business data associated with a business scenario, determining a first category of the business data based on the business scenario, determining a more accurate and scientific second category of the business data based on a significance indicator of the business scenario for the business data belonging to the first category, and obtaining a category determination result of the business data according to the first category and the second category, wherein the second category has higher accuracy than the first category. The method can select more accurate business evaluation references through a more reasonable business data classification method, so as to obtain more scientific and fair business statistical and evaluation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data classification and statistics, and more particularly to methods, apparatus and computer storage media for determining the categories of business data and evaluating business data based on those categories. Background Technology

[0002] Statistical analysis of store business data (such as sales performance) and presentation of it to users and management in the form of settlement reports is a crucial aspect of enterprise management. Business statistics and reports are typically used in platforms and systems for cross-referencing, where data belonging to the same group or category are compared with statistical indicators used as a reference within that group / category to obtain an individual business evaluation corresponding to each data point.

[0003] Typically, business data needs to be stratified or categorized into groups that include all individuals based on selected criteria. However, stratification and categorization rules are often monotonous and crude, failing to consider the multidimensional differences between individuals or units within a stratum, making the group induction and evaluation of the comparison targets neither scientific nor fair.

[0004] Therefore, there is a need to improve the scheme for classifying business data and evaluating business data based on the classification results. Summary of the Invention

[0005] The embodiments of this application propose an improved scheme for the classification and evaluation of business data, which is used to at least partially solve the defects and problems mentioned above. By using a more reasonable business data classification method, a more accurate business evaluation reference can be selected, thereby obtaining more scientific and fair business statistics and evaluation information.

[0006] According to one aspect of this application, a method for determining the category of business data is proposed, comprising: acquiring business data associated with a business scenario; determining a first category of the business data based on the business scenario; determining a second category of the business data for the business data belonging to the first category based on a saliency index of the business scenario, wherein the second category has higher accuracy than the first category; and acquiring the category determination result of the business data based on the first category and / or the second category.

[0007] According to another aspect of this application, a method for evaluating business data is proposed, comprising: determining a first category and a second category of business data according to the method for determining the category of business data as described above; and evaluating business data that belongs to the first category based on the second category.

[0008] According to another aspect of this application, an apparatus for determining the category of business data is proposed, comprising: an acquisition unit configured to acquire business data associated with a business scenario; a classification unit configured to determine a first category of the business data based on the business scenario, determine a second category of the business data for the business data belonging to the first category based on a saliency index of the business scenario, and acquire the category determination result of the business data according to the first category and / or the second category, wherein the second category has higher accuracy than the first category.

[0009] According to another aspect of this application, an apparatus for evaluating business data is provided, comprising: an apparatus for determining the category of business data as described above; and an evaluation unit configured to evaluate business data belonging to a first category based on a second category.

[0010] According to another aspect of this application, a computer-readable storage medium is provided that stores a computer program thereon, the computer program including executable instructions that, when executed by a processor, implement the method described above.

[0011] According to another aspect of this application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the method described above.

[0012] By adopting the business data classification and evaluation scheme proposed in this application, and integrating the practical significance of business scenarios, the algorithm's closed-loop iteration is embedded into the business operation evaluation process and updated in a timely manner. When evaluating the business data of individual objects, the scheme of this application can consider the multi-dimensional characteristic differences and distributions of individual objects at different levels and within the same level, scientifically and comprehensively classifying (e.g., clustering) and grouping business data. It aggregates business data from similar individual objects more quantitatively, and selects appropriate reference values ​​for different evaluation criteria and characteristics. This avoids the unfairness and inaccuracy caused by using fixed evaluation reference values ​​for comparison due to differences in the actual situation and attributes of the data. It can provide more business-valuable classification and evaluation information for subsequent analysis and calculation. Attached Figure Description

[0013] The above and other features and advantages of this application will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0014] Figure 1 This is a schematic logic block diagram illustrating a process for determining the category of business data and evaluating business data according to an embodiment of this application.

[0015] Figure 2This is a schematic flowchart illustrating a method for determining the category of business data and evaluating business data according to an embodiment of this application.

[0016] Figure 3 This is a schematic structural block diagram of a device for determining the category of business data according to an embodiment of this application.

[0017] Figure 4 This is a schematic structural block diagram of an apparatus for evaluating business data according to an embodiment of this application.

[0018] Figure 5 This is a schematic structural block diagram of an electronic device according to an embodiment of the present application. Detailed Implementation

[0019] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided to make the content of this application comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. In the drawings, the dimensions of some elements may be exaggerated or modified for clarity. The same reference numerals in the drawings denote the same or similar structures, and therefore their detailed descriptions will be omitted.

[0020] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details described, or other methods, elements, etc., can be employed. In other instances, well-known structures, methods, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0021] This document uses the classification of restaurant business data (e.g., sales data) and the evaluation of each restaurant's sales performance as an example to introduce embodiments of this application for determining the category of business data and the method and apparatus for evaluating business data based on that category. However, those skilled in the art should understand that this description is merely exemplary and not a limitation of the solutions described in this application. The solutions of the embodiments of this application can be applied not only to the catering industry but also to business scenarios in other industries (e.g., retail, transportation, manufacturing, etc.) that require the classification and / or evaluation of business data (e.g., sales data from retail stores, or business data from other individual objects or other object data).

[0022] Business data typically originates from the stores or units that generate it. In the following text, the source of business data will be referred to as an individual object. Therefore, the source of all acquired business data will be the entirety of business units comprised of these individual objects. Generally, the smallest unit of an individual object is a business unit, and these business units exist in a hierarchical relationship. In this paper, the classification and evaluation of business data is equivalent to the classification and evaluation of the individual objects (business units) that generate the business data. Therefore, the distinction is no longer made between classifying and evaluating business data itself and classifying and evaluating the individual objects or various combinations or sets of individual objects corresponding to the business data. When referring to business data belonging to a certain category, it also refers to the individual objects within the combinations or sets of individual objects belonging to that category.

[0023] When analyzing business data, it is necessary to determine the evaluation characteristics of the business data and the evaluation indicators that serve as reference values ​​for these characteristics. These indicators act as comparison objects for the evaluation characteristic values ​​of each individual entity (e.g., a business unit such as a restaurant). Based on the comparison results between the evaluation characteristic values ​​and evaluation indicators of each business unit, the operating performance of each business unit is obtained and scored and evaluated using preset evaluation / scoring standards. This provides reference information for improving the business unit's products, services, and management.

[0024] This type of statistical analysis of business data, such as settlement reports, is typically implemented using platforms and systems for horizontal comparison. For each business unit within the same group, level, or category, its performance-related statistical data is compared with corresponding reference data. Reference data can be the best-performing data, average data, or data calculated according to other standards within that group, level, or category. Reference data may differ from the business data of all business units; therefore, it is calculated reference data corresponding to a "virtual" business unit.

[0025] Business data typically possesses multiple features, each characterizing an attribute or characteristic of the data in a specific dimension. Furthermore, each dimension of the business data exhibits different feature values. The number and distribution of feature dimensions, as well as the feature values ​​for each dimension, can vary for each piece of business data. Therefore, using different feature dimensions to statistically analyze business data may yield different results, thus affecting the evaluation outcomes.

[0026] When performing statistical analysis on business data of business units by hierarchy, it is generally necessary to stratify these groups of business units according to specific features of interest based on the characteristic dimensions of the business data. The hierarchy is equivalent to the category of the business unit. Then, statistical calculations can be performed based on the business data of the business units within each collected and recorded hierarchy.

[0027] For example, when calculating sales figures for restaurants in the catering industry, a hierarchical standard linked to geographical location can be used to classify restaurants. The classification or categorization of restaurant groups involves categorizing restaurants at different geographical levels at the management level. For instance, in a nationwide catering company, regional managers need to obtain and examine sales data for the entire group of restaurants at the regional level (across provinces); provincial regional managers handle sales data at the provincial level; city regional managers handle sales data at the city level; district / county regional managers handle sales data at the district / county level (e.g., a district within a city); location regional managers handle sales data at the location level (e.g., a commercial area or several streets); and store managers handle sales data at the store level (one or several stores).

[0028] Generally, a restaurant can be the smallest business unit in a hierarchy or category. A collection of restaurants can also serve as the smallest business unit. For example, in an application scenario corresponding to a single store manager overseeing multiple stores in a catering company, the smallest business unit could be a collection of multiple stores. Different collections of restaurants are managed by managers at different levels. In this case, the collection of restaurants corresponds to a category at the geographical region level. For instance, the collection of restaurants for a district / county regional manager is the set of all restaurants at the district / county level. If a district / county regional manager manages all the restaurants managed by area managers within their jurisdiction, then the collection of restaurants managed by a district / county regional manager can also be the sum of the restaurants in the collections managed by those area managers.

[0029] According to statistical requirements, the statistical granularity can be reduced to different levels, and the specified statistical indicators can be calculated layer by layer for the business data of the business units in the statistical level.

[0030] Granularity refers to the level of detail in classifying and statistically analyzing business units / data within a group of business units. In the classification and statistics at the regional manager level, different granularities can be chosen when calculating statistical indicators. For example, sales figures can be calculated down to the lowest level of restaurants, and then the sales data can be statistically analyzed level by level until the regional manager level. Business managers at the district / county level directly manage the statistical data at the location / regional manager level. Their statistical unit can also be the overall statistical data of the set of restaurants corresponding to the location / regional manager level (e.g., examining the sales performance of each location / regional manager within the district / county manager's jurisdiction), without needing to calculate down to the individual restaurant granularity. Similarly, the overall statistical data of the set of restaurants at the location / regional manager level can be calculated from the sales data of the restaurants within that set managed by the location / regional manager (i.e., examining the performance of each restaurant within the jurisdiction). If there are intermediate levels between the area manager level and the restaurant (store manager) level, the overall statistical data for the set of restaurants managed by each level of business managers or store managers can be calculated recursively until the statistical data at the restaurant level is calculated. The integration of overall statistical data between levels will not be discussed in detail in this article.

[0031] Statistical indicators or statistics include, for example, sales revenue of restaurant establishments, sales revenue / inventory ratio of retail stores, and output of manufacturing plants. Here, a statistical indicator or statistic is a characteristic value, parameter value, or variable value of the statistical feature dimension of interest. These values ​​can be count values, ratio values, probability values, and other numerical values. For these statistical indicators, the calculated reference values ​​can be the mean, median, or mode of these statistical indicators or statistics.

[0032] The mode, in statistical analysis, refers to the value that exhibits a clear central tendency in a statistical distribution, representing the general distribution level of the data. The mode can be the number or value that appears most frequently in a set of data. Sometimes, a set of data can have multiple modes. While the reliability of using the mode to characterize a set of data is generally limited, the concept of the mode is unaffected by extreme data and its calculation method is simple. When conducting statistical analysis, if a set of data contains individual data points with significant variations, choosing the mean or median to represent the "central tendency" of the data may be more appropriate. However, when a set of data, or the data points of interest within that set, cannot be compared using common sorting or comparison rules, the mode is particularly suitable for characterizing the central tendency or statistical characteristics of the data, while arithmetic or other means and medians are difficult to define accurately. The mode is more suitable for combinations or sets of non-numerical data.

[0033] In this article, restaurant sales figures can be sales figures related to a specific time period. For example, business data can be statistically analyzed over time periods such as days, weeks, months, quarters, or years. For instance, WPSA (weekly persale) is monthly sales measured in weeks. To avoid discrepancies in monthly sales figures due to variations in the number of days a store operates each month, other data such as last month's WPSA (weekly sales figures for the previous month), last year's WPSA (weekly sales figures for the same store during the same period last year, compared longitudinally), or rolling 12-month WPSA (a type of cumulative monthly sales figure) can also be used as restaurant sales figures. Furthermore, the WPSA family for a business category refers to the corresponding WPSA data within that business category. This business category differs from the store category, and may include different ordering channels such as takeout, dine-in, and in-store ordering machines. Those skilled in the art will recognize that the WPSA family for a business category also includes different segmented sales data such as last month's WPSA, last year's WPSA, and rolling 12-month WPSA.

[0034] After obtaining the business data of the business units in the level of interest or statistics and calculating the corresponding statistical indicators, the overall statistical data of each set of business units in that level (such as a set of catering outlets) is compared with the statistical indicators, and the statistical results are summarized.

[0035] It can be seen that statistically analyzing business data by hierarchy can be understood as classifying business units using a single classification rule (such as the geographical region hierarchy of stores in a nationwide catering enterprise as mentioned above), and then calculating the corresponding evaluation indicators for the business data of business units in each category to obtain the corresponding evaluation and statistical results. This method of statistically analyzing business data under a single classification rule may have several drawbacks.

[0036] First, the classification method based solely on hierarchical relationships is crude and simplistic, failing to consider the multi-dimensional differences between units within and between hierarchies. For example, classifying and evaluating stores based on a single dimension—their geographical location (their geographic region, or, at the management level, their geographic jurisdiction)—does not account for the impact of other dimensions of business data on the classification and evaluation process. For instance, if a nationwide catering company categorizes its stores solely by market and city, the sales performance of a store in Shanghai might be evaluated simply by averaging the sales of all its stores in Shanghai. In this paper, the focus of statistical analysis of business data is on the statistical data of the observed features. Evaluating this statistical data requires calculating statistical indicators for that feature as reference values. Therefore, the evaluation process involves selecting the statistical feature to be evaluated and comparing it with statistical indicators. Statistical features and evaluation features can be understood equivalently, as can statistical indicators and evaluation indicators. Generally, the selection of conventional evaluation indicators for business data is relatively limited and may be more susceptible to other factors, but these factors are not considered during the classification process, leading to unreasonable statistical and evaluation results. Accurate statistical indicators should be associated with the characteristics or attributes of each dimension of the business unit to comprehensively reflect its features across all dimensions. In this paper, statistical indicators or evaluation indicators are business metrics selected by users from the multi-dimensional characteristics of business data to evaluate the data's features. These are statistically calculated indicators (e.g., mean, median, mode, etc.). Business indicators correspond to business scenarios. Different business scenarios can select different business indicators as references for evaluation features. For example, statistics in the catering and sales industries mainly involve sales / marketing scenarios, where business characteristics include sales revenue, sales frequency, and average order value. The corresponding business indicators are reference values ​​for these characteristics, and therefore, the selected evaluation features or statistical features can be referenced accordingly. Different statistical / evaluation features will result in different calculated statistical indicators or evaluation indicators. Therefore, a single classification rule leads to a monotonous stratification approach that is not scientifically sound in its group induction of the individual objects being statistically analyzed and compared, failing to obtain more accurate classification categories.

[0037] Using a single hierarchical relationship for classification is unreasonable. Users are prone to categorizing business data or the hierarchical level of the business unit to which the data belongs, making it impossible to obtain a fair and effective evaluation of business units and data based on unreasonable classification rules. For example, in classifying restaurants by geographical region, the provincial level category for a regional manager could include categories such as Beijing, Shanghai, and Jiangsu Province. This provincial category includes the collection of all restaurants within that province or a collection of restaurants at the prefecture-level city level. However, the operating performance of restaurants or sets of restaurants within each geographical region has inherent advantages and disadvantages due to geographical differences, making it unfair to compare statistical characteristics between restaurants within the same province after classification by province. For example, within a province, due to population differences, restaurants located in the provincial capital generally have higher sales (in this case, sales volume is used as a statistical or evaluative feature) than restaurants in other cities. Using the city's geographical region as the hierarchical level for these restaurants is unfair. Similarly, within the city geographical region level, comparing the sales of restaurants in popular commercial areas and those near residential areas within the same city is also unfair. For restaurants located near residential areas within the same city tier (e.g., provincial capitals of different provinces), such as a restaurant owned by the same company located in Shanghai's Xintiandi and another in Beijing's Sanlitun, classifying them as the same restaurant category allows for a fairer determination of statistical or evaluative characteristics when compiling sales figures. This fairness is reflected in considering not only the geographical location of the business unit but also other dimensions of its characteristics, such as the type of business district (commercial or residential) and city tier (municipality, first-tier city, or second / third-tier city).

[0038] Figure 1 A block diagram illustrating the logical relationship of a process for determining the category of business data for a business unit in statistics, and evaluating the business data (business unit) based on the category, according to an embodiment of this application.

[0039] First, the process acquires business data 101 associated with business scenario 102. Business data 101 is data associated with a business unit and can include information about the business unit itself or information generated by the business unit. Therefore, business data 101 can have multi-dimensional features or attributes associated with the business unit, and can also include data generated by the business unit or scenario data associated with business scenario 102. Broadly speaking, business data 101 can also include various data associated with business scenarios, such as sales figures for restaurant stores. This data is generated by the business unit in its business activities, is influenced by the characteristics of the business unit itself and business scenario 102, and can also be used as classification features, evaluation features, or statistical features in the classification, statistics, and evaluation of business data 101. Business data 101 can be acquired and stored by each business unit, or it can be acquired and centrally stored and processed by the business unit's superior unit or a professional department of the enterprise (e.g., the statistics department of the head office or regional branch of a national catering enterprise). These business data 101 can be stored on local devices or databases of the enterprise and stores, or on servers or databases such as on the network or in the cloud, and can be accessed, retrieved and processed by users through general or dedicated interfaces or interfaces using wired or wireless technologies.

[0040] Business scenario 102 represents the information about the scenarios and environments in which the business unit, or business personnel, performs relevant business activities or completes business processes when conducting statistics, classification, and evaluation of business data 101. For example, business scenario 102 for a restaurant represents the set of scenario information involved in the activities and processes by which the restaurant provides food and food-related services to customers. This information includes, but is not limited to, store-related information (such as store location, food type, and service items), as well as consumption information of customers who dine in the store (such as average transaction value and store sales). Business scenario 102 constrains the dimensions of the numerous dimensions and attributes of business data 101 that are related to the business activities or processes of interest.

[0041] The acquired business data 101 needs to undergo feature engineering processing based on the requirements of business scenario 102. Feature engineering can be done manually, including judging, defining, and selecting multiple business features 111 that are associated with the business scenario 102 that needs attention. These business features 111 represent different dimensions or attributes of the business data 101.

[0042] The number of business features 111 can be adjusted according to the needs of business scenario 102 and statistical evaluation. For example, the business data 101 of a restaurant can include multi-dimensional features of the restaurant, involving the restaurant's city level, business district type (including commercial areas, residential areas, scenic areas, etc.), parent and subsidiary stores (including subsidiary stores, parent stores, etc.), floor (involving scheduling information), number of floors, restaurant type (including city stores, airport stores, etc.), price type, operating market (e.g., markets in Shanghai and non-Shanghai), area, daily business hours (specifically including start and end times), IE date (involving the date of the most recent renovation before opening, and whether that date is more than 3 months from the current statistical date), weather, and other business feature vectors. Furthermore, each of the above-mentioned dimensional feature vectors may have multiple business feature components. For example, the store's area can include the area of ​​each floor, where the area is also related to the number of floors. These dimensional business features and their components are mainly related to the physical and environmental attributes of the store, such as location, type, time, and weather, and are generally unrelated to subjective factors. Therefore, these dimensional features can be selected as multiple business features 111 of the restaurant associated with the sales scenario.

[0043] Among them, some features in business feature 111 are strongly correlated with business scenario 102 and have a significant impact on business activities. These business features 111 can be extracted as initial business features 112 associated with business scenario 102 to facilitate the initial classification of business data 101. Importance, also known as significance, indicates the degree of influence of a feature or attribute in a certain dimension on the business activities of a business unit. The higher the significance, the greater the impact; conversely, the lower the significance, the smaller the impact.

[0044] The process of classifying restaurants based on their geographical location, as mentioned above, can be seen as an example of using the geographical location level of a restaurant as the initial classification of the business feature 112. In the business scenario 102 where restaurants provide food and food-related services, the geographical location of a restaurant has a significant impact on its sales performance. Therefore, dimensions such as city level, business district type, and restaurant type in business feature 111 are business features with high importance (high significance). At least one of these business features can be selected as the initial business feature 112 for the initial classification of business data 101 and its corresponding business units (restaurants). It is evident that business feature 111 includes not only the initial business feature 112 but also at least other business features with more dimensions different from the initial business feature 112. These other business features can be used to determine further classifications with higher accuracy than the initial classification. These other business features used to determine further classifications of business data can be called additional business features to distinguish them from the initial business features.

[0045] The business data 101, after initial classification, is divided into j initial categories, also known as the first category, denoted as Category 1-1, Category 1-2, ..., Category 1-i, ..., Category 1-j, where i and j are positive integers, and i = 1, 2, ..., j. Thus, based on the initial business features 112, the business data 101 associated with the business scenario 102 can be divided into j initial category units, where the business data contained in the i-th initial category unit is labeled as business data 101-1-i, and these business data 101-1-i all belong to category 1-i. For example, for restaurant outlets of a nationwide catering enterprise, city level and business district type can be used as the two initial business features 112 as the basic dimensions for initial classification. When there are 4 city levels and 6 business district types, 4 categories can be obtained based on the combination of city level and business district type. There are 24 initial categories, corresponding to 24 initial category units, i.e., j = 24.

[0046] The initial classification, as the first classification of business data 101, considers only a smaller number of dimensions of features or attributes compared to the number of dimensions of business features 111, resulting in a relatively coarse classification. In subsequent steps of the classification and evaluation scheme of this application, business data 101 or its business units need to be further classified to improve the accuracy and fairness of the classification. Those skilled in the art will understand that a restaurant or a collection of restaurants managed by the same manager, as the smallest business unit, may be adjusted in subsequent classification steps (e.g., clustering classification) to a new category different from the initial category associated with the geographical region at the time of initial classification.

[0047] Feature engineering includes not only judging, defining and selecting business feature 111, but also preprocessing of business feature 111, including but not limited to expanding and filtering business feature 111.

[0048] The expansion process augments and fills in the missing business feature dimensions and data in business data 101, enabling the corresponding business feature 111 to have corresponding feature values ​​and standardizing the format of business feature 111. Before the expansion process, business data 101 does not possess these expanded business features. The purpose of the expansion process is to allow business data 101 lacking the desired business features to participate alongside business data possessing the desired business features in the process of selecting additional business features for further classification. Here, the desired business features typically refer to additional business features. Therefore, the expansion process can augment and fill in one or more missing additional business features for business data 101. For example, if business data 101 from a certain restaurant lacks dimension feature data associated with the city level feature, the feature values ​​associated with the corresponding city level feature for this restaurant can be filled using the mode calculation result of business data 101 from other restaurants, thus completing the value of the business feature in business data 101. For example, the completion operation can use dictionary matching based on city-city level relationships. When matching, if the same city corresponds to multiple city levels, you can take the first city level by mode, or take the mode of the city levels in the province or national data where the restaurant is located for matching.

[0049] The filtering process is used to filter business data 101 from business units that exhibit abnormal states in business scenario 102. This data 101 cannot be used for subsequent clustering and classification, thus preventing the accurate classification of business data 101 and the corresponding business unit (restaurant store), and consequently hindering accurate and fair statistical analysis and evaluation. Filtered business data 101 typically cannot be re-eligible for clustering and classification through extended processing. For example, restaurants with a zero WPSA value for the current month and / or fewer than 21 valid business days in the current month are considered to have abnormal business operations; their missing dimensional features and their values ​​cannot be filled through extended processing, and therefore they will be filtered out.

[0050] Expansion and filtering processes are used to more efficiently determine further classifications of business data; therefore, expansion can be based on additional business features among multiple business characteristics. Filtering primarily targets business data of business units with anomalous states, especially business data with anomalous states in additional business features.

[0051] After feature engineering, business features 111 (including initial business features 112) can be output in the form of a feature wide table, distinguishing between system variables and business feature variables. System variables include, for example, the restaurant's store code, year and month, city level, business district, and business category (e.g., takeout), which are system-related variables shared by all business units (restaurants). Business feature variables include feature variables of dimensions other than the system variables mentioned above, and are relatively more numerous (e.g., generally more than 20). These variables can represent features related to business scenario 102 and the attributes of the business unit (e.g., restaurant). Business feature variables generally change with business scenario 102 and / or time. In the feature wide table, the value of each business feature 111 indicates whether the features of these dimensions exist or the corresponding data they possess. For example, for a numerical business feature 111, its value represents the corresponding numerical value of the business feature 111; for character or other non-numerical business features, they can be represented by 0-1 variable values, where a value of 0 indicates that the corresponding feature value does not exist, and a value of 1 indicates that the feature value of that dimension exists.

[0052] Based on the feature-engineered business features 111, and targeting the significance index 103 associated with business scenario 102, one or more additional business features are selected and determined from the business features 111 that can be used to determine further classification of business data 101. The additional business features include at least those different from the initial business features 112. Among the additional business features that can be used to determine further classification, a first business feature 121 is selected. The first business feature 121 is the most important additional business feature among the business features 111 possessed by all business data 101 that has the strongest correlation or the greatest impact on business scenario 102; that is, a significant feature in the business data set for all business data 101. The significance index 103 is used to determine the relevance, importance, or significance of business features 111, especially additional business features, to the statistical or evaluation objectives of business scenario 102. The selection of the significance index 103 is related to the objects of classification, statistics, and evaluation of business scenario 102 by business personnel. For example, in the statistical evaluation scenario of a restaurant, when sales volume (e.g., WPSA) is used as the statistical and evaluation target, sales volume is selected as the significance indicator 103, and the process of determining the first business characteristic 121 becomes the process of finding the additional business characteristic that has the most significant impact on sales volume or sales value.

[0053] The relevance, importance, or significance of additional business features to significance index 103 can be represented by the significance value of the additional business features. Additional business features with higher significance can be assigned larger significance values. Then, based on these significance values, the additional business features are sorted, and the n most important (highest significance value) first business features 121 for the statistical or evaluation objective under business scenario 102 are selected, where n is a positive integer less than or equal to the total number of business data 101. For example, using WPSA (Monthly Sales Amount) as the statistical or evaluation objective, the significance values ​​of the additional business features of the restaurant business data 101 are calculated. The additional business features whose significance values ​​meet preset conditions (e.g., the significance value exceeds a preset threshold or falls within a preset threshold range) or whose number of significance values ​​meets preset conditions meets a preset quantity condition (e.g., the number of qualified additional business features exceeds a preset quantity threshold or falls within a preset quantity threshold range) are selected as the first business features 121. For example, when business scenario 102 requires at least 30 salient features for statistical analysis and evaluation, 30 primary business features 121 can be selected from multiple additional business features (at least more than 30). During the selection process, a larger sample set of collected data can be obtained by combining all business data 101 (e.g., the nationwide restaurant stores and their business data of a nationwide catering enterprise) as the data to be selected.

[0054] Additional business features and their corresponding significance values ​​can be determined using regression algorithms, and first business features 121 that meet the threshold number requirement (e.g., n) can be further filtered out. Alternatively, a model incorporating regression algorithms can be constructed to determine the first business feature 121 and its corresponding significance value based on the input additional business features. When using WPSA as the significance index 103, the significance value can be calculated for the current month's WPSA or using historical WPSA. Regression calculations are generally performed on a monthly basis, so the current month's WPSA may be a good regression target. The regression target can also be extended to quarterly WPSA. Regression algorithms can include linear regression and nonlinear regression algorithms. Those skilled in the art will recognize that other algorithms can also be used to calculate the significance level data of the additional business features.

[0055] Before performing regression algorithms, the business features 111, especially the feature variables with additional business features, can be screened for discontinuous and continuous variables, transforming discontinuous variables into continuous variables. For continuous variables, truncation can be performed based on a predetermined variable range to eliminate the negative impact of outliers on statistical analysis. Truncation is equivalent to filtering out noise from the input data that may affect the subsequent classification (e.g., clustering) results.

[0056] The process of determining the significance value of the additional business features to determine the first business feature 121 can be implemented through a machine learning model or a neural network model. The input to the model can be the feature vector of the multi-dimensional additional business features processed as continuous variables, and the output of the model is the determined first business feature 121 and its corresponding significance value. The neural network model can further include deep neural networks (DNNs), convolutional neural networks (CNNs), and other neural network model types capable of implementing algorithms such as regression. Before using the model, the model parameters can be pre-trained using a labeled training data sample set.

[0057] When at least one of the business data 101 and business scenario 102 is updated, the model parameters can be updated offline or online in a timely or periodic manner to obtain better model tracking performance. For example, in summer, weather (rainfall) may be one of the business features that has a significant impact on sales, and weather can be included as one of the first business features 121 with a large significance value; while in autumn, less rain reduces the impact of weather on sales, so weather is no longer included as the first business feature 121. Therefore, the model can be updated periodically or regression calculations can be re-performed according to seasonal changes. When the significance indicator 103 selected by business personnel changes (for example, when store sales are no longer considered as the significance indicator 103, but rather store customer satisfaction is considered), different models can also be used or the model parameters can be updated.

[0058] After obtaining the first business feature 121, for each first category 1-i obtained from the first classification, the business data set consisting of business data 101-1-i included in the initial category unit corresponding to category 1-i is again determined and selected from the n first business features 121 based on the significance index 103. This selection is used for a second classification of the business data set of business data 101-1-i, i.e., the second classification. Here, m is a positive integer less than or equal to n. These m second business features 122 are the first business features 121 (which are also business features 111) that have the greatest relevance, are the most important, and are the most significant to the significance index 103 for the business data set consisting of business data 101-1-i in each initial category unit (a subset of all business data 101). By using the second business feature 122 as the classification rule or focus dimension for the second classification, the business data 101-1-i in each initial category unit can be further classified into a more accurate and scientific category. Therefore, the second classification has a higher classification accuracy than the first classification, or in other words, the second category has a higher accuracy than the first category.

[0059] A regression algorithm can be used to select a second business feature 122 and its corresponding significance value from the first business feature 121 for the business data set consisting of business data 101-1-i. Furthermore, a threshold number (e.g., m) of second business features 122 can be selected, where m is a positive integer less than or equal to n. The process of determining the second business feature 122 is used to find additional business features that have a more significant impact on the significance index 103 determined for business scenario 102, based on the first category, in order to optimize the results of the first classification. The second business feature 122 is determined based on the business data set consisting of business data 101-1-i in the first category 1-i of the first classification. Therefore, m second business features 122 can characterize the more targeted relevance or significance of business data 101 in this business data set for business scenario 102, while n first business features 121 are those additional business features of all business data 101 that have a higher relevance or significance for business scenario 102. It can be considered that, for the business data combinations 101-1-i in the first category 1-i, the second business feature 122 is more important than the first business feature 121 when considering the significance index 103. Generally speaking, the second business feature 122 determined based on each business data combination in the first category 1-i is different, because these business data 101 have already been distinguished in the first classification by the initial business features 112 with different feature values.

[0060] Similar to the process of determining the first business feature 121, the process of determining the second business feature 122 can be based on the same or similar significance index 103 (e.g., WPSA). The significance values ​​of the first business feature 121 can be calculated and ranked, and corresponding significance index screening conditions can be set to further filter out m second business features 122 from the first business feature 121. Alternatively, the second business feature 122 and its corresponding significance value can be determined using a regression algorithm, and this process can be implemented using the machine learning model or neural network model mentioned above. Since the continuous variable transformation and / or truncation processing has already been performed in the process of determining the first business feature 121, these processing steps can be implemented or not implemented in the process of determining the second business feature 122, depending on the circumstances.

[0061] For example, when the business data 101 of catering outlets of a nationwide catering enterprise are divided into 24 initial category units according to the first category (e.g., category 1-i) of the first classification, the number of catering outlets and the number of business data 101 in each initial category unit should be less than the total number of outlets and the total number of business data nationwide. In each initial category unit (e.g., represented accordingly by category 1-i), regression calculations are performed again on the significant variables of the n first business features 121 to further determine the m significant feature vectors (i.e., second business features 122) and their significance values ​​of the business data 101-1-i with respect to the significance index 103 in each initial category unit. Compared to catering outlets nationwide, the number of catering outlets and their business data 101 in each initial category unit is reduced, which is why the number m of the newly determined second business features 122 is generally less than n.

[0062] Since the business data set changes when determining the relevance or significance of each business data point 101 to the significance index 103, the significance values ​​of the determined significant feature variables can be further calibrated during the determination of the second business feature 122 to reflect the impact of changes in the business data set on the significance values. For example, the significance values ​​of n first business features 121 can be weighted, and the weights can include the regression coefficients of the first business features 121 corresponding to the significance values. The regression coefficients are usually determined based on the number of business units (restaurants) or business data in the initial category unit and the total number of all business units (e.g., restaurants nationwide) or business data during the determination of the second business feature 122 (i.e., the second regression operation). For example, the regression coefficient can be the ratio of the number of business units (restaurants) or business data in the initial category unit to the total number of all business units (e.g., restaurants nationwide) or business data. Thus, multiplying the significance value of the first business feature 121 by the regression coefficient yields the calibrated relevance or significance of the first business feature 121 in the initial category unit. The second regression is the process of determining the significant characteristic variables in business feature 111 that have a major impact on the classification and calculation of statistical indicators of catering stores, based on the constraints or further limitations of business scenario 102 proposed by the business department. The two regression calculations can use the same or different regression models. Alternatively, model parameters can be determined separately for each regression calculation, or both models can be pre-trained simultaneously. Accordingly, model parameters can be updated separately for each feature determination (regression) process, or the model parameters can be updated as a whole based on the updated business data 101.

[0063] After obtaining m second business features 122, in each initial category unit, the business data combination belonging to the first category 1-i 101-1-i is classified a second time to obtain a more accurate, scientific and fair second category for the business data combination, denoted as category 2-1, 2-2, ..., 2-i, ..., 2-k, where k is a positive integer and i is a positive integer less than or equal to k.

[0064] The m second business features 122 can be used as the dimensional features for clustering, and a clustering method can be used for the second classification. Those skilled in the art will understand that various clustering methods can be used to complete the second classification.

[0065] The k-means clustering method is a clustering algorithm based on Euclidean distance; the closer two targets are, the greater their similarity. The k-means algorithm first uses k data samples as initial cluster centers. Then, for each data sample, it calculates the distance to the k cluster centers and assigns it to the cluster corresponding to the smallest distance. For each cluster, it recalculates the cluster center. A cluster center can be the centroid of all data samples belonging to that cluster. The algorithm repeats the process of assigning data samples to cluster centers and calculating new cluster centers for each assigned cluster until a preset termination condition is met. The termination condition can be the number of iterations, the minimum error change reaching a (maximum or minimum) threshold, or falling within a predetermined range. The k-means clustering method has low computational complexity and good scalability for large datasets. Traditional clustering methods cannot guarantee that data samples assigned to each cluster have the same weight on each dimension feature. Therefore, it is necessary to analyze the influence or role of the corresponding dimension features on each data sample (e.g., business units or corresponding business data) based on their importance. The sorting and filtering operations of significance values ​​when determining the first business feature 121 and the second business feature 122, as well as the calibration operations when determining the second business feature 122, described above can solve the above-mentioned problems of clustering methods (including k-means clustering methods).

[0066] Specifically, during the clustering process of restaurant stores and their business data, if the number of restaurant stores and their corresponding business data in each initial category unit is less than a set threshold (e.g., 30), these restaurant stores and their business data do not participate in clustering but form a separate group. In this case, the business data in the initial category unit forms a separate second category. If the number of restaurant stores and their corresponding business data is greater than the set threshold, then the business data in the initial category unit is clustered to further divide it into multiple second categories 2-i. Each category 2-i includes approximately the set threshold number of business data or the restaurant stores corresponding to that business data. For example, if an initial category unit includes business data of 100 restaurant stores (more than 30), this initial category unit needs to be further subdivided into four second categories (e.g., denoted as categories 2-1, 2-2, 2-3, and 2-4) by clustering, each containing 30, 30, 30, and 10 business data (restaurant stores) respectively. The clustering rules can also include a default threshold of at least 15 business data points in each second category (i.e., the threshold includes both upper and lower limits); otherwise, the second category will be merged into the nearest other second category. Therefore, the second category 2-4, which contains 10 restaurant business data points, will be merged into one or more of the nearest second categories 2-1, 2-2, and 2-3, which contain 30 restaurant business data points.

[0067] After the second classification, the business data combinations consisting of business data 101-1-i belonging to the first category 1-i are further subdivided into multiple more precise second categories 2-i. For each second category 2-i, the multiple business data 101 (or their corresponding business units) belonging to that second category are... Figure 1 The data is labeled as business data 101-2-i. The result of the second classification completed through clustering can be represented as a list of the correspondence between the number of business data 101-2-i (or business unit) and the number or name of the second category 2-i to which it belongs.

[0068] Clustering algorithms can be implemented using separate models, such as the machine learning or neural network models described above, or they can be integrated into the aforementioned models (e.g., integrated into the model that determines the second business feature 122). According to embodiments of this application, the entire functionality, including determining the second business feature 122 through two regression algorithms and performing a second classification based on the second business feature 122 to determine the second category 2-i, can be implemented using a single holistic model. Each algorithm / function can be performed using a sub-model or a different part of the holistic model (e.g., different layers of a neural network model). The holistic model can be pre-trained using a labeled training dataset or updated with newer business data 101 periodically or in real-time.

[0069] The second category 2-i obtained after the second classification is an accurate category that is strongly correlated, important, or significant with the specific business scenario 102. Each piece of business data 101-2-i included in the category has significant business characteristics in multiple dimensions that are the same or similar to those associated with the business scenario 102. Based on the above first category 1-i and second category 2-i, especially the second category 2-i, the category determination result of the data can be obtained as the final category information.

[0070] Based on the evaluation feature 104 selected by the business personnel, statistical analysis and evaluation are performed on the same type of business data 101-2-i belonging to the same second category 2-i. Evaluation feature 104, as a personalized measurement indicator defined by the business personnel and related to the business scenario 102, can be a dimensional feature of the significance indicator 103, or a dimensional feature different from the significance indicator 103. Evaluation feature 104 can be a business feature of a certain dimension of the business data 101-2-i, such as a business feature from a certain dimension of the second business feature 122, or a feature of a dimension different from the additional business features.

[0071] Based on evaluation feature 104, a reference value or benchmark is calculated for the feature values ​​of evaluation feature 104 for all business data 101-2-i in the second category 2-i. The reference value or benchmark can be calculated using the average, median, or mode of each feature value. The reference value or benchmark may differ from the feature value of evaluation feature 104 for each business data 101-2-i; therefore, there may not actually be a business data 101-2-i corresponding to that reference value. This reference value or benchmark can be called a virtual benchmark to distinguish it from the traditional approach of selecting a specific business unit or business data as a reference object for comparison and evaluation of other business units or business data.

[0072] By comparing the feature value of the evaluation feature 104 of each business data 101-2-i with the reference value, the business performance of business data 101-2-i can be obtained, that is, the business performance of the business unit (e.g., a restaurant) corresponding to business data 101-2-i. Based on the business performance of business data 101-2-i or the corresponding business unit, the evaluation 105-2-i of the business data or business unit can be obtained.

[0073] For example, when analyzing the sales performance of restaurants across a nationwide catering enterprise, the restaurants and their business data (data and dimensional features related to sales performance) are first classified into accurate second categories through two classification processes, following the steps described above. If the initial business feature 112 of the first classification is defined as a specific business district at a certain city level, then the restaurants and their business data in the second category should have the same or similar significant business features. For example, the business data of stores located in bustling commercial districts (Sanlitun in Beijing, Xintiandi in Shanghai, etc.) in first-tier provincial capitals (Category 1) within the commercial district type of provincial capitals should correspond to the dimensional features of stores in bustling commercial districts of first-tier provincial capitals (Category 2). If the statistical indicator selected by the business personnel is sales performance (performance), then the evaluation feature 104 can be selected as a feature related to sales performance (e.g., sales revenue, such as WPSA) and a reference value can be calculated based on the overall sales performance level of the restaurants in Category 2-i. Furthermore, for each restaurant, the same store business indicator can be statistically analyzed and evaluated. Same store can include longitudinal comparison data of business indicators for the same restaurant within the same period, such as comparing the previous month's WPSA with the current month's WPSA, or comparing business indicators from 2019 to 2020 and beyond. Longitudinal comparison data includes not only comparisons of characteristic values ​​of business indicators, but also information on changes in characteristic values, such as statistical data like the increase in sales or the percentage increase in sales. Accordingly, based on the overall sales level of the business data (or restaurants) in this second category, the mean, median, or mode of these Same store values ​​are calculated to obtain a Same store reference value that can serve as a virtual benchmark. The Same store value of each restaurant within the same category is then compared with the reference value to evaluate the restaurant's sales performance.

[0074] By adopting the business data classification and evaluation scheme proposed in this application, and integrating the practical significance of business scenarios, the algorithm's closed-loop iteration is embedded into the business operation evaluation process and updated in a timely manner. When evaluating business data of individual objects such as business units, the scheme of this application can consider the multi-dimensional characteristic differences and distributions of individual objects at different levels and within the same level, scientifically and comprehensively considering more business characteristics, such as airport and scenic area mobile displays, business district types, and weather. When classifying (e.g., clustering) and grouping business data, business data from similar individual objects are aggregated together, breaking through the limitations of conventional thinking, allowing more business units (e.g., restaurants) and their business data with similar dimensional characteristics in a quantitative sense to be clustered into the same category for comparison. For example, KFC in Shanghai Xintiandi and KFC in Beijing Sanlitun are compared in the same category, breaking through geographical limitations. For different statistical and evaluation standards, corresponding evaluation features are selected and corresponding reference values ​​are calculated. Scientific classification results can increase the amount of sample data, improve the accuracy of business data and its business units' evaluation, avoid unfairness and inaccuracy when comparing with fixed evaluation reference values ​​due to differences in actual data and attributes, and provide more business-valued classification and evaluation information for subsequent analysis and calculation.

[0075] Figure 2 This application illustrates a method for determining the category of business data according to embodiments thereof, and a method for evaluating the business data based on the determined category of business data. The method incorporates... Figure 1 The same or similar steps in the classification and evaluation process described will not be detailed hereafter.

[0076] The method for determining the category of business data may include step S210 for acquiring business data associated with a business scenario, step S220 for determining a coarse first category of business data based on the business scenario, step S230 for determining a more accurate second category of business data belonging to the first category based on a selected saliency indicator associated with the business scenario, and step S240 for acquiring the category determination result of business data based on the first category and the second category.

[0077] Step S220 uses initial business features 201 associated with the business scenario to complete the first classification process for obtaining the first category. Initial business features 201 are those business features that are easily identified during feature engineering of the business data and are significantly associated with the business scenario. For example, in a business statistics scenario for a nationwide enterprise, initial business features 201 may be associated with geographical location or the level of enterprise management.

[0078] Step S230 further includes a sub-step S231 of determining multiple additional business features associated with the business scenario, a sub-step S232 of selecting a first business feature from the determined multiple additional business features, a sub-step S233 of selecting and determining a second business feature from the first business feature for the business data combination composed of business data in the first category determined in step S220, and a sub-step S234 of determining a second category of business data in the business data combination composed of business data belonging to the first category based on the determined second business feature.

[0079] In sub-step S231, these additional business features include at least those different from the initial business feature 201 used in step S220 to expand the dimensions of focus of the statistical business data. Based on the selected additional business features, feature dimension data can be expanded and filled in, and abnormal state business features can be filtered out, depending on the different situations of the business features possessed by the business data.

[0080] The sub-step S232 of selecting the first business feature further includes determining the relevance 202, such as the significance value, of each business feature, particularly the additional business features, to the significance index selected for the business scenario. Then, those additional business features whose relevance 202 meets preset conditions are selected as the first business features with the greatest impact on the significance index. The screening of the first business features may include sorting the significance values ​​and selecting multiple first business features with the largest significance values. The significance value can be calculated using a regression algorithm. Preprocessing is performed on the additional business features used for regression calculation, including screening, transformation, and truncation of continuous variables.

[0081] Sub-step S233, which determines the second business feature, further includes calculating the correlation 203 of multiple first business features determined in sub-step S232 with the significance index of the business scenario for the business data group composed of business data belonging to the first category. This yields a second business feature with a more significant impact on the business data group within the first category, thereby optimizing the results of the first classification in subsequent classification processes. The determination of the second business feature can also use a regression algorithm. The correlation 203 of the first business feature can be calibrated to obtain a calibrated correlation. Calibration can be achieved through correlation weighting, where the weights can be determined, for example, based on the number of business data in the business data group belonging to the first category and the total number of business data, such as choosing a ratio between the two.

[0082] The regression algorithms in sub-steps S232 and S233 can use the same or different regression algorithms. The regression algorithms can also be implemented using algorithmic models such as machine learning models or neural network models. These models can be pre-trained before use and their parameters can be updated as business data is updated during use, thereby tracking changes in business data and business scenario requirements.

[0083] Sub-step S234 further includes using a clustering method to perform a second classification of the business data groups in the first category using multiple determined second business features, thereby obtaining the second category of the business data. The clustering method can be, for example, various clustering algorithms such as k-means clustering. Sub-step S234 ultimately determines the more accurate, scientific, and fair second category information for each business data point (or business unit associated with the business data) in the first category.

[0084] Finally, in step S240, the category of the business data is determined according to the first category and / or the second category. The second category, as a further subdivision of the first category, has higher accuracy.

[0085] Based on the categorized business data, evaluation features can be selected according to the statistical needs associated with the business scenario. Figure 2 The dashed line in the diagram illustrates step S250, where for each combination of business data in the first category, a reference value or benchmark value for the characteristic of the evaluation feature is calculated based on the overall situation of the business data combination. Each piece of business data is then evaluated based on the comparison between the characteristic value and the benchmark value. Therefore, compared to the method for determining the category of business data, the method for evaluating business data adds step S250 to steps S210 to S240.

[0086] Embodiments of this application also propose a device 300 for classifying business data (e.g., Figure 3 (as shown) and equipment 400 for evaluating business data (such as...) Figure 4 (As shown).

[0087] The device 300 includes an acquisition unit 310 for acquiring business data associated with a business scenario, a classification unit 320 for determining a first category of the business data based on the business scenario, further classifying the business data belonging to the first category based on a significant indicator associated with the business scenario to obtain a more accurate and scientific second category, and acquiring the category determination results of the business data based on the first and / or second categories. The classification unit 320 can also be further used to implement, for example... Figure 2The functions in steps S220 to S240 shown are illustrated. The device 300 may also include an output unit (not shown) for presenting the classification results of business data to the user or providing them to other devices or units in the form of tables, images, etc.

[0088] The device 400 for evaluating business data includes an acquisition unit 410 for acquiring business data associated with business scenarios, a classification unit 420 for determining a first category of business data based on business scenarios, further classifying the business data in the first category based on significant indicators associated with business scenarios to obtain a more accurate and scientific second category, a classification unit 420 for determining the category of business data acquired based on the first and / or second categories, and an evaluation unit 430 for evaluating the business data based on the (first and second) categories of the business data acquired by the acquisition unit 410 and the (first and second) categories of the business data determined by the classification unit 420. Device 300 can also be embedded in device 400 as a sub-device of device 400, performing the functions of the acquisition unit 410 and the classification unit 420. The evaluation unit 430 can acquire the required business data and its categories based on the data interface with device 300. Device 400 may also include an output unit (not shown) to present the statistical and evaluation results of the business data to business personnel in various ways to guide the company's operations.

[0089] It should be noted that although several modules or units for determining the categories of business data and evaluating business data based on the determined categories are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units. Components shown as modules or units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the technical solution of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0090] In exemplary embodiments of this application, a computer-readable storage medium is also provided, on which a computer program is stored. The program includes executable instructions that, when executed by, for example, a processor, can implement the steps of the method for determining the category of business data and evaluating the business data based on the determined category as described in any of the above embodiments. In some possible implementations, various aspects of this application can also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in this specification for determining the category of business data and evaluating the business data based on the determined category, according to various exemplary embodiments of this application.

[0091] The program product for implementing the above-described method according to embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0092] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0093] The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0094] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0095] In an exemplary embodiment of this application, an electronic device is also provided, which may include a processor and a memory for storing executable instructions of the processor. The processor is configured to perform the steps of the method for determining the category of business data and evaluating the business data based on the determined category in any of the above embodiments by executing the executable instructions.

[0096] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0097] The following reference Figure 5 To describe an electronic device 500 according to this embodiment of the present application. Figure 5 The electronic device 500 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0098] like Figure 5 As shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including storage unit 520 and processing unit 510), a display unit 540, etc.

[0099] The storage unit stores program code that can be executed by the processing unit 510, causing the processing unit 510 to perform the steps described in the methods for determining the category of business data and evaluating business data based on the determined category, according to various exemplary embodiments of this application. For example, the processing unit 510 can perform actions such as... Figure 2 The steps are shown in the figure.

[0100] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 5201 and / or a cache storage unit 5202, and may further include a read-only memory unit (ROM) 5203.

[0101] The storage unit 520 may also include a program / utility 5204 having a set (at least one) program module 5205, such program module 5205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0102] Bus 530 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0103] Electronic device 500 can also communicate with one or more external devices 600 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 500, and / or with any device that enables electronic device 500 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 550. Furthermore, electronic device 500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 560. Network adapter 560 can communicate with other modules of electronic device 500 via bus 530. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0104] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes several instructions to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the method according to the embodiments of this application for determining the category of business data and evaluating the business data based on the determined category.

[0105] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

Claims

1. A method for determining the category of business data, comprising: Obtain business data related to the business scenario; Determining a first category of the business data based on the business scenario, wherein determining the first category of the business data based on the business scenario further includes: determining the first category based on an initial business feature among multiple business features of the business data associated with the business scenario; For business data belonging to the first category, a second category is determined based on a saliency index of the business scenario, wherein the second category has higher accuracy than the first category, and determining the second category of the business data includes dynamically determining the second category of the business data in response to changes in the saliency index; and Based on the first category and the second category, obtain the category determination result of the business data. The second category of the business data, determined based on the significance indicators of the business scenario, includes: Determine multiple additional business features from the plurality of business features, wherein the plurality of additional business features include at least business features that are different from the initial business features; The relevance of the additional business features is determined using a regression algorithm; A first business feature is selected from the additional business features based on their relevance. For each of the first categories, in the business data belonging to the first category, determine the relevance of the first business feature of the business data to the significance index of the business scenario; select a second business feature from the first business feature based on the relevance of the first business feature; The business data is clustered based on the second business characteristic to determine the second category of the business data belonging to the first category. Specifically, dynamically determining the second category of the business data in response to changes in the significance index includes redetermining the correlation through the regression algorithm and dynamically filtering the second business features to re-execute the clustering process.

2. The method according to claim 1, characterized in that, Determining the additional business features among the multiple business features further includes: The business characteristics of the business data are extended to give the business data the determined additional business characteristics.

3. The method according to claim 1, characterized in that, Determining the additional business features among the multiple business features further includes: The business data is filtered out if at least one of the multiple business characteristics is abnormal.

4. The method according to claim 1, characterized in that, Selecting the second business feature from the first business feature based on the relevance of the first business feature further includes: The correlation of the first business feature is calibrated to obtain the calibrated correlation; The second service feature is selected based on the calibrated correlation.

5. The method according to claim 1, characterized in that, The calibration of the relevance of the first business feature further includes: The relevance of the first business feature is calibrated based on the number of business data belonging to the first category and the total number of business data.

6. The method according to claim 1, characterized in that, The calibration of the relevance of the first business feature further includes: The calibrated relevance of the first business feature is the product of the relevance of the first business feature and the ratio of the number of business data belonging to the first category to the total number of business data.

7. The method according to claim 1, characterized in that, The clustering includes k-means clustering.

8. The method according to claim 1, characterized in that, The regression algorithm is implemented using a machine learning model or a neural network model.

9. The method according to claim 1, characterized in that, It also includes updating the correlation based on updates to the business data.

10. The method according to any one of claims 1 to 9, characterized in that, The business scenario includes catering establishments, and the initial business characteristics of the business data are associated with the geographical location of the catering establishments.

11. The method according to claim 10, characterized in that, The significance index includes the sales revenue of the restaurant.

12. A method for evaluating business data, comprising: The method according to any one of claims 1 to 11 determines the first category and the second category of the business data; For the business data belonging to the first category, the business data is evaluated based on the second category.

13. The method according to claim 12, characterized in that, The evaluation of the business data based on the second category further includes: Select the evaluation features for the business scenario; For each of the first categories, in the business data belonging to the first category Determine reference values ​​for the evaluation characteristics of the business data belonging to the second category; Compare the values ​​of the evaluation features of the business data with the reference values ​​of the evaluation features.

14. The method according to claim 13, characterized in that, The reference value for the evaluation feature includes one of the average, median, and mode values ​​of the evaluation feature of the business data.

15. An apparatus for determining the category of business data, comprising: The acquisition unit is configured to acquire business data related to the business scenario. as well as A classification unit is configured to determine a first category of the business data based on the business scenario, determine a second category of the business data for business data belonging to the first category based on a saliency index of the business scenario, and obtain a category determination result of the business data based on the first category and the second category, wherein the second category has higher accuracy than the first category, and determining the second category of the business data includes dynamically determining the second category of the business data in response to changes in the saliency index. The determination of the first category of the business data based on the business scenario further includes: determining the first category based on initial business features among multiple business features of the business data associated with the business scenario. The second category of the business data, determined based on the significance indicators of the business scenario, includes: Determine multiple additional business features from the plurality of business features, wherein the plurality of additional business features include at least business features that are different from the initial business features; The relevance of the additional business features is determined using a regression algorithm; A first business feature is selected from the additional business features based on their relevance. For each of the first categories, in the business data belonging to the first category, determine the relevance of the first business feature of the business data to the significance index of the business scenario; select a second business feature from the first business feature based on the relevance of the first business feature; The business data is clustered based on the second business characteristic to determine the second category of the business data belonging to the first category. Specifically, dynamically determining the second category of the business data in response to changes in the significance index includes redetermining the correlation through the regression algorithm and dynamically filtering the second business features to re-execute the clustering process.

16. An apparatus for evaluating business data, comprising: The device according to claim 15; as well as The evaluation unit is configured to evaluate the business data belonging to the first category based on the second category.

17. A computer-readable storage medium having a computer program stored thereon, the computer program including executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 14.

18. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Smart supply chain system

    JP2021163485A