Grouping method and grouping system for convenience stores in energy industry and processor
By combining internal and external indicators, using contour coefficient evaluation and clustering algorithms, grouping convenience stores in the energy industry, solving the problem of inaccurate grouping results in the existing technology, and achieving more efficient management and decision-making support.
Patent Information
- Application Number
- CN202411925684.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to group convenience stores in the energy industry accurately and effectively, resulting in complex management processes and inefficient efficiency.
By obtaining the first indicator related to the internal sales data of the convenience store and the second indicator related to the external industry data, the outline coefficient is used to evaluate the number of targets, and grouping convenience stores based on the clustering algorithm.
It improves the accuracy and practicality of grouping results of convenient stores, improves the accuracy of grouping results, and thus better serves business decision-making and development.
Smart Images

Figure CN120031607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of retail management, and in particular to a grouping method, a grouping system and a processor for convenience stores in the energy industry. Background Art
[0002] For convenience stores in the energy industry, in order to achieve efficient management of convenience stores at all levels in provinces, cities and municipalities across the country, and to effectively evaluate the operating assessment of subordinate units by each management level and the top-down disassembly of the plan, it is necessary to solve the problem of grouping convenience stores to simplify the management process and improve management efficiency. For example, in the annual plan split scenario, it is necessary to set sales targets for the next year at the end of each year, so it is necessary to group convenience stores by province, city and city, so as to set similar sales plan targets for similar convenience stores; in the business supervision scenario, it is necessary to group convenience stores by province, city and city, so as to conduct similar performance management and evaluation scores for the operating conditions of similar convenience stores, and then give diagnostic suggestions; in the commodity operation scenario, it is necessary to configure and distribute similar commodity combinations based on similar convenience stores, so as to improve the level of commodity operation. It can be seen that in order to meet the needs of benchmarking and analysis in different business scenarios, there is a need for grouping convenience stores at all levels of provinces and cities.
[0003] However, the inventors of this application found in the process of realizing the present invention that the current method of grouping convenience stores usually only selects internal enterprise data and conventional retail indicators. Therefore, the selection of characteristic indicators is difficult to reflect the characteristics of the external industry environment and the actual situation of this application field, and the accuracy and practicality of the grouping results need to be improved. In addition, it is difficult to quantitatively determine the specific number of groups in advance, and often only manually assign values based on experience, which is not enough to support the refined management of convenience stores at the provincial, municipal and city levels. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a grouping method, grouping system and processor for convenience stores in the energy industry, so as to ensure the accuracy and practicality of the grouping results of convenience stores in the energy industry, improve the accuracy of the grouping results, and better serve business decision-making and development.
[0005] In order to achieve the above-mentioned purpose, an embodiment of the present invention provides a method for grouping convenience stores in the energy industry, the grouping method comprising: obtaining a first indicator related to internal sales data and a second indicator related to external industry data of the convenience store, wherein the first indicator includes a surrounding retail indicator of the convenience store and an indicator related to oil consumption; determining a target number of groupings of the convenience stores by means of contour coefficient evaluation; and grouping the convenience stores in a clustering manner according to the first indicator, the second indicator and the target number.
[0006] Optionally, the grouping types of the convenience stores include: provincial grouping, city grouping and store grouping, wherein the store grouping includes inter-provincial store grouping and same-city store grouping.
[0007] Optionally, the number of the second indicators of the store grouping, the province grouping, and the city grouping decreases, and the number of the second indicators of the cross-provincial store grouping is greater than the number of the second indicators of the same-city store grouping.
[0008] Optionally, the surrounding retail indicators include: setting commodity sales revenue and gross profit indicators; the indicators related to oil consumption include: oil non-conversion rate, pure non-oil revenue per ton of oil, and gasoline-diesel ratio; the second indicator includes: the surrounding population and age attributes of the convenience store.
[0009] Optionally, after obtaining a first indicator related to internal sales data and a second indicator related to external industry data of the convenience store, the grouping method further includes: preprocessing the first indicator and the second indicator, wherein the preprocessing includes automatic detection of outliers and / or cross-domain data standardization.
[0010] Optionally, the automatic outlier detection includes a non-proportional interquartile range method, and the cross-domain data standardization includes unification of data units and nonlinear normalization of data values.
[0011] Optionally, grouping the convenience stores in a clustering manner according to the first indicator, the second indicator and the target number includes: grouping the convenience stores according to the first indicator, the second indicator and the target number by using any one of K-means clustering, hierarchical clustering, DBSCAN, spectral clustering, and Gaussian mixture model.
[0012] Optionally, after the convenience stores are grouped in a clustering manner according to the first indicator, the second indicator and the target number, the grouping method further includes: performing multi-classification on the clustering grouping results using the XGBoost algorithm, and determining the contribution of each of the first indicator and the second indicator to the clustering grouping results through supervised learning; setting a weight for each indicator according to the contribution of each indicator; and updating the grouping of the convenience stores in a clustering manner according to the first indicator, the second indicator, the weight of each indicator and the target number.
[0013] On the other hand, the present invention provides a grouping system for convenience stores in the energy industry, the grouping system comprising: an indicator acquisition device for acquiring a first indicator related to internal sales data and a second indicator related to external industry data of the convenience store, wherein the first indicator includes a peripheral retail indicator of the convenience store and an indicator related to oil consumption; a quantity determination device for determining a target number of groupings of the convenience stores by means of contour coefficient evaluation; and a clustering grouping device for grouping the convenience stores in a clustered manner according to the first indicator, the second indicator and the target number.
[0014] On the other hand, the present invention provides a processor for running a program, wherein the program, when run, is used to execute: the method for grouping convenience stores in the energy industry as described above.
[0015] Through the above technical solution, the present application can firstly select different internal and external indicators in combination with the business characteristics of the convenience stores in the energy industry; secondly, evaluate the clustering effect through the silhouette coefficient, measure the compactness within the same cluster and the separation between different clusters, and confirm the target number of groups; finally, perform grouping based on the clustering algorithm. The present invention can ensure the accuracy and practicality of the grouping results of convenience stores in the energy industry, effectively improve the accuracy of the grouping results, and thus better serve business decision-making and development.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0018] Figure 1 It is a flowchart of a method for grouping convenience stores in the energy industry provided according to an embodiment of the present invention;
[0019] Figure 2 is a schematic diagram of factors that should be considered in a grouping model provided according to an embodiment of the present invention;
[0020] Figure 3 is a schematic diagram of a correlation heat map between indicators provided according to an embodiment of the present invention;
[0021] Figure 4 It is a structural diagram of a grouping system for convenience stores in the energy industry provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0023] The present invention first provides a method for grouping convenience stores in the energy industry, such as Figure 1 As shown, the grouping method of the present invention may include steps S110-S130.
[0024] Step S110, obtaining a first indicator related to the internal sales data of the convenience store and a second indicator related to the external industry data.
[0025] In one embodiment, the grouping types of convenience stores may include: provincial grouping, prefecture-level city grouping, and store grouping, etc., wherein the store grouping includes cross-provincial store grouping and same-city store grouping. That is to say, provincial grouping is to group several similar provinces into one group, and prefecture-level city grouping is to group several similar cities into one group. It can be seen that neither of these two groupings involves specific stores, and only store grouping will specifically group stores. Among them, cross-provincial store grouping is to group all stores across the province, which may involve tens of thousands of stores; while same-city store grouping is to group stores in the same city, which involves a small number of stores. The present invention can be used through big data analysis, through the relevant attributes and indicators within the enterprise (i.e., the first indicator), and at the same time combined with the distribution types of provincial and regional companies, prefecture-level city companies, and the objective environment of each convenience store. These external indicators (i.e., the second indicator), conduct in-depth mining and analysis to find the laws and patterns in the indicator data, which can provide data support for subsequent grouping models.
[0026] in, Figure 2It shows some factors that should be considered in the grouping model of this application. In particular, the first indicator may include the peripheral retail indicators of convenience stores and indicators related to oil consumption; the second indicator may include public data provided by the National Bureau of Statistics, and some key indicator data may come from objective data provided by external manufacturers, such as the commercial Nielsen database (data of different dimensions can be purchased according to demand). In one embodiment, considering the complexity of grouping, the number of second indicators of store grouping, provincial grouping, and prefecture-level grouping can decrease in sequence. Specifically, the parameter range involved in store grouping is the widest, including macro and micro levels, so the most external factors need to be considered, while provincial grouping and prefecture-level grouping mainly involve parameters at the macro level, so relatively speaking, there is no need to consider so many external factors, and the second indicator required can decrease in sequence. In addition, it can be understood that the number of second indicators of cross-provincial store grouping should be greater than the number of second indicators of same-city store grouping. This is because for a large number of stores, more external conditions are bound to be required for classification, so cross-provincial store grouping requires more external factors than same-city store grouping. It should be emphasized that the selection of these indicator data is a unique application of the algorithm. By using both internal and external indicators for group calculations, the calculated group results can be made more scientific and comparable.
[0027] In summary, one of the innovative points of the present invention is reflected in the selection of characteristic values. For the convenience stores in the energy industry, especially the characteristics of most of them being together with gas stations, in addition to selecting the retail characteristics around conventional convenience stores, the present invention specially selects some indicators related to oil consumption as characteristic values for grouping clustering as important parameters for grouping, thus greatly improving and helping the scientificity and interpretability of the clustering algorithm. For example, in the actual classification index weight calculation, the applicant found that some industry-specific indicators have a greater impact on the grouping results, such as oil non-conversion rate, non-oil income per ton of oil, and gasoline-diesel ratio and other energy industry characteristic indicators. Including some indicator requirements for the future energy industry in urban planning, it also has a certain impact on the grouping results of provinces, regions and cities.
[0028] Specifically, the present invention particularly selects features related to the characteristics of convenience stores in the energy industry. Considering that the users generally go to convenience stores to consume because they need to refuel their cars, the present invention adds some indicators related to oil-non-business for some business scenarios of oil-non-interaction. For example, the first indicator may include the peripheral retail indicators of the convenience store and indicators related to oil consumption. Among them, the peripheral retail indicators may include: setting commodity sales revenue and gross profit indicators, etc.; indicators related to oil consumption may include: oil-non conversion rate (oil-non conversion rate), non-oil revenue per ton of oil, gasoline-diesel ratio, etc. Among them, the oil-non conversion rate refers to the consumption of non-oil commodities driven by gasoline and diesel consumption, which can be calculated by the following formula: oil-non conversion rate = number of non-oil customers / number of oil customers or number of non-oil orders / number of oil orders. And the non-oil revenue per ton of oil refers to the non-oil sales revenue indicator brought by oil sales, which can be calculated by the following formula: non-oil revenue per ton of oil = non-oil sales revenue / retail sales volume of oil products. It can be seen that the above two indicators can be used as grouping models for calculation, which has a good recognition for store group grouping. After actual testing, the results of store grouping were basically recognized by the demand side and achieved very good results.
[0029] At the same time, the second indicator can include: the surrounding population and age attributes of the convenience store. This is mainly aimed at another situation related to the characteristics of the convenience stores in the energy industry. That is, because of the large scale of convenience stores in the energy industry, the use of centralized procurement has a large bargaining space for commodity prices, and can ensure the authenticity of the commodities, such as high-end wine and high-end cigarettes. Therefore, for the selection of sales revenue and gross profit indicators of certain key commodities, calculation based on the convenience store grouping results is also very helpful. Since the selection of these indicators is determined by the special business characteristics of convenience stores in the energy industry, the scientific nature of the grouping results can be improved. Therefore, attributes such as the population and age around the store can be selected at the same time as the characteristics of the store-related consumer population, and can also be considered as important grouping parameters.
[0030] It can be seen that different grouping models can be applied to different business scenarios, and the selection of indicators is critical to the scientific nature of the grouping results. Therefore, different indicators may be used to calculate grouping data for different application scenarios. The applicant found that the clustering effect was best achieved when internal enterprise indicator data and Nielsen external data were used to jointly model the clustering effect. For example, in the feature selection of one of the application scenarios, features covering 25 dimensions in total can be selected, including the operating conditions of the store, the characteristics of the city where it is located, the relevant attributes of the population, as well as the geographical location and business district characteristics. Such a feature combination is not only comprehensive, but also can accurately reflect the key factors affecting business development, providing strong data support for subsequent clustering analysis.
[0031] Examples of four different classification models are given below.
[0032] In one embodiment, for the inter-provincial store grouping in the store grouping, please refer to Table 1, and the indicators used may include: an internal first indicator with a total of 11 dimensions of data, and an external second indicator with a total of 130 dimensions of data.
[0033] Table 1, Grouping of stores across provinces
[0034]
[0035]
[0036]
[0037]
[0038] In addition, for the grouping of stores in the same city in the store grouping, please refer to Table 2, and the indicators used may include: the internal first indicator with a total of 7 dimensions of data, and the external second indicator with a total of 30 dimensions of data.
[0039] Table 2, Grouping of stores in the same city
[0040] Indicator ID Indicator name Is it internal 1 Number of convenience stores yes 2 Total building area yes 3 Canopy area yes 4 Business hall area yes 5 Convenience store area yes 6 Number of fuel dispensers yes 7 Number of categories of merchandise that can be traded yes 8 Population attributes_Education level_Working population_College no 9 Population attributes_Age_Working population_25-29 no 10 Population attributes_Age_Working population_30-34 no 11 Population attributes_Age_Working population_40-44 no 12 Population attributes_Age_Working population_45-49 no 13 Population attributes_Age_Resident population_18 and below no 14 Population attributes_age_resident population_19-24 no 15 Population attributes_age_resident population_35-39 no 16 Population attributes_age_resident population_40-44 no 17 Population attributes_age_resident population_45-49 no 18 Population attributes_Life stage_Resident population_Having children_No children no 19 Population attributes_Life stage_Resident population_Child status_Children aged 0-6 years old no 20 Population attributes_Life stage_Resident population_Child status_Children aged 7-18 no 21 Population attributes_Income level_Resident population_Wealth index_Medium no 22 Population attributes_Consumption level_Working population_Online consumption ability_Low no 23 Population attributes_Gender_Resident population_Female no 24 Population attributes_Asset status_Working population_Probability of owning a car_Low no 25 Business District Data_Population_Working Population no 26 Business District Data_Population_Residential Population no 27 Business District Data_Population_Total Population no 28 Canopy area no 29 Accommodation services_hotels and guest houses no 30 Office Building no 31 Urban Rail Transit no 32 shopping mall no 33 Health no 34 Total consumer population (total population) no 35 Resident population no 36 Working people no 37 Number of stores no
[0041] In addition, for the provincial grouping, please refer to Table 3. The internal first indicator has 14 dimensions of data; the external second indicator involves 35 dimensions of data, of which 15 dimensions are required for modeling.
[0042] Table 3, Provincial and Regional Store Grouping
[0043] Serial number Indicator name Is it internal 1 Number of convenience stores that can be operated yes 2 Gas station operation data yes 3 Stores with non-oil revenue of one million yuan yes 4 Store with non-oil income of 500,000 yuan yes 5 Non-store with revenue of 300,000 yuan yes 6 Urban planning factors yes 7 income yes 8 gross profit yes 9 Average order value yes 10 Key product revenue yes 11 Ton of oil pure gun non-oil income yes 12 Non-oil gross profit per ton of oil yes 13 Inventory turnover days yes 14 Sales rate yes 15 Per capita income no 16 Number of full-caliber gas stations in the province (including social stations) no 17 Percentage of PetroChina gas stations no 18 Number of cars in the province (including new energy vehicles, 10,000) no 19 Number of new energy vehicles in the province (10,000 vehicles) no 20 Provincial GDP (trillion yuan) no 21 Population of the province (10,000) no 22 Area (10,000 km2) no 23 Number of fuel dispensers per unit area no 24 Number of stores no 25 FMCG Category Sales Index no 26 Importance of urban sales no 27 Channel modernization no 28 MT chain concentration (TOP10 chains) sales contribution no 29 Urban Rail Transit no
[0044] In addition, for the grouping of cities, please refer to Table 4, and the indicators used may include: the internal first indicator has 6 dimensions of data, and the external second indicator has 12 dimensions of data.
[0045] Table 4, Grouping of stores in cities
[0046] Serial number Indicator name Is it internal 1 Number of convenience stores yes 2 Total building area yes 3 Canopy area yes 4 Business hall area yes 5 Convenience store area yes 6 Number of fuel dispensers yes 7 Importance of urban sales no 8 Channel modernization no 9 Sales contribution of MT chain concentration in urban areas no 10 Community no 11 Office Building no 12 Urban Rail Transit no 13 shopping mall no 14 Health no 15 Total consumer population (total population) no 16 Resident population no 17 Working people no 18 Number of stores no
[0047] Step S120, determining the target number of convenience store groups by means of silhouette coefficient evaluation.
[0048] Step S130, grouping the convenience stores in a clustering manner according to the first indicator, the second indicator and the target quantity.
[0049] Among them, clustering algorithm is a machine learning method, the core idea of which is to divide the data set by measuring the similarity or distance between objects, so that the objects in the same group have a higher similarity, while the similarity between different groups is lower. This algorithm is widely used in data mining and pattern recognition to discover potential patterns and structures in data sets.
[0050] Clustering algorithms have a wide range of applications, including but not limited to e-commerce, social networking, financial services, healthcare, bioinformatics, and image processing. For example, in e-commerce, clustering algorithms can be used to provide personalized recommendations and user grouping services; in the healthcare field, clustering algorithms can be used to group patients with diseases to provide personalized treatment plans. In general, clustering algorithms are a powerful data analysis tool that can help people better understand and explore the inherent structure and patterns of data, thereby supporting better decision-making and actions.
[0051] Specifically, the clustering algorithms used in the present invention may include K-means clustering (dividing similar provincial and municipal stores into the same group based on distance measurement), hierarchical clustering (forming a tree-like clustering structure by merging or decomposing), DBSCAN, spectral clustering, Gaussian mixture model, etc. In other words, step S130 may include: according to the first indicator, the second indicator and the target number, using any one of K-means clustering, hierarchical clustering, DBSCAN, spectral clustering, and Gaussian mixture model to group the convenience stores.
[0052] Among them, K-means clustering is one of the most powerful and widely used clustering algorithms, which can be used to divide data into a predefined number of clusters. However, the existing K-means clustering on the market also has some disadvantages, such as being sensitive to outliers and noise points, and requiring the number of clusters to be specified in advance. In this regard, this model is optimized based on the above-mentioned disadvantages, wherein, for the sensitivity to outliers and noise points, an outlier detection algorithm (Local Outlier Factor LOF) can be used to identify and remove outliers in the data set, and the data is cleaned (removing duplicate values, filling missing values, etc.). In addition, the present invention can pre-specify the number of clusters, for example, the clustering effect of different numbers of classifications can be evaluated by the silhouette coefficient evaluation method in step S120, and the best number of classifications can be selected. Alternatively, a hierarchical distance auxiliary determination method can also be used to assist in determining a suitable number of clusters. Then, based on the selected optimal number of classifications, a K-means algorithm can be used for clustering operations.
[0053] In summary, this application can firstly combine the business characteristics of convenience stores in the energy industry, select different internal and external indicators for grouping calculations; secondly, evaluate the clustering effect through the silhouette coefficient, measure the compactness within the same cluster and the separation between different clusters, and confirm the target number of groups; finally, based on the K-means clustering algorithm and PCA (Principal Components Analysis) dimensionality reduction technology (reduce the dimension of the data and extract the most important influencing factors), effectively improve the accuracy of various grouping models.
[0054] For the use of specific provincial and municipal store clustering grouping models, there are four major algorithms, including provincial grouping, prefectural-level city grouping, same-city store grouping, and cross-provincial store grouping. Internal + external data sources are used in the grouping model. According to different algorithms, the data dimensions used are different. Among them, some external data comes from public data such as the National Bureau of Statistics, and some key indicator data comes from objective data provided by external manufacturers (external Nielsen data). At the same time, combined with the internal indicator data of the enterprise, the national provincial grouping calculation is completed, and the function has been launched, with good application feedback. In addition, for prefectural-level city grouping and convenience store grouping, the POC test of the grouping model has been completed (taking a provincial company in the group as an example), and the test results basically meet user expectations.
[0055] In one embodiment, after step S110, the grouping method of the present invention may further include:
[0056] Step S140, preprocessing the first indicator and the second indicator.
[0057] Among them, data preprocessing is a key step to ensure data quality. Data preprocessing can include operations such as data cleaning, data integration, and data transformation to ensure the accuracy and consistency of data and improve the accuracy of the grouping model.
[0058] In one embodiment, the model proposes an intelligent outlier detection method and a cross-domain data normalization method, that is, the preprocessing may include (intelligent) automatic outlier detection and / or cross-domain data normalization.
[0059] For intelligent outlier detection methods, because there are many features involved in the grouping model, data anomalies often occur, such as missing values, outliers, data inconsistency, business logic errors, etc. In this regard, automatic outlier detection can include the non-proportional interquartile range method, and then combine the features with business understanding to automatically identify outliers in the data.
[0060] Specifically, the existing IQR (Interquartile Range) method on the market is to sort a set of data from small to large and divide it into four equal parts (or as close to four equal parts as possible), each part contains about one-quarter of the data. However, this model divides each part in different proportions according to the characteristics of the analyzed data, that is, the non-proportional interquartile range method (non-proportional IQR). Taking the store revenue indicator as an example, the first quartile is 0, that is, 0 is used as the dividing line, and the revenue less than 0 is divided into the lower quartile; the second quartile value is the median. For the third quartile, it is necessary to combine the top 10% of the sales value of all stores in the city, and then average it, and compare the average value of the top 10% with the highest sales of the store. If it is higher than the highest sales of the store, the average value is the third quartile. If it is less than or equal to the highest sales of the store, there is no third quartile value. This non-proportional IQR based on combined business can ensure the accuracy and reliability of the remaining data.
[0061] The cross-domain data standardization method can cope with the challenges of different cross-domain data, such as inconsistent data definitions, inconsistent data units, and diverse data formats. That is, in order to improve the quality and consistency of data, a customized cross-domain data standardization method is proposed in this model. Among them, cross-domain data standardization can include the unification of data units and the nonlinear normalization of data values.
[0062] 1) Data unit standardization: This grouping model customizes a set of data unit requirements. When collecting data from each convenience store, data units are verified and converted to convert data in different units into unified units.
[0063] 2) Data value standardization: Use data nonlinear normalization in combination with business requirements to ensure data value consistency. Common normalization in the market is to transform data to a specific range (usually between 0 and 1) through linear transformation. For example, through the formula: Calculate, where X ′ is the normalized target value, X is the original data value, and X nax is the maximum value in the series data, X min Is the minimum value in the series data values.
[0064] This model uses nonlinear normalization. Different normalization methods such as logarithmic transformation and Box-Cox transformation are used for different features according to the business, which can more flexibly process data with different distribution characteristics.
[0065] The specific data preprocessing steps are: missing value detection, data cleaning, and data distribution analysis. Among them, data distribution analysis can be performed based on the bar charts and box plots of all feature distributions to truncate outliers. Then, the correlation coefficient heat map between each indicator can be used to calculate the correlation coefficient between each indicator. Figure 3 That is to say, if the correlation between two indicators is high (for example, greater than 0.85), it can be understood that the two indicators are related to each other, so one of the indicators can be removed to ensure the independence of the indicators involved in the calculation.
[0066] In one embodiment, after step S130, the grouping method of the present invention may further include:
[0067] Step S151, using the XGBoost algorithm to perform multi-classification on the clustering grouping results, and determining the contribution of each of the first indicator and the second indicator to the clustering grouping results by supervised learning;
[0068] Step S152, setting a weight for each indicator according to the contribution degree of each indicator;
[0069] Step S153, updating the grouping of convenience stores in a clustering manner according to the first indicator, the second indicator, the weight of each indicator and the target quantity.
[0070] The purpose of this step is to combine the results of cluster analysis with business reality in order to further optimize and adjust the grouping results after grouping the convenience stores in the energy industry. That is, to conduct an in-depth analysis of the clustering results in combination with the business background, so as to explore the importance of each feature in the clustering results. Specifically, the results of cluster analysis can be compared and calibrated with the actual needs of the business department, and the model indicators and weights can be adjusted and optimized accordingly based on the feedback and suggestions of the business party. Doing so can not only ensure the accuracy and practicality of the clustering results, but also further improve the accuracy of the clustering effect, so as to better serve business decision-making and development.
[0071] Specifically, the XGBoost algorithm can be used to perform multi-classification on the clustering results, and the importance score of each feature can be obtained through supervised learning. Finally, the features are sorted according to these scores and displayed through visualization to more intuitively understand the contribution of each feature in the clustering model. On the one hand, the stability and generalization ability of the model can be evaluated by dividing the data set multiple times, and the simulation accuracy of the model can be cross-validated. On the other hand, the parameters of the clustering algorithm can be adjusted or a more appropriate algorithm can be selected based on the evaluation results, and the grouping model can be fed back, evaluated and optimized.
[0072] Based on the grouping model above, the grouping result data can be used in business supervision analysis, annual plan splitting, product management and other scenarios. Through actual testing, it can achieve good results in various business scenarios (such as annual plan splitting, business supervision, product operation, etc.), for example:
[0073] (1) Annual plan split scenario: The head office that manages convenience stores needs to set sales targets for the next year at the end of each year. Based on the provincial grouping model, sales plan targets can be set for provinces in the same group.
[0074] Specifically, in the annual plan splitting scenario, when splitting from top to bottom, when splitting from the whole country to the regional companies, it is necessary to split the revenue and gross profit according to different proportions based on the provincial and regional groups; when splitting from the provincial and regional companies to the lower-level municipal companies, it is necessary to split the indicators according to the different proportions of data calculated by different groups. Similarly, when splitting the plan of the municipal company to the store, it is necessary to decompose the plan of the municipal company into each store according to different store groups. The process flow between the business center and the algorithm, the data middle station obtains Nielsen external indicator data through the interface, combines the obtained provincial and municipal store economic environment indicator data and the provincial and municipal store business indicator data as the algorithm input, and obtains the provincial and municipal store grouping through calculation, which is used for the commodity group model of the commodity center and the indicator diagnosis and analysis basis of the business supervision and inventory evaluation. At the same time, the annual plan is decomposed and issued according to the grouping situation and coefficient. The grouping result of the present invention can ensure the rationality and executability of the plan splitting, and the assessment of the plan completion rate is more targeted.
[0075] (2) Operation Supervision Scenario: This scenario is based on the results of national provincial and regional groupings, prefecture-level city groupings, and convenience store groupings in the same city. It is used to evaluate and score the operating conditions of provinces, prefecture-level cities, and convenience stores in the same city, provide diagnostic suggestions, and calculate the average value based on each grouping as the median value for comparison to help improve the level of operations.
[0076] Specifically, in terms of business analysis, in the business evaluation of provincial companies, the basis of the evaluation is based on the basic data of the province. Mainly based on the selected various indicator data, the position of the provincial companies in the same group, and according to the different comparison results of high, medium and low, various diagnostic suggestions are put forward in a targeted manner to guide the improvement and enhancement of the business indicators of provincial companies. Similarly, based on the grouped data, the business conditions of different levels such as prefecture-level companies and convenience stores in the same city are evaluated and compared, and various indicators of organizations in the same group are evaluated and diagnosed according to the development level. The proposed diagnostic conclusions are more scientific and operational.
[0077] (3) Product operation scenarios: Product operations combine and distribute products in different ways based on the characteristics of different store grouping types to improve the level of product operations.
[0078] Specifically, in terms of merchandise management and the combination of merchandise packages, different merchandise combinations need to be adopted according to different store groupings and the characteristics of stores in different groups, so as to increase store sales revenue and gross profit, improve indicators such as inventory turnover rate, and improve operating levels.
[0079] Through the above technical solution, the present application can firstly select different internal and external indicators in combination with the business characteristics of the convenience stores in the energy industry; secondly, evaluate the clustering effect through the silhouette coefficient, measure the compactness within the same cluster and the separation between different clusters, and confirm the target number of groups; finally, perform grouping based on the clustering algorithm. The present invention can ensure the accuracy and practicality of the grouping results of convenience stores in the energy industry, effectively improve the accuracy of the grouping results, and thus better serve business decision-making and development.
[0080] On the other hand, the present invention also provides a grouping system for energy industry convenience stores, such as Figure 4 As shown, the grouping system may include:
[0081] The indicator acquisition device 210 is used to acquire a first indicator related to internal sales data of the convenience store and a second indicator related to external industry data, wherein the first indicator may include a peripheral retail indicator of the convenience store and an indicator related to oil product consumption;
[0082] The quantity determining means 220 is used to determine the target quantity of the convenience stores grouped by means of silhouette coefficient evaluation; and
[0083] The clustering grouping device 230 is used to group the convenience stores in a clustering manner according to the first indicator, the second indicator and the target quantity.
[0084] Through the above technical scheme, the present application can firstly select different internal and external indicators in combination with the business characteristics of the convenience stores in the energy industry; secondly, evaluate the clustering effect through the silhouette coefficient, measure the compactness within the same cluster and the separation between different clusters, and confirm the target number of groups; finally, grouping is performed based on the clustering algorithm. The present invention can ensure the accuracy and practicality of the grouping results of convenience stores in the energy industry, effectively improve the accuracy of the grouping results, and thus better serve business decision-making and development.
[0085] Among them, the grouping system also includes a processor and a memory. The above-mentioned indicator acquisition device 210, quantity determination device 220, and cluster grouping device 230 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0086] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be provided, and the purpose of the present invention is achieved by adjusting kernel parameters.
[0087] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0088] An embodiment of the present invention provides a storage medium having a program stored thereon, which, when executed by a processor, implements the method for grouping convenience stores in the energy industry.
[0089] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes the method for grouping convenience stores in the energy industry when running.
[0090] The embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor, and when the processor executes the program, each step of the above-mentioned method for grouping convenience stores in the energy industry is implemented. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0091] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing the various steps of the above-mentioned method for grouping convenience stores in the energy industry.
[0092] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0093] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0094] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0096] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0097] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0098] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0099] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0100] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for grouping convenience stores in the energy industry, characterized in that: The grouping method comprises: Acquire a first indicator related to internal sales data of the convenience store and a second indicator related to external industry data, wherein the first indicator includes a peripheral retail indicator of the convenience store and an indicator related to oil product consumption; Determining a target number of groups for the convenience stores by means of silhouette coefficient evaluation; and The convenience stores are grouped in a clustering manner according to the first indicator, the second indicator, and the target quantity.
2. The grouping method according to claim 1, characterized in that: The grouping types of the convenience stores include: provincial grouping, city grouping and store grouping, wherein the store grouping includes inter-provincial store grouping and same-city store grouping.
3. The grouping method according to claim 2, characterized in that: The number of the second indicators of the store grouping, the province grouping, and the city grouping decreases, and the number of the second indicators of the cross-provincial store grouping is greater than the number of the second indicators of the same-city store grouping.
4. The grouping method according to claim 1, characterized in that: The peripheral retail indicators include: setting commodity sales revenue and gross profit indicators; The indicators related to oil consumption include: oil non-conversion rate, non-oil income per ton of oil, and gasoline-diesel ratio; The second indicator includes: the surrounding population and age attributes of the convenience store.
5. The grouping method according to claim 1, characterized in that: After obtaining the first indicator related to the internal sales data and the second indicator related to the external industry data of the convenience store, the grouping method further includes: The first indicator and the second indicator are preprocessed, wherein the preprocessing includes automatic outlier detection and / or cross-domain data standardization.
6. The grouping method according to claim 5, characterized in that: The automatic outlier detection includes a non-proportional interquartile range method, and the cross-domain data standardization includes unification of data units and non-linear normalization of data values.
7. The grouping method according to claim 1, characterized in that: The grouping the convenience stores in a clustering manner according to the first indicator, the second indicator, and the target quantity includes: According to the first indicator, the second indicator and the target quantity, the convenience stores are grouped by using any one of K-means clustering, hierarchical clustering, DBSCAN, spectral clustering and Gaussian mixture model.
8. The grouping method according to claim 1, characterized in that: After grouping the convenience stores in a clustering manner according to the first indicator, the second indicator, and the target quantity, the grouping method further includes: Perform multi-classification on the clustering grouping results by using the XGBoost algorithm, and determine the contribution of each of the first indicator and the second indicator to the clustering grouping results by means of supervised learning; According to the contribution degree of each indicator, a weight is set for each indicator; The grouping of the convenience stores is updated in a clustering manner according to the first indicator, the second indicator, the weight of each indicator, and the target quantity.
9. A grouping system for convenience stores in the energy industry, characterized in that: The grouping system comprises: An indicator acquisition device, used to acquire a first indicator related to internal sales data of the convenience store and a second indicator related to external industry data, wherein the first indicator includes a peripheral retail indicator of the convenience store and an indicator related to oil product consumption; a quantity determining device for determining a target quantity for grouping the convenience stores by means of silhouette coefficient evaluation; and A clustering grouping device is used to group the convenience stores in a clustering manner according to the first indicator, the second indicator and the target quantity.
10. A processor, characterized in that: Used to run a program, wherein the program, when run, is used to execute: a method for grouping convenience stores in the energy industry as described in any one of claims 1-8.