Data management platform and data management method
By analyzing the sales volume and user trajectory data of retail outlets, combined with geographic distance, sales volume is predicted and data risks are assessed. This solves the problem of lagging real-time risk assessment in data governance in the retail industry, and achieves real-time supervision of data quality and improved efficiency of big data analysis.
Patent Information
- Application Number
- CN202510788870.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-19
AI Technical Summary
Under the chain operation model of the retail industry, traditional data governance methods are unable to reflect the data risks of retail outlets in real time, resulting in delayed data quality assessment and affecting the efficiency and accuracy of big data analysis.
By collecting sales data and user trajectory data from retail outlets, analyzing user distribution similarity and sales similarity, and combining geographical distance, sales are predicted and data risks are assessed, and real-time data risk assessment and governance are performed using the data governance platform.
It has improved the efficiency of data quality supervision in retail outlets, optimized business strategies, improved the efficiency of big data processing and the accuracy of sales forecasts, and reduced data anomalies.
Smart Images

Figure CN120672372A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data risk governance, and in particular to a data governance platform and a data governance method. Background Art
[0002] The retail chain operation model is a business organization structured around the unified management and operation of multiple retail outlets. It employs standardized merchandise displays and service specifications, creating economies of scale to reduce costs, increase brand awareness, and enhance market reach. Data governance within this retail chain model primarily involves collecting large amounts of data from retail outlets and building a data cloud platform to conduct data risk assessment, data quality management, and data lifecycle management. This data is then mined through big data technologies to obtain deeper insights and optimize retail outlets' operational strategies.
[0003] Terminal retail outlets in the retail industry generally require a large amount of manual operation, which may lead to various risks of data tampering. Traditional data quality assessment of retail outlets is mainly achieved through manual inspection, rule verification, regular audits, etc. These methods have a high lag in data risk assessment and cannot reflect the data risk situation of retail outlets in real time; moreover, the data governance platform uses the same calculation strategy for the data generated by each retail outlet, which seriously affects the efficiency and accuracy of the platform's subsequent big data analysis. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a data governance platform and a data governance method.
[0005] According to a first aspect of an embodiment of the present invention, a data governance method is provided, and the technical solutions adopted are as follows: Collect sales data and user trajectory data from retail outlets; Analyzing the user distribution around each of the retail outlets based on the user trajectory data to obtain the user distribution similarity between the current retail outlet and other retail outlets; Obtaining a sales volume similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets; Based on the sales data, analyzing the historical sales similarity between the retail outlets, and combining the sales similarity change rate to obtain the predicted sales similarity between the current retail outlet and other retail outlets; Based on the sales data, analyzing the sales stability of other retail outlets, and combining the predicted sales similarity to obtain the predicted sales of the current retail outlet; Analyze the difference between the predicted sales and actual sales of current retail outlets, determine the data risk level of current retail outlets, and complete data governance.
[0006] In some embodiments of the present invention, based on the user trajectory data, analyzing the user distribution around each of the retail outlets to obtain the user distribution similarity between the current retail outlet and other retail outlets includes: Get the radius of each retail outlet r a planned area within the range, and calculating a second distance between a center point of the planned area and the retail outlet; Based on the user trajectory data, evaluating user portrait characteristics of the planning area to obtain a user group parameter vector; Obtaining user distribution around each of the retail outlets based on the user group parameter vector and the second distance; Analyze the differences in user distribution between the current retail outlet and other retail outlets to obtain the similarity of user distribution between the current retail outlet and other retail outlets.
[0007] In some embodiments of the present invention, the user group parameter vector includes average passenger flow and average consumption level.
[0008] In some embodiments of the present invention, analyzing the difference in user distribution between the current retail outlet and other retail outlets to obtain the similarity of user distribution between the current retail outlet and other retail outlets includes: The cosine similarity between the user group parameter vectors of the current retail outlet and all the planned areas corresponding to other retail outlets is analyzed, as well as the second distance difference between the current retail outlet and all the planned areas corresponding to other retail outlets is analyzed to obtain the user distribution similarity between the current retail outlet and other retail outlets.
[0009] In some embodiments of the present invention, obtaining the sales volume similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets includes: Sort the user distribution similarities between the current retail outlet and the other retail outlets in ascending order based on the first distances between the current retail outlet and the other retail outlets, to obtain a user distribution similarity sequence corresponding to the current retail outlet; Calculating a difference between a previous element and a next element in the user distribution similarity sequence and a third distance between corresponding retail outlets, and counting the number of positive values of the difference; The sales volume similarity change rate of the current retail outlet is obtained by combining the absolute values of the differences corresponding to all adjacent elements in the user distribution similarity sequence, the third distance, and the number of positive differences.
[0010] In some embodiments of the present invention, based on the sales data, analyzing the historical sales similarity between the retail outlets, and combining the sales similarity change rate to obtain the predicted sales similarity between the current retail outlet and other retail outlets, includes: Analyze the difference between the sales data of the current retail outlet and other retail outlets for the same product number, traverse all products with the same number, and obtain the historical sales similarity between the current retail outlet and other retail outlets; Obtaining the sales similarity between the current retail outlet and the other retail outlets based on the sales similarity change rate of the current retail outlet and the first distance between the current retail outlet and the other retail outlets; The predicted sales similarity between the current retail outlet and other retail outlets is obtained by combining the sales similarity of the current day and the historical sales similarity.
[0011] In some embodiments of the present invention, based on the sales data, analyzing the sales stability of other retail outlets, and combining the predicted sales similarity to obtain the predicted sales of the current retail outlet, includes: Based on the sales data, analyzing the variance of the historical sales data of other individual retail outlets and the variance of the historical overall sales data of all retail outlets to obtain the sales stability of other retail outlets; According to the sales stability, combined with the predicted sales similarity and the sales data of other retail outlets on the same day, all other retail outlets are traversed to obtain the predicted sales of the current retail outlet.
[0012] In some embodiments of the present invention, the difference between the predicted sales volume and the actual sales volume of the current retail outlet is analyzed to obtain the data risk level of the current retail outlet and complete data governance, including: Analyze the absolute value of the difference between the predicted sales volume and the actual sales volume of all types of goods at the current retail outlets to obtain the data risk level of the retail outlets; Set data risk level thresholds; Determining whether the data risk level is greater than the data risk level threshold; If so, there is data anomaly in the retail outlet corresponding to the data risk level, and an anomaly notification is sent to complete data governance.
[0013] According to a second aspect of an embodiment of the present invention, a data governance platform is provided, comprising: a retail outlet data collection module, a data risk assessment module, and a big data processing module; wherein: Retail outlet data collection module, used to collect sales data and user trajectory data of retail outlets; Data risk assessment module, used to analyze the real-time data risk level of retail outlets and complete data governance; The big data processing module, including the big data analysis model, is used to perform big data processing on the managed retail outlet data.
[0014] In some embodiments of the present invention, the data risk assessment module includes: A user distribution similarity analysis unit is used to analyze the user distribution around each of the retail outlets based on the user trajectory data, and obtain the user distribution similarity between the current retail outlet and other retail outlets; a sales similarity analysis unit configured to obtain a sales similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets; and to analyze the historical sales similarity between the retail outlets based on the sales data, and obtain a predicted sales similarity between the current retail outlet and other retail outlets based on the sales similarity change rate; A predicted sales acquisition unit, configured to analyze the sales stability of other retail outlets based on the sales data, and obtain the predicted sales of the current retail outlet in combination with the predicted sales similarity; The data risk level analysis unit is used to analyze the difference between the predicted sales volume and the actual sales volume of the current retail outlets to obtain the data risk level of the current retail outlets.
[0015] Compared with the existing technology, the data governance platform and data governance method provided by the present invention have the following beneficial effects: 1. This invention uses the sales data of other retail outlets to estimate the sales of a single current retail outlet, and obtains data anomalies based on the difference between the predicted sales and the actual sales. This can effectively improve the efficiency of data quality supervision of retail outlets, thereby improving the efficiency of platform big data processing and optimizing the operating strategies of retail outlets.
[0016] 2. The present invention analyzes the user distribution around each of the retail outlets to obtain the similarity of user distribution between the current retail outlet and other retail outlets. The user distribution around a retail outlet determines the sales volume of the retail outlet. Therefore, the more consistent the user distribution shown by the surrounding environment of two retail outlets, the higher the possibility that the two retail outlets have similar sales volume, providing a data basis for subsequent sales volume prediction of retail outlets.
[0017] 3. The present invention takes into account that similar user groups at different geographical distances may have inconsistent consumption habits, but user groups with closer distances are more likely to have consistent consumption habits. Therefore, it can be assumed that the closer the distance between two retail outlets, the more the similarity of the user distribution around them can represent the similarity of the users' consumption behavior habits, and thus the similarity of the sales volume of goods at the two retail outlets can be estimated; that is, the present invention corrects the weight of user distribution in predicting the sales volume of retail outlets by the influence of geographical distance, thereby improving the reliability of retail outlet sales prediction.
[0018] 4. The present invention also takes into account the stability of sales at other retail outlets. When the sales at other retail outlets are relatively stable, the higher the credibility value and the greater the reference value of the prediction using similarity assessment, the more important it is when predicting sales at the current retail outlet. That is, the present invention uses the sales stability of other retail outlets as the weight for predicting sales at the current retail outlet, thereby further improving the accuracy of predicting sales at the current retail outlet. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 A schematic diagram of the basic process of a data governance method provided by one embodiment of the present invention; Figure 2 A schematic diagram of the distribution of retail outlets in a planned area provided by an embodiment of the present invention; Figure 3 A schematic diagram of the basic composition of a data governance platform provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a data governance platform and data governance method proposed in accordance with the present invention, its specific implementation method, structure, features and effects. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. Terms such as "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a circuit structure, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such article or device. In the absence of further limitations, the phrase "comprising a ..." to define an element does not preclude the presence of other identical elements in the article or device comprising the element.
[0023] The following describes in detail a specific scheme of a data governance method provided by the present invention with reference to the accompanying drawings.
[0024] See also Figure 1 , which shows the basic process of a data governance method provided by an embodiment of the present invention.
[0025] like Figure 1 As shown, an embodiment of the present invention provides a data governance method, which specifically includes: S100: Collect sales data and user trajectory data of retail outlets.
[0026] Data collection terminals are installed at retail chains of a certain scale to collect data from these outlets and transmit it to a cloud computing platform via wireless or wired channels. For a single retail outlet, the types of data required include sales and user trajectory data. This involves using intelligent cash registers to obtain daily sales data for the retail outlet. Furthermore, using a GIS (Geographic Information System) system, combined with geo-fencing technology, and obtaining anonymized signaling data from base station operators, user trajectory data around the retail outlet is collected.
[0027] At this point, sales data and user trajectory data of retail outlets have been obtained.
[0028] Sales volume data in the retail industry is key data for analyzing data risks at retail outlets. It can clearly show sales trends, sales structure, and the preferences of users near retail outlets. The distribution of users around a retail outlet determines the sales volume of that retail outlet. Therefore, the more consistent the user distribution around two retail outlets, the more likely the two retail outlets are to have similar sales volumes.
[0029] Furthermore, when it comes to geographic distance, similar user groups may have different consumption habits, but user groups that are closer together are more likely to have similar consumption habits. Therefore, the closer two retail outlets are to each other, the more similar the user distribution around them is, representing the similarity of their consumer behavior habits. This, in turn, can be used to estimate the similarity of the sales volume at the two retail outlets. Accordingly, if two retail outlets have relatively similar user distributions and are closer together, the likelihood that this similarity in user distribution represents similar sales volume is greater. Therefore, for a retail outlet, the likelihood of similar sales volume at multiple other retail outlets can be used to derive estimated sales volume. If the estimated sales volume differs significantly from the actual sales volume, the risk of data anomalies is high.
[0030] Therefore, first, the degree of user distribution similarity between any two retail outlets is calculated. Second, the likelihood of sales similarity between the current retail outlet and other retail outlets is analyzed by calculating the distribution of distance changes between retail outlets, thereby obtaining the predicted sales value of the current retail outlet. Finally, the difference between the estimated sales and actual sales of the current retail outlet is calculated to obtain the data anomaly risk level of the current retail outlet. This specifically includes steps S200 to S600.
[0031] S200: Analyze the user distribution around each retail outlet based on the user trajectory data to obtain the user distribution similarity between the current retail outlet and other retail outlets.
[0032] Based on user trajectory data, analyze the user distribution around each retail outlet to obtain the user distribution similarity between the current retail outlet and other retail outlets. Further analysis includes: First, obtain the radius of each retail outlet r The specific implementation method is: using the GIS system to obtain all planning areas within a 2km radius of each retail outlet, and calculate the second distance between the center point of the planning area and the retail outlet, wherein the second distance is the straight-line distance between the geographical location of the center point of the planning area and the geographical location of the retail outlet. Figure 2 As shown, retail outlets The planning area within a radius of 2km is obtained as the center, and the white circle represents the center point of the planning area. The surrounding users can be expressed as (primary school, 100m), (residential area, 300m), and (office area, 200m).
[0033] Then, based on the user trajectory data, the user profile characteristics of the planning area are evaluated to obtain the user group parameter vector. Furthermore, the user group parameter vector includes the average passenger flow and the average consumption level. The specific implementation method is as follows: First, the analysis range of the planning area is set to 300 meters radiating outward from the planning area, such as Figure 2 As shown. For the average passenger flow of the planning area, the average passenger flow of the planning area can be represented by counting the number of user trajectories within the analysis range of the planning area. For the average consumption level of the planning area, firstly, the user trajectory data within 300 meters around the planning area in one day is obtained through mobile phone signaling data, and the consumption places within the planning area are obtained through the GIS system. The user trajectory data is combined to determine the length of time the user stays at the consumption place. When the length of time the user stays at the consumption place is greater than the preset stay threshold (generally the stay threshold is set to 5 minutes), it indicates that the user has made consumption there, and then the consumption situation of the user is represented according to the average consumption level of this type of consumption place; then the sum of the user's consumption in different consumption places is combined to obtain the user's daily consumption situation; finally, the average consumption level of the planning area is obtained based on the consumption situation of multiple users.
[0034] Then, based on the user group parameter vector and the second distance, the user distribution around each retail outlet is obtained. The user distribution around is expressed as: [(user group parameter vector 1 of planning area 1, second distance 1 corresponding to planning area 1), (user group parameter vector 2 of planning area 2, second distance 2 corresponding to planning area 2), ..., (planning area The user group parameter vector , planning area The corresponding second distance )],in, Retail outlets The number of planning areas within a radius of 2 km, and the user group parameter vector is expressed as (average passenger flow, average consumption level).
[0035] Finally, the difference in user distribution between the current retail outlet and other retail outlets is analyzed to obtain the similarity of user distribution between the current retail outlet and other retail outlets. The specific implementation method is: the more similar the user group parameter vectors of the planning area are, the higher the possibility that users will make the same consumption. At the same time, the second distance between the planning area and the retail outlet also affects the user's consumption willingness. When the second distance is closer, the user's willingness to go there for consumption is higher. Therefore, the higher the similarity of the user group parameter vectors of the two planning areas corresponding to the current retail outlet and other retail outlets, and the closer the second distance from the retail outlets, the higher the similarity of user distribution around the two retail outlets. Therefore, the difference in user distribution between the current retail outlet and other retail outlets is analyzed to obtain the similarity of user distribution between the current retail outlet and other retail outlets, specifically including: analyzing the cosine similarity between the user group parameter vectors of all planning areas corresponding to the current retail outlet and other retail outlets, and analyzing the second distance difference between all planning areas corresponding to the current retail outlet and other retail outlets, to obtain the similarity of user distribution between the current retail outlet and other retail outlets. Construct the current retail outlet With other retail outlets The formula for calculating the user distribution similarity is: ; Where, Indicates current retail outlets With other retail outlets User distribution similarity; represents the cosine similarity function; Indicates current retail outlets No. The user group parameter vector corresponding to the planning area; Indicates other retail outlets No. The user group parameter vector corresponding to the planning area; Indicates current retail outlets No. The center point of the planning area and the current retail outlets The second distance between Indicates other retail outlets No. planning areas and other retail outlets The second distance between Indicates current retail outlets The number of planning areas within a 2km radius; Indicates other retail outlets The number of planning areas within a 2km radius; Represents natural base An exponential function with base .
[0036] Cosine similarity between the parameter vectors of user groups in the planning areas of two retail outlets Indicates the similarity between the user group parameter vectors of the planning areas of the two retail outlets, The larger the value, the greater the similarity of user distribution between the two retail outlets; It represents the second distance difference between the planned area of the current retail outlet and the planned areas of other retail outlets. The smaller the value, the more similar the impact of the second distance on the user distribution similarity is, and the greater the user distribution similarity between the two retail outlets.
[0037] S300: Obtaining a sales volume similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets.
[0038] The above steps calculate the similarity between two retail outlets based on the user distribution around them. However, even if the user distribution around two retail outlets is the same, the same user group may have different consumption habits due to differences in their regions. That is, if the distance between two retail outlets is significantly different, the actual consumption habits in the two regions may be significantly different. Therefore, for two retail outlets, similar user distributions are more likely to result in similar sales. However, as the distance between the two retail outlets increases, the likelihood of similar user distributions resulting in similar sales gradually decreases. Therefore, for the current retail outlet, if the trend of the user distribution similarity of other retail outlets gradually decreasing with increasing distance is more obvious, it indicates that the sales similarity between other retail outlets and the current retail outlet is lower at greater distances.
[0039] Based on the above analysis, in an embodiment of the present invention, the sales volume similarity change rate of the retail outlets is obtained based on the user distribution similarity and the first distance between the retail outlets. Further including: First, according to the first distance between the current retail outlet and other retail outlets, the user distribution similarities between the current retail outlet and other retail outlets are sorted in ascending order to obtain the user distribution similarity sequence corresponding to the current retail outlet. The first distance to all other retail outlets, and then based on the current retail outlet The first distance between the current retail outlet and all other retail outlets is in ascending order. Sort by user distribution similarity with all other retail outlets to get the current retail outlet The corresponding user distribution similarity sequence is recorded as ( 、 、 ,…, ,…, ),in, Indicates current retail outlets Similarity sequence with user distribution The user distribution similarity between other retail outlets, Represents the number of all other retail outlets.
[0040] Then, the difference between the previous element and the next element in the user distribution similarity sequence and the third distance between the corresponding retail outlets are calculated, and the number of positive differences is counted.
[0041] Finally, the sales similarity change rate of the current retail outlets is obtained by combining the absolute value of the difference and the third distance of all adjacent elements in the user distribution similarity sequence and the number of positive differences. The formula for calculating the sales similarity change rate caused by user distribution is: ; Where, Indicates current retail outlets The rate of change in sales similarity due to user distribution; represents the number of all other retail outlets; Indicates current retail outlets Similarity sequence with user distribution The similarity of user distribution among other retail outlets; Indicates current retail outlets Similarity sequence with user distribution The similarity of user distribution among other retail outlets; The first The other retail outlets corresponding to the first element The third distance between other retail outlets corresponding to the elements; express The value is the number of positive values; represents the linear normalization function.
[0042] The number of positive values The larger the number, the more it indicates that the As the first distance gradually increases, the trend of user distribution similarity gradually decreasing becomes more obvious, indicating that the first distance between the two retail outlets is likely to have a greater impact on sales, indicating that the current retail outlets The greater the rate of change of sales similarity affected by distance, the greater the absolute value of the difference in user distribution similarity corresponding to adjacent elements in the user distribution similarity sequence. The larger the value is and the smaller the corresponding third distance is, the greater the impact of distance on the similarity of user distribution is, indicating that the current retail outlets The greater the rate of change in sales similarity affected by distance.
[0043] S400: Based on the sales data, the historical sales similarity between retail outlets is analyzed, and combined with the sales similarity change rate, the predicted sales similarity between the current retail outlet and other retail outlets is obtained.
[0044] The above steps are all calculated based on the data of the retail outlets on that day, and the current retail outlets are obtained. The sales similarity change rate caused by the user distribution on that day; it is necessary to further analyze the historical sales similarity between retail outlets based on sales data. Specifically, analyze the difference between the sales data of the same numbered product at the current retail outlet and other retail outlets, traverse all the same numbered products, and obtain the historical sales similarity between the current retail outlet and other retail outlets. Construct the current retail outlet With other retail outlets The calculation formula for the historical sales similarity between the two is: ; Where, Indicates current retail outlets With other retail outlets Similarity of historical sales between them; Indicates current retail outlets and other retail outlets The number of types of products with the same product number; Indicates current retail outlets No. The sales quantity of the commodity on the previous day; Indicates other retail outlets No. The sales quantity of the commodity on the previous day; Represents natural base An exponential function with base ; represents the linear normalization function.
[0045] By comparing the sales volume of each product in two retail outlets on the previous day, the historical sales similarity between the two retail outlets can be obtained.
[0046] Combined with the historical sales similarity and the first distance between the current retail outlet and other retail outlets and the sales similarity change rate of the current retail outlet, the predicted sales similarity of the current retail outlet and other retail outlets is obtained. Specifically, based on the sales similarity change rate of the retail outlets and the first distance between the current retail outlet and other retail outlets, the sales similarity of the current retail outlet and other retail outlets on the same day is obtained; combined with the sales similarity of the same day and the historical sales similarity, the predicted sales similarity of the current retail outlet and other retail outlets on the same day is obtained. Construct the current retail outlet on the same day With other retail outlets The formula for calculating the similarity of predicted sales volume is: ; Where, Indicates the current retail outlets on that day With other retail outlets Similarity of predicted sales volume; Indicates the current retail outlets on the previous day With other retail outlets Similarity of historical sales between them; Indicates current retail outlets The rate of change in sales similarity due to user distribution; Indicates current retail outlets With other retail outlets The first distance between represents the maximum value function; represents the linear normalization function.
[0047] For current retail outlets , other retail outlets The first distance The larger his retail outlets With current retail outlets Similarity of historical sales volume to current retail outlets The smaller the contribution of the predicted sales similarity; As the distance increases, its value becomes smaller. When the distance is large, the contribution of the subsequent attenuation is greater than the historical sales similarity. The value of will be negative, and the value is 0, so use Indicates that, that is, from 0 and The maximum value is selected as the predicted sales similarity.
[0048] S500: Based on the sales data, analyze the sales stability of other retail outlets, and combine the predicted sales similarity to obtain the predicted sales of the current retail outlet.
[0049] For current retail outlets , based on the current retail outlets estimated by other retail outlets When predicting sales volume, it is necessary to consider the stability changes of sales volume at other retail outlets. When the sales volume at other retail outlets is relatively stable, the higher the credibility value of the prediction using similarity assessment, the greater the reference value, and the more reliable the prediction is for the current retail outlets. The higher the degree of importance in prediction.
[0050] Based on the above analysis, in an embodiment of the present invention, based on sales data, the sales stability of other retail outlets is analyzed, and combined with the predicted sales similarity, the predicted sales of the current retail outlet is obtained. Further including: First, when the sales data change difference of a single retail outlet is small, it means that the sales data is more stable; but at the same time, for the entire retail industry, there may be a decline or increase in overall sales. Therefore, the smaller the difference between the sales data change of a single retail outlet and the overall change, the higher the stability of the sales data of the retail outlet. Therefore, based on the sales data, the variance of the historical sales data of other single retail outlets and the variance of the historical overall sales data of all retail outlets are analyzed to obtain the sales stability of other retail outlets. No. The sales stability calculation formula of a product is: ; Where, Indicates other retail outlets No. The sales stability of a product, that is, predicting the current retail outlets No. Other retail outlets when selling this product No. The importance of the commodity; Indicates other retail outlets No. The variance of the historical sales data of a product; Indicates all retail outlets The variance of the historical overall sales data of a product; represents the linear normalization function.
[0051] Then, based on the sales stability, combined with the predicted sales similarity and the sales data of other retail outlets on the same day, all other retail outlets are traversed to obtain the predicted sales of the current retail outlet on the same day. No. The formula for calculating the predicted sales volume of a product on that day is: ; Where, Indicates current retail outlets No. The forecast sales volume of the product for the day; represents the number of all other retail outlets; Indicates the current retail outlets on that day With other retail outlets Similarity of predicted sales volume; Indicates the current retail outlets on that day The sum of the predicted sales similarities with all other retail outlets; Indicates other retail outlets No. The daily sales data of the commodity; Indicates other retail outlets No. The sales stability of this product indicates that other retail outlets No. The sales volume of this product is important for predicting the current retail outlets No. The importance of sales of a product.
[0052] Indicates the normalized predicted sales similarity. The larger the value, the higher the similarity of other retail outlets. Sales data for current retail outlets The greater the influence weight of the predicted sales volume, the right Weighted, the current retail outlets are finally obtained Forecasted sales for the day.
[0053] S600: Analyze the difference between the predicted sales and actual sales of the current retail outlets, obtain the data risk level of the current retail outlets, and complete data governance.
[0054] The above steps use the sales data of other retail outlets on the same day to obtain the sales forecast for the current retail outlet. Generally, in the entire data platform, retail outlets with data anomalies are rare. When calculating the data risk level of retail outlets, there are two cases: First, the current retail outlet is the one with data anomalies, while the other retail outlets are normal. In this case, the sales forecast is based on normal data. Second, the current retail outlet is the one with normal data, while the abnormal retail outlets are other retail outlets. However, because the abnormal retail outlets are rare, the role of the abnormal data is diluted when the forecast is combined with the abnormal retail outlets, and the impact can be ignored.
[0055] Therefore, it is more likely that the sales volume of the current retail outlet predicted by other retail outlets is normal sales volume. Therefore, if the predicted sales volume of the current retail outlet is significantly different from the actual sales volume of the current retail outlet, it means that the data risk level of the current retail outlet is relatively high.
[0056] Based on the above analysis, in an embodiment of the present invention, by analyzing the difference between the predicted sales volume and the actual sales volume of the current retail outlets, the data risk level of the current retail outlets is obtained, and data governance is completed. Further including: First, we analyze the absolute value of the difference between the predicted sales and actual sales of all types of goods in the current retail outlets to obtain the data risk level of the retail outlets. The formula for calculating the data risk level is: ; Where, Indicates current retail outlets the degree of data risk; Indicates current retail outlets No. The actual sales volume of the commodity on that day; Indicates current retail outlets No. The forecast sales volume of the product for the day; Indicates current retail outlets and other retail outlets The number of types of products with the same product number; This is to prevent the denominator from being 0.
[0057] The governance platform's computing resource allocation is then adjusted based on the proportional value of the data risk level. Specifically, the data governance platform initially allocates computing resources to retail outlets equally. After determining the data risk level of each retail outlet, the governance platform's computing resource allocation is adjusted based on the proportional value of the data risk level, i.e., retail outlets with greater data risk levels are allocated more computing resources.
[0058] Finally, set the data risk level threshold, which can be 0.65; determine whether the data risk level is greater than the data risk level threshold; if so, there is data anomaly in the retail outlets corresponding to the data risk level, and send an anomaly notification to notify the relevant processing personnel that more attention needs to be paid, and in the subsequent processing of the data governance platform, more resources should be allocated to finally complete data governance.
[0059] It should be noted that the embodiment of the present invention uses the current retail outlets as the analysis object for analysis and description. The same method can be used to analyze all retail outlets, and ultimately the data risk level of all retail outlets can be obtained, completing the data governance of all retail outlets.
[0060] Based on the same inventive concept as the above method, this embodiment also provides a data governance platform.
[0061] See also Figure 3 , which shows the basic composition of a data governance platform provided by an embodiment of the present invention.
[0062] like Figure 3 As shown, a data governance platform includes: a retail outlet data collection module 10, a data risk assessment module 20 and a big data processing module 30; wherein: Retail outlet data collection module 10, used to collect sales data and user trajectory data of retail outlets; Data risk assessment module 20, used to analyze the real-time data risk level of retail outlets and complete data governance; The big data processing module 30 includes a big data analysis model, which is used to perform big data processing on the managed retail outlet data.
[0063] Furthermore, the data risk assessment module 20 includes: a user distribution similarity analysis unit 21, a sales similarity analysis unit 22, a predicted sales acquisition unit 23 and a data risk degree analysis unit 24. Among them: A user distribution similarity analysis unit 21 is used to analyze the user distribution around each retail outlet based on the user trajectory data, and obtain the user distribution similarity between the current retail outlet and other retail outlets; The sales similarity analysis unit 22 is configured to obtain a sales similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets; and to analyze the historical sales similarity between the retail outlets based on the sales data, and obtain the predicted sales similarity between the current retail outlet and other retail outlets based on the sales similarity change rate; The predicted sales volume acquisition unit 23 is used to analyze the sales volume stability of other retail outlets based on the sales volume data, and obtain the predicted sales volume of the current retail outlet in combination with the predicted sales volume similarity; The data risk level analysis unit 24 is configured to analyze the difference between the predicted sales volume and the actual sales volume of the current retail outlet, and obtain the data risk level of the current retail outlet.
[0064] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0065] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A data governance method, characterized in that: The method comprises: Collect sales data and user trajectory data from retail outlets; Analyzing the user distribution around each of the retail outlets based on the user trajectory data to obtain the user distribution similarity between the current retail outlet and other retail outlets; Obtaining a sales volume similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets; Based on the sales data, analyzing the historical sales similarity between the retail outlets, and combining the sales similarity change rate to obtain the predicted sales similarity between the current retail outlet and other retail outlets; Based on the sales data, analyzing the sales stability of other retail outlets, and combining the predicted sales similarity to obtain the predicted sales of the current retail outlet; Analyze the difference between the predicted sales and actual sales of current retail outlets, determine the data risk level of current retail outlets, and complete data governance.
2. The data governance method according to claim 1, characterized in that: Analyze the user distribution around each of the retail outlets based on the user trajectory data to obtain the user distribution similarity between the current retail outlet and other retail outlets, including: Get the radius of each retail outlet r a planned area within the range, and calculating a second distance between a center point of the planned area and the retail outlet; Based on the user trajectory data, evaluating user portrait characteristics of the planning area to obtain a user group parameter vector; Obtaining user distribution around each of the retail outlets based on the user group parameter vector and the second distance; Analyze the differences in user distribution between the current retail outlet and other retail outlets to obtain the similarity of user distribution between the current retail outlet and other retail outlets.
3. The data governance method according to claim 2, characterized in that: The user group parameter vector includes average passenger flow and average consumption level.
4. The data governance method according to claim 3, characterized in that: Analyze the differences in user distribution between the current retail outlet and other retail outlets to obtain the similarity of user distribution between the current retail outlet and other retail outlets, including: The cosine similarity between the user group parameter vectors of the current retail outlet and all the planned areas corresponding to other retail outlets is analyzed, as well as the second distance difference between the current retail outlet and all the planned areas corresponding to other retail outlets is analyzed to obtain the user distribution similarity between the current retail outlet and other retail outlets.
5. The data governance method according to claim 1, wherein: Obtaining a sales volume similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets includes: Sort the user distribution similarities between the current retail outlet and the other retail outlets in ascending order based on the first distances between the current retail outlet and the other retail outlets, to obtain a user distribution similarity sequence corresponding to the current retail outlet; Calculating a difference between a previous element and a next element in the user distribution similarity sequence and a third distance between corresponding retail outlets, and counting the number of positive values of the difference; The sales volume similarity change rate of the current retail outlet is obtained by combining the absolute values of the differences corresponding to all adjacent elements in the user distribution similarity sequence, the third distance, and the number of positive differences.
6. The data governance method according to claim 1, characterized in that: Based on the sales data, the historical sales similarity between the retail outlets is analyzed, and combined with the sales similarity change rate, the predicted sales similarity between the current retail outlet and other retail outlets is obtained, including: Analyze the difference between the sales data of the same numbered product at the current retail outlet and other retail outlets, traverse all products with the same number, and obtain the historical sales similarity between the current retail outlet and other retail outlets; Obtaining the sales similarity between the current retail outlet and the other retail outlets based on the sales similarity change rate of the current retail outlet and the first distance between the current retail outlet and the other retail outlets; The predicted sales similarity between the current retail outlet and other retail outlets is obtained by combining the sales similarity of the current day and the historical sales similarity.
7. The data governance method according to claim 1, wherein: Based on the sales data, the sales stability of other retail outlets is analyzed, and combined with the predicted sales similarity, the predicted sales of the current retail outlet is obtained, including: Based on the sales data, analyzing the variance of the historical sales data of other individual retail outlets and the variance of the historical overall sales data of all retail outlets to obtain the sales stability of other retail outlets; According to the sales stability, combined with the predicted sales similarity and the sales data of other retail outlets on the same day, all other retail outlets are traversed to obtain the predicted sales of the current retail outlet.
8. The data governance method according to claim 1, wherein: Analyze the discrepancy between the predicted sales and actual sales of current retail outlets, determine the data risk level of current retail outlets, and complete data governance, including: Analyze the absolute value of the difference between the predicted sales volume and the actual sales volume of all types of goods at the current retail outlets to obtain the data risk level of the retail outlets; Set data risk level thresholds; Determining whether the data risk level is greater than the data risk level threshold; If so, there is data anomaly in the retail outlet corresponding to the data risk level, and an anomaly notification is sent to complete data governance.
9. A data governance platform, characterized in that: The governance platform includes: a retail outlet data collection module, a data risk assessment module, and a big data processing module; wherein: Retail outlet data collection module, used to collect sales data and user trajectory data of retail outlets; Data risk assessment module, used to analyze the real-time data risk level of retail outlets and complete data governance; The big data processing module, including the big data analysis model, is used to perform big data processing on the managed retail outlet data.
10. The data governance platform according to claim 9, characterized in that: The data risk assessment module includes: A user distribution similarity analysis unit is used to analyze the user distribution around each of the retail outlets based on the user trajectory data, and obtain the user distribution similarity between the current retail outlet and other retail outlets; a sales similarity analysis unit configured to obtain a sales similarity change rate of the retail outlets based on the user distribution similarity and the first distance between the retail outlets; and to analyze the historical sales similarity between the retail outlets based on the sales data, and obtain a predicted sales similarity between the current retail outlet and other retail outlets based on the sales similarity change rate; A predicted sales acquisition unit, configured to analyze the sales stability of other retail outlets based on the sales data, and obtain the predicted sales of the current retail outlet in combination with the predicted sales similarity; The data risk level analysis unit is used to analyze the difference between the predicted sales volume and the actual sales volume of the current retail outlets to obtain the data risk level of the current retail outlets.
Citation Information
Patent Citations
Data security management method and system, terminal equipment and storage medium
CN116049859A
Data management method and system based on big data analysis
CN118797528A
Retail data analysis method and device for retail store and medium
CN118822586A
Real-time processing method and system for online shop sales data
CN120123413A