An audit risk prediction method and system based on big data
By using big data-based audit risk prediction methods, a quantitative data structure and individual business benchmark values are generated, and weighting coefficients are dynamically adjusted. This solves the problem of misjudgment in store risk identification by traditional audit methods and achieves more accurate risk assessment and early warning.
Patent Information
- Application Number
- CN202610440296.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-07
AI Technical Summary
Traditional auditing methods struggle to identify unusual transactions in smaller stores that, while involving smaller sums, pose higher actual risks when dealing with stores of varying sizes within large chain retail enterprises, leading to reduced accuracy in risk warnings.
The big data-based audit risk prediction method generates a quantitative data structure and individual business benchmark values, divides groups based on similarity, dynamically adjusts weight coefficients according to the length of operation, calculates a comprehensive risk index, and outputs risk warning signals.
It improved the accuracy of audit risk warnings for stores of different operating sizes, reduced the misjudgment rate caused by traditional single-dimensional assessments, and enhanced the accuracy and comprehensiveness of risk assessments.
Smart Images

Figure CN122347478A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing systems or methods specifically applicable to administrative, commercial, financial, management, supervisory or predictive purposes, and in particular relates to an audit risk prediction method and system based on big data. Background Technology
[0002] With the rapid development of information technology and the increasing complexity of economic activities, traditional auditing methods have shown significant limitations when faced with massive amounts of data. Traditional auditing work mainly relies on the experience and judgment of auditors and manual sampling inspection. This approach is not only inefficient but also easily influenced by subjective human factors, making it difficult to detect potential audit risks in a timely manner, thus failing to effectively guarantee audit quality.
[0003] In related technologies, a data mining-based auditing system can be employed. This system automatically analyzes various types of financial data collected during the audit process by establishing data models and performs risk assessments based on preset risk thresholds, thereby helping auditors quickly identify abnormal transactions and suspicious data. This approach improves the efficiency and accuracy of auditing work to a certain extent.
[0004] However, large chain retailers need to audit thousands of their stores annually. When using the aforementioned system to assess the risk of sales data from these stores, the system's uniform assessment standards result in the same risk threshold being applied to large stores in first-tier cities and small stores in third- and fourth-tier cities. In this situation, the system struggles to identify abnormal transactions in smaller stores that, while involving smaller sums, actually carry higher risks. This reduces the accuracy of risk warnings and makes it difficult to allocate audit resources to truly high-risk areas. Summary of the Invention
[0005] This application provides a big data-based audit risk prediction method and system to improve the accuracy of audit risk warnings for stores of different operating sizes.
[0006] Firstly, this application provides a big data-based audit risk prediction method, which generates a quantitative data structure and individual operating benchmark value corresponding to each store based on the operating data of each store in multiple historical operating cycles. Calculate the similarity between the quantitative data structures of each store, and classify stores with similarity higher than a preset threshold into the same group; Based on the operating data of each store in each group during the current audit period, a group operating benchmark value corresponding to each group is generated. Calculate the deviation rate between the operating data of each store in the current audit period and the group operating benchmark value of the group to which it belongs, and obtain the group deviation value; calculate the deviation rate between the operating data of each store in the current audit period and the individual operating benchmark value, and obtain the individual deviation value; Determine the operating duration of each store; When the opening duration is less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the first weight coefficient and the weight coefficient of the individual deviation value is set as the second weight coefficient; when the opening duration is not less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the second weight coefficient and the weight coefficient of the individual deviation value is set as the first weight coefficient, with the first weight coefficient being greater than the second weight coefficient. The group deviation value and individual deviation value corresponding to each store are multiplied by their respective weight coefficients and summed to obtain a comprehensive risk index corresponding to each store. If the comprehensive risk index exceeds the preset risk threshold, a risk warning signal will be output.
[0007] By adopting the above technical solution, a quantitative data structure and individual operating benchmark values are established through the analysis of historical store operating data. Group segmentation based on similarity ensures that risk assessment considers both group and individual characteristics. The weighting coefficients of group and individual deviation values are dynamically adjusted according to the length of operation. Newly opened stores are given more consideration to the operating performance of similar stores, while stores with longer operating histories are given more consideration to their own historical performance. This differentiated weighting improves the accuracy of risk assessment. By calculating a comprehensive risk index and setting risk thresholds, the misjudgment rate caused by traditional single-dimensional assessments is reduced, and the accuracy of audit risk warnings for stores of different operating sizes is improved.
[0008] In conjunction with some implementation methods of the first aspect, in some implementation methods, based on the operating data of each store over multiple historical operating cycles, a quantitative data structure and individual operating benchmark value corresponding to each store are generated, specifically including: Obtain event records for each store across multiple historical operating periods. The event records include the event type and the time of the event. Based on the time of the event, multiple historical operating cycles are divided into event-impact cycles and normal cycles. The event-impact cycle is the operating cycle that includes the time of the event. Calculate the deviation between the operating data of each store during the event impact period and the average operating data during the normal period; The deviation values are classified according to the event type, and the median of all deviation values corresponding to each event type is used as the standard influence of the corresponding event type. Subtract the standard impact amount of the corresponding event type from the operating data of each store during the event impact period to obtain the corrected operating data for the event impact period. Based on the revised operational data during the event impact period and the operational data during the normal period, the individual operational benchmark value for each store is calculated.
[0009] By adopting the above technical solution, and identifying and quantifying the impact of special events in historical operating cycles, a more accurate benchmark for individual businesses was established. The operating cycle was divided into event-impact cycles and normal cycles. By calculating the deviation between the event-impact cycles and normal cycles, the standard impact of various events was obtained and used to correct the operating data in the event-impact cycles. This reduced the interference of special events on the calculation of the benchmark and improved its representativeness. Analyzing the corrected historical data to calculate the individual business benchmark enhanced the benchmark's reflection of the store's normal operating status.
[0010] In conjunction with some implementation methods of the first aspect, in some implementation methods, the similarity between the quantitative data structures of each store is calculated, specifically including: Calculate the difference in the same dimension of data in the quantitative data structure of any two stores; Divide the difference by the standard deviation of each dimension's data across all stores to obtain the standardized difference value for each dimension; Sum the squared standardized differences of each dimension to obtain the sum, and then take the square root of the sum to obtain the similarity between any two stores.
[0011] By employing the above technical solution and using standardized processing to calculate the similarity between stores, the impact of differences in the magnitude of data across different dimensions is eliminated by dividing the differences in data across each dimension by the standard deviation. The Euclidean distance, obtained by performing a square root operation on the sum of squares of the standardized differences, is used as a measure of similarity, improving the scientific rigor of the similarity calculation. This calculation method considers the comprehensive influence of multi-dimensional data, reducing the impact of fluctuations in single-dimensional data on similarity judgment. This standardized similarity calculation method enhances the rationality of store group segmentation.
[0012] In conjunction with some implementation methods of the first aspect, in some implementation methods, based on the operating data of each store within each group during the current audit period, a group operating benchmark value corresponding to each group is generated, specifically including: Calculate the first mean of the operating data of all stores within the group during the current audit period; Calculate the difference between the operating data of each store in the group and the first mean; The operating data of stores whose differences exceed the preset deviation value are removed to obtain the corrected group; Calculate the second mean of the operating data of stores in the corrected group, and use the second mean as the group's operating benchmark value.
[0013] By adopting the above technical solution, the first mean of all stores within the group is calculated. Samples with excessively large discrepancies are then eliminated by comparing this mean with a preset deviation value, forming a corrected group. A second mean is then calculated for the corrected group as the final group operating benchmark, improving the representativeness of the benchmark. This progressive calculation method reduces the interference of abnormal operating data from individual stores on the group benchmark, enhancing its stability. By establishing a more accurate group operating benchmark, the accuracy of risk assessment based on group comparisons is improved.
[0014] In conjunction with some implementations of the first aspect, in some implementations, after outputting the risk warning signal, the method further includes: The risk coverage rate is obtained by counting the number of stores that triggered risk warning signals during the current audit period and calculating the ratio of the number of stores to the total number of stores. If the risk coverage rate exceeds the preset coverage rate threshold, risk characteristics are extracted for each store that triggers the risk warning signal. The risk characteristics include group affiliation information and abnormal operation dimensions. Calculate the percentage of stores that triggered risk warning signals within each group to obtain the group risk density; Groups whose risk density exceeds a preset density threshold are classified as high-risk groups; For each high-risk group, stores that trigger risk warning signals are clustered based on the abnormal operation dimension to determine a subset of stores with the same abnormal operation dimension; When the proportion of stores in the determined store subset to the total number of stores in the high-risk group that triggered risk warning signals exceeds a preset consistency threshold, the high-risk group is determined to have systemic risk. Extracting common abnormal operational dimensions from a subset of stores as sources of systemic risk for high-risk groups; When the number of high-risk groups identified as having sources of systemic risk exceeds a preset group threshold, a global systemic risk is determined to exist. It outputs systemic risk warning signals, sources of systemic risk, and identification information of high-risk groups.
[0015] By employing the aforementioned technical solution, and through statistical analysis of the number of stores triggering risk warning signals and calculation of risk coverage, combined with the calculation of group risk density and the identification of high-risk groups, the system can elevate risk analysis from individual risk warnings to the group level. Based on this, cluster analysis is performed on stores within the high-risk group that share the same abnormal operational dimensions, and a consistency threshold is used to determine whether systemic risk exists. This allows for the identification of the sources of systemic risk and ultimately, the assessment of whether global systemic risk exists. This multi-layered risk analysis method improves the accuracy of risk identification, expands risk warnings from individual stores to the group and global levels, and enhances the comprehensiveness of risk warnings.
[0016] In conjunction with some implementation methods of the first aspect, in some implementation methods, after extracting the same abnormal operational dimensions of a subset of stores as the source of systemic risk for high-risk groups, the method further includes: Identify the systemic risk sources for each high-risk group, with each systemic risk source containing one or more dimensions of abnormal operations; For different high-risk groups, pairwise comparisons of systemic risk sources are made, and the correlation between any two systemic risk sources is calculated. When the correlation exceeds the preset correlation threshold, the two corresponding systemic risk sources are merged into the same risk root cause, and all abnormal business dimensions contained in the two systemic risk sources are taken as the abnormal dimension set of the risk root cause. All sources of systemic risk are iteratively merged until no source of systemic risk with a correlation degree exceeding a preset correlation threshold is found. The number of risk root causes after iterative merging is determined. When the number of risk root causes does not exceed a preset root cause threshold and the number of high-risk groups with systematic risk sources exceeds a preset group threshold, it is determined that there is a global systematic risk.
[0017] By employing the aforementioned technical solution, and through correlation analysis and iterative merging of systemic risk sources from different high-risk groups, the system can group highly correlated risk sources into a single risk root cause. During the merging process, by combining the abnormal operational dimensions of related risk sources into a set of abnormal dimensions for the risk root cause, the complete information of the risk sources is preserved. Combining preset root cause thresholds and group thresholds for global systemic risk determination reduces redundant calculations and judgments of risk root causes. This correlation-based risk root cause merging method improves the accuracy of systemic risk identification and reduces the fragmentation of risk root causes.
[0018] In conjunction with some implementation methods of the first aspect, in some implementation methods, the correlation between any two sources of systemic risk is calculated, specifically including: Extract the abnormal operational dimensions contained in the first source of systemic risk to obtain the first set of abnormal dimensions; extract the abnormal operational dimensions contained in the second source of systemic risk to obtain the second set of abnormal dimensions; Calculate the intersection of the first set of anomalous dimensions and the second set of anomalous dimensions to obtain the common set of anomalous dimensions, and count the number of dimensions in the common set of anomalous dimensions to obtain the number of intersections; calculate the union of the first set of anomalous dimensions and the second set of anomalous dimensions to obtain the complete set of anomalous dimensions, and count the number of dimensions in the complete set of anomalous dimensions to obtain the number of unions. The ratio of the number of intersections to the number of unions is used to obtain the correlation between the first and second sources of systematic risk.
[0019] By adopting the above technical solution, a quantifiable method for assessing the correlation of risk sources is established by calculating the intersection and union of the abnormal operational dimensions contained in two systemic risk sources and using the ratio of the intersection quantity to the union quantity to determine the correlation degree. This correlation calculation method based on set operations considers the degree of overlap of abnormal operational dimensions, making the correlation calculation more objective and reasonable. By including both the commonalities and differences of abnormal operational dimensions in the calculation, the distinguishability of the correlation index is improved, making it easier to judge the correlation strength between risk sources, thereby improving the accuracy of risk root cause attribution.
[0020] Secondly, embodiments of this application provide an audit risk prediction system based on big data. The audit risk prediction system based on big data includes: one or more processors and a memory; the memory is coupled to one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and one or more processors call the computer instructions to cause the system to perform the method described in the first aspect and any possible implementation of the first aspect.
[0021] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a system, cause the system to perform the method described in the first aspect and any possible implementation thereof.
[0022] Fourthly, embodiments of this application provide a computer program product that, when run on a system, causes the system to execute the method described in any possible implementation of the first aspect.
[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. This application provides a big data-based audit risk prediction method. It establishes a quantitative data structure and individual operating benchmarks by analyzing historical operating data of stores, and segments stores based on similarity, ensuring that risk assessment considers both group and individual characteristics. The weighting coefficients of group and individual deviation values are dynamically adjusted based on the length of operation. Newly opened stores are given more consideration to the operating performance of similar stores, while stores with longer operating histories are given more consideration to their own historical performance. This differentiated weighting improves the accuracy of risk assessment. By calculating a comprehensive risk index and setting risk thresholds, the misjudgment rate caused by traditional single-dimensional assessments is reduced, and the accuracy of audit risk warnings for stores of different operating sizes is improved.
[0024] 2. This application provides a big data-based audit risk prediction method. By statistically analyzing the number of stores triggering risk warning signals and calculating the risk coverage rate, combined with the calculation of group risk density and the identification of high-risk groups, the system can elevate risk analysis from individual risk warnings to the group level. Based on this, cluster analysis is performed on stores within the high-risk group that share the same abnormal operational dimensions, and a consistency threshold is used to determine whether systemic risk exists. This identifies the sources of systemic risk and ultimately determines whether global systemic risk exists. This multi-level risk analysis method improves the accuracy of risk identification, expands risk warnings from individual stores to the group and global levels, and enhances the comprehensiveness of risk warnings. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating an audit risk prediction method based on big data in an embodiment of this application.
[0026] Figure 2 This is another flowchart illustrating a big data-based audit risk prediction method in the embodiments of this application.
[0027] Figure 3 This is a schematic diagram of the physical device structure of an audit risk prediction system based on big data, provided in an embodiment of this application. Detailed Implementation
[0028] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0029] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0030] The following example is used in conjunction with Figure 1 This application describes a big data-based audit risk prediction method in its embodiments: Please see Figure 1 This is a flowchart illustrating a big data-based audit risk prediction method in an embodiment of this application.
[0031] S101. Based on the operating data of each store in multiple historical operating cycles, generate a quantitative data structure and individual operating benchmark value that correspond one-to-one with each store. Based on the operating data of each store across multiple historical operating cycles, the system generates a quantitative data structure and individual operating benchmark value corresponding to each store. Specifically, the system acquires event records for each store across multiple historical operating cycles, including event type and event occurrence time. Based on the event occurrence time, the system divides the multiple historical operating cycles into event-impact cycles and normal cycles, with the event-impact cycle being the operating cycle including the event occurrence time. The system calculates the deviation between the operating data of each store within the event-impact cycle and the average operating data within the normal cycle. The deviation values are categorized according to event type, and the median of all deviation values corresponding to each event type is used as the standard impact value for that event type. The system subtracts the standard impact value for the corresponding event type from the operating data of each store within the event-impact cycle to obtain the corrected event-impact cycle operating data. Based on the corrected event-impact cycle operating data and the operating data within the normal cycle, the system calculates the individual operating benchmark value for each store.
[0032] A quantitative data structure refers to a multi-dimensional vector used to describe the core operational characteristics of a single store. Its dimensions may include, but are not limited to, financial and operational indicators such as store sales, customer traffic, gross profit margin, inventory turnover rate, average customer transaction amount, and member repurchase rate. It may also include relatively static attributes such as store area, type of business district, and surrounding population density, aiming to comprehensively depict a store's profile. An individual operating benchmark value is an expected value representing a store's "normal" or "expected" level, calculated based on its long-term past performance, excluding the interference of special events. An event record identifies specific situations that may affect store operations; it includes at least two attributes: event type and event time. Event type categorizes the nature of the event, such as "mall promotion," "store renovation," "competitor opening," and "public holiday." An operating cycle is a fixed unit of time, such as a day, a week, or a month. An event impact cycle refers to the operating cycles that include the event's occurrence time, while a normal cycle refers to all operating cycles that do not include any known events. The standard impact measure is the generalized and standardized magnitude of the impact of a specific type of event on operating data. It is calculated by statistically analyzing the data deviations of all stores that have experienced similar events. The median is used to eliminate the interference of individual extreme cases and make it more representative. The corrected operating data for the event impact period is obtained by subtracting the standard impact measure from the original data, thereby restoring the operating level that the store might have achieved if the event had not occurred.
[0033] The first approach is based on time series decomposition and statistical correction. The system first extracts daily or weekly operating data for each store from the database over the past few years (e.g., 3 years), forming a time series. Simultaneously, the system maintains an event log recording the start and end times of various events. The system iterates through the time series of each store, dividing it into event-impact periods and normal periods based on the event log. For each event type, the system collects operating data from all stores within the corresponding event-impact period and calculates the deviation from the average data within the normal period. For example, for the "May Day Promotion" event, the system calculates the increase in sales for all stores during the May Day period compared to their average daily sales in the normal period. Then, the system uses a median function to calculate these increases, obtaining the standard impact of the "May Day Promotion." Next, the system corrects the historical data for each store: subtracting the standard impact of the corresponding event type from the data points within the event-impact period. Finally, based on the corrected complete historical data (including corrected event impact cycle data and original normal cycle data), the system can use time series forecasting algorithms such as Exponential Smoothing or ARIMA (Autoregressive Integral Moving Average) to generate the individual operating benchmark value for the store in the next audit cycle. The second approach is based on machine learning-based anomaly detection and data imputation. The system aggregates the normal cycle data from all stores and trains an unsupervised anomaly detection model, such as Isolation Forest or Variational Autoencoder (VAE). Then, this model is applied to the event impact cycle data; the model's output anomaly score or reconstruction error can be considered as the deviation value caused by the event. Similarly, the median of the deviation value is calculated by event type as the standard impact, and the data is corrected. After obtaining the corrected complete historical data, the system can train a supervised learning regression model, such as Gradient Boosting Decision Trees or Long Short-Term Memory (LSTM) networks, where historical data serves as features and future operating data serves as labels. The predicted value obtained by using this model to predict the current audit cycle is the individual operating benchmark value for that store.
[0034] S102. Calculate the similarity between the quantitative data structures of each store, and classify stores with similarity higher than a preset threshold into the same group; The system calculates the similarity between the quantitative data structures of each store, specifically including: calculating the difference of the same dimension data in the quantitative data structures of any two stores; dividing the difference by the standard deviation of each dimension data in all stores to obtain the standardized difference value of each dimension; squaring the standardized difference values of each dimension and summing them to obtain the sum value, and taking the square root of the sum value to obtain the similarity between any two stores, and classifying stores with similarity values higher than a preset threshold into the same group.
[0035] Similarity is a quantitative metric used to measure how closely two stores are similar in their operational characteristics. Standardized difference is a crucial step; it eliminates the influence of differences in units and numerical ranges across different dimensions by dividing the difference in raw data between two stores on a particular dimension by the standard deviation of that dimension across all stores. For example, sales revenue is much larger than inventory turnover rate; without standardization, the difference in sales revenue will dominate the overall similarity calculation. A preset threshold is a critical value used to determine whether two stores are sufficiently similar to be classified into the same group. Setting this threshold requires a balance between group purity and size. Too high a similarity requirement (i.e., too low a distance threshold) may result in an excessively small group size, while too low a requirement may introduce dissimilar stores, reducing the reference value of the group benchmark.
[0036] The first approach is based on hierarchical clustering. The system first constructs a quantified data structure vector for each store. Then, it calculates the standardized Euclidean distance between all stores, forming an N×N distance matrix (N is the total number of stores). Next, the system uses an agglomerative hierarchical clustering algorithm, such as Ward's linkage method, which tends to merge clusters that minimize the increase in intra-cluster variance. This process generates a tree-like cluster structure diagram. Audit analysts can visually observe this tree diagram and set a distance threshold as a cutting line based on their business understanding; horizontally cutting the tree diagram yields the final store group division. The second approach is based on density clustering. The system also first calculates the quantified data structure vector for all stores. Then, it uses the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm for clustering. DBSCAN requires two parameters: neighborhood radius (Eps) and minimum number of neighboring samples (MinPts). Here, Eps is the preset threshold. MinPts defines the minimum number of stores required to form a dense area (i.e., a group). The DBSCAN algorithm automatically groups stores with closely spaced densities into the same group and effectively identifies outlier stores that do not belong to any group.
[0037] S103. Based on the operating data of each store in each group during the current audit period, generate a group operating benchmark value that corresponds one-to-one with each group. Based on the operating data of each store in each group during the current audit period, the system generates a group operating benchmark value that corresponds one-to-one with each group. Specifically, this includes: calculating the first mean of the operating data of all stores in the group during the current audit period; calculating the difference between the operating data of each store in the group and the first mean; removing the operating data of stores whose difference is greater than a preset deviation value to obtain the corrected group; calculating the second mean of the operating data of stores in the corrected group, and using the second mean as the group operating benchmark value.
[0038] The first mean is the simple arithmetic average of a certain operating metric (such as sales revenue) for all stores within the group. It is a preliminary, uncorrected measure of group centrality. The preset deviation is a threshold used to identify extreme individuals within the group. If a store's operating data deviates from the group's first mean by more than this threshold, it is considered an outlier. These outliers are removed to prevent individual extreme cases (such as a store closing due to a fire resulting in zero sales, or a customer making an excessively large purchase due to winning a lottery) from unduly influencing the calculation of the group baseline and thus compromising its representativeness. The corrected group refers to the set of stores remaining after removing these extreme outlier stores from the original group. The second mean is the arithmetic average calculated based on this corrected group.
[0039] The first method is outlier removal based on standard deviation. For a group, the system first calculates the first mean (μ) and standard deviation (σ) of a certain operating metric (e.g., "average daily sales") for all its stores within the current audit period. Then, the system defines a preset deviation value as a multiple of the standard deviation, such as k*σ, where k is typically 2 or 3. The system iterates through each store within the group, calculating the absolute value of the difference between its operating data and the first mean μ. If this absolute difference is greater than k*σ, the store is marked as an outlier and removed. After removing all outlier stores, the system recalculates the mean of the operating data for the remaining stores in the corrected group; this result is the second mean, which is the operating baseline for the group. The second method is outlier removal based on interquartile range (IQR). The system first calculates the first quartile (Q1, the 25th percentile) and the third quartile (Q3, the 75th percentile) of the operating data for all stores within the group. Then, the interquartile range (IQR) is calculated as IQR = Q3 - Q1. The preset deviation value is implicit here, implemented by defining the boundaries of outliers. Typically, a value less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR is considered an outlier. The system identifies and removes all store data that meet this condition, forming a corrected group. Finally, the mean of the remaining store data in the corrected group is calculated as the group's operating benchmark.
[0040] S104. Calculate the deviation rate between the operating data of each store in the current audit period and the group operating benchmark value of the group to which it belongs, and obtain the group deviation value; calculate the deviation rate between the operating data of each store in the current audit period and the individual operating benchmark value, and obtain the individual deviation value. Group deviation measures a store's performance difference compared to its peers. A high positive value indicates that the store significantly outperforms its similar group, while a large negative value indicates that it performs far worse than its peers; the latter is usually a more concerning risk signal. Individual deviation measures a store's performance difference compared to its own historical performance. A high deviation (regardless of whether it's positive or negative) indicates that the store's current operating status has deviated from its inherent, stable pattern, and some unknown changes may have occurred. Deviation rate is a standardized metric, typically calculated as (actual value - baseline value) / baseline value.
[0041] S105. Determine the opening duration of each store; Direct queries based on master data management are possible. In a well-managed enterprise information system, there is typically a store master data table or service. This table stores various static information for each store, which invariably includes an opening date field. When the risk prediction process is initiated, the system only needs to execute a simple database query or API call for each store to obtain its opening date. Then, by subtracting the opening date from the current system date, the operating duration can be obtained in days, months, or years.
[0042] S106. When the opening duration is less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the first weight coefficient and the weight coefficient of the individual deviation value is set as the second weight coefficient; when the opening duration is not less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the second weight coefficient and the weight coefficient of the individual deviation value is set as the first weight coefficient. The preset duration is a benchmark set by business experts based on industry experience and company characteristics. It represents the approximate time required for a new store to stabilize operations, establish a stable customer base, and develop its own operational patterns after opening. For example, a fast-food restaurant might need 6 months to stabilize, while a large supermarket might need 18 to 24 months. The preset duration serves as the dividing line between new and established stores. New stores, due to insufficient historical data accumulation, have unstable and less valuable individual performance benchmarks; while established stores possess sufficiently long historical data, and their individual benchmarks more reliably reflect their operational capabilities. The first and second weighting coefficients are two preset values. The first weighting coefficient is greater than the second weighting coefficient. For example, the first weighting coefficient can be set to 0.7, and the second weighting coefficient to 0.3. Furthermore, to ensure the normalization of the total weight, their sum is usually equal to 1. The logic behind this step is: for stores with an opening duration shorter than the preset duration (i.e., new stores), the system considers their own historical data too short to be of limited reference value, while comparing their performance with carefully selected similar groups better illustrates their operational health. Therefore, a higher weight is assigned to group deviation values (first weight coefficient), and a lower weight is assigned to individual deviation values (second weight coefficient). Conversely, for stores with an operating duration of at least a preset duration (i.e., mature stores), the system considers them to have accumulated sufficient historical data, formed a stable operating model, and their individual operating benchmark values are highly reliable. In this case, deviations from their own historical data are more likely to reveal potential problems, so a higher weight is assigned to individual deviation values (first weight coefficient), while the weight of group deviation values is correspondingly reduced (second weight coefficient).
[0043] S107. Multiply the group deviation value and individual deviation value corresponding to each store by the corresponding weight coefficient and sum them to obtain the comprehensive risk index corresponding to each store. The overall risk index is the final risk measure for each store; it is a numerical value. The higher the index, the greater the overall abnormality exhibited by the store within the current audit period, and the higher its potential audit risk. The audit team can use this index to prioritize and screen stores requiring special attention. The formula is: Overall Risk Index = (Group Deviation Value × Group Deviation Value Weighting Coefficient) + (Individual Deviation Value × Individual Deviation Value Weighting Coefficient).
[0044] S108. If the comprehensive risk index exceeds the preset risk threshold, output a risk warning signal.
[0045] The preset risk threshold is a critical score used to distinguish between high-risk and normal stores. This threshold is set by the audit department based on its risk tolerance, audit resources, and historical experience. It can be a fixed value or dynamically adjusted. When the overall risk index exceeds this threshold, it means that the store's abnormality has reached a level requiring manual intervention. The risk warning signal is a system-generated notification designed to inform relevant personnel of potential risks at a specific store. This signal should at least include the store's identifier, the overall risk index score, the group deviation and individual deviation values that constitute that score, and relevant operational data to enable auditors to quickly understand the situation and conduct subsequent analysis.
[0046] In the above embodiments, a quantitative data structure and individual operating benchmark values are established by analyzing historical operating data of stores, and groups are segmented based on similarity, so that risk assessment considers both group characteristics and retains individual characteristics. The weighting coefficients of group deviation values and individual deviation values are dynamically adjusted according to the length of operation. Newly opened stores are given more reference to the operating performance of similar store groups, while stores with longer operating histories are given more consideration to their own historical performance. This differentiated weighting allocation improves the accuracy of risk assessment. By calculating a comprehensive risk index and setting risk thresholds, the misjudgment rate caused by traditional single-dimensional assessment is reduced, and the accuracy of audit risk warnings for stores of different operating sizes is improved.
[0047] The above embodiments detail the basic process of audit risk prediction based on big data, achieving risk assessment and early warning at the individual store level. However, in practical applications, in addition to focusing on the risk status of individual stores, it is also necessary to identify and warn of potential group and systemic risks from a more macro perspective. The following will combine... Figure 2 Another big data-based audit risk prediction method is described in the embodiments of this application: Please see Figure 2 This is another flowchart illustrating a big data-based audit risk prediction method in this application embodiment.
[0048] S201. Count the number of stores that triggered risk warning signals during the current audit period, and calculate the ratio of the number of stores to the total number of stores to obtain the risk coverage rate. This is a warning message generated by the system when the overall risk index of a single store exceeds a preset risk threshold. Stores that trigger this signal are considered high-risk stores. The total number of stores refers to all stores operated by the company. Risk coverage is a percentage that quantifies what proportion of all stores exhibit abnormalities and trigger risk warnings. This metric reflects the breadth of risk events and is the first signal to determine whether a problem is an isolated phenomenon or a widespread trend. A low risk coverage may indicate that the problem is limited to a few stores, while an unusually high risk coverage suggests a potentially broader, shared factor affecting multiple stores.
[0049] S202. If it is determined that the risk coverage rate exceeds the preset coverage rate threshold, risk characteristics are extracted from each store that triggered the risk warning signal. When the system determines that the risk coverage rate exceeds a preset coverage threshold, it extracts risk characteristics from each store that triggered the risk warning signal. These risk characteristics include group affiliation information and abnormal operational dimensions. The preset coverage threshold is a pre-defined critical percentage, such as 3% or 5%. When the risk coverage rate exceeds this value, the system considers the prevalence of the risk to have reached a level requiring group analysis. Risk characteristics are key information used to describe the specific profile of a risky store, and they include at least two aspects: group affiliation information and abnormal operational dimensions. Group affiliation information indicates which similar store group the risky store belongs to, as defined in step S102, which is the basis for group analysis. Abnormal operational dimensions specifically indicate which operational indicators deviated from the target store's warning, such as "excessive individual deviation in sales volume" or "excessively low group deviation in inventory turnover rate."
[0050] The first approach is based on secondary parsing of warning logs. In the first embodiment, when the system generates a risk warning signal, detailed diagnostic information can be stored in a structured format (such as JSON) in a log file or database. This information includes store ID, comprehensive risk index, group deviation value, individual deviation value, group ID, and the top N abnormal operating dimensions that contribute the most to the risk index and their deviation values. When S202 is triggered, the system only needs to traverse the warning logs of all risky stores within the current audit period, parse this structured data, and extract the group ID as group affiliation information and the list of abnormal operating dimensions as abnormal operating dimensions to form a new analysis dataset. The second approach is real-time correlation query and caching. The system can maintain a near real-time cache (e.g., using Redis) to store detailed risk calculation results for each store, including its group ID and deviations of each dimension. When the risk coverage exceeds the threshold, the system obtains a list of all store IDs that triggered the warning, and then uses this list to batch query and correlate the group affiliation information and the abnormal operating dimensions that caused the warning from the cache or backend database.
[0051] S203. Calculate the percentage of stores that trigger risk warning signals within each group to obtain the group risk density; The group risk density is calculated by dividing the number of stores within the group that triggered risk warning signals by the total number of stores in that group. This metric can be seen as the risk coverage rate within a specific group. It contrasts with the global risk coverage rate, helping analysts determine whether the current overall high risk is caused by a small number of risky stores across all groups, or by a large-scale risk outbreak within a few groups. A high-risk-density group means that the stores within that group are more likely to have been affected by some common negative factor.
[0052] S204. Groups whose group risk density exceeds a preset density threshold are classified as high-risk groups; High-risk groups refer to the set of stores whose probability of experiencing problems is far higher than the average, and are the most likely source of systemic risk. This step, through a simple comparative judgment, narrows the scope of analysis from all groups to a few high-risk groups, which is a prerequisite for subsequent in-depth root cause analysis.
[0053] S205. For each high-risk group, cluster the stores that trigger risk warning signals based on the abnormal operation dimension to determine the subset of stores with the same abnormal operation dimension. The abnormal operation dimension refers to the specific deviation of indicators extracted from S202 that caused the store to trigger an alert. Clustering here is a broad concept, referring to the process of grouping objects with the same attributes into one category. In this scenario, the clustering is based on the abnormal operation dimension. A subset of stores with the same abnormal operation dimension refers to the set of all stores within the high-risk group that triggered alerts due to completely identical abnormal operation indicators (e.g., all due to "low individual sales" and "low group inventory turnover"). Each such subset of stores points to a potential systemic problem affecting the group.
[0054] There are two main ways to implement this step. The first is precise grouping based on hash mapping. The system first iterates through each high-risk group. For each risky store within the group, the system obtains its list of abnormal operating dimensions. To ensure comparability, the system first sorts this dimension list (e.g., alphabetically) and then converts it into a unique string or hash value. The system uses a hash table (dictionary) with this unique string or hash value as the key and a list of store IDs as the value. The system iterates through the risky stores within the high-risk group again, adding the store IDs to the corresponding value list in the hash table based on the keys generated from their abnormal dimension lists. After the iteration, each key-value pair in the hash table represents a subset of stores with the same abnormal operating dimension. The second method is approximate clustering based on feature vectors. When there are many abnormal operating dimensions or subtle differences, exact matching may be too strict. In this case, the system can convert the abnormal operating dimensions of each risky store into a high-dimensional binary feature vector, where each bit represents a possible abnormal dimension, with a value of 1 indicating the presence of the abnormality and 0 indicating its absence. Then, for these feature vectors within each high-risk group, the system can apply clustering algorithms, such as K-Means or hierarchical clustering, using Hamming distance as the distance metric. After clustering, each cluster represents a subset of stores whose abnormal operating dimensions are very similar but not necessarily identical.
[0055] S206. When the proportion of the number of stores in the determined store subset to the total number of stores in the high-risk group that triggered risk warning signals exceeds a preset consistency threshold, it is determined that the high-risk group has systemic risks. The preset consistency threshold is a percentage, such as 30% or 50%, which defines a sufficiently significant standard. The higher this percentage, the stronger the homogeneity of the problems exhibited by the risky stores within the group. Systemic risk is an important determination made in this step. When the above percentage exceeds the consistency threshold, the system determines that the high-risk group has systemic risk, meaning that what affects the group is not just individual poor management, but one or more common, deep-seated factors at play.
[0056] S207. Extract the common abnormal operation dimensions of a subset of stores as the source of systemic risk for high-risk groups; After extracting the common abnormal operation dimensions of a subset of stores as the systematic risk sources of high-risk groups, the method further includes: determining the systematic risk sources of each high-risk group, where each systematic risk source contains one or more abnormal operation dimensions; performing pairwise comparisons of the systematic risk sources of different high-risk groups to calculate the correlation between any two systematic risk sources; when the correlation exceeds a preset correlation threshold, merging the corresponding two systematic risk sources into the same risk root cause, and using all abnormal operation dimensions contained in the two systematic risk sources as the abnormal dimension set of the risk root cause; iteratively merging all systematic risk sources until there are no systematic risk sources with a correlation exceeding the preset correlation threshold; determining the number of risk root causes after iterative merging, and determining the existence of a global systematic risk when the number of risk root causes does not exceed a preset root cause threshold and the number of high-risk groups with systematic risk sources exceeds a preset group threshold.
[0057] The calculation of the correlation between any two sources of systemic risk specifically includes: extracting the abnormal business dimensions contained in the first source of systemic risk to obtain a first set of abnormal dimensions; extracting the abnormal business dimensions contained in the second source of systemic risk to obtain a second set of abnormal dimensions; calculating the intersection of the first and second sets of abnormal dimensions to obtain a common set of abnormal dimensions, and counting the number of dimensions in the common set of abnormal dimensions to obtain the number of intersections; calculating the union of the first and second sets of abnormal dimensions to obtain a complete set of abnormal dimensions, and counting the number of dimensions in the complete set of abnormal dimensions to obtain the number of unions; and calculating the ratio of the number of intersections to the number of unions to obtain the correlation between the first and second sources of systemic risk.
[0058] Systemic risk sources are the specific causes leading to systemic risk in a high-risk group. They are directly constituted by the common abnormal operational dimensions of the subset of stores exceeding the consistency threshold in S206; it is a set of one or more abnormal dimensions. Correlation is an indicator used to measure the similarity between two different systemic risk sources. According to the detailed scheme, it is obtained by calculating the ratio of the intersection to the union of the abnormal dimension sets of the two sources. The preset correlation threshold is a value between 0 and 1. When the correlation between two systemic risk sources exceeds this threshold, they are considered to essentially point to the same problem and can be merged. Risk root causes are the final data after iterative merging, representing a class of essentially the same systemic risks. Their abnormal dimension set is the sum of the dimensions of all merged systemic risk sources. The preset root cause threshold and preset group threshold are two standards used to determine whether global systemic risk exists: the former limits the number of core problem types ultimately identified, and the latter requires that the number of groups affected by these core problems be sufficiently large.
[0059] The first approach is an iterative merging algorithm based on the adjacency matrix. 1. The system first extracts the systematic risk sources for all high-risk groups, with each source being a set of anomalous dimensions. 2. The system constructs an N×N correlation matrix, where N is the number of systematic risk sources, and each element (i, j) of the matrix is the Jaccard similarity between the i-th and j-th risk sources. 3. The system enters a loop: it searches for the maximum correlation value in the matrix. If this value is greater than a preset correlation threshold, the corresponding two risk sources are merged (creating a new set of dimensions, the union of the two source dimensions), and these two old sources are removed from the pending list and added to the new source. Then, the correlation matrix is recalculated based on the new source list. 4. This loop continues until the maximum correlation value in the matrix no longer exceeds the threshold. At this point, the remaining risk sources in the list are the final root causes of risk. 5. The system counts the number of root causes and how many different high-risk groups these root causes initially involved. 6. Finally, the system makes a judgment: if the number of root causes of risk is not greater than a preset root cause threshold, and the number of affected high-risk groups exceeds a preset group threshold, then a global systemic risk is determined to exist. The second method is a graph-based community detection algorithm. 1. The system treats each source of systemic risk as a node in a graph. 2. The system calculates the correlation (Jaccard similarity) between any two nodes. If the correlation exceeds a preset correlation threshold, an edge is connected between the two nodes. 3. This constructs a risk source correlation graph. 4. The system runs a community detection algorithm (such as the Louvain algorithm or connected component algorithm) on this graph. Each community (or connected component) in the graph directly corresponds to a root cause of risk. The set of abnormal dimensions of this root cause is the union of the dimensions of all nodes (risk sources) within the community. 5. Subsequent statistical and final judgment steps are the same as in the first method. This graph method is generally more efficient than the iterative matrix method, especially when the number of risk sources is large.
[0060] S208. When the number of high-risk groups with identified sources of systemic risk exceeds a preset group threshold, it is determined that there is a global systemic risk. The preset group threshold is an integer representing the minimum number of affected groups that management deems necessary to trigger the highest level of alert. For example, the threshold could be set to 5, meaning that if five or more different types of store groups simultaneously experience their own systemic problems, it should be considered a global risk. Even if the problems (sources of systemic risk) of each group are not identical, when the number of affected groups reaches a certain scale, this multi-point concurrent phenomenon itself constitutes a significant risk signal at the macro level. It may indicate widespread pressure or problems in the company's overall management framework, market adaptability, or operational support system.
[0061] S209. Output systemic risk warning signals, sources of systemic risk, and identification information of high-risk groups.
[0062] The first approach is to generate structured risk reports and visualize them. The system can integrate the analysis results into a multi-level interactive dashboard. The top layer of the dashboard can be a company map, highlighting the geographical areas where high-risk groups are concentrated with different colors and providing a conclusion on the overall systemic risk. Users can click on a high-risk group to drill down to the second layer to view the group's risk density, systemic risk sources (i.e., common anomaly dimensions), and a list of affected stores. At this layer, charts can also be used to show the specific deviations of these anomaly dimensions. If there are global risk root causes, the dashboard will have a dedicated module that uses Sankey diagrams or relationship diagrams to show how different sources of systemic risk are grouped into a few risk root causes. The second approach is to trigger automated notification and task allocation processes. When the system determines that there is systemic risk or global systemic risk, it can trigger a series of automated actions through an API interface according to preset rules. For example, the system can automatically generate a summary email containing all key information and send it to C-level management (such as the chief audit officer or chief operating officer). Meanwhile, for each identified source of systemic risk, the system can automatically create an investigation task in a project management tool (such as JIRA or Asana) and assign the task to the special audit team responsible for that business area (such as supply chain or marketing), attaching a list of high-risk groups and detailed data links in the task description.
[0063] In the above embodiments, by statistically analyzing the number of stores triggering risk warning signals and calculating the risk coverage rate, combined with the calculation of group risk density and the identification of high-risk groups, the system can elevate risk analysis from individual risk warnings to the group level. Based on this, cluster analysis is performed on stores within the high-risk group that share the same abnormal operational dimensions, and a consistency threshold is used to determine whether systemic risk exists, thereby identifying the source of systemic risk and ultimately determining whether global systemic risk exists. This multi-level risk analysis method improves the accuracy of risk identification, expands risk warnings from individual stores to the group and global levels, and enhances the comprehensiveness of risk warnings.
[0064] The system in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference needed]. Figure 3 This is a schematic diagram of the physical device structure of an audit risk prediction system based on big data, provided in an embodiment of this application.
[0065] It should be noted that, Figure 3The structure of the system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0066] like Figure 3 As shown, the system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 302 or a program loaded from storage portion 308 into Random Access Memory (RAM) 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.
[0067] The following components are connected to I / O interface 305: input section 306 including a camera, infrared sensor, etc.; output section 307 including a liquid crystal display (LCD) and speakers, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0068] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.
[0069] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. The transmitted data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
[0070] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0071] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or it may exist independently and not assembled into the system. The storage medium carries one or more computer programs that, when executed by a processor of a system, cause the system to implement the methods provided in the above embodiments.
[0072] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0073] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0074] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0075] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for predicting audit risks based on big data, characterized in that, include: Based on the operating data of each store in multiple historical operating cycles, a quantitative data structure and individual operating benchmark value corresponding to each store are generated. Calculate the similarity between the quantitative data structures of each store, and classify stores with similarity values higher than a preset threshold into the same group; Based on the operating data of each store in each of the aforementioned groups during the current audit period, a group operating benchmark value corresponding to each of the aforementioned groups is generated; Calculate the deviation rate between the operating data of each store in the current audit period and the group operating benchmark value of the group to which it belongs, and obtain the group deviation value; calculate the deviation rate between the operating data of each store in the current audit period and the individual operating benchmark value, and obtain the individual deviation value; Determine the operating duration of each of the aforementioned stores; When the opening duration is less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the first weight coefficient and the weight coefficient of the individual deviation value is set as the second weight coefficient; when the opening duration is not less than the preset duration, the weight coefficient of the group deviation value of the corresponding store is set as the second weight coefficient and the weight coefficient of the individual deviation value is set as the first weight coefficient, wherein the first weight coefficient is greater than the second weight coefficient. The group deviation value and individual deviation value corresponding to each store are multiplied by their respective weight coefficients and summed to obtain a comprehensive risk index corresponding to each store. If the comprehensive risk index exceeds a preset risk threshold, a risk warning signal is output.
2. The method according to claim 1, characterized in that, The process of generating a quantitative data structure and individual business benchmark value corresponding to each store based on the operating data of each store over multiple historical operating cycles specifically includes: Obtain event records for each of the aforementioned stores within the multiple historical operating cycles, wherein the event records include event type and event occurrence time; Based on the time of the event, the multiple historical operating cycles are divided into event-impact cycles and normal cycles, wherein the event-impact cycle is the operating cycle that includes the time of the event. Calculate the deviation between the operating data of each store during the event impact period and the average operating data during the normal period; The deviation values are classified according to the event type, and the median of all deviation values corresponding to each event type is used as the standard influence quantity for the corresponding event type. Subtract the standard impact amount of the corresponding event type from the operating data of each store during the event impact period to obtain the corrected operating data for the event impact period. Based on the revised operational data during the event impact period and the operational data during the normal period, the individual operational benchmark value for each store is calculated.
3. The method according to claim 1, characterized in that, The calculation of the similarity between the quantitative data structures of each store specifically includes: Calculate the difference in the same dimension of data in the quantitative data structure of any two of the stores; Divide the difference by the standard deviation of each dimension's data across all stores to obtain the standardized difference value for each dimension; The standardized difference values of each dimension are squared and summed to obtain a sum value. The square root of the sum value is then taken to obtain the similarity between any two stores.
4. The method according to claim 1, characterized in that, The process of generating a group operating benchmark value corresponding to each of the aforementioned groups based on the operating data of each store within the current audit period specifically includes: Calculate the first mean of the operating data of all stores within the group during the current audit period; Calculate the difference between the operating data of each store in the group and the first mean; The operating data of stores whose differences exceed the preset deviation value are removed to obtain the corrected group; Calculate the second mean of the operating data of stores in the corrected group, and use the second mean as the operating benchmark value of the group.
5. The method according to claim 1, characterized in that, After outputting the risk warning signal, the method further includes: The number of stores that triggered the risk warning signal during the current audit period is counted, and the ratio of the number of stores to the total number of stores is calculated to obtain the risk coverage rate. If the risk coverage rate exceeds a preset coverage threshold, risk features are extracted for each store that triggered the risk warning signal. The risk features include group affiliation information and abnormal operation dimensions. Calculate the percentage of stores within each group that triggered the risk warning signal to obtain the group risk density; A group whose risk density exceeds the preset density threshold is defined as a high-risk group; For each of the aforementioned high-risk groups, the stores that triggered the risk warning signal are clustered based on the abnormal operation dimension to determine a subset of stores with the same abnormal operation dimension; When the proportion of the number of stores in the aforementioned store subset to the total number of stores in the high-risk group that triggered the risk warning signal exceeds a preset consistency threshold, it is determined that the high-risk group has systemic risk. Extracting the common abnormal operational dimensions of the aforementioned subset of stores as the source of systemic risk for the high-risk group; When the number of high-risk groups with the aforementioned sources of systemic risk exceeds a preset group threshold, a global systemic risk is determined to exist. Output systemic risk warning signals, the sources of the systemic risks, and the identification information of the high-risk groups.
6. The method according to claim 5, characterized in that, After extracting the common abnormal operational dimensions of the subset of stores as the source of systemic risk for the high-risk group, the method further includes: Identify the systemic risk sources for each of the aforementioned high-risk groups, with each systemic risk source comprising one or more dimensions of abnormal operations; For different high-risk groups, pairwise comparisons of systemic risk sources are made, and the correlation between any two systemic risk sources is calculated. When the correlation is determined to exceed a preset correlation threshold, the two corresponding systemic risk sources are merged into the same risk root cause, and all abnormal business dimensions contained in the two systemic risk sources are taken as the abnormal dimension set of the risk root cause. All sources of systemic risk are iteratively merged until no source of systemic risk with a correlation degree exceeding the preset correlation threshold is found. The number of risk root causes after iterative merging is determined. When the number of risk root causes does not exceed a preset root cause threshold and the number of high-risk groups with the source of the systemic risk exceeds a preset group threshold, it is determined that the global systemic risk exists.
7. The method according to claim 6, characterized in that, The calculation of the correlation between any two sources of systemic risk specifically includes: Extract the abnormal operational dimensions contained in the first source of systemic risk to obtain the first set of abnormal dimensions; extract the abnormal operational dimensions contained in the second source of systemic risk to obtain the second set of abnormal dimensions; Calculate the intersection of the first abnormal dimension set and the second abnormal dimension set to obtain the common abnormal dimension set, and count the number of dimensions in the common abnormal dimension set to obtain the number of intersections; calculate the union of the first abnormal dimension set and the second abnormal dimension set to obtain the complete abnormal dimension set, and count the number of dimensions in the complete abnormal dimension set to obtain the number of unions; The ratio of the number of intersections to the number of unions is calculated to obtain the correlation between the first source of systemic risk and the second source of systemic risk.
8. A big data-based audit risk prediction system, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the system to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the system, the system performs the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product is run on the system, the system performs the method as described in any one of claims 1-7.