Tobacco Supply Compliance Supervision System Based on Consumer Characteristics Big Data
By combining federated collaborative risk control and privacy set intersection technology with graph neural networks, encrypted interaction and joint modeling of cross-departmental data are achieved, solving the problems of data silos and privacy protection in the tobacco regulatory system and improving the accuracy and security of regulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA NAT TOBACCO CORP JIANGSU PROVINCE CO
- Filing Date
- 2026-03-12
- Publication Date
- 2026-05-26
AI Technical Summary
The existing tobacco regulatory system is constrained by the data silo effect and privacy protection regulations, making it difficult to achieve cross-departmental data integration. This results in the inability to accurately identify deeply hidden illegal groups, and the existing regulatory models are prone to false alarms or failure when faced with legitimate transfers and sparse scenarios.
A compliance supervision system for tobacco supply distribution based on big data of consumer characteristics is adopted. Through a federal collaborative risk control module, a data collection and preprocessing module, a data fusion module, and anomaly detection and risk classification module, encrypted interaction and joint modeling of cross-departmental data are achieved. Combined with privacy set intersection strategy and graph neural network, joint feature vectors are generated for anomaly detection.
It enables accurate identification of abnormal signals during the distribution of goods without disclosing the original data, reducing false alarm rates, improving regulatory efficiency and security, blocking attacks in sparse data scenarios, and ensuring privacy and security while meeting data compliance requirements.
Smart Images

Figure CN122088979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco supply compliance supervision technology, and in particular to a tobacco supply compliance supervision system based on big data of consumer characteristics. Background Technology
[0002] Currently, tobacco monopoly supervision mainly relies on internal data such as orders, logistics, and monopoly licenses within the tobacco industry, using single-dimensional statistical analysis to identify abnormal operators. However, with the improvement of criminals' counter-investigation capabilities, tobacco-related illegal activities are becoming more covert, networked, and cross-regional, often involving abnormal fund flows, tax declaration fraud, or collusion between personnel from different departments. The existing regulatory system is limited by the "data silo" effect, making it difficult for tobacco authorities to obtain high-value data from external departments such as public security, taxation, and market supervision in real time. This results in the inability to build a complete chain of regulatory evidence and accurately identify deeply hidden illegal groups.
[0003] More importantly, in attempting to break down data silos and facilitate cross-departmental collaboration, there is a significant technical conflict between the need for data fusion and privacy protection regulations:
[0004] On the one hand, in order to accurately combat the outflow and illegal hoarding of genuine cigarettes, it is necessary to introduce multi-source heterogeneous data for joint modeling; on the other hand, the data held by various departments contains highly sensitive citizens' privacy and trade secrets, and laws and regulations strictly prohibit the direct exchange of raw plaintext data.
[0005] Furthermore, existing regulatory models often employ fixed physical thresholds, such as fixed distances or fixed sales fluctuations. This leads to a large number of false alarms when dealing with legitimate cross-regional transfers in chain operations. Moreover, when faced with differential attacks or piecemeal evasion of regulation using small sample sparse scenarios, they fail due to insufficient privacy protection or rigid rules, making it impossible to achieve dynamic and accurate compliance regulation while ensuring data security. Summary of the Invention
[0006] This invention aims to at least partially address one of the technical problems in related technologies. Therefore, the objective of this invention is to propose a tobacco supply compliance monitoring system based on big data analysis of consumer characteristics, in order to improve regulatory efficiency and security.
[0007] To achieve the above objectives, a first aspect of the present invention proposes a tobacco supply compliance supervision system based on big data of consumer characteristics, comprising:
[0008] The federated collaborative risk control module is configured to build a distributed monitoring network that includes federated core nodes and federated collaborative nodes. The federated center performs identity authentication and encrypted interaction of model parameters between the federated core nodes and the federated collaborative nodes, and restricts the federated core nodes and the federated collaborative nodes to only exchange encrypted intermediate parameters and not exchange original plaintext data.
[0009] The data acquisition and preprocessing module is configured to acquire multi-source heterogeneous data from the federated collaborative nodes and convert the multi-source heterogeneous data into unified federated standard semantic data based on preset semantic mapping rules.
[0010] The data fusion module is configured to align the regulatory objects of the federal core node and the federal collaborative node based on a privacy set intersection strategy, and to concatenate the encrypted tobacco end features with cross-departmental features to generate a joint feature vector;
[0011] The anomaly detection and risk classification module is configured to input the joint feature vector into a preset collaborative detection model, identify abnormal signals in the cargo delivery process and calculate the risk classification result, and generate corresponding control instructions based on the risk classification result.
[0012] To achieve the above objectives, a second aspect of this invention proposes a method for regulating the compliance of tobacco supply based on big data of consumer characteristics. This method is applied in a regulatory system comprising a federal collaborative risk control module, a data collection and preprocessing module, a data fusion module, and an anomaly detection and risk classification module. The method includes:
[0013] The federated collaborative risk control module constructs a distributed monitoring network containing federated core nodes and federated collaborative nodes. The federated center authenticates each node and controls the federated core node and the federated collaborative node to only exchange encrypted model parameters or intermediate parameters, without exchanging original plaintext data.
[0014] The data acquisition and preprocessing module collects multi-source heterogeneous data from the federated collaboration node and executes preset semantic mapping rules to convert the multi-source heterogeneous data into unified federated standard semantic data.
[0015] The data fusion module executes a privacy set intersection strategy to align the common regulatory objects of the federated core nodes and the federated collaborative nodes, and concatenates the encrypted tobacco end features with cross-departmental features to generate a joint feature vector.
[0016] The anomaly detection and risk classification module inputs the joint feature vector into a preset collaborative detection model to identify abnormal signals in the cargo delivery process, calculates the risk classification results, and generates corresponding control instructions based on the risk classification results.
[0017] To achieve the above objectives, a third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, it implements the above-described method for regulating the compliance of tobacco supply based on big data of consumer characteristics.
[0018] The tobacco supply compliance supervision system based on consumer characteristic big data, as described in this invention, effectively solves the problems of privacy leakage and insufficient regulatory accuracy in cross-departmental data fusion. Its main beneficial effects are as follows:
[0019] First, by leveraging federated collaborative risk control and privacy set intersection technology, multi-party joint modeling was achieved without exchanging original plaintext data. This not only met data compliance requirements but also fully explored the correlation value of cross-departmental data. Second, through three-flow consistency verification, graph neural network correlation detection, and a dynamic exemption mechanism based on topological stability, complex commercial disguises can be penetrated to accurately distinguish between legitimate chain operation scheduling and illegal off-site laundering, significantly reducing the false alarm rate. Finally, a dynamic noise injection mechanism based on privacy budgets was introduced to completely block differential attacks and reconstruction attacks targeting sparse data scenarios. This ensured the transparency of regulatory data while achieving high-strength privacy and security protection, resulting in a dual improvement in regulatory efficiency and security. Attached Figure Description
[0020] Figure 1 This is a schematic diagram illustrating the implementation of the tobacco supply compliance supervision system based on big data of consumer characteristics provided by this invention.
[0021] Figure 2 This is a scatter plot of tobacco-related keyword clustering based on information entropy in the tobacco supply compliance supervision system based on consumer characteristic big data provided by this invention;
[0022] Figure 3 This is a graph neural network dimensionality reduction and clustering effect diagram of the cross-departmental joint feature vector in the tobacco supply compliance supervision system based on consumer characteristic big data provided by this invention;
[0023] Figure 4 This is a three-dimensional nonlinear mapping surface diagram of the link stability coefficient, time span, and transaction frequency in the tobacco supply compliance supervision system based on big data of consumer characteristics provided by this invention.
[0024] Figure 5 This is a histogram comparing the skewed distribution of end-consumer probability between normal retail sales and abnormal centralized distribution in the tobacco supply compliance supervision system based on big data of consumer characteristics provided by this invention.
[0025] Figure 6 This is a Laplace dynamic noise probability density function evolution curve driven by different scale parameters in the tobacco supply compliance supervision system based on consumer characteristic big data provided by this invention;
[0026] Figure 7 This is a curve of the scale parameter of the tobacco supply compliance supervision system based on consumer characteristic big data, which is resistant to differential attacks and adapts to the query frequency.
[0027] Figure 8 This is a flowchart illustrating the method for regulating the compliance of tobacco supply based on big data of consumer characteristics provided by this invention.
[0028] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0029] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0030] The following describes, with reference to the accompanying drawings, an embodiment of the present invention: a tobacco supply compliance supervision system, method, and electronic device based on big data of consumer characteristics.
[0031] Example 1:
[0032] This embodiment provides a tobacco supply compliance supervision system based on big data of consumer characteristics. The system is deployed in a cloud server cluster with high-performance computing capabilities and large-scale data storage capabilities. The cloud server cluster includes multiple distributed computing nodes and physically isolated secure storage arrays.
[0033] To address the real-time processing needs of massive amounts of tobacco consumption characteristic data, the computing nodes are interconnected via a high-speed fiber optic network, and a load balancing strategy is employed to dynamically allocate computing tasks. The overall architecture is designed to completely break down data silos between different regulatory departments, while strictly adhering to national laws and regulations regarding data security and personal information protection.
[0034] like Figure 1As shown in this embodiment, the tobacco supply compliance supervision system based on consumer characteristic big data provides includes a federal collaborative risk control module, a data collection and preprocessing module, a data fusion module, and an anomaly detection and risk classification module. These four core modules constitute the basic technical foundation of this system, and through their collaborative interaction, they achieve compliance tracking and intelligent auditing of the entire lifecycle of tobacco supply distribution.
[0035] For example, the federated collaborative risk control module is configured to build a distributed monitoring network comprising federated core nodes and federated collaborative nodes. In this distributed monitoring network, each node has an independent computing engine and a trusted execution environment. The federated center performs identity authentication and encrypted interaction of model parameters between the federated core nodes and the federated collaborative nodes. The identity authentication mechanism relies on a public key infrastructure (PKI) system, with an authoritative digital certificate authority issuing a unique digital certificate to each node joining the network.
[0036] Once node authentication is successful, an encrypted communication tunnel based on a transport layer security protocol is established between the nodes. The core security mechanism of the system is to restrict the exchange between the federal core node and the federal collaborative nodes to only encrypted intermediate parameters and not to the exchange of original plaintext data. This architecture fundamentally cuts off the possibility of original sensitive data flowing through the network. The federal core node is only responsible for coordinating global regulatory tasks and distributing initialized model weights, while the federal collaborative nodes perform local calculations within their respective security boundaries using locally held tobacco-related regulatory data, and finally submit gradient update data or encrypted statistical summaries to the federal center, thereby achieving cross-departmental collaborative risk control calculations while ensuring that data does not leave the domain.
[0037] For example, the data acquisition and preprocessing module is configured to collect multi-source heterogeneous data from the federated collaboration node and convert the multi-source heterogeneous data into unified federated standard semantic data based on preset semantic mapping rules. Because different regulatory departments have adopted different database standards and data dictionaries in their long-term information technology development, the collected multi-source heterogeneous data exhibits significant differences in field definitions, data formats, and storage encoding.
[0038] The data acquisition and preprocessing module first connects to the data front-end machine of the federated collaboration node through a secure application programming interface (API) to perform timed batch extraction of incremental or full data. The extracted data enters a high-speed cache, and the system then invokes preset semantic mapping rules. These preset semantic mapping rules are essentially a vast multi-dimensional ontology knowledge graph, containing alignment standards for industry terms covering multiple fields such as public security, taxation, market supervision, and tobacco monopoly. Through deep analysis by the mapping engine, the system automatically cleans, transforms, and reconstructs the chaotic, multi-source heterogeneous data, ultimately generating federated standard semantic data with a completely unified format and unambiguous semantics, laying a solid data foundation for subsequent advanced analysis and hidden feature extraction.
[0039] Specifically, the data acquisition and preprocessing module is configured to perform the following steps when extracting data elements:
[0040] First, a sensitive word cloud is constructed using an inverted index database, and natural language segmentation algorithms are employed to extract and cluster keywords. The inverted index database enables rapid full-text structured retrieval of massive amounts of unstructured text, such as law enforcement records, public complaint letters, and cross-departmental investigation notices. The system then extracts frequently occurring tobacco-related sensitive words to form a multi-dimensional sensitive word cloud.
[0041] The system then introduces a natural language segmentation algorithm to accurately segment long texts based on a tobacco industry-specific thesaurus, remove meaningless auxiliary stop words, and perform high-dimensional data clustering on keywords with similar parts of speech or highly related semantics.
[0042] Next, the information entropy of the keyword is calculated, whereby the information entropy is calculated based on the probability of changes in the keyword's left or right neighboring words in the corpus. The specific formula for calculating the information entropy is:
[0043] ;
[0044] In the formula, This represents the information entropy value of the extracted keywords. Represents the specific keyword being calculated itself. This represents the statistical probability of different derivative words appearing in the left or right adjacent words of a specific keyword across the entire corpus.
[0045] By calculating information entropy, we can measure with great accuracy the semantic richness of a keyword in its context and the degree of freedom in information combination.
[0046] Subsequently, the system determines whether the information entropy of the keyword exceeds a preset entropy clustering threshold. If so, the keyword is extracted as a tobacco-related data element, and the unstructured text data is transformed into structured feature data. The preset entropy clustering threshold is a high-precision empirical value derived from training with massive amounts of historical regulatory text. Keywords with information entropy greater than this threshold typically mean that they, as independent tobacco-related data elements, possess an extremely large information capacity. Accurately extracting these data elements and transforming them into structured feature data greatly improves the computational efficiency and early warning accuracy of subsequent machine learning models processing complex text data.
[0047] like Figure 2 The figure illustrates the specific mathematical judgment process for feature data cleaning and extraction in the data acquisition and preprocessing module. The horizontal axis represents the sequence number of each independent keyword sample initially segmented and extracted from massive unstructured text by the natural language segmentation algorithm, while the vertical axis represents the actual information entropy value of each corresponding keyword calculated by the system's backend.
[0048] As can be seen from the data distribution pattern in the figure, the full set of keyword samples exhibits a significant two-layer discrete distribution in the two-dimensional space. A horizontal dashed line at a value of 0.5 is used as the clustering segmentation threshold baseline. This horizontal dashed line sharply divides the scatter plot into two feature regions of different colors. The bright red scatter points above the horizontal baseline represent high-value tobacco-related feature words with a calculation result greater than 0.5. These red scatter points are often accompanied by extremely rich word combination variations in context, meaning they carry highly targeted core business information such as case investigation or illegal logistics. The system will accurately extract and purify these into structured tobacco-related data elements for subsequent machine learning.
[0049] Conversely, scattered blue dots below the horizontal baseline represent ordinary interference words with a calculation result of less than or equal to 0.5. These blue dots are usually common words with no actual business significance and have extremely low information carrying capacity. The system will automatically remove them as invalid noise.
[0050] The clear separation effect of red and blue scatter points on both sides of the preset clustering baseline is achieved through the effect of red and blue scatter points. Figure 2 This study demonstrates that using information entropy theory for dimensionality reduction and key feature extraction of tobacco-related text data is highly scientific and accurate. It overcomes, to some extent, the technical shortcomings of traditional keyword matching methods that easily introduce a large number of invalid interference terms, laying an extremely pure and high-quality data foundation for subsequent cross-departmental feature fusion calculations in the entire distributed regulatory network.
[0051] For example, the data fusion module is configured to align the regulatory objects of the federal core node and the federal collaborative node based on a privacy set intersection strategy, and to concatenate the encrypted tobacco end features with cross-departmental features to generate a joint feature vector.
[0052] In real-world cross-departmental joint law enforcement and supervision scenarios, accurately identifying suspects of common concern to multiple departments while absolutely avoiding the exposure of their unique customer lists is an extremely challenging engineering problem. The privacy set intersection strategy perfectly solves this core contradiction.
[0053] The system utilizes privacy-preserving set intersection algorithms based on inadvertent transmission encryption protocols or homomorphic hash algorithms. This ensures that each server node participating in collaborative computation can only compute the common intersection portion of its own datasets, while any anonymized data outside the intersection remains completely unknown. After successfully and securely aligning the jointly monitored objects, each business node extracts local behavioral features specific to that entity. To prevent unauthorized reverse engineering attacks during network splicing and transmission, all extracted feature sequences undergo extremely rigorous asymmetric encryption. The data fusion module receives these multi-source encrypted features and performs high-level dimensional splicing between the encrypted tobacco-end features and cross-departmental features, ultimately generating a joint feature vector representing a comprehensive behavioral profile of the monitored object. This joint feature vector directly serves as the core input for subsequent multi-layer deep machine learning models.
[0054] It is also important to note that the data fusion module is configured to perform both horizontal and vertical federated fusion. In the complex business context of national tobacco supply compliance supervision, a single-dimensional federated learning architecture often cannot adequately address the data silos across different administrative levels and business departments. Therefore, in horizontal federated fusion, the core federated nodes at each level train their own neural network sub-models locally, and the global model parameters are aggregated and updated using a secure aggregation strategy.
[0055] In practical implementation, each municipal or county-level tobacco supervision and management department acts as a distributed federated core node, independently training its underlying neural network sub-model using homogeneous attribute data accumulated from its local operations. After training iterations, each business node never transmits any plaintext data of business features to the external network. Instead, it simply adds random noise masks to the locally calculated model gradient weight parameters before sending them to the upper-level federated central node. The federated central node uses a multi-party secure aggregation strategy to perform a global weighted average calculation on all reported gradient weight parameters in a completely encrypted state, thereby securely aggregating and updating the global model parameters. The updated parameters are then distributed to all participating nodes in real time, thoroughly realizing the collaborative evolution of anomaly detection performance across all levels of regulatory nodes.
[0056] Furthermore, in the vertical federated fusion, a joint feature vector is constructed, comprising tobacco-related features of the first preset dimension and cross-departmental features of the second preset dimension. The data distribution at this point reveals different attribute characteristics of the same group of regulated entities held by different national functional departments. The tobacco-related features are generated based on retailers' ordering, delivery, and inventory data. This first-hand data is authoritatively provided by the tobacco department's internal monopoly system, accurately depicting the daily operating scale, cash flow, and inventory digestion capacity of each retailer. The cross-departmental features are generated based on tobacco-related case data, unlicensed operation data, and tax record data. This cross-border data is encrypted and provided by collaborative nodes such as public security organs, market supervision administrations, and national tax authorities, providing crucial anti-counterfeiting judgment criteria from external regulatory dimensions such as criminal records, business qualifications, and compliant tax fund flows. By deeply integrating the tobacco-related features of the first preset dimension and the cross-departmental features of the second preset dimension, the vertical federated fusion process thoroughly pieces together a complete picture of the regulated entities' behavior, leaving no blind spots.
[0057] For example, the anomaly detection and risk classification module is configured to input the joint feature vector into a pre-set collaborative detection model, identify abnormal signals during the tobacco supply process, calculate risk classification results, and generate corresponding control instructions based on the risk classification results. The pre-set collaborative detection model is a composite high-level discriminant model that deeply integrates multiple cutting-edge artificial intelligence algorithm architectures. When the joint feature vector from the high-dimensional dense feature space is input into this discriminant model, the multi-layer fully connected perceptron and long short-term memory temporal analysis network within the model simultaneously and in parallel perform deep logical mining of the hidden feature interactions in the vector sequence. Whether it's deeply hidden illicit fund repatriation variations or cleverly disguised abnormal cross-regional logistics transportation trajectories, the collaborative detection model can effectively separate and distinguish them through the hyperplane classification boundary in the high-dimensional feature space, thereby accurately identifying any subtle abnormal signals during the tobacco supply process.
[0058] The subsequent calculation of risk classification results is not a simple, rigid judgment based on a single rule threshold. Instead, it involves an extremely complex comprehensive weighted assessment and scoring based on the statistical confidence level of the extracted abnormal signals, the scope of potential social impact, and the final severity of harm from similar historical illegal cases. Ultimately, based on the risk classification results (high, medium, low, etc.) calculated from the assessment, the system's automated rule engine automatically generates corresponding control and execution instructions. These instructions cover a variety of immediate and effective practical response measures, such as automatically suspending the merchant's subsequent allocation of scarce goods, automatically issuing on-site inspection and verification task orders with GPS location tracking, or directly pushing red alert letters to the corresponding public security economic investigation and supervision departments.
[0059] For example, the anomaly detection and risk classification module includes a three-flow consistency verification unit, which is configured to execute the following logic. Three-flow consistency verification is the core data detection method for identifying merchants' fraudulent order-brushing transactions and large-scale cross-provincial distribution of genuine cigarettes. Specifically:
[0060] The system first acquires the retailer's cash flow data, logistics and distribution data, and information flow order data. These three types of core underlying data represent the online banking payment and settlement trajectory, the spatial and geographical transfer trajectory of physical cigarette goods, and the electronic ordering demand trajectory sent by lower-level merchants to the provincial platform, respectively, during the closed-loop process of commercial transactions.
[0061] Subsequently, the system executes an automated inspection program to detect whether funds have been received but there is no logistics confirmation record. This information gap usually strongly suggests a money laundering activity disguised as a normal tobacco order but actually involving illegal dark web fund settlements, or a purely financial fund transfer scheme.
[0062] Furthermore, the system calculates the geographical distance between the member's consumption address and the actual delivery address in a spatial dimension, and determines whether this geographical distance exceeds a preset distance threshold. The preset distance threshold is the maximum tolerable distance in kilometers pre-set by the geographic information system based on the administrative geographical boundaries of the merchant's city and the reasonable local logistics and distribution coverage area for residents. If the registered consumer address or the location address of the frequently used consumption device that purchased this batch of tobacco is extremely far from the retailer's actual delivery address and spans multiple provinces and cities, it often indicates serious illegal activities by criminals using shell companies registered in different locations to conduct cross-regional sales and stockpiling.
[0063] Furthermore, the system performs cross-validation to detect situations where the information flow shows frequent purchases but the corresponding cash flow has no corresponding payment records. This serious discrepancy reveals that illegal underground groups may be using a large number of fake orders to maliciously obtain scarce cigarette supplies from tobacco companies, while the actual black market transaction funds are all used for illegal hedging operations through offline covert channels or the dark web of digital cryptocurrencies.
[0064] If the system detects that a target entity meets any of the above conditions, it determines that its transaction behavior has a serious logical flaw and immediately generates a consistency anomaly judgment signal. At the same time, it significantly increases the risk classification calculation weight of the retailer within the risk control model. This rigorous, multi-dimensional data cross-comparison logic makes any high-IQ illegal behavior that attempts to superficially disguise itself in a single logistics or financial data chain instantly exposed, greatly enhancing the overall compliance and regulatory system's robustness in the face of unknown risks.
[0065] It should also be noted that this system includes a verifiable chain of evidence module, which is configured to perform the following anti-tampering record operation mechanism to ensure that all intelligent regulatory conclusions given by the system have irrefutable and immutable legal effect at the judicial review level. Specifically:
[0066] First, the evidence chain fusion unit integrates the joint feature vector, model inference path, and verification results fed back by the federated collaborative nodes to generate an evidence package. In a fully automated, unattended monitoring environment, since the judgment and decision-making process for massive amounts of data is completed rapidly by complex, black-box machine learning models, the ability to provide a highly interpretable and transparent chain of evidence to courts or auditing institutions is particularly crucial. The evidence chain fusion unit silently binds in the background the original structured business features that triggered this risk control anomaly, the core inference path such as the decision tree node branching or the attention highlight weight allocation matrix of the neural network during the deep model's calculation process, and the digitized, anti-counterfeiting copies of conclusive investigation documents fed back by external collaborative nodes such as public security and tax authorities, forming a comprehensive evidence package that is logically closed-loop and mutually corroborating.
[0067] Next, the hash solidification unit extracts the digital digest of key information from the evidence package and generates a unique hash value. These hash values are then concatenated in chronological order to form a hash chain. The digital digest extraction process employs a high-strength, military-grade secure hash algorithm by default. Even the slightest change to a single punctuation mark or pixel in the complete evidence package will immediately and drastically alter the unique hash value generated by the algorithm. These timestamped hash values are then tightly linked together end-to-end using underlying cryptographic hash pointers according to the precise logical order of events, thus forming an unbreakable, tamper-proof hash chain on the physical storage medium.
[0068] Finally, the system stores the hash chain in a pre-built blockchain network via a smart contract protocol for law enforcement supervision and post-event traceability. The pre-built blockchain network employs a highly decentralized Byzantine fault-tolerant consensus mechanism to maintain the operation of all network nodes. Any malicious host attempting to tamper with data, or any unauthorized human intervention by internal personnel, will be instantly rejected by all other normal, honest nodes in the network, triggering an automatic repair mechanism. When frontline law enforcement personnel or higher-level disciplinary committee auditors conduct subsequent review and evidence collection, they only need to recalculate the hash value of the locally retrieved electronic evidence package using the same algorithm and compare it with the hash chain permanently recorded in the distributed blockchain network ledger. This mathematical principle ensures the original consistency and absolute authenticity of the entire electronic evidence.
[0069] Specifically, this system also includes a privacy protection and access control module, which is configured to perform tiered protection based on data sensitivity levels. Faced with the exponentially growing volume of extremely sensitive cross-departmental core government data, adopting a rigid, one-size-fits-all data security strategy would not only significantly slow down the entire cluster's processing efficiency but also make it fundamentally difficult to effectively balance the sharp contradiction between the business value generated by the free flow of data and the protection of individual privacy and security.
[0070] Therefore, for datasets with different sensitivity levels, the module employs highly differentiated cutting-edge modern cryptographic technologies in its underlying engine. For criminal suspect identity information and extremely sensitive corporate tax records, which are strictly marked as core sensitive data, the system uses homomorphic encryption for end-to-end encryption, combined with irreversible desensitization processing using partial field hiding rules. The brilliance of homomorphic encryption lies in its ability to allow distributed federated computing nodes to perform complex addition or multiplication operations directly in an unreadable ciphertext state. The entire computation process requires no key, and the result obtained by decryption at the center after the computation is complete is perfectly consistent with the result obtained by performing the same rule-based computation on the plaintext beforehand. This means that core-level sensitive data, throughout its entire long lifecycle of flowing and jointly computing across multiple network layers, will never expose its true plaintext content to any interceptor.
[0071] Meanwhile, by incorporating some field hiding rules into the data visualization output layer, such as replacing multiple sensitive digits in the middle of a citizen's real ID number with invisible asterisk characters through character masks, or forcibly physically truncating a user's detailed delivery address to a vague area street level, the security risk of core internal data being illegally spied on and stolen by relevant operators is further reduced from the perspective of physical management.
[0072] On the other hand, for daily logistics delivery schedules and store inventory records marked as general sensitive data, the system employs a differential privacy strategy to add noise, while controlling the noise intensity within a preset threshold range. The core principle of the differential privacy strategy lies in automatically injecting random numerical noise following a specific probability distribution model into the statistical query results of each response from the backend. This ensures that no matter how much external background knowledge an attacker outside the network possesses, they will absolutely be unable to accurately infer the actual single-store inventory quantity or specific logistics delivery trajectory of any particular physical retailer from the overall statistical results returned by the system.
[0073] By strictly controlling the overall injected noise intensity within a preset noise intensity threshold range, the system ensures, on a rigorous mathematical level, that the added fuzzy noise is sufficient to completely mask the real data differences and fluctuations of minute individuals, while absolutely not causing devastating disruption to the macro-level provincial and overall statistical trends and sales warning patterns. Thus, through this rigorous hierarchical protection system, the system only outputs high-dimensional macro-statistical results during cross-network segment data fusion and computation collaboration without disclosing any sensitive privacy information involving underlying individuals, perfectly achieving the highest level of security governance goal advocated by the state: data is usable but invisible.
[0074] Optionally, the system also includes a dynamic threshold and feedback learning module, which is configured to execute the following closed-loop iterative process. Because rigid and static fixed regulatory judgment rules are often extremely easy for highly intelligent counter-surveillance criminals to gradually discover the system's security warning and evasion boundaries through long-term, probing low-frequency transactions, this level of intelligent regulatory system must possess deep learning capabilities for self-evolution and dynamic adversarial adaptation at its underlying architecture.
[0075] During continuous daily operation, the system interface receives real-time verification feedback signals for the risk classification results. These feedback signals include verified violation tags confirmed by manual verification or false alarm tags with no evidence. These valuable corrective feedback signals typically originate directly from on-site investigation reports by frontline law enforcement and inspection personnel or case review conclusions by senior business experts. If the feedback signal is a false alarm tag, it fully indicates that the current version of the system's early warning rules are set too harshly and sensitively, or that the algorithm model has failed to consider the normal lag fluctuations in legitimate business operations caused by inconvenient transportation in specific remote administrative areas. In this case, the system's automatic correction engine will accurately extract the multi-dimensional spatiotemporal features of this specific false alarm scenario and store them in a specially maintained normal fluctuation feature library, and automatically increase the anomaly detection trigger threshold for this specific business scenario according to a preset adjustment ratio. The preset adjustment ratio is a learning rate penalty parameter that dynamically changes based on time decay logic. By gradually increasing the anomaly detection threshold in a stepwise manner on the mathematical model, the system can intelligently and adaptively relax the data tolerance for large-scale commercial promotions during national statutory and compliant holidays or normal year-end group purchases by large enterprises and institutions. This significantly reduces the needless waste and consumption of already strained grassroots law enforcement resources caused by massive invalid and false warnings.
[0076] Conversely, if the manual feedback signal is a genuine violation label, it strongly proves that the system has successfully and accurately captured potential serious illegal and criminal activities in the vast sea of data. To further consolidate and expand this hard-won detection achievement, the system's data flow orchestration engine will immediately label the joint feature vector corresponding to the entity with a positive sample and supplement it to the core training library to significantly strengthen the connection feature weights of the underlying recognition model. Subsequently, through the federated core node, the updated and optimized sensitive threshold parameters are sent in real time to every node terminal in the massive distributed monitoring network using a high-priority channel. This non-stop, continuously optimized feedback training closed-loop mechanism enables the collaborative detection model deployed across the network to continuously learn highly adversarial game experience from the latest investigation and handling cases day and night. Its detection accuracy for deep-seated and hidden abnormal behaviors and its keen sense of cracking down on tobacco-related illegal and criminal activities will inevitably show an exponential explosive increase over time, thoroughly helping the tobacco monopoly governance capabilities to achieve a leap from passive defense to proactive intelligent closed-loop management.
[0077] For example, the anomaly detection and risk classification module further includes a graph neural network association detection unit, which is configured to construct a cross-departmental heterogeneous graph. The cross-departmental heterogeneous graph includes nodes of retailer, suspect, taxpayer and delivery node types, as well as edges representing the association, registration matching and delivery relationships between nodes.
[0078] In reality, the intricate and highly organized counter-surveillance networks of tobacco-related criminal gangs often do not appear directly. Instead, they secretly control dozens or even hundreds of seemingly unrelated retail entities holding legitimate licenses to collect and resell goods in small, piecemeal fashion. Traditional risk control models based on one-dimensional linear statistical regression or isolated account auditing are simply ineffective in detecting these structural criminal variations hidden within deep-seated social relationships and financial networks.
[0079] The graph neural network association detection unit introduces topological concepts, constructing a high-dimensional cross-departmental heterogeneous graph in memory space to thoroughly visualize and mathematically represent the complex business relationships that were originally scattered across multiple different government databases. In this computationally massive graph parallel database, a vast number of physical tobacco retailers, suspects with any criminal registration records on the national public security police platform, various corporate entities that have filed tax returns in the industrial and commercial tax system, and backbone distribution and warehousing nodes responsible for all inter-provincial logistics transit all objectively exist as indispensable vertex entities in the multi-dimensional graph structure network.
[0080] The various interconnected behaviors that these entities exhibit in the real world, such as hidden kinship social connections, complex and continuous abnormal fund transfers, overlapping and shared addresses of registered legal entities and tax registration information, and a high degree of overlap in the time and space routes of a large number of abnormal package logistics and delivery, constitute the connecting edges used to represent the relationships between nodes, registration matching relationships, and business delivery relationships.
[0081] After the server cluster successfully completes the underlying memory construction of the aforementioned ultra-large-scale graph structure, the system main process calls specialized graph computation operators to perform high-concurrency graph convolution matrix operations on the cross-departmental heterogeneous graph. This is used to deeply explore the complex and hidden connections between multiple retailers who are essentially sharing the same behind-the-scenes suspects, or to intelligently identify and expose serious structural gaps in the information, such as the tax entity's declared scale being extremely inconsistent with the actual registered physical operating area of the underlying retailers. The core graph convolution operation continuously aggregates the high-dimensional attribute features of all surrounding neighboring nodes with different hop counts for each specific node in a multi-dimensional space, enabling local minor anomalies to propagate deeply along the connected edges in a multi-hop ripple manner throughout the entire macro graph topology.
[0082] After feature enhancement and filtering through multi-layer graph deep convolutional networks, the illegal activities of criminals who attempt to deliberately conceal the true identity of their actual criminal controllers through extremely complex cross-shareholding business nominee structures or high-frequency, secretive shell company fund transfers will be instantly exposed.
[0083] For example, the system's computing engine can extremely accurately uncover dozens of grassroots retailers scattered in remote mountainous areas of different provinces, whose geographical locations have no overlap. The accounts of the ultimate controllers of the multi-layered financial flows behind these retailers all mysteriously point to the same fugitive police officer who had a major tobacco smuggling conviction ten years ago. This highly concealed and time-spanning deep network connection path is something that cannot be discovered by simply relying on traditional manual reports and visually reviewing them line by line.
[0084] Similarly, if the graph convolution matrix operation discovers during iteration that the absolute scale of the hidden cash flow under a certain core taxpayer entity has a serious structural discrepancy of orders of magnitude with the throughput capacity of the small shops indicated by the business registration information of the corresponding multiple remote physical retail households, it strongly suggests that there is an extremely high probability that the node is using tobacco transactions as a cover to carry out the crime of issuing false value-added tax invoices at exorbitant prices or committing huge-scale illegal money laundering.
[0085] Ultimately, based on the hidden correlation paths mined by graph computing algorithms and the massive structural anomalies, the system automatically merges alarms to generate a group-linked anomaly alarm signal targeting the entire network of interests. This high-level alarm signal will directly prompt national high-level regulatory departments to completely upgrade from isolated fines and crackdowns targeting individual end-user retailers to joint, cluster-style, targeted destruction operations against the entire cross-provincial organized crime black and gray industry chain, fundamentally and thoroughly enhancing the macro-level control and precise destruction effectiveness of the national tobacco monopoly compliance regulatory system.
[0086] like Figure 3 This paper demonstrates the core working mechanism and technical verification effect of the graph neural network association detection unit in the anomaly detection and risk classification module for mining group linkage anomalies. The horizontal and vertical axes in the figure represent the dimensionality-reduced principal components one and two, respectively, mapped from the high-dimensional joint feature vector extracted after graph convolution matrix operations to the two-dimensional planar space.
[0087] In the feature coordinate system, the gray scattered points represent a large number of regular and compliant retail nodes. These gray scattered points exhibit a highly dispersed and irregular diffuse reflection distribution throughout the feature space, objectively reflecting the normal commercial retail behavior of legitimate merchants who are independent and do not interfere with each other.
[0088] In stark visual contrast and data comparison, there is a cluster of bright red dots located in the upper right region of the coordinate system, near the horizontal axis value of eight and the vertical axis value of seven. These red dots represent a group of retailers who may be geographically distant and seemingly have no commercial overlap. However, in the cross-departmental heterogeneous graph constructed by the system, they secretly share the same entity with a criminal record or anomaly in the structured financial transactions at the underlying level. The multi-hop feature aggregation effect of the graph convolution algorithm generates a powerful topological centripetal force in the mathematical dimension, forcibly pulling these deeply hidden anomalous nodes tightly together and forming an extremely dense cluster of interconnected anomalous criminal nodes.
[0089] Meanwhile, the red dotted-line curve surrounding this high-risk node cluster clearly delineates the hidden topological boundary. This continuous closed boundary precisely severs the organized crime's profit-driven black and gray industry chain from innocent, legitimate businesses outside. This significant dimensionality reduction and clustering spatial transformation effect demonstrates the technical advantages of this system in using graph neural network topology to process multi-source heterogeneous government data. It overcomes the technical shortcomings of traditional isolated account review models, which cannot detect shell company structures and cross-regional, fragmented crime operations. At the system's underlying algorithmic logic, it achieves dimensionality reduction detection and precise strikes against complex tobacco-related illegal networks.
[0090] Example 2:
[0091] Based on Embodiment 1, this embodiment further elaborates on the detailed working mechanism and internal logic of the core component in the anomaly detection and risk classification module: the topology anomaly filtering unit.
[0092] In the practice of regulating the distribution of tobacco products, regulatory systems often face extremely complex business scenarios. The most typical pain point is how to accurately distinguish between legitimate modern chain business operations and illegal cross-provincial distribution of genuine cigarettes. Traditional regulatory rules are usually based on a single physical distance threshold. For example, when the geographical distance between a member's consumption address and the actual delivery address exceeds a preset distance threshold, such as 50 kilometers, the system will mechanically determine it as a violation.
[0093] However, this rigid judgment logic is highly susceptible to generating massive false alarms when dealing with large chain convenience stores or supermarkets that have unified ordering and settlement but distributed across different locations, severely disrupting normal commercial circulation. At the same time, simply removing this distance restriction would allow criminals to quickly register numerous shell chain companies, using false headquarters ordering guise to hoard and resell scarce goods in different locations, creating a regulatory loophole.
[0094] To completely resolve this technical paradox, this embodiment introduces a topology anomaly filtering unit. This unit is configured to execute a rigorous false alarm suppression process when the geographical distance exceeds a preset distance threshold. This process no longer relies solely on the absolute distance in physical space, but delves into the information topology of transaction relationships and the discrete entropy value dimension of end-sales. Through multi-dimensional dynamic verification, it achieves intelligent exemption for legitimate chain operation activities and precise detection of hidden violations.
[0095] For example, the first stage of the false alarm suppression process is to construct a payment and delivery relationship subgraph. When the system's underlying screening engine detects that a retailer's order has triggered a geographical distance exceeding the limit alarm, the topology anomaly filtering unit does not immediately issue a penalty instruction, but instead activates the graph data mining engine. This engine uses the retailer as an anchor point and performs deep backtracking in a massive historical transaction database to extract all payment and delivery records for that retailer within a preset historical period. The preset historical period is typically set to the past twelve months or longer to ensure coverage of complete business quarterly and annual cycles. The system treats each extracted payment transaction as a payment node and each delivery receipt location as a delivery node, establishing a directed relationship from payment nodes to delivery nodes. This directed relationship represents the essential business logic of "who is buying" and "who is receiving the goods."
[0096] By aggregating these points and edges, the system constructs a payment and delivery relationship subgraph in memory, centered on the retailer. Within this subgraph, the system can clearly identify the specific payment and delivery connection path that triggered the anomaly and mark it as a risk edge to be verified. This step restores the isolated single-violation alarm to its overall business transaction network context, providing the necessary data structure foundation for subsequent stability analysis.
[0097] The second stage of the process involves calculating the link stability coefficient, a crucial mathematical tool for distinguishing between temporary resale relationships and long-term, stable business partnerships. In real-world business logic, legitimate transfer relationships between the headquarters and branches of a chain enterprise typically exhibit long-term continuity and regularity, while shell links used for illegal resale are often short-lived or exhibit highly unstable, pulse-like transaction frequencies. Based on this pattern, the unit module calculates the link stability coefficient based on the time span of the connection between the payment node and the delivery node, as well as the historical transaction frequency.
[0098] In this calculation model, the link stability coefficient is positively correlated with the time span, meaning the longer the cooperation period, the higher the trust level. At the same time, the coefficient has a non-linear mapping relationship with the frequency of historical transactions to prevent the system from being deceived by simply high-frequency order-brushing behavior.
[0099] Specifically, this embodiment uses the following formula to define and calculate the link stability coefficient. :
[0100] ;
[0101] In the formula, This represents the final calculated link stability coefficient; the higher the value, the more stable the business credibility of the payment and delivery relationship. This represents the number of natural days that have passed since the system first recorded the establishment of the connection between the payment node and the delivery node, using a logarithmic function. The physical meaning of processing it is to recognize the value of time accumulation (positive correlation), but as time goes on indefinitely, the marginal trust gain it brings will gradually level off, which is in line with the accumulation law of business reputation. It is a natural constant; This represents the convergence factor, a preset model hyperparameter used to control the steepness of the function curve, i.e., the system's sensitivity to changes in trading frequency. This represents the total number of transactions successfully completed between the nodes within the historical period. This represents the preset minimum number of trust transactions threshold, which is the minimum number of interactions the system believes are required to establish a stable business relationship.
[0102] And the denominator part This constitutes a typical variation of the Sigmoid function, which serves to perform a non-linear mapping of transaction frequency:
[0103] When historical transaction count Much smaller than the threshold When the denominator approaches infinity, the overall coefficient becomes... Approaching zero means that occasional, low-frequency cross-regional delivery activities will be judged as extremely unstable; when the number of transactions... Once the threshold is exceeded, the denominator quickly converges to 1, at which point the overall coefficient... It is mainly determined by the time span within the molecule.
[0104] This design cleverly filters out abnormal reselling behavior that involves frantically placing orders in a short period but lacks a long-term cooperative history. The system calculates... After determining the value, compare it with the preset stability threshold.
[0105] like Figure 4 This demonstrates the underlying mathematical logic and dynamic judgment effect of the topology anomaly filtering unit in assessing the stability of business partnerships. The horizontal axis in the figure represents the total number of successful transactions completed by payment and delivery nodes within a historical period; the vertical axis represents the number of natural days elapsed since the initial connection transaction; and the vertical axis represents the link stability coefficient ultimately calculated by the system.
[0106] From the spatial waveform transformation and color distribution of the 3D surface, it can be clearly observed that the entire mapped surface exhibits a composite feature of significant step-like and smooth growth. In the range of low transaction counts on the horizontal axis, i.e., below the preset minimum trust transaction count threshold (as shown by the horizontal axis value 15 in the attached figure), the surface appears dark blue and has an extremely low height, almost completely touching the bottom plane. This objectively indicates that if the shipping and receiving parties are only engaging in occasional or low-frequency cross-regional delivery activities, regardless of the time span of their connection establishment, the system will utilize the step-like characteristics of the nonlinear mapping to forcibly suppress its stability coefficient to an extremely low level, thus extremely sensitively filtering out those shell links that are frantically generating orders in a short period or have been dormant for a long time but suddenly engage in illegal reselling.
[0107] Once the number of transactions on the horizontal axis exceeds this trust threshold, the waveform of the surface rapidly rises in an extremely steep manner, and the color of the surface smoothly transitions from dark blue to light blue, yellow, and finally to dark red at the highest point. At this point, the time span represented by the vertical axis begins to play a core and dominant role.
[0108] As the surface extends along the vertical axis, it can be seen that the height of the red surface, representing the stability coefficient, exhibits a smooth logarithmic growth trend as the number of days increases. In the early stages of business cooperation, trust accumulates rapidly, resulting in a steeper upward slope for the red surface. However, as the cooperation continues indefinitely, due to the diminishing marginal gain of commercial reputation, the upward trend at the top of the surface gradually flattens out.
[0109] Figure 4 This three-dimensional curved surface nonlinear mapping mechanism, driven by both time span and transaction frequency, not only closely aligns with the accumulation and sedimentation patterns of genuine business reputation, but also provides data verification support for the system to accurately distinguish between impulsive illegal reselling and long-term stable compliant chain transfers through highly intuitive graphical color and curvature waveform indicators. This enhances the regulatory system's ability to prevent false alarms and its precision strike effectiveness when facing highly disguised and complex business scenarios.
[0110] like If the value is less than the preset stability threshold, it indicates that the relationship is fragile and suspicious. The system will directly maintain the consistency anomaly judgment and will not perform further calculations, thereby saving computing resources.
[0111] Optionally, if the link stability coefficient If the value is greater than or equal to the preset stability threshold, the process enters a more in-depth third stage, namely, calculating the end-consumer entropy. This is the core of this system's function to prevent shell chains from engaging in fraudulent activities.
[0112] Even if the payment and delivery relationship appears stable in the long term, it's possible that criminals might register a legitimate shell company specifically for receiving tobacco from other locations and conducting illegal wholesale. To uncover this deception, the system must track the destination of the goods after they arrive. If the goods are received by a compliant chain store, they will inevitably be sold piecemeal to a large number of different ordinary consumers in the surrounding area; if they are received by a resale gang, the goods will inevitably flow to a few specific downstream buyers.
[0113] Therefore, the system tracks the terminal sales data associated with the delivery nodes and calculates the terminal consumption entropy based on the distribution of the number of independent consumers within a preset monitoring period. The terminal consumption entropy is used to quantitatively characterize the dispersion of the sales of goods.
[0114] Specifically, this embodiment uses a formula based on an improved version of the Shannon entropy principle to calculate the end-consumer entropy. :
[0115] ;
[0116] In the formula, This represents the calculated end-consumer entropy value; This represents the total number of independent consumer entities that make a purchase within a preset monitoring period after the goods are received at the delivery address. The counting variable in the summation operation represents the first... An independent consumer; Representing the The specific quantity of tobacco purchased by each consumer during this period; This represents the total quantity of tobacco sold by this distribution node during the specified period. Represents the natural logarithm operation; in the formula This item actually calculates the first... The probability of the proportion of goods purchased by a consumer to the total sales volume.
[0117] According to the principle of entropy increase, when goods are sold to a large number of different consumers in an extremely uniform and scattered manner, that is... Very large, and each The values are all small and close, the probability distribution tends to be uniform, and the calculated entropy value... It will tend towards the maximum value, which is very much in line with the business characteristics of real retail stores that accumulate small amounts into large ones.
[0118] Conversely, when goods are sold in bulk to a few large customers, that is... Very small, or a few If the value is particularly large, the probability distribution will exhibit extreme skewness, leading to a change in the calculated entropy value. The sharp decline and near-zero levels exhibit typical characteristics of large-scale wholesale or reselling.
[0119] Through the above formula, the system successfully transforms the abstract concept of retail authenticity into a calculable and comparable numerical indicator.
[0120] like Figure 5 This visually demonstrates the core judgment principle of this system in preventing shell chain merchants from illegally hoarding and reselling goods in different locations within the topology anomaly filtering unit. The horizontal axis in the figure represents the percentage of tobacco purchased by a single independent consumer within a preset monitoring period relative to the total sales volume of that delivery node, while the vertical axis represents the probability density distribution of this purchase ratio in the overall statistics.
[0121] from Figure 5The waveform transformation and color contrast of the graph clearly reveal two distinct characteristics of business behavior. The blue histogram and its corresponding dark blue fitted curve represent the end-consumer distribution of compliant and authentic chain retail branches. This blue curve exhibits a typical high-dispersion, long-tail distribution, with its probability density peak concentrated in the extremely low proportion range of 1% to 2% on the horizontal axis. This objectively reflects the characteristic of most ordinary citizens in real retail scenarios engaging in small, scattered purchases that accumulate over time. At this point, the end-consumer entropy value calculated by the system through the underlying algorithm approaches its maximum.
[0122] In stark contrast is the red histogram and its corresponding dark red fitted curve, which represents the consumer distribution of illegal merchants suspected of abnormal distribution or laundering. Its waveform exhibits an extreme unimodal skewed distribution, with the probability density peak abruptly concentrated in the high proportion range of about 20% on the horizontal axis.
[0123] This proves that the batch of tobacco goods did not actually flow into the retail terminal after arriving at the so-called branch, but was instead monopolized and purchased by a very small number of specific large-scale buyers, exhibiting typical characteristics of bulk wholesale and illegal resale. At this time, the end-consumer entropy value calculated by the system will experience a precipitous drop.
[0124] The above two probability density distribution curves demonstrate a significant spatial separation effect between low-proportion retail areas and high-proportion monopoly areas on the horizontal axis. Figure 5 The unique end-consumer dispersion verification mechanism of the proof system possesses extremely rigorous mathematical logic. This mechanism successfully transforms the abstract retail authenticity into a clearly visible probability skewed waveform indicator, thereby giving the compliance supervision system a powerful technical capability to accurately identify complex empty shell disguises and implement dynamic exemptions or severe crackdowns.
[0125] For example, the final stage of the process is to perform a dynamic exemption determination. The system will calculate the end-consumer entropy. A final comparison is performed with a preset dispersion threshold. This dispersion threshold is a baseline derived from training on the average operating data of normal retail stores in the region. If the end-consumer entropy... If the dispersion exceeds the preset threshold, the system derives the following conclusion from the data: Although the goods underwent a long-distance cross-regional physical displacement and the payer and the consignee were in different locations, their transaction relationship was long-standing and stable. More importantly, the goods were indeed sold to local ordinary people in a real and scattered manner, and no illegal secondary distribution occurred.
[0126] Based on this conclusive evidence, the system determines that the current geographical distance deviation is a normal logistics scheduling behavior in compliant chain operations, and then performs a whitelist operation, automatically cancels the consistency anomaly judgment signal, and may mark the merchant as a low-risk compliant chain merchant, and give it more lenient credit policies in future supervision.
[0127] It should also be noted that if the terminal consumption entropy If the deviation is less than or equal to the preset dispersion threshold, the conclusion is quite the opposite: although the merchant operates under the guise of a "chain store" and maintains the illusion of operation for a considerable period, the goods received are not actually sold at retail; instead, they are transferred in whole orders or in large quantities, posing a very high risk of off-site hoarding or resale. In this case, the so-called stable relationship becomes evidence of organized crime. Therefore, the system will not only not exempt the merchant from the anomaly, but will also maintain the aforementioned consistent anomaly judgment signal and, based on the severity of the excessively low entropy value, increase the risk classification weight by one or more levels, directly triggering a more severe on-site inspection or business suspension order.
[0128] Through the four closely linked steps described above, the topology anomaly filtering unit described in detail in this embodiment constructs a closed-loop verification logic that considers both historical relationships and current results. This greatly reduces compliance costs for compliant enterprises while exposing illegal activities that attempt to use complex business structures to conceal their true reselling intentions.
[0129] Example 3:
[0130] Based on the above Embodiment 1 and Embodiment 2, this embodiment further elaborates on the advanced core components responsible for the underlying data security defense line in the entire regulatory system.
[0131] In existing digital supervision practices of tobacco monopoly, although the system has adopted basic differential privacy strategies to add random noise to generally sensitive data, this static, fixed-intensity noise injection mechanism often exposes serious security vulnerabilities when faced with extremely complex geographical distributions of businesses. Especially in typical small-sample, sparse scenarios such as remote mountainous areas or highly dispersed rural commercial outlets, a specific geographical grid may contain only a very small number of tobacco retailers. In such cases, the small, fixed-intensity random noise is simply insufficient to effectively conceal the true logistics and delivery records and accurate inventory levels of individual merchants. Attackers or unauthorized internal personnel can directly reconstruct the absolute trade secrets of these micro-level individuals through simple algebraic elimination, thereby triggering incalculable privacy leakage risks and legal compliance crises.
[0132] To fundamentally resolve the serious technical conflict arising from the excessive weakening of fixed-difference privacy noise in sparse data scenarios, leading to the failure of privacy defenses, this system's privacy protection and access control module also includes a dynamic privacy budget management unit. This unit is specifically configured to execute the following rigorous dynamic noise injection process. This process no longer treats privacy protection as a static data post-processing filtering mechanism, but creatively abstracts individual privacy as a digitally constrained resource, similar to a financial account balance, that can be precisely quantified, allocated, and consumed. Through comprehensive budget management throughout the entire lifecycle, it ensures the high availability of macro-level statistical laws while achieving absolute mathematical security for micro-level individual data.
[0133] For example, the first critical initial step in this dynamic noise injection process is to initialize the privacy budget. The system's master scheduling service automatically scans the topology of the entire network's monitored areas and allocates a total privacy budget for a preset period to data subjects within those areas. In this process, the so-called data subjects do not refer to a general, chaotic mass of underlying data streams, but are strictly defined as a collection of independent, protected objects consisting of one or more tobacco retail entities with related attributes, logistics transit warehouses in specific areas, or consumer groups in specific business districts.
[0134] The preset cycle is a cyclical time window that the system pre-sets based on the statutory assessment cycle or financial audit settlement cycle of the State Tobacco Monopoly Administration, such as a complete natural day, a natural assessment week, or a natural financial month.
[0135] The total privacy budget, theoretically speaking, represents the maximum mathematical upper limit to which the system can tolerate the exposure of a specific data subject's privacy information to the outside world within the aforementioned complete preset cycle. At the absolute zero point of each new preset cycle, the system automatically injects a complete total privacy budget into the secure isolation sandbox of each independent data subject and initializes a dedicated privacy consumption ledger in the system's underlying distributed key-value pair in-memory database, strictly resetting the initial cumulative value of the currently consumed budget to zero.
[0136] This initialization mechanism, similar to the periodic allocation of fiscal funds, ensures that even in the most severe data query attack environment, the system will only passively limit the potential privacy leakage risk to a single preset period. Once the period naturally changes, the privacy protection armor of all individuals will instantly return to the highest full-load defense state, laying a solid logical foundation for subsequent high-concurrency query verification.
[0137] After budget initialization is completed, the process immediately moves on to the second core step: calculating sensitivity and consumption. When the system's data service bus receives a macro-statistical query request from a law enforcement terminal or a higher-level audit dashboard in real time, the query parsing engine within the privacy budget dynamic management unit immediately intercepts the request and performs deep syntax tree parsing and feature object extraction to accurately calculate the global sensitivity of the query.
[0138] Global sensitivity is the core mathematical benchmark in differential privacy theory. Its physical meaning refers to the maximum possible change in the final output statistical results of a specific query request caused by arbitrarily adding or deleting even a single retailer's business record in the entire massive underlying tobacco retail database.
[0139] For example, if the current query request is to calculate the total monthly inventory of premium cigarettes at all retailers in a core business district, and the system's preset rules limit the maximum allowed inventory per store to one thousand cartons, then the global sensitivity value of this query request is strictly locked at one thousand. This is because any extreme change in the data of a single store will never have a maximum impact on the overall total volume exceeding this limit. After accurately determining the global sensitivity, the system does not blindly allocate resources uniformly. Instead, it extracts the digital identity certificate and authorization level of the system operator who initiated the query request in parallel, and allocates the privacy budget consumption for each query based on the query permissions.
[0140] In the system's permission matrix mapping table, when core data analysis experts from national or provincial tobacco monopoly bureaus with high-level authorization conduct macroeconomic analysis of the entire province, the system allocates a larger single privacy budget expenditure. This means that the subsequent fuzzy noise injected by the system is relatively small, enabling the presentation of more granular and high-fidelity statistical trends. Conversely, if grassroots grid inspectors conduct general cross-regional data access, the system allocates an extremely small single privacy budget expenditure, resulting in data with a strong fuzzy random mask. This innovative mechanism, which directly links digital identity permissions with mathematical privacy consumption, perfectly balances the efficient and compliant flow of data elements with the absolute and robust protection of the privacy of underlying individuals.
[0141] Optionally, after the sensitivity and single-use cost calculations are completed, the process irreversibly proceeds to the third line of defense: budget checking. This is the highest-priority veto interception mechanism in the entire dynamic noise injection process. The privacy budget dynamic management unit securely reads the accumulated consumed budget value for the target data subject from the distributed memory database using a read-only lock mechanism. Subsequently, the system's central arithmetic logic unit performs a rigorous parallel addition and comparison operation to determine whether the sum of the consumed budget and the current privacy budget cost exceeds the total privacy budget. If it does, the query is rejected.
[0142] This seemingly simple algebraic interception logic actually embodies the highest security philosophy for combating continuous data snooping attacks. In actual attack-defense network scenarios, malicious attackers often use low-privilege accounts to perform massive, intensive, high-frequency automated replay query requests. They attempt to completely cancel out the superimposed random noise by mathematically averaging the thousands of identical results returned by the system, which contain slight noise. In this way, they can perfectly reconstruct the underlying sensitive inventory or transaction data through the law of large numbers. This is the infamous averaging differential attack in the industry.
[0143] However, due to the rigidity of the budget check process, every probing query by a malicious attacker will irreversibly and permanently deduct the privacy budget balance from the data subject's account. The moment the cumulative deduction from frequent malicious queries approaches or exceeds the total privacy budget security threshold, the system immediately triggers a circuit breaker mechanism. This not only returns a cold, denied access error code to the caller, refusing to provide any misleading or forged data, but also simultaneously writes the source network protocol address and digital signature of the abnormal request into the system's underlying security alert log, directly blocking all subsequent query permissions from that request source within the current preset period. This robust, rigid circuit breaker mechanism completely shatters any illusions of illegal mathematical reconstruction that attempt to exploit time to gain privacy data.
[0144] Specifically, if the budget check successfully confirms sufficient balance, the system will perform the formal budget deduction and then transfer the processing pipeline to the fourth data synthesis stage, which generates dynamic noise. Traditional static data anonymization often uses fixed-bit truncation or fixed pseudo-random number replacement, which is easily cracked by reverse engineering. The randomization engine in this unit, however, uses high-dimensional mathematical theory and the Laplace mechanism to generate dynamic noise values and add them to the query results. The Laplace mechanism is a noise generator built on an absolutely rigorous mathematical probability distribution. The noise data it generates is not random jumble, but strictly follows a Laplace continuous probability density distribution with an expected value centered at zero.
[0145] To achieve accurate generation of dynamic noise, the system incorporates the following Laplace probability density formula:
[0146] ;
[0147] In the formula, This represents the random noise variable that the system will inject into the original query results. Its value range covers the entire real number axis. This means that the injected noise may be a positive value that causes the statistical results to be artificially high, or a negative value that causes the statistical results to be artificially low. This ensures that the mathematical expectation of the macro statistical total always approaches the absolute true value in a large number of independent queries. Represents the random noise variable The absolute probability density distribution value of the occurrence; This represents performing absolute value mathematical operations on random noise variables to ensure that the decay characteristics of the exponential function exhibit perfect axial symmetry on both sides of the positive and negative number axes. This is the core mathematical variable controlling the steepness or flatness of the entire Laplace distribution, i.e., the scale parameter used to generate the dynamic noise value. It is the dynamic change of this scale parameter that determines the breadth and masking ability of each injected noise.
[0148] like Figure 6 This demonstrates the underlying mathematical mechanism by which the privacy protection and access control module automatically adjusts the strength of data masking in the face of different query adversarial environments. The horizontal axis in the figure represents the random noise values that the system will inject into the original macroscopic statistical results, and the vertical axis represents the probability density distribution of these random noise values.
[0149] In the feature coordinate system, three colored probability density distribution curves with distinctly different sharpness and coverage are presented. The dark blue distribution curve represents the noise generation state in a normal legitimate query scenario. In this case, due to sufficient privacy budget expenditure per query and extremely low historical query frequency, the scale parameter calculated by the system is extremely small. The dark blue curve exhibits an extremely steep peak shape, with most probability densities tightly converged and concentrated near the value of zero on the horizontal axis. This objectively indicates that the noise interference injected by the system is extremely slight, ensuring high fidelity of normal macroscopic statistical trends.
[0150] The orange distribution curve represents a transitional defense state when query sensitivity is moderate or query frequency begins to rise. Its waveform peaks begin to decline significantly, and the tails on both sides gradually expand outwards on the horizontal axis. The dark red distribution curve represents the highest defense meltdown state when the system encounters extremely high-frequency, intensive probing or suspected differential attacks. At this point, the attack penalty mechanism within the system is instantly triggered, the scale parameter expands exponentially, and the waveform of the dark red curve undergoes a drastic flattening transformation. The original spikes completely collapse, and its probability density extends extremely far to both sides of the horizontal axis in an extremely gentle manner.
[0151] This waveform transformation powerfully demonstrates that, under extreme scenarios of continuous differential attacks, the dynamic noise generated by the system has a very high probability of randomly taking huge values of tens or even hundreds, thereby thoroughly polluting and overwhelming the real inventory or outbound data of a specific single store by actively generating extremely exaggerated and violent fluctuations.
[0152] The above three curves demonstrate significant morphological evolution effects driven by different scale parameters. Figure 6 This demonstrates that the dynamic noise injection mechanism possesses extremely sensitive anti-attack adaptive expansion capabilities, which to some extent overcomes the fatal flaw of static fixed noise failing to defend against sparse data scenarios.
[0153] The scaling parameter used to generate the dynamic noise value is calculated based on the global sensitivity, the single privacy budget cost, and the query frequency. The system will then generate the final random noise variable. The algebraic addition is directly superimposed onto the original statistical values calculated by the database engine. Then, the superimposed contamination result, which is protected by multiple mathematical layers, is returned to the query requester terminal as the final security response message, completely cutting off any physical possibility of the original data being exposed.
[0154] For example, in order to make the above-mentioned dynamic noise generation process truly intelligent and resilient against attacks, the most core technical breakthrough of this embodiment lies in redefining the key factors that determine the noise intensity.
[0155] The calculation logic of the scale parameter is configured such that the scale parameter is positively correlated with the global sensitivity, negatively correlated with the single privacy budget cost, and increases as the data subject is queried more frequently within a preset time period.
[0156] In order to convert the above logical literals into machine instructions that can be used by the central processing unit for high-speed floating-point operations, the system internally hard-coded the following formula for calculating the scale parameter:
[0157] ;
[0158] In the formula, It accurately represents the global sensitivity of a specific query operation, calculated by the system through rigorous evaluation during the query parsing step. It accurately represents the amount of privacy budget consumed per instance during the authentication phase, which is precisely divided and allocated based on the permission level of the operator requesting the terminal. This represents the attack penalty coefficient pre-configured in the global security policy configuration file. It is a positive-zero empirical constant weight specifically used to adjust the sensitivity of the system's defense and counterattack. This represents the total count of query frequencies of the target data set, which is recorded and dynamically updated in real time by the system memory state machine, within a preset time period by various concurrent requests.
[0159] From the construction of this complete formula, a rigorous mathematical mapping logic can be clearly derived: the numerator of the formula contains... This determines that the scale parameter is positively correlated with the global sensitivity, meaning that if the fluctuation limit of the queried business indicator is larger, the system will intelligently amplify the absolute range of the generated noise to ensure the effectiveness of data masking; the denominator of the formula includes This determines that the scale parameter is negatively correlated with the single privacy budget expenditure; that is, the smaller the single privacy expenditure allocated by the system, the lower the authorization level. In this case, the denominator becomes smaller, leading to a decrease in the overall scale parameter. It expands rapidly, generating extremely loud noise that interferes with the vision of those with lower privileges.
[0160] The multiplier factor in the above formula includes This determines that the scaling parameter increases with the frequency of querying the data subject within a preset time period. If a target retailer or specific grid encounters frequent and continuous queries within a short period, this multiplier factor will surge exponentially, instantly increasing the scaling parameter. Magnified to astronomical figures.
[0161] like Figure 7 This provides an intuitive and insightful understanding of the underlying defense and counterattack mechanism of the privacy budget dynamic management unit in the face of malicious differential attacks. The horizontal axis in the figure represents the total count of queries made by external systems or operators to the target data set within a preset time period, while the vertical axis represents the core scale parameter calculated by the system to generate dynamic Laplace noise. The absolute value of this parameter directly determines the intensity of the random noise masking superimposed on the actual statistical results.
[0162] In the feature coordinate system, the smooth curve in dark blue represents the baseline of parameter evolution under normal legal query scenarios. When the query frequency of the horizontal axis is in the low-frequency safe range of less than ten times, the dark blue curve is almost completely close to the bottom of the coordinate axis and maintains horizontal extension. This means that the noise injected by the system at this time is extremely weak, thereby ensuring the high fidelity of macro statistical data and business availability.
[0163] However, when the system encountered a ensemble differential attack by a malicious attacker attempting to average out noise through massive, dense queries, the deep red highlighted curve displayed a stunning dynamic expansion defense waveform. As the query frequency on the horizontal axis continued to accumulate and exceeded the tolerance range of normal business operations, the deep red curve instantly broke away from its initial flat state and erupted with an extremely dramatic exponential upward surge.
[0164] When the number of queries on the horizontal axis climbs to between forty and fifty, the scale parameter value corresponding to the dark red curve on the vertical axis exhibits a geometric progression of hundreds or even nearly two thousand times, forming an extremely steep defensive barrier on the right side of the graph. This nonlinear waveform expansion effect, based on adaptive amplification of query frequency, demonstrates that this monitoring system possesses a highly sensitive anti-attack and countermeasure capability. By generating extremely large and irregular dynamic noise in an instant, it completely contaminates and destroys any mathematical possibility for malicious attackers attempting to steal micro-level individual business secrets by subtracting multiple query results, thus fundamentally constructing an unbreakable data privacy security defense.
[0165] It is also important to note that this method of using scale parameters... The dynamic decay mathematical architecture, which is forcibly bound to and intertwined with historical query frequency, has its fundamental defensive mission as blocking differential attacks. In high-level network espionage countermeasures, in addition to the averaged differential attack mentioned above, attackers often use an attack strategy called set difference subtraction.
[0166] For example, an attacker might first use a legitimate account to query the total monthly cigarette shipments of all tobacco stores in a county-level city. The system returns an overall statistical result with minimal baseline noise. Then, within a short time, the attacker queries again, requesting the total monthly cigarette shipments of all other stores in the same county-level city, excluding a specific designated store. If the system still uses static, fixed noise, the attacker can simply perform a very simple algebraic subtraction on their local terminal to precisely extract and steal the core commercial shipment secrets of that single, specific merchant within a very small, fixed error range.
[0167] However, under the dynamic computing logic deployed in this system, when an attacker initiates a second precise differential query request, the memory state machine instantly detects a sharp increase in the query frequency index for that specific region, and the attack penalty mechanism is instantly triggered. The central arithmetic logic unit will then adjust the scale parameter of the second query request according to the aforementioned mathematical formula. This noise can be amplified tens or even hundreds of times. As a result, the Laplace dynamic noise injected into the second query result will exhibit extremely exaggerated and violent fluctuations, with the absolute value of the noise potentially far exceeding the actual outbound data volume of that specific merchant. When the attacker attempts to subtract the two query results locally again, the second query result has been thoroughly and deeply contaminated and distorted by the enormous dynamic noise. The final result of the subtraction will be a jumble of mathematical garbage that completely defies business logic and has no reference value whatsoever.
[0168] The system utilizes a technique that uses dynamic noise adaptive expansion to cause the result of differential subtraction to completely lose its mathematical and statistical significance. Under the friendly premise of not closing the normal and legitimate macro-statistical query channel, it extremely accurately, powerfully, and elegantly blocks any form of micro-set differential attack attempt to steal secrets. This elevates the overall data security protection level of the tobacco supply compliance supervision system based on consumer characteristic big data to a new, indestructible technical level.
[0169] Example 4:
[0170] like Figure 8 As shown, this embodiment further provides a method for regulating the compliance of tobacco supply based on big data of consumer characteristics. This method is fully integrated into a regulatory system that includes a federal collaborative risk control module, a data collection and preprocessing module, a data fusion module, and an anomaly detection and risk classification module. Addressing the long-standing limitations of tobacco monopoly regulation due to data silos and the serious technical conflicts between data fusion requirements and privacy protection regulations during cross-departmental collaboration, as mentioned in the background, this embodiment provides a comprehensive solution. This regulatory method is logically consistent with the three system embodiments described above, and its implementation process includes the following four core execution steps:
[0171] The first step involves constructing a distributed regulatory network and implementing encrypted interaction. In this initial step, the system builds a distributed regulatory network comprising a federal core node and federal collaborative nodes through the federal collaborative risk control module. Specifically, the federal core node acts as the business hub, coordinating the overall compliance review of tobacco supply distribution; while the federal collaborative nodes are widely distributed among external government agencies such as public security, taxation, and market supervision that possess key tobacco-related information. To ensure absolute access security for this cross-departmental regulatory network, the system must immediately conduct strict digital identity authentication for each node through the federal center. After a secure handshake and authentication, the system controls the interaction between the federal core node and the federal collaborative nodes from the underlying network protocol stack, allowing only encrypted model parameters or intermediate parameters to be exchanged, without exchanging raw plaintext data. This operating mechanism is highly consistent with the absolute mathematical security concept emphasized in Example 3, completely eliminating the risk of illegal interception or unauthorized retention of citizens' privacy and trade secrets during cross-network segment plaintext transmission, perfectly resolving the core contradiction between data fusion and privacy protection at the methodological and process level.
[0172] The system then performs multi-source data acquisition and semantic mapping. Due to the varying ages and technical standards of information technology development across different tobacco regulatory departments, the underlying stored data exhibits strong fragmentation and heterogeneity. Therefore, the system collects multi-source heterogeneous data from the federated collaboration nodes through the data acquisition and preprocessing module. This multi-source heterogeneous data includes both structured high-frequency tax invoicing records and unstructured economic crime investigation transcripts. After physical acquisition, the system does not directly feed this dirty data to the upper-layer algorithm but instead enforces preset semantic mapping rules. These rules embed industry-wide unified terminology alignment standards covering multiple dimensions of government affairs, thereby converting the multi-source heterogeneous data into unified federated standard semantic data. This conversion step is crucial; it not only completely eliminates semantic ambiguity during cross-system data interaction but also paves the way for the subsequent generation of a high-dimensional feature space.
[0173] The system then proceeds to the crucial steps of privacy set intersection and feature fusion. Faced with a massive and opaque list of cross-departmental regulatory targets, a direct full comparison would inevitably expose a large amount of innocent information about uninvolved individuals. Therefore, the system employs a privacy set intersection strategy through the data fusion module. This strategy utilizes advanced cryptographic algorithms to precisely align the common regulatory targets of the federal core nodes and federal collaborative nodes within the ciphertext space, while ensuring the absolute anonymity of any entity list outside the intersection. After successfully identifying the common regulatory targets, the system extracts the business characteristics of these targets from different departments and concatenates the encrypted tobacco-related features with the cross-departmental features to generate a joint feature vector. This generated joint feature vector is a highly condensed and layered encrypted secure data matrix, essentially converging the cross-border correlation patterns of merchants' cash flow, logistics, and information flow, providing a powerful data weapon for subsequently penetrating the deep-seated commercial disguises of criminals.
[0174] Finally, the system executes anomaly identification and closed-loop control steps. The system inputs the joint feature vector into a pre-set collaborative detection model through the anomaly detection and risk grading module. This model incorporates a three-flow consistency check and a high-dimensional graph neural network topology algorithm. Upon receiving the joint feature vector, the model comprehensively scans and identifies abnormal signals during the cargo delivery process. Whether it's deeply hidden illegal fund repatriation or cleverly disguised abnormal cross-regional logistics transportation, the system can accurately quantify and determine the risk by calculating the risk grading results through a multi-dimensional weighting system.
[0175] During this process, the system also integrates topology anomaly filtering logic to accurately distinguish between legitimate chain operation scheduling and illegal off-site laundering and hoarding activities based on the dispersion of end-consumer spending, thereby significantly reducing the false alarm rate for frontline law enforcement. Finally, the system generates corresponding control instructions based on the risk classification results. These automatically generated control instructions will directly flow to the business execution layer, driving automatic reduction of supply allocation, dispatching on-site manual inspection work orders, or initiating cross-departmental joint enforcement applications, forming a perfect regulatory closed loop from anomaly detection to business disruption.
[0176] In summary, the regulatory method described in this embodiment complements and is deeply integrated with the aforementioned system hardware architecture, topology false alarm prevention design, and dynamic privacy budget technology. This method not only completely breaks down data barriers in cross-departmental collaboration and achieves multi-party security joint modeling using federated collaboration and privacy set intersection technology, but also accurately penetrates the concealed and networked disguises of tobacco-related illegal activities through intelligent anomaly detection algorithms.
[0177] Example 5:
[0178] Corresponding to the above embodiments, the present invention also proposes an electronic device.
[0179] like Figure 9 The diagram shows a structural schematic of an electronic device according to the present invention. The electronic device 100 includes a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, for example, via a bus 102. Optionally, the electronic device 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one unit, and the structure of this electronic device 100 does not constitute a limitation on the embodiments of the present invention.
[0180] Processor 101 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure. Processor 101 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0181] Bus 102 may include a pathway for transmitting information between the aforementioned components. Bus 102 may be a PCI bus or an EISA bus, etc. Bus 102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0182] The memory 103 stores a computer program corresponding to the tobacco supply compliance supervision method based on consumer characteristic big data according to the above embodiments of the present invention. This computer program is controlled and executed by the processor 101. The processor 101 executes the computer program stored in the memory 103 to implement the content shown in the aforementioned method embodiments.
[0183] Among them, electronic devices 100 include, but are not limited to: mobile terminals such as laptops and PADs (tablet computers) and fixed terminals such as desktop computers. Figure 9 The electronic device 100 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.
[0184] Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art will understand the scope of the present invention.
[0185] The above embodiments can be varied, modified, replaced, and modified.
Claims
1. A tobacco source supply compliance monitoring system based on consumption characteristics big data, characterized in that, Comprise: A federal collaborative risk control module configured to build a distributed supervision network comprising a federal core node and a federal collaborative node, perform identity authentication on the federal core node and the federal collaborative node through a federal center, and perform encrypted interaction of model parameters between the federal core node and the federal collaborative node, and limit the interaction between the federal core node and the federal collaborative node to only intermediate parameters after encryption without exchanging original plaintext data; A data acquisition and preprocessing module configured to acquire multi-source heterogeneous data from the federal collaborative node, and convert the multi-source heterogeneous data into unified federal standard semantic data based on a preset semantic mapping rule; A data fusion module configured to align the supervision objects of the federal core node and the federal collaborative node based on a privacy set intersection strategy, and splice encrypted tobacco end features and cross-department features to generate a joint feature vector; An anomaly detection and risk grading module configured to input the joint feature vector into a preset collaborative detection model, identify abnormal signals in the source delivery process and calculate risk grading results, and generate corresponding control instructions according to the risk grading results.
2. The system of claim 1, wherein, When performing data element extraction, the data acquisition and preprocessing module is configured to perform the following steps: Sensitive word clouds are constructed using an inverted index database, and natural language segmentation algorithms are used to extract and cluster keywords; The information entropy of the keyword is calculated, wherein the information entropy is calculated based on the change probability of the left or right adjacent words of the keyword in the corpus; Determine whether the information entropy of the keyword exceeds a preset entropy clustering threshold, if so, extract the keyword as tobacco-related data elements, and convert unstructured text data into structured feature data.
3. The system of claim 1, wherein, The data fusion module is configured to perform horizontal federal fusion and vertical federal fusion: In horizontal federal fusion, control the local training of neural network sub-models by each level of the federal core node, and update the global model parameters through a secure aggregation strategy; In vertical federal fusion, the joint feature vector comprising tobacco end features of a first preset dimension and cross-department features of a second preset dimension is constructed; Wherein, the tobacco end features are generated based on ordering data, distribution data and inventory data of retailers, and the cross-department features are generated based on tobacco-related case data, unlicensed operation data and tax record data.
4. The system of claim 1, wherein, The anomaly detection and risk grading module includes a three-flow consistency verification unit, which is configured to perform the following logic: Obtain the fund flow data, logistics distribution data and information flow order data of the retailer; Detect whether there is a state where the fund has been credited but there is no logistics signing record; Calculate the geographic distance between the member consumption address and the actual distribution address, and determine whether the geographic distance exceeds a preset distance threshold; Detect whether there is a state where the information flow shows high-frequency purchases but the fund flow has no corresponding payment record; If any of the above conditions is met, a consistency anomaly determination signal is generated, and the risk grading calculation weight of the retailer is increased.
5. The system of claim 1, wherein, Further comprising a verifiable evidence chain module, the verifiable evidence chain module is configured to: The evidence chain fusion unit integrates the joint feature vector, the model inference path, and the verification result fed back by the federal collaborative node to generate an evidence package; The hash solidification unit extracts a digital digest of the key information in the evidence package and generates a unique hash value, and concatenates the hash values in chronological order to form a hash chain; The hash chain is stored in a preset blockchain network for supervision and traceability.
6. The system of claim 1, wherein, A privacy protection and permission control module is also included, which is configured to perform hierarchical protection according to the data sensitivity level: For the information of the person involved and the tax records marked as core sensitive data, a homomorphic encryption strategy is used for encryption, and a partial field hiding rule is used for desensitization processing; For the logistics distribution and inventory records marked as general sensitive data, a differential privacy strategy is used to add noise, and the noise intensity is controlled within a preset noise intensity threshold range; Through the hierarchical protection, only statistical results are output during data fusion calculation without leaking individual privacy information.
7. The system of claim 1, wherein, A dynamic threshold and feedback learning module is also included, which is configured to perform the following closed-loop iteration process: Receive a verification feedback signal for the risk classification result, the feedback signal including a true violation label or a false positive label; If the feedback signal is a false positive label, the features of the false positive scenario are extracted and stored in a normal fluctuation feature library, and the abnormal detection threshold under the scenario is automatically increased by a preset adjustment ratio; If the feedback signal is a true violation label, the corresponding joint feature vector is supplemented to the training set to strengthen the model weight, and the updated threshold parameter is sent to the distributed supervision network through the federal core node.
8. The system of claim 1, wherein, The anomaly detection and risk classification module also includes a graph neural network correlation detection unit, which is configured to: Construct a cross-department heterogeneous graph, which includes nodes of retail outlets, persons involved, tax subjects, and distribution node types, and edges representing the association between nodes, registration matching relationships, and distribution relationships; Perform graph convolution operations in the cross-department heterogeneous graph to mine association paths of multiple retail outlets sharing the same person involved, or structural deviations of tax subjects and retail outlets registration information; Based on the mined association paths and structural deviations, generate a group linkage anomaly signal.
9. The system of claim 4, wherein, The anomaly detection and risk classification module also includes a topological anomaly filtering unit for executing a false positive suppression process when the geographic distance exceeds a preset distance threshold, the process including: Construct a payment and distribution association subgraph: take the retail outlet as an anchor point, extract payment records and distribution records in a preset historical period, and establish a directed relationship from the payment node to the distribution node; Calculate the link stability coefficient: based on the time span and historical transaction frequency of the association between the payment node and the distribution node, calculate the link stability coefficient, wherein the link stability coefficient is positively correlated with the time span, and has a non-linear mapping relationship with the historical transaction frequency; The end consumption entropy is calculated: if the link stability coefficient meets the preset stability requirement, terminal sales data associated with the distribution node is tracked, and the end consumption entropy is calculated according to the number distribution of independent consumers in a preset monitoring period, the end consumption entropy being used to represent the discrete degree of the sales of the goods source; The dynamic exemption determination is performed: if the end consumption entropy is greater than a preset discrete degree threshold, it is determined that the current geographic distance deviation belongs to a compliant chain operation behavior, and the consistency anomaly determination signal is revoked.
10. The system of claim 6, wherein, The privacy protection and permission control module further includes a privacy budget dynamic management unit, which is configured to perform the following dynamic noise injection process: Initialize the privacy budget: allocate a total privacy budget in a preset period for data subjects in a regulatory area; Calculate sensitivity and consumption: when a query request is received, the global sensitivity of the query is calculated, and a single privacy budget consumption is allocated according to the query permission; Perform budget check: determine whether the sum of the consumed budget and the current privacy budget consumption exceeds the total privacy budget, and if so, reject the query; Generate dynamic noise: generate a dynamic noise value using the Laplace mechanism and superimpose it on the query result; The scale parameter used to generate the dynamic noise value is calculated based on the global sensitivity, the single privacy budget consumption, and the query frequency; The calculation logic of the scale parameter is configured such that the scale parameter is positively correlated with the global sensitivity, negatively correlated with the single privacy budget consumption, and increases with the increase of the query frequency of the data subject within a preset time period, so as to block differential attacks.