A tobacco retail market network case investigation method and related device
By employing a hierarchical heterogeneous data collection and credibility analysis approach, a semi-supervised learning model is constructed for investigating online cases in the tobacco retail market. This solves the problems of dynamic profile updates and self-iterative adjustments in real-time incoming data in existing technologies, achieving highly adaptable and accurate case investigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG TOBACCO SHANWEI CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies, such as manual comparison and single-modal monitoring on internet channels, are insufficient to capture the evolution of black market chains, cannot achieve real-time dynamic profile updates of cross-platform tobacco-related business data, and lack self-iterative adjustment capabilities.
By acquiring multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, credibility analysis and time series alignment are performed. A semi-supervised learning model is constructed for anomaly identification. Based on dynamic network profiling, the core nodes and critical paths of the gang are analyzed to generate reconnaissance analysis results. The model is then optimized and iterated based on actual reconnaissance results.
It enables real-time processing of dynamically flowing data, breaks through the limitations of static knowledge reasoning, significantly improves the adaptability and accuracy of case investigation, and forms a closed-loop system with self-iterative adjustment capabilities.
Smart Images

Figure CN122114941A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and related equipment for investigating cases involving tobacco retail market networks. Background Technology
[0002] Traditional tobacco retail regulation focuses more on the compliance of offline licenses and permits, and the consistency of terminal sales volume and logistics distribution. However, with the deepening penetration of Internet channels, the concealment of cross-platform resales, and the spread of gray logistics jump points, manual comparison and single-modal monitoring are difficult to capture the evolution of the black chain. This makes external delivery disguise, virtual identity transit, and multi-account control difficult points for law enforcement identification. Therefore, there is an urgent need for a new reconnaissance system with cross-source data collaborative perception, dynamic threshold identification of abnormal indicators, and topological tracing of gang relationships.
[0003] Existing technologies utilize case entity relationships to construct knowledge graphs and enhance semantic connections through attention mechanisms. These graphs, combined with ConvKB models, enable reasoning and calculations between individuals and assets, offering valuable assistance in understanding and organizing cases with broad scope and highly concealed relationships. However, these approaches still rely on fixed graph modeling structures, requiring predefined entity boundaries and exhibiting lag in relationship expansion. Particularly in scenarios involving cross-platform tobacco operations, online resale, and sudden changes in logistics, they cannot dynamically update profiles based on real-time data inflows and lack self-iterative adjustment capabilities. Furthermore, these methods lean towards static knowledge structure reasoning, exhibiting poor evolutionary capabilities. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, electronic device, storage medium, and program product for investigating cases in the tobacco retail market network, aiming to solve at least one problem of the prior art.
[0005] To achieve the above objectives, one aspect of this invention proposes a method for investigating cases related to tobacco retail market networks, the method comprising: Multi-source heterogeneous data is acquired through a hierarchical heterogeneous acquisition architecture. The credibility analysis and time-series alignment of the multi-source heterogeneous data are performed to obtain raw data labeled with data credibility. Among them, the multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from online public platforms, and logistics and delivery related data. The raw data is preprocessed and then correlated and fused under a unified spatiotemporal framework to obtain fused data for each time series. A semi-supervised learning model is constructed based on fused data, and then processed sequentially to obtain anomaly identification data for each round of time series; the anomaly identification data includes anomaly index and judgment threshold. A dynamic network profile of suspected illegal business groups is constructed based on anomaly identification data; the dynamic network profile includes the strength of the group structure evolution in each time series and the group coreness of each node under a unified spatiotemporal framework; Based on dynamic network profiling analysis, the core nodes and key paths of the gang are identified, thereby generating reconnaissance analysis results; among which, the reconnaissance analysis results include suggestions on priority reconnaissance order and evidence collection strategies; The actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model.
[0006] In some embodiments, each piece of collected data in multi-source heterogeneous data is marked with a unique timestamp and source identifier. The process of performing credibility analysis and time-series alignment on the multi-source heterogeneous data to obtain original data marked with credibility includes the following steps: Based on the source identifier, the source level, collection method, and degree of structuring of the collected data are determined, and then the source level score, collection method score, and degree of structuring score of the collected data are mapped and determined. The data credibility of the collected data is obtained by weighting and summing the scores of source level, collection method, and structure degree. Based on a unique timestamp, each collected data from multiple heterogeneous sources is linked together around the same timeline to generate a data chain, resulting in time-aligned raw data.
[0007] In some embodiments, the fused data includes a standardized representation of each node in a unified spatiotemporal framework. The original data is preprocessed, and then correlated and fused within the unified spatiotemporal framework to obtain fused data for each time series. This includes the following steps: The raw data is cleaned, anomaly logic is removed, and cross-modal alignment and standardization are performed to obtain preprocessed data. The preprocessed data is divided according to time sequence to obtain the original data value corresponding to each time window, and then the abnormal difference factor corresponding to each time window is obtained by consistency verification and labeling. The corrected expression is obtained by multiplying the preset artifact suppression strength factor, the confidence coefficient, and the abnormal difference factor; whereby the confidence coefficient is obtained by summing the data confidence of all data within the time window. Based on the difference between the original data value and the corrected expression, the net data expression corresponding to each time window is obtained, and then the real content vector of each node in the unified spatiotemporal framework is obtained through vectorization processing. The true content representation of each node in the unified spatiotemporal framework is obtained by multiplying the true content vector and the credibility weight; whereby the credibility weight is obtained by summing the data credibility of all data in the corresponding node. Based on the ratio of the actual content representation to the sum of the credibility weights of all nodes in the unified spatiotemporal framework, the standardized representation of each node in the unified spatiotemporal framework is determined.
[0008] In some embodiments, the fused data includes a standardized representation of each node in a unified spatiotemporal framework. A semi-supervised learning model is constructed based on the fused data and then processed sequentially to obtain anomaly identification data for each round of the time series, including the following steps: Perform statistical analysis on the fused data for each time series to determine the behavioral distribution center for each time series; Take the previous time series as the first time series and the current time series as the second time series; Obtain the preset threshold benchmark; Based on the difference between the behavior distribution center in the first time series and the behavior distribution center in the second time series, the behavior difference data is determined; The threshold correction value is obtained by multiplying the behavioral difference data, the credibility weighting parameter, and the preset convergence control coefficient; wherein, the credibility weighting parameter is obtained by weighted summation of the data credibility corresponding to all data in the unified spatiotemporal framework under the second time series; The sum of the threshold baseline and the threshold correction value is used as the decision threshold for the first time series; Based on the ratio of the Euclidean norm of the difference between the behavior distribution center of the standardized expression and the second time series to the judgment threshold, the abnormality index of the behavior corresponding to each node in the unified spatiotemporal framework under the second time series is obtained. Using the judgment threshold as the threshold benchmark, the second time sequence as the first time sequence, and the next time sequence as the second time sequence, the process returns to the step of determining the behavior difference data based on the difference between the behavior distribution center of the first time sequence and the behavior distribution center of the second time sequence, and processes the abnormal identification data of each time sequence in chronological order.
[0009] In some embodiments, constructing a dynamic network profile of a suspected illegal business group based on anomaly identification data includes the following steps: The gang core degree of each node is obtained by multiplying the anomaly index and credibility weight of each node under the unified spatiotemporal framework with the preset node association link strength ratio coefficient. The credibility weight is obtained by summing the data credibility of all data in the corresponding node. The evolution strength of each node is obtained by multiplying the coreness of the group with the increase or decrease of the associated links of the corresponding node within a preset time window. The evolution intensity of all nodes within a preset time window is accumulated to obtain the gang structure evolution intensity corresponding to each round of time sequence.
[0010] In some embodiments, the core nodes and key paths of a gang are analyzed based on dynamic network profiling to generate reconnaissance and analysis results, including the following steps: The reconnaissance risk priority score of each node is obtained by weighting and summing the credibility weight, anomaly index and gang core degree corresponding to each node. The credibility weight is obtained by summing the data credibility of all data in the corresponding node. The reconnaissance risk priority scores of all nodes are ranked to determine the priority order for reconnaissance. The risk priority scores of all nodes in the preset path are summed to obtain the cumulative risk score. The reconnaissance value index of the preset path is obtained by the ratio of the cumulative risk score to the link length of the preset path. Based on the reconnaissance risk priority score for each node and the reconnaissance value index for each path, various types of output content are generated to suggest evidence collection strategies.
[0011] In some embodiments, the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model, including the following steps: Risk scoring deviations are determined based on actual reconnaissance results, and intensity weights are determined based on strike effectiveness. The threshold optimization value is obtained by multiplying the risk score bias, the intensity weight, and the preset feedback convergence step size. Based on the difference between the decision threshold and the optimized threshold, a new iteration threshold is determined, and the decision threshold of the semi-supervised learning model is updated using the new iteration threshold. The optimal index value is determined by multiplying the preset gain adjustment coefficient with the gang core degree and the gang structure disintegration intensity increment; wherein, the gang structure disintegration intensity increment is determined by the change in the gang structure evolution intensity corresponding to the reconnaissance analysis results during the execution phase. The abnormality index is determined by the sum of the abnormality index and the index optimization value, and then the abnormality index is updated using the abnormality index after iteration.
[0012] To achieve the above objectives, another aspect of the present invention provides a device for investigating cases related to tobacco retail market networks, the device comprising: The second module preprocesses the raw data and then performs correlation and fusion under a unified spatiotemporal framework to obtain fused data for each time series. The third module constructs a semi-supervised learning model based on the fused data and then processes it sequentially along the time sequence to obtain anomaly identification data for each round of time sequence; among which, the anomaly identification data includes anomaly index and judgment threshold; The fourth module constructs a dynamic network profile of suspected illegal business groups based on anomaly identification data. The dynamic network profile includes the strength of the group structure evolution in each time series and the group coreness of each node under a unified spatiotemporal framework. The fifth module analyzes the core nodes and key paths of the gang based on dynamic network profiling, and then generates reconnaissance analysis results, including suggestions on priority reconnaissance order and evidence collection strategies. The sixth module uses the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results as negative feedback labels to optimize and iterate the semi-supervised learning model.
[0013] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.
[0014] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0015] To achieve the above objectives, another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0016] The embodiments of this invention include at least the following beneficial effects: This invention provides a method, apparatus, electronic device, storage medium, and program product for investigating online cases in the tobacco retail market. This solution acquires multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, performs credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtains raw data marked with data credibility. The multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from publicly available online platforms, and logistics and delivery-related data. The raw data undergoes data preprocessing, and then is correlated and fused within a unified spatiotemporal framework to obtain fused data for each time series. A semi-supervised learning model is constructed based on the fused data. The model is then processed sequentially to obtain anomaly identification data for each time series. The anomaly identification data includes anomaly indices and judgment thresholds. A dynamic network profile of suspected illegal business groups is constructed based on the anomaly identification data. This dynamic network profile includes the group structure evolution intensity for each time series and the group coreness of each node within a unified spatiotemporal framework. The core nodes and critical paths of the group are analyzed based on the dynamic network profile, generating reconnaissance analysis results. These results include suggestions for priority reconnaissance order and evidence collection strategies. The actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model. This invention overcomes the shortcomings of existing technologies, such as reliance on fixed graphs, predefined entity boundaries, delayed relationship expansion, and poor evolutionary capabilities. Specifically, it ensures the timeliness and quality of multi-source data through hierarchical heterogeneous data collection and credibility analysis; it performs temporal fusion and anomaly identification within a unified spatiotemporal framework, enabling real-time processing of dynamically flowing data; it constructs and continuously updates dynamic network profiles based on anomaly identification results, achieving evolutionary analysis of gang structures and core nodes, thus breaking through the limitations of static knowledge reasoning; finally, it uses actual investigation results as negative feedback labels to optimize the model, forming a closed-loop system with self-iterative adjustment capabilities, significantly improving the adaptability and accuracy of case investigation. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an implementation environment for a method for investigating online cases in the tobacco retail market provided in this embodiment of the invention; Figure 2 This is a flowchart illustrating a method for investigating cases related to a tobacco retail market network, as provided in an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating an example of a hierarchical heterogeneous data acquisition architecture provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the data preprocessing and fusion process provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the adaptive analysis model training and updating process provided in an embodiment of the present invention; Figure 6This is a schematic diagram of the process for generating dynamic network profiles and reconnaissance strategies provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the overall process of the tobacco retail market network case investigation method provided in the embodiments of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0019] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”
[0020] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0022] First, it should be noted that, for the sake of understanding the technical solution of this invention, the technical terms that may be involved in the technical solution of this invention will be explained first: 1. Multi-source heterogeneous data: refers to the collection of data from different sources and of different types collected by the system, including tobacco retail terminal transaction data, text and image data from publicly available online platforms, and logistics and delivery related data.
[0023] 2. Adaptive Analysis Model: Based on semi-supervised learning, this analysis model can autonomously learn normal retail behavior patterns even with missing data or noise interference, and can dynamically update the anomaly judgment threshold according to data flow and behavior changes.
[0024] 3. Dynamic Gang Network Profiling: A visualized network topology constructed using abnormal nodes and relationships, containing information such as core nodes and related links of suspected illegal business gangs. It can be updated in real time as new data flows in to depict the gang's evolution process.
[0025] 4. Closed-loop feedback and model iteration: The actual reconnaissance results and strike effects are used as negative feedback labels to feed back into the adaptive analysis model, driving the model to adjust parameters and optimize thresholds, thus realizing a mechanism for continuous evolution of recognition capabilities.
[0026] 5. Credibility Coefficient: A parameter calculated based on the data source level, collection method, and content structuring level. It is used to quantify the credibility of the original data and provide a basis for weight allocation in the multi-source data fusion stage.
[0027] The related technologies tend to focus on static knowledge structure reasoning, and have not built a system that integrates transaction data, delivery trajectories and online dissemination behavior. They do not support the presentation of the behavioral evolution of organizational networks, nor do they have the functions of automated reconnaissance strategies and closed-loop execution feedback.
[0028] In view of this, this invention provides a method and related equipment for investigating cases involving tobacco retail market networks. This method acquires multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, performs credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtains raw data labeled with data credibility. The multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from publicly available online platforms, and logistics and delivery-related data. The raw data undergoes data preprocessing, and then is correlated and fused within a unified spatiotemporal framework to obtain fused data for each time series. A semi-supervised learning model is constructed based on the fused data and then processed sequentially along the time series. Anomaly identification data for each time series is obtained; the anomaly identification data includes anomaly index and judgment threshold; a dynamic network profile of suspected illegal business gangs is constructed based on the anomaly identification data; the dynamic network profile includes the gang structure evolution intensity for each time series and the gang core degree of each node under a unified spatiotemporal framework; the core nodes and critical paths of the gang are analyzed based on the dynamic network profile, thereby generating reconnaissance analysis results; the reconnaissance analysis results include priority reconnaissance order and evidence collection strategy suggestions; the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model. This invention overcomes the shortcomings of existing technologies, such as reliance on fixed graphs, predefined entity boundaries, delayed relationship expansion, and poor evolutionary capabilities. Specifically, it ensures the timeliness and quality of multi-source data through hierarchical heterogeneous data collection and credibility analysis; it performs temporal fusion and anomaly identification within a unified spatiotemporal framework, enabling real-time processing of dynamically flowing data; it constructs and continuously updates dynamic network profiles based on anomaly identification results, achieving evolutionary analysis of gang structures and core nodes, thus breaking through the limitations of static knowledge reasoning; finally, it uses actual investigation results as negative feedback labels to optimize the model, forming a closed-loop system with self-iterative adjustment capabilities, significantly improving the adaptability and accuracy of case investigation.
[0029] It is understood that the tobacco retail market network case investigation method provided by this invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet, laptop, or desktop computer, but it is not limited to these.
[0030] like Figure 1The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0031] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0032] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0033] Terminal 102 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.
[0034] For example, based on Figure 1 The implementation environment shown in this embodiment of the invention provides a method for investigating cases in a tobacco retail market network. The following description uses the application of this method in server 101 as an example. It can be understood that this method can also be applied in terminal 102.
[0035] Reference Figure 2 , Figure 2 This is an optional flowchart of the tobacco retail market network case investigation method provided in the embodiments of the present invention. The executing entity of the tobacco retail market network case investigation method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S600.
[0036] Step S100: Obtain multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, perform credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtain the original data marked with data credibility. Among them, multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from publicly available online platforms, and logistics and delivery related data; It should be noted that each piece of collected data in the multi-source heterogeneous data is marked with a unique timestamp and source identifier. In some embodiments, step S100 may include the following steps: determining the source level, collection method, and degree of structuring of the collected data based on the source identifier, and then mapping and determining the source level score, collection method score, and degree of structuring score corresponding to the collected data; performing a weighted summation of the source level score, collection method score, and degree of structuring score to obtain the data credibility of the collected data; and generating a data chain around the same time axis for each piece of collected data in the multi-source heterogeneous data based on the unique timestamp to obtain time-aligned original data.
[0037] Exemplary examples, such as in some specific implementations, Figure 3 As shown, the data sources exhibit multi-dimensional characteristics and varying collection frequencies. To ensure the consistency of subsequent analysis model training logic, a layered heterogeneous collection architecture is adopted in the data acquisition process. This involves real-time monitoring of tobacco retail terminal transaction behavior, periodic crawling of text and image information from publicly available online platforms, and associated collection of logistics and delivery trajectories to achieve synchronous coverage of the behavioral chain, propagation chain, and physical flow chain. During the collection process, the system generates a unique timestamp and source identifier for each piece of raw data and introduces an adaptive collection frequency adjustment mechanism based on sampling dynamics to prevent abnormal behavior from affecting recognition accuracy due to data delays and collection window jumps during short-term, high-frequency operations. Furthermore, to address the differences in collection dimensions, granularity, and temporal sequence among multi-source data, the system establishes a standardized coding system for different data streams and defines an information credibility coefficient for weight control in the subsequent fusion stage. This coefficient is calculated based on the data source permission level, collection method, and content structuring level, and is expressed as follows:
[0038] in, Indicates the first Credibility of the original data Indicates the source level score. The score indicates the data collection method. The score represents the degree of structure, and the weight coefficients of the three parameters satisfy the following: The purpose of this formula is to assign weights to content from different sources, so that the subsequent fusion and analysis processes avoid the impact of data noise equalization on the stability of the results. Through credibility calculation, the system can prioritize the participation of data from the transaction system and logistics supervision field in the core modeling, ensuring the credibility and traceability of the data foundation. Due to the asynchronous nature of the collected data, the system introduces a time series alignment algorithm to construct a cross-source number index, so that text, images, transaction behavior, and logistics trajectory form a dynamic and traceable data chain around a unified time axis. This mechanism provides a complete spatiotemporal correlation foundation for subsequent data preprocessing and fusion, and constructs the data flow logic for step S200.
[0039] Step S200: Perform data preprocessing on the raw data, and then perform correlation and fusion under a unified spatiotemporal framework to obtain fused data for each time series; It should be noted that the fused data includes the standardized representation of each node in the unified spatiotemporal framework. In some embodiments, step S200 may include the following steps: performing data cleaning, anomaly logic removal, and cross-modal alignment standardization on the original data to obtain preprocessed data; dividing the preprocessed data according to time sequence to obtain the original data value corresponding to each time window, and then obtaining the anomaly difference factor corresponding to each time window through consistency verification annotation; obtaining the corrected representation based on the product of the preset artifact suppression strength factor, the confidence coefficient, and the anomaly difference factor; wherein, the confidence coefficient is based on the data confidence of all data corresponding to the time window. The data is summarized; the summarization operation may include weighted summation or average operation; based on the difference between the original data value and the corrected expression, the net data expression corresponding to each time window is obtained, and then the true content vector of each node in the unified spatiotemporal framework is obtained through vectorization processing; based on the product of the true content vector and the credibility weight, the true content expression of each node in the unified spatiotemporal framework is obtained; the credibility weight is obtained based on the sum of the data credibility of all data in the corresponding node; based on the ratio of the true content expression to the sum of the credibility weights of all nodes in the unified spatiotemporal framework, the standardized expression of each node in the unified spatiotemporal framework is determined.
[0040] Exemplary examples, such as in some specific implementations, Figure 4 As shown, based on the multi-source credibility identifier and time series alignment results formed in step S100, the original data undergoes deep cleaning, anomaly logic removal, and cross-modal alignment standardization. An automated consistency verification mechanism identifies missing items, conflicting fields, and forgery traces, thereby establishing a noise-free unified feature representation before entering the model stage. Addressing the structural differences in multi-source data, the system constructs a cross-domain mapping matrix for transaction logs, public text, image semantic features, and logistics trajectory signals. Through hierarchical feature regularization, different modal data are pulled into a unified spatiotemporal index structure, with credibility as the basis for fusion weights. Simultaneously, abnormal fluctuations caused by tampering, delays, or forgery are suppressed. To avoid weight distortion in the fusion stage, the system constructs an adaptive artifact cancellation function based on the credibility coefficient in step S100. By applying nonlinear suppression to suspicious interference information, the true business trajectory is highlighted, as expressed below:
[0041] in, Indicates in Net data representation after time-series point correction. This represents the original value of the data collected within the corresponding time window. The confidence coefficient is calculated based on the data confidence level in step S100. These are the abnormal difference factors marked by consistency verification. This is the artifact suppression strength factor, used to accurately correct fluctuations in abnormal structured behavior, ensuring that group tampering and fictitious transactions fail in the feature representation dimension. To further improve cross-source alignment stability and mitigate temporal drift caused by different sampling periods, the system constructs a unified temporal fusion function to continuously map text, image, and logistics flow vectors, achieving smooth fusion of cross-source behavior nodes along the time axis.
[0042] in, For the fusion of the first Standardized representation of nodes within a unified spatiotemporal framework This is the cleaned vector of the actual content of this node. To obtain the corresponding credibility weights, This ensures that the fusion results are not biased towards a single source due to differences in data volume and collection frequency, thus maintaining a complete and abstract representation of market behavior trajectories. Through this series of cleaning, noise reduction, and weighted fusion processes, data features are represented without disturbance and possess cross-modal consistency before entering the analysis module, providing stable input and laying the foundation for dynamic recognition for the construction of the adaptive analysis model in step S300.
[0043] Step S300: Construct a semi-supervised learning model based on the fused data and then process it sequentially along the time sequence to obtain the anomaly identification data for each round of time sequence; The anomaly identification data includes anomaly index and judgment threshold; It should be noted that the fused data includes the standardized representation of each node in the unified spatiotemporal framework. In some embodiments, step S300 may include the following steps: performing statistical analysis on the fused data of each time series to determine the behavioral distribution center of each time series; taking the previous time series as the first time series and the current time series as the second time series; obtaining a preset threshold benchmark; determining behavioral difference data based on the difference between the behavioral distribution center of the first time series and the behavioral distribution center of the second time series; obtaining a threshold correction value based on the product of the behavioral difference data, the confidence weighting parameter, and the preset convergence control coefficient; wherein, the confidence weighting parameter is based on the unified spatiotemporal framework under the second time series. The weighted sum of the data credibility corresponding to all data in the empty frame is obtained; the sum of the threshold baseline and the threshold correction value is used as the judgment threshold of the first time series; based on the ratio of the Euclidean norm of the difference between the behavior distribution centers of the standardized expression and the second time series to the judgment threshold, the abnormality index of the behavior corresponding to each node in the unified spatiotemporal frame under the second time series is obtained; the judgment threshold is used as the threshold baseline, the second time series is used as the first time series, the next time series is used as the second time series, and the steps of determining the behavior difference data based on the difference between the behavior distribution centers of the first time series and the behavior distribution centers of the second time series are returned to be executed, and the abnormal identification data of each round of time series is obtained by processing along the time series order.
[0044] Exemplary examples, such as in some specific implementations, Figure 5 As shown, based on the spatiotemporal unified expression vector formed in step S200 and the confidence-weighted feature input, a semi-supervised learning model with environmental adaptability is constructed. A baseline pattern is formed through the long-term behavioral trajectory of normal retail business, and the model maintains stable convergence even under conditions of missing data or noise. The system divides the training paths for labeled and unlabeled samples for this model, and uses gradient control to give data with high confidence a stronger supervisory influence, enabling the model to continuously absorb patterns in the real business feature distribution and gradually solidify the boundaries of normal behavioral features. To avoid model convergence deviation caused by abnormal transactions due to sparse scale, abrupt frequency changes, or disguised behavior, the system introduces a dynamic boundary judgment mechanism. A self-updating threshold strategy is constructed through the abnormal offset vector, and this threshold is adjusted in real time according to data flow and changes in behavioral patterns. The core dynamic threshold update process of the model is generated based on the semi-supervised convergence factor, confidence weight, and the degree of behavioral distribution offset, as expressed below:
[0045] in, The threshold for determining the current time series. Based on the previous round's threshold benchmark, For convergence control coefficient, The confidence-weighted parameters are derived from the fusion process in step S200. and Corresponding to the distribution centers of current and previous temporal behaviors in the fusion feature space, the model, through this dynamic update mechanism, adaptively and sensitively adjusts to overall market trends, local abnormal structural changes, and the evolution of potential gang behaviors, thereby effectively identifying covert bulk cigarette purchases, abnormal logistics concentration points, and duplicate account behavior camouflage patterns. To further enhance the model's identification ability with unsupervised samples, the system introduces a self-correcting feature clustering representation, realizing the dynamic labeling process of abnormal cluster nodes through feature distance offset, expressed as follows:
[0046] in, For the first The abnormal index corresponding to each behavior, For the standardized spatiotemporal fusion vector obtained in step S200, when As the time series continues to rise, the model includes the behavior and its associated nodes in the potential abnormal network structure labeling interval. Through this series of self-learning, self-correction, and dynamic threshold update mechanisms, the model forms a behavior recognition system with anomaly self-detection capabilities, providing basic data structure support for generating an evolving gang network profile in step S400.
[0047] Step S400: Construct a dynamic network profile of a suspected illegal business group based on anomaly identification data; Among them, the dynamic network profile includes the strength of gang structure evolution in each round of time sequence and the gang core degree of each node under the unified spatiotemporal framework; It should be noted that in some embodiments, step S400 may include the following steps: obtaining the gang core degree of each node based on the product of the anomaly index corresponding to each node under the unified spatiotemporal framework, the credibility weight, and the preset node association link strength ratio coefficient; wherein, the credibility weight is obtained based on the sum of the data credibility corresponding to all data in the corresponding node; obtaining the evolution strength of each node based on the product of the gang core degree and the increase or decrease of the association link of the corresponding node within the preset time window; and accumulating the evolution strength of all nodes within the preset time window to obtain the gang structure evolution strength corresponding to each round of time sequence.
[0048] Exemplary examples, such as in some specific implementations, Figure 6As shown, based on the abnormal index sequence and dynamic threshold boundary output in step S300, potential illegal business entities and their transaction, delivery, and propagation links are mapped into computable network units. A gang relationship topology graph with continuous evolution characteristics is constructed through cross-modal association weights. Network nodes consist of actors, logistics touchpoints, and virtual accounts. Edges are generated by frequency density, transaction intensity, same-direction delivery trajectories, and image semantic similarity. The topology structure does not rely on static modeling but exhibits data flow-driven active emergence characteristics. To mitigate the structural noise impact of high-frequency but non-critical nodes on the network morphology, the system constructs a gang core evaluation function based on the credibility weights formed in step S200 and the abnormal distribution center offset obtained in step S300, and performs periodic reordering of the graph structure to identify key nodes in the architecture that have a controlling influence on the business black chain. This function is expressed as follows:
[0049] in, Indicates the first The core strength of a node within a group. The credibility weights after fusion in step S200 The anomaly index obtained in step S300. The node association link strength ratio coefficient is used as the core evaluation result of the function for topology center extraction and peripheral node folding, preventing the graph structure from becoming unstable due to sudden data increases during updates. To present the evolution trend of network topology in the temporal dimension, the system introduces a gang structure transition tensor to quantify the changes in inter-node dependency strength and control path with data increments. This tensor is defined as follows:
[0050] in, for Intensity of gang structure evolution within the time window Let i be the core degree weight of node i (i∈[1,n], where n is the total number of nodes). For the link associated with node i By adjusting the increment or decrement within the window, the system can identify changes in the control hierarchy and external expansion behaviors within the group, thus presenting the replacement of virtual accounts, the proliferation of logistics jump points, and the fission trend of the distribution chain as a dynamic network diffusion pattern. Through the aforementioned core node calibration and topology evolution calculation, the system achieves a panoramic presentation of the group structure based on time sequence at the visualization level, and forms a quantitative representation of organizational hierarchy, command flow, and resource allocation density, providing an executable path basis for step S500 to automatically generate reconnaissance strategies.
[0051] Step S500: Based on dynamic network profiling analysis, identify the core nodes and key paths of the gang, and then generate reconnaissance analysis results; The reconnaissance analysis results include recommendations on priority reconnaissance order and evidence collection strategies; It should be noted that, in some embodiments, step S500 may include the following steps: weighted summation of the credibility weight, anomaly index, and gang core degree corresponding to each node to obtain the reconnaissance risk priority score of the corresponding node; wherein, the credibility weight is obtained based on the summary of the data credibility corresponding to all data in the corresponding node; sorting the reconnaissance risk priority scores of all nodes to determine the priority reconnaissance order; accumulating the reconnaissance risk priority scores of all nodes in the preset path to obtain the cumulative risk score; obtaining the reconnaissance value index of the preset path based on the ratio of the cumulative risk score to the link length of the preset path; and generating various types of output content for evidence collection strategy suggestions based on the reconnaissance risk priority score of each node and the reconnaissance value index of each path.
[0052] For example, in some specific implementations, based on the dynamic gang topology structure formed in step S400, the reconnaissance priority is calculated through three dimensions: behavioral dominance characteristics, link control degree, and propagation efficiency. Core nodes, key transit points, and behavioral diffusion nodes are then arranged in a weighted sequence to form a reconnaissance strategy set with execution order and action direction. To ensure that the strategy generation process is not affected by local noise or abnormal short-term behavioral fluctuations, a comprehensive risk judgment function is introduced. This function jointly quantifies the credibility weight of step S200, the anomaly index of step S300, and the gang coreness of step S400, enabling the model to stably express the true threat level of nodes. This function is expressed as follows:
[0053] in, Indicates the first Node reconnaissance risk priority scoring, The credibility weight after fusion This is a behavioral abnormality index. As a core indicator of the gang, the weighting coefficient The scoring system exhibits characteristics of changing over time and evolving with data input, thus enabling the system's reconnaissance ranking to be non-static and automatically adjusted as group behavior develops. The system generates key reconnaissance paths based on the scoring sequence, forms the strategy foundation by searching for the minimum control chain set involving high-risk nodes, and introduces an evidence chain closure evaluation mechanism to quantify path integrity. The process is expressed as follows:
[0054] in, For path The reconnaissance value index This is the cumulative risk score of all nodes in the path. The formula, which defines the link length corresponding to this path, aims to avoid bias towards lengthy but weakly contributing links and prioritize key paths that can form a closed chain of evidence and trigger the exposure of the group's true organizational structure. The system is based on... and The dual indicators generate three types of output content, including the first type of action path suggestion for pointing to the priority strike node sequence list, the second type of evidence collection strategy for matching image source tracing, fund flow reconstruction and logistics trajectory verification methods, and the third type of risk diffusion prediction for early warning of potential gang migration, alternative entities and hidden link behavior trends. Finally, a complete reconnaissance strategy structure with coherence, execution logic and related proof chain is formed, and the execution effect is put into step S600 to form a closed-loop feedback update model and strategy system.
[0055] Step S600: The actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model. It should be noted that in some embodiments, step S600 may include the following steps: determining the risk score deviation based on the actual reconnaissance results, and determining the intensity weight based on the strike effect; obtaining the threshold optimization value based on the product of the risk score deviation, the intensity weight, and the preset feedback convergence step size; determining the new iteration threshold based on the difference between the judgment threshold and the threshold optimization value, and updating the judgment threshold of the semi-supervised learning model using the new iteration threshold; determining the index optimization value based on the product of the preset gain adjustment coefficient, the gang core degree, and the gang structure disintegration intensity increment; wherein, the gang structure disintegration intensity increment is determined based on the change in the gang structure evolution intensity corresponding to the execution stage implemented by the reconnaissance analysis results; determining the iterative abnormal index based on the sum of the abnormal index and the index optimization value, and updating the abnormal index using the iterative abnormal index.
[0056] For example, in some specific implementations, the model's cognitive bias, the accuracy of gang node identification, and the effectiveness of the evidence chain are retrospectively calibrated based on the reconnaissance results after step S500. The actual effectiveness of the crackdown is injected into the semi-supervised model training loop of step S300 through negative feedback labels, causing the dynamic threshold and anomaly index to converge again to the distribution of real criminal behavior, thereby avoiding the risk of pattern solidification or high-frequency misjudgment after long-term model operation. The negative feedback label consists of two parts: one is a node detection effectiveness indicator, and the other is an evidence chain closure quality score. These two parts jointly determine the gradient direction and weight decay magnitude of the model in the next round of parameter updates. During the feedback calculation process, the system constructs an adaptive iterative function, jointly correcting the judgment error, the degree of deviation of node risk weights, and the misidentification rate of gang structure evolution, so that the anomaly boundary threshold is updated with the convergence of real judicial results. This process is expressed as follows:
[0057] in, The threshold for the new iteration, The threshold of the previous round, To provide feedback on the convergence step size, The risk score deviation in step S500 reconnaissance path, To identify intensity weights for accurate strike results, this function helps the model threshold converge towards the precise identification range and suppresses duplicate labeling of low-risk nodes. To address the lag error in monitoring gang topology evolution, the system quantifies the degree of organizational disintegration, the interruption of node control chains, and the effectiveness of logistics jump point blocking in the strike results using labels. This forms an iterative gain adjustment parameter, which directionally scales the anomaly index, as shown below:
[0058] in, The anomaly index after iteration. For the current index, This is the gain adjustment coefficient. The core degree weights formed in step S400 To enhance the intensity of gang disintegration during the execution phase, this approach aims to increase the sensitivity of core nodes to anomalies and weaken the influence of peripheral, disguised nodes, thereby enabling the model to focus on the long-term stable expression of real criminal organization characteristics. Through the combined effects of threshold callback, anomaly weight gain, and structural feedback enhancement mechanisms, the system forms a closed-loop recursion of reconnaissance execution results, model judgment mechanisms, and topological evolution expression. This continuously reduces false alarms and missed detections throughout the entire process, ultimately achieving self-continuous optimization and iteration of strategies, profiles, and recognition capabilities, forming a terminal recursive stable convergence structure for a complete intelligent reconnaissance system.
[0059] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0060] First, it should be noted that the technical solution of this invention relates to the field of intelligent tobacco retail supervision and online tobacco crime investigation technology. It belongs to the intersection of retail supervision data governance and cross-domain illegal business behavior identification. Through multi-source heterogeneous data collection, dynamic gang profile construction and adaptive reconnaissance strategy deduction, it realizes intelligent identification and organizational structure dismantling of tobacco resale chains.
[0061] In some specific implementations, such as Figure 7 As shown, the method for investigating cases related to tobacco retail market networks provided by this invention can be implemented through the following steps: Step 1, Acquisition of multi-source heterogeneous data: The system simultaneously accesses and collects transaction data from tobacco retail terminals, text and image data from publicly available online platforms, and logistics and delivery-related data in real time; Step 2, Data Preprocessing and Fusion: The raw data collected in Step 1 is automatically cleaned and standardized. The missing, contradictory and abnormal patterns in the data are identified and marked. The impact of human tampering with the data is reduced by the algorithm to remove false data and the multi-source data is correlated and fused under a unified spatiotemporal framework. Step 3: Establish an adaptive analysis model: Based on the fused data processed in Step 2, train a semi-supervised learning model. This model can autonomously learn normal retail behavior patterns and dynamically update the anomaly judgment threshold under conditions of partial data loss or noise interference. Step 4, Dynamic Gang Network Profile Generation: Using the abnormal nodes and relationships output by the model in Step 3, the network topology of suspected illegal business gangs is automatically constructed and visualized. This structure can be dynamically updated as new data flows in to depict its evolution process. Step 5, Automatic Generation of Reconnaissance Strategy: Based on the dynamic network profile generated in Step 4, the core nodes and critical paths of the gang are automatically analyzed, and suggestions for priority reconnaissance order and evidence collection strategy are generated accordingly. Step Six, Closed-Loop Feedback and Model Iteration: The actual reconnaissance results and strike effects implemented according to the strategy in Step Five are used as negative feedback labels and fed back into the adaptive analysis model in Step Three to drive the model to perform the next round of optimization iteration.
[0062] The principles, logic, and explanations of each step from step one to step six correspond to the specific implementation details of steps S100 to S600, and will not be repeated here.
[0063] Specifically, the application product of the technical solution of this invention is an intelligent tobacco-related case investigation system for tobacco regulatory departments. It can integrate multi-source data such as tobacco retail transactions, publicly available online information, and logistics delivery. Through intelligent analysis, dynamic profiling, and strategy generation functions, it provides intelligent support for the accurate identification, gang tracing, and efficient crackdown of illegal online tobacco-related business activities, and helps to build a technical guarantee system for the standardized supervision of the tobacco retail market.
[0064] In summary, this invention overcomes the limitations of existing technologies, such as static map modeling, single data sources, lack of real-time update mechanisms, and insufficient reconnaissance strategy support. It constructs a unified data foundation through hierarchical collection and reliable fusion of multi-source heterogeneous data, and achieves dynamic adjustment of abnormal thresholds and accurate identification of abnormal behaviors based on a semi-supervised adaptive analysis model. It generates a dynamic gang network profile that can be updated in real time by combining core degree evaluation and structural evolution tensor, automatically deduce priority reconnaissance sequences and evidence collection strategies, and uses actual strike results to form a closed-loop feedback to drive the model to continuously iterate. This not only enables cross-domain tracing and organizational visualization of illegal tobacco business activities, but also allows the identification capability to evolve synchronously with the gang's disguise upgrades, achieving precise detection and control throughout the entire cycle, the entire link, and all nodes.
[0065] Specifically, the core logic for implementing the technical solution of this invention includes, but is not limited to: 1. A hierarchical acquisition and trusted fusion mechanism for multi-source heterogeneous data: Through standardized coding, trusted coefficient calculation and time series alignment algorithm, it realizes unified spatiotemporal correlation and false-proof processing of cross-source data such as retail transactions, network information and logistics trajectories.
[0066] 2. An adaptive analysis model based on semi-supervised learning is constructed, which has the ability to learn normal retail behavior patterns autonomously under the conditions of missing data or noise interference, and achieves accurate identification of abnormal behavior through dynamic threshold updates and self-correcting feature clustering.
[0067] 3. Dynamic gang network profiling generation technology: Through core degree evaluation function and structural evolution tensor calculation, it constructs a gang topology that can be updated in real time with the inflow of data, and quantitatively presents organizational hierarchy, control path and evolution trend.
[0068] 4. An automatic generation mechanism for reconnaissance strategies based on dynamic profiling, through comprehensive risk assessment and evidence chain closure evaluation, outputs priority reconnaissance sequences, evidence collection methods, and risk diffusion predictions, thereby achieving optimal allocation of reconnaissance resources.
[0069] 5. A closed-loop iterative system driven by reconnaissance results transforms actual strike effects into negative feedback labels. Through threshold callbacks and abnormal exponential gain adjustments, it drives the continuous optimization and evolution of model recognition capabilities and reconnaissance strategies.
[0070] Compared with the prior art, the technical solution of the present invention has at least the following beneficial effects: 1. The technical solution of this invention achieves synchronous online acquisition of the tobacco retail transaction chain, the postal logistics chain and the online dissemination chain through a multi-source heterogeneous acquisition architecture, forming a unified data foundation across platforms, behaviors and scenarios, making covert reselling behaviors and cross-regional gang structures traceable, effectively breaking through the limitations of traditional investigations that rely on single transaction data or manual investigation.
[0071] 2. The technical solution of this invention constructs a unified spatiotemporal expression matrix by combining credibility parameters with a fusing algorithm for removing false information and preserving truth. This enables structured cleaning of feature pollution caused by malicious forgery, data delay superposition, and cross-subject spoofing, allowing the true business behavior trajectory to stand out from the high-noise mixed data and improving the data purity and interference resistance of network smoke detection.
[0072] 3. The technical solution of this invention dynamically adjusts the anomaly threshold through a semi-supervised adaptive analysis model, so that the model can maintain convergence and stability under conditions of missing data, noise accumulation and gang disguise evolution, and has the self-learning feature of anomaly identification, effectively reducing the false alarm rate and false negative rate caused by reliance on manual modeling and feature solidification.
[0073] 4. The technical solution of this invention constructs and quantifies the evolution of the dynamic gang network profile, and presents the logistics jump point proliferation, capital chain concealment, virtual identity replacement and multi-node distribution and diffusion behavior in real time in the form of a graph, so as to realize the visible identification of the organizational relationship, control path and command center of the black tobacco chain.
[0074] 5. The technical solution of this invention uses an automatic reconnaissance strategy deduction mechanism to transform high-risk core nodes, key circulation links and evidence closure paths into priority investigation sequences, thereby optimizing the allocation of reconnaissance resources and maximizing the benefits of operations, and reducing the missed opportunities and insufficient strike accuracy caused by traditional manual judgment.
[0075] 6. The technical solution of this invention drives model threshold callback, abnormal weight scaling and topology stability reconstruction through execution result feedback, realizes data-identification-reconnaissance-feedback-re-identification loop, enables the system to continuously evolve identification accuracy, strike effectiveness and disguise cracking ability in long-term operation, and finally form a continuous intelligent reconnaissance system that can cope with the evolving trend of new network tobacco-related crimes.
[0076] This invention also provides a device for investigating cases in the tobacco retail market network, which can implement the above-described method. This device may include: The first module is used to acquire multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, perform credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtain raw data marked with data credibility; among them, multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from online public platforms, and logistics and delivery related data; The second module preprocesses the raw data and then performs correlation and fusion under a unified spatiotemporal framework to obtain fused data for each time series. The third module constructs a semi-supervised learning model based on the fused data and then processes it sequentially along the time sequence to obtain anomaly identification data for each round of time sequence; among which, the anomaly identification data includes anomaly index and judgment threshold; The fourth module constructs a dynamic network profile of suspected illegal business groups based on anomaly identification data. The dynamic network profile includes the strength of the group structure evolution in each time series and the group coreness of each node under a unified spatiotemporal framework. The fifth module analyzes the core nodes and key paths of the gang based on dynamic network profiling, and then generates reconnaissance analysis results, including suggestions on priority reconnaissance order and evidence collection strategies. The sixth module uses the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results as negative feedback labels to optimize and iterate the semi-supervised learning model.
[0077] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0078] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0079] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0080] like Figure 8 As shown, Figure 8 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0081] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0082] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0083] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0084] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0085] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0086] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0087] The invention provides a method, apparatus, electronic device, storage medium, and program product for investigating online cases in the tobacco retail market. It acquires multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, performs credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtains raw data labeled with data credibility. The multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from publicly available online platforms, and logistics and delivery-related data. The raw data undergoes data preprocessing, and then is correlated and fused within a unified spatiotemporal framework to obtain fused data for each time series. A semi-supervised learning model is constructed based on the fused data and then processed sequentially along the time series. The process obtains anomaly identification data for each time series, including anomaly indices and judgment thresholds. A dynamic network profile of suspected illegal business groups is constructed based on this data. This dynamic network profile includes the group structure evolution intensity for each time series and the group coreness of each node within a unified spatiotemporal framework. The core nodes and critical paths of the group are analyzed based on the dynamic network profile, generating reconnaissance analysis results. These results include suggestions for priority reconnaissance order and evidence collection strategies. The actual reconnaissance results and strike effectiveness corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model. This invention overcomes the shortcomings of existing technologies, such as reliance on fixed graphs, predefined entity boundaries, delayed relationship expansion, and poor evolutionary capabilities. Specifically, it ensures the timeliness and quality of multi-source data through hierarchical heterogeneous data collection and credibility analysis; it performs temporal fusion and anomaly identification within a unified spatiotemporal framework, enabling real-time processing of dynamically flowing data; it constructs and continuously updates dynamic network profiles based on anomaly identification results, achieving evolutionary analysis of gang structures and core nodes, thus breaking through the limitations of static knowledge reasoning; finally, it uses actual investigation results as negative feedback labels to optimize the model, forming a closed-loop system with self-iterative adjustment capabilities, significantly improving the adaptability and accuracy of case investigation.
[0088] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. A method for investigating cases involving tobacco retail market networks, characterized in that, The method includes the following steps: Multi-source heterogeneous data is acquired through a hierarchical heterogeneous acquisition architecture. The credibility analysis and time-series alignment of the multi-source heterogeneous data are then performed to obtain raw data marked with data credibility. The multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from online public platforms, and logistics and delivery related data. The original data is preprocessed and then correlated and fused under a unified spatiotemporal framework to obtain fused data for each time series. A semi-supervised learning model is constructed based on the fused data, and then processed sequentially to obtain anomaly identification data for each round of time sequence; wherein, the anomaly identification data includes an anomaly index and a judgment threshold; Based on the anomaly identification data, a dynamic network profile of suspected illegal business groups is constructed; wherein, the dynamic network profile includes the strength of the group structure evolution in each round of time sequence and the group core degree of each node under the unified spatiotemporal framework; Based on the dynamic network profiling analysis of the gang's core nodes and key paths, reconnaissance analysis results are generated; wherein, the reconnaissance analysis results include priority reconnaissance order and evidence collection strategy suggestions; The actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results are used as negative feedback labels to optimize and iterate the semi-supervised learning model.
2. The method according to claim 1, characterized in that, Each piece of collected data in the multi-source heterogeneous data is marked with a unique timestamp and source identifier. The process of performing credibility analysis and time-series alignment on the multi-source heterogeneous data to obtain the original data marked with credibility includes the following steps: Based on the source identifier, the source level, collection method, and structure degree of the collected data are determined, and then the source level score, collection method score, and structure degree score of the collected data are mapped and determined. The data credibility of the collected data is obtained by weighting and summing the source level score, the collection method score, and the structure degree score. Based on the unique timestamp, each of the collected data from the multi-source heterogeneous data is linked together around the same time axis to generate a data chain, thereby obtaining the time-aligned original data.
3. The method according to claim 1, characterized in that, The fused data includes a standardized representation of each node in the unified spatiotemporal framework. The process of preprocessing the original data and then performing correlation and fusion within the unified spatiotemporal framework to obtain fused data for each time series includes the following steps: The original data is cleaned, anomaly logic is removed, and cross-modal alignment and standardization are performed to obtain preprocessed data; The preprocessed data is divided according to time sequence to obtain the original data value corresponding to each time window, and then the abnormal difference factor corresponding to each time window is obtained by consistency verification labeling. The corrected expression is obtained by multiplying the preset artifact suppression strength factor, the confidence coefficient, and the abnormal difference factor; wherein the confidence coefficient is obtained by summing the data confidence of all data within the time window; Based on the difference between the original data value and the corrected expression, the net data expression corresponding to each time window is obtained, and then the real content vector of each node in the unified spatiotemporal framework is obtained through vectorization processing. The true content representation of each node in the unified spatiotemporal framework is obtained by multiplying the true content vector with the credibility weight; wherein, the credibility weight is obtained by summing the data credibility of all data in the corresponding node. The standardized expression of each node in the unified spatiotemporal framework is determined based on the ratio of the actual content expression to the sum of the credibility weights of all nodes in the unified spatiotemporal framework.
4. The method according to claim 1, characterized in that, The fused data includes the standardized representation of each node in the unified spatiotemporal framework. The step of constructing a semi-supervised learning model based on the fused data and then processing it sequentially to obtain anomaly identification data for each round of time series includes the following steps: Statistical analysis is performed on the fused data for each time series to determine the behavioral distribution center for each time series; Take the previous time series as the first time series and the current time series as the second time series; Obtain the preset threshold benchmark; Based on the difference between the behavior distribution center of the first time series and the behavior distribution center of the second time series, behavior difference data is determined; The threshold correction value is obtained by multiplying the behavioral difference data, the credibility weighting parameter, and the preset convergence control coefficient; wherein, the credibility weighting parameter is obtained by weighted summation of the data credibility corresponding to all data in the unified spatiotemporal framework under the second time series; The sum of the threshold benchmark and the threshold correction value is used as the determination threshold of the first time series; The anomaly index of the behavior corresponding to each node in the unified spatiotemporal framework under the second time series is obtained by the ratio of the Euclidean norm of the difference between the standardized expression and the behavior distribution center of the second time series to the judgment threshold. Using the determination threshold as the threshold benchmark, the second time sequence as the first time sequence, and the next time sequence as the second time sequence, the process returns to the step of determining the behavior difference data based on the difference between the behavior distribution center of the first time sequence and the behavior distribution center of the second time sequence, and processes the abnormal identification data of each round of time sequence in chronological order.
5. The method according to claim 1, characterized in that, The process of constructing a dynamic network profile of a suspected illegal business group based on the anomaly identification data includes the following steps: The gang core degree of each node is obtained by multiplying the anomaly index and credibility weight of each node under the unified spatiotemporal framework with the preset node association link strength ratio coefficient. The credibility weight is obtained by summing the credibility of all data in the corresponding node; The evolution strength of each node is obtained by multiplying the core strength of the group with the increase or decrease of the associated links of the corresponding node within a preset time window. The evolution intensity of all nodes within the preset time window is accumulated to obtain the gang structure evolution intensity corresponding to each round of time sequence.
6. The method according to claim 1, characterized in that, The process of analyzing the core nodes and key paths of the gang based on the dynamic network profiling to generate reconnaissance and analysis results includes the following steps: The reconnaissance risk priority score of the corresponding node is obtained by weighting and summing the credibility weight, the anomaly index and the core degree of the gang for each node; The credibility weight is obtained by summing the credibility of all data in the corresponding node; The reconnaissance risk priority scores of all nodes are sorted to determine the priority reconnaissance order; The reconnaissance risk priority scores of all nodes in the preset path are summed to obtain the cumulative risk score value; The reconnaissance value index of the preset path is obtained by the ratio of the cumulative risk score to the link length of the preset path. Based on the reconnaissance risk priority score for each node and the reconnaissance value index for each path, various types of output content are generated to suggest the evidence collection strategy.
7. The method according to claim 1, characterized in that, The step of using the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results as negative feedback labels to optimize and iterate the semi-supervised learning model includes the following steps: The risk score deviation is determined based on the actual reconnaissance results, and the intensity weight is determined based on the strike effect. The threshold optimization value is obtained by multiplying the risk score deviation, the intensity weight, and the preset feedback convergence step size. Based on the difference between the judgment threshold and the optimized threshold, a new iteration threshold is determined, and the judgment threshold of the semi-supervised learning model is updated using the new iteration threshold. The optimal index value is determined by multiplying the preset gain adjustment coefficient by the core strength of the gang and the increment of the gang structure disintegration intensity; wherein, the increment of the gang structure disintegration intensity is determined by the change in the gang structure evolution intensity in the execution phase corresponding to the reconnaissance analysis results; Based on the sum of the anomaly index and the optimized index value, the iterative anomaly index is determined, and the anomaly index is updated using the iterative anomaly index.
8. A device for investigating cases involving tobacco retail market networks, characterized in that, The device includes: The first module is used to acquire multi-source heterogeneous data through a hierarchical heterogeneous acquisition architecture, perform credibility analysis and time-series alignment on the multi-source heterogeneous data, and obtain raw data marked with data credibility; wherein, the multi-source heterogeneous data includes tobacco retail terminal transaction data, text and image data from online public platforms, and logistics and delivery related data; The second module preprocesses the raw data and then performs correlation and fusion under a unified spatiotemporal framework to obtain fused data for each time series. The third module constructs a semi-supervised learning model based on the fused data and then processes it sequentially to obtain anomaly identification data for each round of time sequence; wherein, the anomaly identification data includes an anomaly index and a judgment threshold; The fourth module constructs a dynamic network profile of suspected illegal business groups based on the anomaly identification data; wherein, the dynamic network profile includes the strength of the group structure evolution in each round of time sequence and the group core degree of each node under the unified spatiotemporal framework; The fifth module analyzes the core nodes and key paths of the gang based on the dynamic network profile, and then generates reconnaissance analysis results; wherein, the reconnaissance analysis results include priority reconnaissance order and evidence collection strategy suggestions; The sixth module uses the actual reconnaissance results and strike effects corresponding to the reconnaissance analysis results as negative feedback labels to optimize and iterate the semi-supervised learning model.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.