Transaction process-oriented risk sample adaptive selection and labeling method

By constructing a multimodal spatiotemporal transaction feature matrix and a business anomaly exposure index, and combining density clustering and dynamic quota game mechanism, the problems of false positives and redundant labeling of risk samples in the industrial manufacturing supply chain are solved, and the accurate screening and efficient review of high-risk candidate samples are achieved.

CN121765387BActive Publication Date: 2026-05-01BAIWEIJINKE (SHANGHAI) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BAIWEIJINKE (SHANGHAI) INFORMATION TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing risk assessment models in the industrial manufacturing supply chain rely on single-dimensional data distribution entropy values ​​or fixed confidence thresholds for sample screening and labeling. This results in an excessive number of false positives and redundant labeling during large-scale anomalies, exceeding the cognitive processing capacity of the review team and leading to the omission or delay in the processing of genuine high-risk samples.

Method used

By acquiring the transaction flow sequence of procurement nodes and the spatial flow logs of physical warehousing and distribution nodes, a multimodal spatiotemporal transaction feature matrix is ​​constructed. The relative entropy of the graph structure feature vector and the physical link hysteresis damping coefficient are quantified to generate a business anomaly exposure index. Combined with density clustering and dynamic quota game mechanism, high-entropy refined samples are adaptively selected and pushed to the review terminal.

Benefits of technology

Accurately intercept false positive interference, remove homogeneous and redundant samples, achieve a closed-loop human-machine collaboration, ensure the physical authenticity of high-risk candidate samples and the accuracy of screening, and avoid review congestion and backlog.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765387B_ABST
    Figure CN121765387B_ABST
Patent Text Reader

Abstract

The application discloses a risk sample adaptive selection and labeling method for transaction process, and particularly relates to the technical field of industrial risk control data processing, and is used for solving the problems of business exception redundancy push and audit overload. First, a multi-modal space-time transaction feature matrix is constructed by fusing transaction flow and physical flow log, high-frequency deviation samples are extracted by combining graph structure relative entropy and lag damping coefficient, and false positive interference is excluded; then, isomorphic risk clusters are analyzed based on density clustering, homogenization redundancy is stripped by using cluster information entropy and centrality, representative samples are extracted to generate a high-entropy refined subset, and audit redundancy is reduced; finally, a dynamic quota game mechanism is established by monitoring the concurrent throughput and cognitive load threshold, the optimal sample slice is intercepted to the terminal discrimination labeling, and a closed-loop system from feature fusion to adaptive push is constructed, which provides scientific support for industrial transaction risk control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial risk control data processing technology, specifically to an adaptive selection and labeling method for risk samples oriented towards transaction processes. Background Technology

[0002] In the process of digital transformation of the industrial manufacturing supply chain, the bulk procurement and warehousing processes generate massive amounts of transaction node flows and spatial movement logs. These high-frequency transactions constitute the core supply chain network of industrial enterprises. As the interactions between business entities become increasingly complex, abnormal risks such as defaults, fraud, or supply chain disruptions hidden in the context of normal transactions are difficult to detect intuitively. Accurately identifying and labeling these hidden risk events is a key link in ensuring the stable operation of the industrial supply chain and preventing systemic operational crises.

[0003] Existing risk assessment models rely entirely on a single-dimensional data distribution entropy value or a fixed confidence threshold to determine sample recommendation strategies during sample screening and labeling. This static feature extraction method severs the entity mapping relationship between financial cash flow and physical warehousing and logistics, resulting in the extraction of a large number of false positive data that are abnormal only in financial data but stable and normal in physical flow. At the same time, when faced with homogeneous, large-scale, explosive business anomalies, existing random sampling or recommendation mechanisms will instantly flood the manual review terminal with a massive number of anomaly records with the same triggers, causing redundant and duplicate labeling, which seriously exceeds the cognitive processing capacity of the review team, and consequently, causes genuine high-risk samples to be missed or delayed in processing while waiting in the queue. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an adaptive selection and labeling method for risk samples oriented towards the transaction process, thus solving the problems mentioned in the background.

[0005] To achieve the above objectives, this invention employs the following technical solution: an adaptive selection and labeling method for risk samples in a transaction process, comprising the following steps: S1. Obtaining the transaction flow sequence of procurement nodes and the spatial flow logs of corresponding physical warehousing and distribution nodes in an industrial manufacturing supply chain scenario; performing heterogeneous data fusion on the transaction node flow sequence and spatial flow logs according to timestamp granularity and business entity identifiers to construct a multimodal spatiotemporal transaction feature matrix; S2. Mapping the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space; extracting graph structure feature vectors representing the flow of funds and materials; quantifying the relative entropy of the graph structure feature vectors deviating from the historical baseline distribution to represent the data concept drift; parsing the spatial flow logs to extract the physical link hysteresis damping coefficient; fusing the data concept drift and the physical link hysteresis damping coefficient to generate a business anomaly exposure index; and extracting the data concept drift from the multimodal spatiotemporal transaction feature matrix. S3. High-frequency deviation samples are used to construct a high-risk candidate sample pool; S4. Business anomaly exposure index is introduced as a distance metric weight, and density-based feature space clustering is performed on the high-risk candidate sample pool to form multiple isomorphic risk clusters. The industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster are calculated. Homogeneous redundant samples with low information gain are stripped, and representative transaction samples with centrality peaks and marginal free distributions are extracted to generate a high-entropy refined sample subset; S5. The concurrent throughput of the industrial transaction cycle and the real-time cognitive load threshold of the manual review resource pool are monitored. The high-entropy refined sample subset is sorted in descending order according to the business anomaly exposure index to generate candidate sample slices. A dynamic quota game mechanism is established between the candidate sample slices and the real-time cognitive load threshold. The optimal sample slice within the cognitive load carrying boundary is selected and pushed to the review terminal for manual identification to generate a labeled industrial transaction risk sample set.

[0006] Further, step S1 includes the following steps: acquiring purchase order details and fund settlement records from the supply chain enterprise resource planning system and integrating them into a purchase transaction node flow sequence; collecting vehicle trajectory coordinates and warehouse throughput node scanning records uploaded by IoT devices and converting them into spatial flow logs of physical warehousing and distribution nodes; extracting globally unique order tracking codes from the purchase transaction node flow sequence and spatial flow logs as business entity identifiers; performing time dimension slicing on the purchase transaction node flow sequence and spatial flow logs according to a preset time window; and tensor concatenating data segments with the same globally unique order tracking code and the same time dimension slice to output a multimodal spatiotemporal transaction feature matrix.

[0007] Furthermore, the process of mapping the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space, extracting graph structure feature vectors representing fund flows and material circulation, and quantifying the relative entropy of the graph structure feature vectors deviating from the historical benchmark distribution to represent the data concept drift is as follows: The transaction participants and warehousing operation areas in the multimodal spatiotemporal transaction feature matrix are analyzed as graph nodes, and the fund allocation path and logistics transportation route are analyzed as directed edges to construct a business entity relationship graph; the business entity relationship graph is mapped to a high-dimensional topological space, and the graph node centrality aggregation feature and directed edge weight feature are extracted and merged to generate graph structure feature vectors representing fund flows and material circulation; the weighted graph structure benchmark feature vectors within the historical normal transaction cycle are extracted to generate a historical benchmark distribution, and the relative entropy between the probability density function of the current graph structure feature vector and the probability density function of the historical benchmark distribution is calculated as the data concept drift.

[0008] Furthermore, the specific process of extracting the physical link hysteresis damping coefficient from the spatial flow log, fusing the data concept drift degree and the physical link hysteresis damping coefficient to generate the business anomaly exposure index, and extracting high-frequency deviation samples from the multimodal spatiotemporal transaction feature matrix to construct a high-risk candidate sample pool is as follows: Extract the standard expected delivery timestamp and the actual node receipt timestamp from the spatial flow log, and calculate the time difference between the standard expected delivery timestamp and the actual node receipt timestamp; normalize the ratio of the time difference with the preset industry benchmark grace period to output the physical link hysteresis damping coefficient; obtain the feature weights corresponding to the data concept drift degree and the physical link hysteresis damping coefficient, perform a linear combination superposition on the weighted data concept drift degree and the physical link hysteresis damping coefficient, and output the business anomaly exposure index; track the business anomaly exposure index of each transaction sample in the multimodal spatiotemporal transaction feature matrix according to the time sliding window, extract the transaction samples corresponding to the abnormal fluctuation peak points exceeding the dynamic statistical confidence interval, mark them as high-frequency deviation samples, and merge the high-frequency deviation samples to generate a high-risk candidate sample pool.

[0009] Furthermore, a business anomaly exposure index is introduced as a distance metric weight. Density-based feature space clustering is performed on the high-risk candidate sample pool to form multiple isomorphic risk clusters. The specific process for calculating the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster is as follows: Extract the graph structure feature vector and corresponding business anomaly exposure index of each sample in the high-risk candidate sample pool. Convert the business anomaly exposure index into a spatial distance penalty coefficient. Combine the graph structure feature vector to calculate the weighted spatial distance between each sample. Perform density reachability clustering based on the weighted spatial distance to divide the high-risk candidate sample pool into multiple isomorphic risk clusters. Calculate the topological in-degree and out-degree of the industrial circulation nodes contained in each isomorphic risk cluster to generate the industrial node influence centrality. Calculate the distribution dispersion of the graph structure feature vectors of the samples in each isomorphic risk cluster and output the intra-cluster information entropy.

[0010] Furthermore, the specific process of generating a high-entropy refined sample subset by stripping away homogeneous redundant samples with low information gain and extracting representative transaction samples with centrality peaks and marginal free distributions is as follows: Set a baseline entropy value, remove all samples in homogeneous risk clusters whose information entropy is lower than the baseline entropy value, and complete the stripping of homogeneous redundant samples. In the remaining homogeneous risk clusters, extract the samples in which the influence centrality of industrial nodes reaches the highest value range as centrality peak samples; calculate the physical topological distance of each sample in the remaining homogeneous risk clusters from the corresponding cluster center, extract the sample with the largest physical topological distance as marginal free distribution samples, and perform tensor merging of the centrality peak samples and marginal free distribution samples to obtain representative transaction samples, outputting a high-entropy refined sample subset.

[0011] Furthermore, the specific process of monitoring the concurrent throughput of the industrial transaction cycle and the real-time cognitive load threshold of the manual review resource pool, and generating candidate sample slices by sorting the high-entropy refined sample subset in descending order according to the business anomaly exposure index is as follows: The order creation frequency and fund settlement frequency within the industrial transaction cycle are collected and combined to generate concurrent throughput; the number of online review specialists, average historical review time, and backlog of work orders in the manual review resource pool are obtained, and the real-time cognitive load threshold is generated through weighted combination calculation; the business anomaly exposure index corresponding to each sample in the high-entropy refined sample subset is obtained; the high-entropy refined sample subset is sorted in descending order according to the business anomaly exposure index; and the high-entropy refined sample subset after descending order is segmented and truncated according to a preset step size to generate multiple candidate sample slices.

[0012] Furthermore, the high-entropy refined sample subset is sorted in descending order according to the business anomaly exposure index to generate candidate sample slices. A dynamic quota game mechanism is established between the candidate sample slices and the real-time cognitive load threshold. The optimal sample slice within the cognitive load carrying boundary is selected and pushed to the review terminal for manual identification. The specific process of generating a labeled industrial transaction risk sample set is as follows: The number of samples contained in the candidate sample slice is converted into the expected review time. The expected review time and the real-time cognitive load threshold are input into the supply and demand matching function to construct a dynamic quota game mechanism. In the dynamic quota game mechanism, candidate sample slices with expected review times close to and not exceeding the real-time cognitive load threshold are extracted, and the extracted candidate sample slices are marked as the optimal sample slices. The transaction records in the optimal sample slice and the corresponding spatial flow logs are packaged and pushed to the review terminal. The risk qualitative feedback instruction returned by the review terminal based on manual identification is received. The risk qualitative feedback instruction is converted into a classification label and bound to the corresponding sample in the optimal sample slice, and the labeled industrial transaction risk sample set is output.

[0013] The present invention has the following beneficial effects:

[0014] (1) An adaptive selection and labeling method for risk samples oriented towards the transaction process is proposed. This method heterogeneously integrates the flow sequence of procurement transaction nodes with the spatial flow logs of physical warehousing and distribution nodes to construct a multimodal spatiotemporal transaction feature matrix. Furthermore, it jointly quantifies the relative entropy of graph structure feature vectors and the physical link hysteresis damping coefficient in a high-dimensional topological space. This step breaks through the limitations of pure financial data analysis. Through bidirectional cross-validation of capital flow and material flow, the generated business anomaly exposure index can objectively reflect the disturbances in the real physical world, thereby accurately intercepting false positive interference that only stays at the data level. This significantly improves the physical authenticity and screening accuracy of the high-risk candidate sample pool.

[0015] (2) An adaptive selection and labeling method for risk samples in the transaction process is proposed. This method uses the business anomaly exposure index as a weight to perform feature space clustering to remove homogeneous and redundant samples, and establishes a dynamic quota game mechanism between candidate sample slices and the real-time cognitive load threshold of the manual review resource pool. This step effectively intercepts the explosive repeated push caused by homogeneous risk clusters, enabling the system to adaptively extract and push the most information-gaining high-entropy refined sample slices based on the actual carrying capacity boundary of the current review team. This completely eliminates the waste of computing power and the backlog of review caused by invalid and redundant labeling, and realizes a human-machine collaborative closed loop that fits the business rhythm.

[0016] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0017] Figure 1 This is a flowchart of the risk sample adaptive selection and labeling method for the transaction process according to the present invention.

[0018] Figure 2 Flowchart for generating the business anomaly exposure index and constructing a high-risk candidate sample pool.

[0019] Figure 3 Flowchart for generating high-entropy refined sample subsets.

[0020] Figure 4 A flowchart illustrating the dynamic quota game mechanism and human-machine collaboration. Detailed Implementation

[0021] This application's embodiments address the problems of existing technologies, such as high false positive rates due to the separation of financial and physical connections in risk sample screening, and the problems of homogeneous and redundant samples crowding out review resources and causing review overload caused by static push mechanisms, through an adaptive selection and labeling method for risk samples oriented towards transaction processes.

[0022] The overall concept of the solution in this application embodiment is as follows:

[0023] First, heterogeneous data from the supply chain procurement transaction node sequence and the spatial flow log of the corresponding physical warehousing and distribution nodes are fused using timestamps and business entity identifiers to construct a multimodal spatiotemporal transaction feature matrix. Second, this matrix is ​​mapped to a high-dimensional topological space, and a business anomaly exposure index is generated by combining the relative entropy representing the conceptual drift of capital and material flow data and the physical link hysteresis damping coefficient. This index is used to extract high-frequency deviation samples and construct a high-risk candidate sample pool. Next, density-based feature space clustering is performed on the candidate pool to remove redundant samples with low information gain and extract representative transaction samples with centrality peaks and marginal free distributions to generate a high-entropy refined sample subset. Finally, a dynamic quota game mechanism is constructed by monitoring concurrent throughput and real-time cognitive load thresholds. Within the cognitive load carrying boundary, the optimal sample slice is extracted and pushed to the review terminal for manual identification, ultimately generating a labeled industrial transaction risk sample set.

[0024] Please see Figure 1 This invention provides a technical solution: an adaptive selection and labeling method for risk samples in a transaction process, comprising the following steps: S1. Obtaining the transaction flow sequence of procurement nodes and the spatial flow logs of corresponding physical warehousing and distribution nodes in an industrial manufacturing supply chain scenario; performing heterogeneous data fusion on the transaction node flow sequence and spatial flow logs according to timestamp granularity and business entity identifiers to construct a multimodal spatiotemporal transaction feature matrix; S2. Mapping the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space; extracting graph structure feature vectors representing the flow of funds and materials; quantifying the relative entropy of the graph structure feature vectors deviating from the historical baseline distribution to represent the data concept drift; parsing the spatial flow logs to extract the physical link hysteresis damping coefficient; fusing the data concept drift and the physical link hysteresis damping coefficient to generate a business anomaly exposure index; and extracting high-frequency deviations from the multimodal spatiotemporal transaction feature matrix. S3. Construct a high-risk candidate sample pool; S4. Introduce the business anomaly exposure index as a distance metric weight, perform density-based feature space clustering analysis on the high-risk candidate sample pool to form multiple isomorphic risk clusters, calculate the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster, remove homogeneous redundant samples with low information gain, and extract representative transaction samples with centrality peaks and marginal free distributions to generate a high-entropy refined sample subset; S5. Monitor the concurrent throughput of the industrial transaction cycle and the real-time cognitive load threshold of the manual review resource pool, sort the high-entropy refined sample subset in descending order according to the business anomaly exposure index to generate candidate sample slices, establish a dynamic quota game mechanism between candidate sample slices and the real-time cognitive load threshold, extract the optimal sample slice within the cognitive load carrying boundary and push it to the review terminal for manual identification to generate a labeled industrial transaction risk sample set.

[0025] In this implementation plan, step S1 functions to acquire the transaction flow sequence of procurement nodes and the spatial flow logs of corresponding physical warehousing and distribution nodes in the industrial manufacturing supply chain scenario. It then performs heterogeneous data fusion on the transaction node flow sequence and spatial flow logs according to timestamp granularity and business entity identifiers to construct a multimodal spatiotemporal transaction feature matrix. Heterogeneous data fusion refers to the process of combining financial billing details with vehicle trajectory records or warehouse barcode scanning records from the physical world, using a unified timeline and globally unique business order numbers for underlying splicing and parameter alignment. The multimodal spatiotemporal transaction feature matrix is ​​a three-dimensional data set that simultaneously includes financial amount-based cash flow information and information on the corresponding goods' physical spatial flow status. The technical advantage of this step is that it breaks the limitation of traditional risk control that only focuses on financial system data, forcibly binding online financial transactions and offline physical logistics to the same dimension. This provides a solid foundational data base with both temporal sequence and spatial location for subsequent investigation of fraudulent transactions or logistics anomalies, avoiding information blind spots caused by a single data source from the outset. Step S2 maps the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space, extracts graph structure feature vectors representing fund flows and material circulation, quantifies the relative entropy of the graph structure feature vectors deviating from the historical baseline distribution to represent data concept drift, analyzes spatial circulation logs to extract physical link hysteresis damping coefficients, integrates data concept drift and physical link hysteresis damping coefficients to generate a business anomaly exposure index, and extracts high-frequency deviation samples from the multimodal spatiotemporal transaction feature matrix to construct a high-risk candidate sample pool. Here, high-dimensional topological space mapping refers to the process of transforming tabular transaction details into a network structure composed of participating enterprise nodes and logistics route connections; the physical link hysteresis damping coefficient is a quantitative indicator of the delay penalty when the actual time spent on bulk goods in warehousing and transportation exceeds the standard expected delivery time. The technical function of this step is to perform bidirectional cross-validation between the relative entropy of pure data distribution anomalies and the actual logistics delays in the physical world. Only when the financial flow network structure undergoes a sudden change and logistics nodes show significant lags will a high-value business anomaly exposure index be output. This accurately extracts high-frequency deviation samples that actually indicate business disruptions to form a candidate pool, directly filtering out false positives that are merely fluctuations in financial data while logistics remain normal. Step S3 introduces the business anomaly exposure index as a distance metric weight, performs density-based feature space clustering analysis on the high-risk candidate sample pool to form multiple isomorphic risk clusters, calculates the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster, removes homogeneous redundant samples with low information gain, and extracts representative transaction samples with centrality peaks and marginal detached distributions to generate a high-entropy refined sample subset.Homogeneous risk clusters refer to a large number of abnormal transaction data sets with extremely similar characteristics caused by the same physical reason, such as a power outage or blockade of the same core port, within a high-risk sample pool. Marginalized distributions refer to extremely rare, novel, and unknown abnormal samples that do not belong to any large cluster and are scattered and fragmented. The technical role of this step is to solve the problem of data redundancy labeling caused by explosive system anomalies. By calculating the information entropy within the cluster, hundreds of redundant samples with completely identical causes are identified and discarded, retaining only the most representative core collapse node samples and rare novel abnormal samples. This compresses massive amounts of high-risk data into a small but highly valuable high-entropy subset, significantly reducing ineffective review actions. Step S4 monitors the concurrent throughput of the industrial transaction cycle and the real-time cognitive load threshold of the manual review resource pool. It sorts the high-entropy refined sample subset in descending order according to the business anomaly exposure index to generate candidate sample slices. A dynamic quota game mechanism is established between the candidate sample slices and the real-time cognitive load threshold. The optimal sample slice within the cognitive load carrying boundary is selected and pushed to the review terminal for manual identification, generating a labeled industrial transaction risk sample set. The real-time cognitive load threshold refers to the maximum physical limit of the workload that the current online risk control review specialists can handle, considering the number of staff on duty, the average review time per case in the past, and the current backlog of pending work orders. The dynamic quota game mechanism refers to the process by which the system continuously balances the mental resources required for the samples to be reviewed with the actual remaining mental resources of the current human team. The technical role of this step is to completely eliminate the bottleneck between machine initial screening and human final review, enabling it to adaptively adjust the push threshold according to the actual fatigue level and idle computing power of the risk control team. During peak business periods, only the top core high-risk samples are pushed, ensuring that the human team will not collapse due to being overwhelmed by massive alerts. Ultimately, this achieves a perfect rhythm match between machine computing power and human computing power, and high-quality closed-loop risk labeling.

[0026] Specifically, step S1 includes the following steps: acquiring purchase order details and fund settlement records from the supply chain enterprise resource planning system and integrating them into a purchase transaction node flow sequence; collecting vehicle trajectory coordinates and warehouse throughput node scanning records uploaded by IoT devices and converting them into spatial flow logs of physical warehousing and distribution nodes; extracting globally unique order tracking codes from the purchase transaction node flow sequence and spatial flow logs as business entity identifiers; performing time dimension slicing on the purchase transaction node flow sequence and spatial flow logs according to a preset time window; and tensor concatenating data segments with the same globally unique order tracking code and the same time dimension slice to output a multimodal spatiotemporal transaction feature matrix.

[0027] This implementation plan first connects to the supply chain enterprise resource planning (ERP) system via an enterprise-level business interface to extract purchase order details, including material categories, transaction prices, and payment terms, along with corresponding fund settlement records. These purely financial, discrete events are then integrated into a purchase transaction node sequence based on their chronological order. Simultaneously, it connects to the underlying IoT device management platform to collect continuous trajectory coordinates uploaded by the GPS of transport vehicles and RFID scan check-in records from various levels of warehousing throughput nodes. This data, depicting the real movement trajectory of bulk goods in the physical world, is transformed into spatial flow logs for physical warehousing and distribution nodes. To break down data silos between financial accounts and the physical system, the system automatically extracts the globally unique order tracking code shared by the purchase transaction node sequence and the spatial flow logs as a business entity identifier, establishing a core anchor point for aligning heterogeneous data across systems. Next, time-dimensional slicing is performed on the purchase transaction node sequence and spatial flow logs according to a preset time window. Considering the varying lengths of industrial bulk transaction cycles and the phased characteristics of physical transportation, precise time-dimensional slicing is a prerequisite for subsequently capturing local spatiotemporal business fluctuations. Specifically, the global observation time span is set as follows: The preset time window step size is Then the first The start timestamp of each time dimension slice With end timestamp The calculation relationship is expressed as follows: ; ;in, Global observation time span; : The step size of the preset time window; : The ascending integer index of the time slice; :No. The starting timestamp of each time dimension slice; : Initial reference timestamp of the observation period; :No. The end timestamp of each time dimension slice. For the preset time window step size. The method for determining the turnover cycle is based on the dynamic average method of historical logistics nodes. Specifically, it extracts the average actual turnover time of all bulk orders of the same product category between adjacent key warehousing nodes within a complete past business year as the average turnover cycle. The value of is chosen to ensure that the granularity of the time slice perfectly matches the actual industrial physical flow rhythm, avoiding feature sparsity due to overly fine slices or the masking of instantaneous anomalies due to overly coarse slices. Finally, data segments with the same globally unique order tracking code and the same time dimension slice are tensor-concatenated to output a multimodal spatiotemporal transaction feature matrix. This crucial step transforms isolated financial data streams and physical trajectory streams into dense tensor structures that can be directly processed in parallel by computers, achieving deep fusion of multidimensional heterogeneous features within the topological space. A specific business entity identifier is set. In the The feature vector of the procurement transaction node flow sequence within each time dimension slice is: The corresponding spatial flow log feature vector is Then the multimodal spatiotemporal transaction feature matrix Local tensor blocks under a specific entity and a specific time slice The structure is as follows: ;in, : Specific business entity identifier sequence number; : Feature vector of the procurement transaction node sequence; :Feature vector of spatial flow log; Multimodal spatiotemporal transaction feature matrix; Local tensor block; : A non-linear activation function, used here to eliminate the surge in dimensional differences when merging heterogeneous data such as financial amounts and physical distances; Financial feature mapping weight matrix; Tensor cascade splicing operator; Physical feature mapping weight matrix; The total number of business entity identifiers; Total number of time slices; Tensor aggregation stacking operator. Regarding the financial feature mapping weight matrix. Mapping weight matrix with physical features In this embodiment, principal component analysis is used to extract the principal component contribution rates in the original financial feature space and the original physical feature space, respectively. The contribution rates of each feature are transformed into normalized diagonal matrices as the corresponding mapping weight matrices. This allows for adaptive retention of the core business features with the largest variance during the tensor splicing stage, suppressing edge noise interference and providing a high-quality data source for the subsequent construction of a high-dimensional relationship graph.

[0028] Please see Figure 2Specifically, the process of mapping the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space, extracting graph structure feature vectors representing fund flows and material circulation, and quantifying the relative entropy of the graph structure feature vectors deviating from the historical benchmark distribution to represent the data concept drift is as follows: The transaction participants and warehousing operation areas in the multimodal spatiotemporal transaction feature matrix are analyzed as graph nodes, and the fund allocation path and logistics transportation route are analyzed as directed edges to construct a business entity relationship graph; the business entity relationship graph is mapped to a high-dimensional topological space, and the graph node centrality aggregation feature and directed edge weight feature are extracted and merged to generate graph structure feature vectors representing fund flows and material circulation; the weighted graph structure benchmark feature vectors within the historical normal transaction cycle are extracted to generate a historical benchmark distribution, and the relative entropy between the probability density function of the current graph structure feature vector and the probability density function of the historical benchmark distribution is calculated as the data concept drift.

[0029] In this implementation scheme, firstly, matrix dimensionality reduction and relation extraction techniques are used to analyze the transaction participants and warehousing operation areas in the multimodal spatiotemporal transaction feature matrix as graph nodes, and simultaneously analyze the fund transfer path and logistics transportation route as directed edges, thereby constructing a business entity relationship graph reflecting the topology of the upstream and downstream interactive network of the supply chain in computer memory. Then, the business entity relationship graph is mapped to a high-dimensional topological space, and graph neural networks or random walk algorithms are used to extract the graph node centrality aggregation features and the directed edge weight features, merging them to generate graph structure feature vectors representing the flow of funds and materials. This step restores the originally fragmented transaction records into a three-dimensional supply chain transaction ecosystem network, allowing abnormal operations of local nodes to be exposed along the directed edges. Next, the weighted graph structure baseline feature vector within the historical normal transaction cycle is extracted to generate a historical baseline distribution. To accurately quantify the deviation of the current transaction batch network structure from the historical normal pattern, the system calculates the relative entropy between the probability density function of the current graph structure feature vector and the probability density function of the historical baseline distribution, using this as the data concept drift degree. By calculating relative entropy, systemic abrupt changes in trading patterns can be keenly detected. Specifically, the total dimension of the graph structure feature vector is set to... , in the The probability density function of the current graph structure feature vector across the feature dimensions is expressed as: The probability density function of the historical baseline distribution is expressed as: The relative entropy, which characterizes the drift of data concepts, is... The calculation logic is as follows: ;in, The relative entropy value of the data concept drift; The total number of dimensions of the feature vectors of the graph structure; Integer ascending index of the feature dimension; : The probability density function of the eigenvectors of the current graph structure; No. Specific feature sampling values ​​for each dimension; Logarithmic function; The probability density function of the historical baseline distribution. For the probability density function... and This implementation scheme discloses a determination method, namely, using kernel density estimation to smooth and fit the graph structure feature vector sequence within a preset time sliding window to generate a continuous probability density curve, thereby avoiding computational crashes caused by discrete sampling points. By calculating relative entropy, the system can not only detect anomalies in single transaction amounts, but also detect abnormalities in organized network structures such as money laundering or chain-reaction fraudulent trade.

[0030] Specifically, the process of extracting the physical link hysteresis damping coefficient from the spatial flow log, fusing the data concept drift degree and the physical link hysteresis damping coefficient to generate the business anomaly exposure index, and extracting high-frequency deviation samples from the multimodal spatiotemporal transaction feature matrix to construct a high-risk candidate sample pool is as follows: Extract the standard expected delivery timestamp and the actual node receipt timestamp from the spatial flow log, and calculate the time difference between them; normalize the ratio of the time difference to a preset industry benchmark grace period to output the physical link hysteresis damping coefficient; obtain the feature weights corresponding to the data concept drift degree and the physical link hysteresis damping coefficient, perform a linear combination superposition on the weighted data concept drift degree and the physical link hysteresis damping coefficient, and output the business anomaly exposure index; track the business anomaly exposure index of each transaction sample in the multimodal spatiotemporal transaction feature matrix according to a time sliding window, extract the transaction samples corresponding to the abnormal fluctuation peak points exceeding the dynamic statistical confidence interval, mark them as high-frequency deviation samples, and merge the high-frequency deviation samples to generate a high-risk candidate sample pool.

[0031] In this implementation plan, the standard estimated delivery timestamp and the actual node receipt timestamp are extracted from the spatial flow log, and the time difference of logistics delivery is obtained by subtracting the timestamps. Then, the time difference is normalized by ratio calculation with a preset industry benchmark grace period, and the physical link hysteresis damping coefficient is output. This operation filters out routine congestion or minor delays, only penalizing and amplifying severely delayed physical stagnation. Specifically, the standard estimated delivery timestamp is set as follows: The actual node receipt timestamp is The preset industry benchmark grace period is Then the physical link hysteresis damping coefficient The calculation logic is as follows: ;in, Physical link hysteresis damping coefficient; An exponential function with the natural constant as its base; : Non-negative sensitivity modulator; : The function to find the maximum value; : Actual delivery time stamp; Standard estimated delivery timestamp; :Preset industry benchmark grace period. For sensitivity adjustment factors The optimal value for the index is determined by publicly employing a grid search method combined with historical actual default rate curves for fitting and verification. The feature weights corresponding to the data concept drift degree and the physical link hysteresis damping coefficient are obtained, and a linear combination is performed on the weighted data concept drift degree and the physical link hysteresis damping coefficient to output the business anomaly exposure index. This step deeply cross-references network anomalies in the digital space with logistical stagnation in the physical space. Specifically, the business anomaly exposure index is set as... The feature weights of data concept drift are: The characteristic weights of the physical link hysteresis damping coefficient are: The fusion calculation logic is as follows: ;in, Business anomaly exposure index; Feature weights for data concept drift; The characteristic weights of the physical link hysteresis damping coefficient. For the above characteristic weights... and This implementation scheme discloses a determination method, namely, using the inverse variance allocation method of historical data, assigning higher weights to feature dimensions with lower historical volatility, ensuring that the weight allocation is not influenced by human experience. Finally, according to the business anomaly exposure index of each transaction sample in the multimodal spatiotemporal transaction feature matrix tracked by a time sliding window, the transaction samples corresponding to the abnormal volatility peak points exceeding the dynamic statistical confidence interval are extracted and marked as high-frequency deviation samples. Specifically, the upper limit of the dynamic statistical confidence interval is set as follows: The average business anomaly exposure index within the sliding time window is The standard deviation is Then control the upper limit The delineation logic is as follows: ;in, The upper limit of the control for dynamic statistical confidence intervals; Mean of business anomaly exposure index; : Confidence multiplier; Standard deviation of the business anomaly exposure index. For the confidence multiplier... The determination method in this implementation scheme uses Chebyshev's inequality to back-calculate based on the maximum tolerable false alarm rate. Through the above steps, the system will determine the business anomaly exposure index. Strictly greater than the control limit The system accurately extracts samples and merges high-frequency deviation samples to generate a high-risk candidate sample pool. This process completely abandons the empirically-based fixed threshold approach, allowing the screening process to perfectly adapt to the rhythm of macroeconomic cycles and the peak and off-peak seasons of industrial production, ensuring that the candidate samples output to subsequent processing modules have extremely high quality.

[0032] Please see Figure 3 Specifically, the business anomaly exposure index is introduced as a distance metric weight. Density-based feature space clustering is performed on the high-risk candidate sample pool to form multiple isomorphic risk clusters. The specific process for calculating the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster is as follows: Extract the graph structure feature vector and the corresponding business anomaly exposure index of each sample in the high-risk candidate sample pool. Convert the business anomaly exposure index into a spatial distance penalty coefficient. Combine the graph structure feature vector to calculate the weighted spatial distance between each sample. Perform density reachability clustering based on the weighted spatial distance to divide the high-risk candidate sample pool into multiple isomorphic risk clusters. Calculate the topological in-degree and out-degree of the industrial circulation nodes contained in each isomorphic risk cluster to generate the industrial node influence centrality. Calculate the distribution dispersion of the graph structure feature vectors of the samples in each isomorphic risk cluster and output the intra-cluster information entropy.

[0033] In this implementation scheme, the graph structure feature vectors of each sample in the high-risk candidate sample pool and the corresponding business anomaly exposure index generated by the previous steps are first extracted. To overcome the limitation of traditional geometric distance in reflecting the severity of industrial business, the system converts the business anomaly exposure index into a spatial distance penalty coefficient and calculates the weighted spatial distance between samples by combining the graph structure feature vectors. This operation amplifies the feature differences between high-risk samples, forcing highly consistent anomaly events with similar business causes to converge rapidly within the topological space. Specifically, the sorting labels of any two transaction samples in the high-risk candidate sample pool are set as follows: and Then the weighted spatial distance between samples The calculation logic is as follows: ;in, Spatial distance penalty coefficient; : Penalty adjustment constant; : No. Business anomaly exposure index for each sample; No. Business anomaly exposure index for each sample; : The integer sorting index of the sample; : The integer sorting index of the comparison sample; Weighted spatial distance; The total feature dimension of the graph structure feature vector; : Sequence number of the feature dimension; No. The feature vector of the sample graph structure at the th sample graph structure is in the th... The possible values ​​for the dimension; :No. The feature vector of the sample graph structure at the th sample graph structure is in the th... The value of the dimension. Regarding the penalty adjustment constant. This implementation scheme discloses a method for determining the optimal constant value, which is to maximize the ratio of intra-class aggregation to inter-class dispersion of historically labeled similar risk samples and then adaptively optimize using grid search. Next, density reachability clustering is performed based on weighted spatial distance to divide the high-risk candidate sample pool into multiple homogeneous risk clusters. This aims to automatically aggregate a large number of similar abnormal transactions caused by the same underlying physical failure, such as a core port being closed due to a typhoon, into an independent set, facilitating subsequent batch analysis and redundancy elimination. Subsequently, the topological in-degree and out-degree of the industrial circulation nodes contained in each homogeneous risk cluster are calculated to generate the influence centrality of the industrial nodes. This step quantifies the pivotal position of the core physical node that caused the risk cluster in the entire supply chain network. The specific calculation logic is as follows: ;in, : Influence centrality of industrial nodes; Classification labels for isomorphic risk clusters; : No. The total number of core industrial circulation nodes contained within a homogeneous risk cluster; : Sequence number of the core industrial circulation node; : No. The topological in-degree of each core industrial circulation node; : No. The topological out-degree of each core industrial circulation node; Node flow energy level weights. (This refers to the weights for node flow energy levels.) This implementation scheme discloses an objective weighting method that calculates the percentage of bulk cargo throughput of a specific node in a complete historical business year relative to the total network throughput. Finally, it calculates the distribution dispersion of the sample graph structure feature vectors within each isomorphic risk cluster, outputting the cluster information entropy. This step is used to measure the diversity of data representation within this type of risk. The calculation formula is as follows: ;in, Intra-cluster information entropy; : No. The total number of samples within a homogeneous risk cluster; : No. The probability of the occurrence of a feature within a cluster of a sample graph structure feature vector; : Natural logarithm function.

[0034] Specifically, the process of generating a high-entropy refined sample subset by stripping away homogeneous redundant samples with low information gain and extracting representative transaction samples with centrality peaks and marginal free distributions is as follows: A baseline entropy value is set, and all samples within homogeneous risk clusters whose information entropy is lower than the baseline entropy value are removed, thus completing the stripping of homogeneous redundant samples. In the remaining homogeneous risk clusters, samples whose industrial node influence centrality reaches the highest value range are extracted as centrality peak samples. The physical topological distance of each sample within the remaining homogeneous risk cluster from its corresponding cluster center is calculated, and the sample with the largest physical topological distance is extracted as a marginal free distribution sample. The centrality peak sample and the marginal free distribution sample are tensor-merged to obtain representative transaction samples, outputting the high-entropy refined sample subset.

[0035] In this implementation plan, to maximize the productivity of the manual risk control review team, the system first sets a baseline entropy value. It then directly removes all samples from homogeneous risk clusters whose information entropy is below this baseline, thoroughly eliminating redundant samples. Low entropy means that hundreds or thousands of abnormal alarms within the cluster are repeated manifestations of the same single trigger; removing these effectively prevents invalid alarm storms from crowding out computing power. Regarding the baseline entropy value... The method for determining the risk clusters is as follows: This implementation plan discloses the extraction of the entropy distribution sequence of risk clusters over a full year, and takes the lower quartile of this sequence as the dynamic baseline. Subsequently, in the retained homogeneous risk clusters, samples with the highest centrality of industrial node influence are extracted as centrality peak samples. This ensures that the selected samples are necessarily the most fatal and destructive typical cases that destroy the core hubs of the supply chain. At the same time, in order to prevent the risk control model from having blind spots in its understanding of new and unknown risks, the system calculates the physical topological distance of each sample in the retained homogeneous risk clusters from its corresponding cluster center, and extracts the sample with the largest physical topological distance as the marginal free distribution sample. This calculation accurately captures rare variant samples that are floating on the edge of large-scale risk clusters. Specifically, the calculation logic of the physical topological distance is as follows: ;in, Physical topological distance; Topological space mapping constant; : No. Cluster center feature column vectors of isomorphic risk clusters; : Matrix transpose symbol. Ultimately, the system performs tensor merging of the centrality peak samples representing core business breakout points with the marginal free distribution samples representing potential evolutionary trends to obtain representative transaction samples, outputting a high-entropy refined sample subset. This operation, while drastically compressing the volume of data to be reviewed, perfectly preserves the highest-dimensional anomaly diversity and core representativeness, achieving a crucial purification and transformation of high-risk data into high-value business knowledge.

[0036] Please see Figure 4 Specifically, the process of monitoring the concurrent throughput of industrial transaction cycles and the real-time cognitive load threshold of the manual review resource pool, and generating candidate sample slices by sorting the high-entropy refined sample subset in descending order according to the business anomaly exposure index, is as follows: The order creation frequency and fund settlement frequency within the industrial transaction cycle are collected and combined to generate the concurrent throughput; the number of online review specialists, average historical review time, and backlog of work orders in the manual review resource pool are obtained, and the real-time cognitive load threshold is generated through weighted combination calculation; the business anomaly exposure index corresponding to each sample in the high-entropy refined sample subset is obtained; the high-entropy refined sample subset is sorted in descending order according to the business anomaly exposure index; and the high-entropy refined sample subset after descending order is segmented and truncated according to a preset step size to generate multiple candidate sample slices.

[0037] In this implementation plan, the order creation frequency and fund settlement frequency within the industrial transaction cycle are first collected, and then smoothly combined to generate concurrent throughput. This step aims to monitor the current business pressure on the industrial system in real time, as transaction peaks often occur in the early morning of commodity trading or at the end of logistics delivery. Next, to prevent blind pushes from a purely machine-centric perspective from overwhelming the risk control team with abnormal alarms, the system obtains the number of online review specialists in the human review resource pool, the average historical review time, and the number of backlogged work orders. These are then weighted and combined to generate a real-time cognitive load threshold. The real-time cognitive load threshold represents the maximum remaining available man-hours that the current human team can provide without crashing or serious misjudgments. Specifically, the real-time cognitive load threshold is set as follows: The calculation logic is as follows: ;in, Real-time cognitive load threshold; : Specialist fatigue attenuation factor; Number of online review specialists; : Current monitoring time window length; : Weighting of urgency penalty for backlogged work orders; Number of backlogged work orders; Average historical review time. (For specialist fatigue attenuation factor) The method for determining the factor involves using a nonlinear logistic regression function based on continuous on-duty time for fitting the output. The longer the team's continuous shift work time, the smaller this factor becomes, thus objectively reflecting the physiological decline in human brain computing power. Subsequently, the system obtains the business anomaly exposure index corresponding to each sample in the high-entropy refined sample subset output from the previous steps, and sorts the high-entropy refined sample subset in descending order according to the business anomaly exposure index. This operation ensures that the top high-risk samples with the greatest disruptive impact on the supply chain and the widest business impact are always at the forefront of the queue. Finally, the high-entropy refined sample subset, after being sorted in descending order, is segmented and truncated according to a preset step size, generating multiple candidate sample slices. The preset step size can be considered as the standard capacity benchmark for each data packet sent to the human terminal. Through segmentation and truncation, the originally large high-risk subset is cut into hierarchical task slices, laying a data structure foundation with appropriate granularity for subsequent human-machine supply and demand game.

[0038] Specifically, the process of generating a labeled industrial transaction risk sample set involves sorting a high-entropy refined sample subset in descending order according to the business anomaly exposure index to generate candidate sample slices, establishing a dynamic quota game mechanism between the candidate sample slices and the real-time cognitive load threshold, and extracting the optimal sample slice within the cognitive load carrying boundary and pushing it to the review terminal for manual identification. The specific steps are as follows: The number of samples contained in the candidate sample slices is converted into expected review time; the expected review time and the real-time cognitive load threshold are input into a supply-demand matching function to construct a dynamic quota game mechanism; in the dynamic quota game mechanism, candidate sample slices whose expected review time is close to and does not exceed the real-time cognitive load threshold are extracted, and these extracted candidate sample slices are marked as optimal sample slices; the transaction records and corresponding spatial flow logs within the optimal sample slices are packaged and pushed to the review terminal; the risk qualitative feedback instructions returned by the review terminal based on manual identification are received, the risk qualitative feedback instructions are converted into classification labels and bound to the corresponding samples within the optimal sample slices, and the labeled industrial transaction risk sample set is output.

[0039] In this implementation plan, the number of samples contained in the candidate sample slice is first converted into the expected review time. Because the complexity of the upstream and downstream supply chain networks involved in different abnormal transaction samples varies, the system cannot simply push samples by number; instead, it precisely measures the demand based on the expected human mental effort required. Specifically, the plan sets the... The number of samples contained in each candidate sample slice is The time required for basic manual identification of a single high-risk sample is The overall difficulty coefficient of the candidate slice samples is The expected review time for that slice is... The calculation logic is as follows: ;in, Expected review time; : The ascending sequence number of the candidate sample slice; : No. The number of samples contained in each candidate sample slice; The time required for basic manual identification of a single high-risk sample; : Sample overall difficulty coefficient. (Regarding the sample overall difficulty coefficient...) The method for determining the number of branches is disclosed in this implementation plan by using a linear mapping to determine the average number of branch connections of graph nodes in the topological space for all samples within the slice. A denser number of branch connections indicates a greater number of business entities requiring verification, thus increasing the difficulty level. Next, the expected review time is compared with the real-time cognitive load threshold calculated in the previous step. A dynamic quota game mechanism is constructed by inputting a supply-demand matching function. Essentially, this mechanism is a constrained resource optimization process aimed at finding a perfect balance between machine-pushed risk confirmation demands and the current human-capable computing power supply. Within this dynamic quota game mechanism, the system calculates the slice's carrying capacity reserve. ;in, : Slice bearing capacity; Real-time cognitive load threshold; Expected review time. From all slices in the set that satisfy the condition that the slice's capacity margin is greater than or equal to zero, the system extracts the slice whose expected review time most closely approximates the real-time cognitive load threshold, thus obtaining the slice's capacity margin. The system obtains the candidate sample slice with the smallest non-negative value and marks it as the optimal sample slice. This calculation step completely eliminates the system congestion and paralysis problem easily caused by traditional static thresholds, achieving an adaptive dynamic balance between computing power saturation and utilization while avoiding overload. Finally, the transaction records in the optimal sample slice and the corresponding spatial flow logs are packaged and pushed to the review terminal. Risk control specialists directly combine the cross data of financial statements and logistics trajectories for comprehensive manual judgment. The system receives the risk qualitative feedback instructions returned by the review terminal based on manual identification, converts the risk qualitative feedback instructions into machine-readable classification labels and binds them to the corresponding samples in the optimal sample slice, outputting the final labeled industrial transaction risk sample set. These accurately labeled data, which are accompanied by experts' high-level business insights, are subsequently re-fed into the underlying industrial basic model, thereby constructing an intelligent risk control closed loop of model self-evolution and continuously expanding business cognition boundaries.

[0040] In summary, this application has at least the following effects:

[0041] An adaptive selection and labeling method for risk samples in the transaction process is proposed. This method integrates multimodal heterogeneous data by combining the flow sequence of supply chain procurement transaction nodes with the spatial flow logs of physical warehousing and distribution nodes. Within a high-dimensional topological space, it generates a business anomaly exposure index by combining data concept drift degree and physical link hysteresis damping coefficient. This fundamentally breaks the limitations of a single financial data perspective and effectively filters out false positive anomalies that are not present in normal physical flows. Furthermore, by using density-based feature space clustering analysis and intra-cluster information entropy calculation, it accurately removes homogeneous redundant samples with low information gain and purifies a high-entropy refined sample subset containing centrality peaks and marginal detached distributions. Finally, it innovatively introduces a real-time cognitive load threshold for the manual review resource pool to construct a dynamic quota game mechanism. This enables the precise and timely delivery of optimal sample slices on demand and in appropriate quantities. This not only completely eliminates the congestion and computing power waste in the manual review system caused by explosive homogeneous anomalies, but also constructs an efficient human-machine collaborative risk control closed loop that closely matches the real industrial transaction turnover rhythm.

[0042] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0043] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0044] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0045] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0046] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0047] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An adaptive selection and labeling method for risk samples oriented towards transaction processes, characterized in that, Includes the following steps: S1. Obtain the transaction flow sequence of procurement nodes and the spatial flow log of corresponding physical warehousing and distribution nodes in the industrial manufacturing supply chain scenario. Perform heterogeneous data fusion on the transaction node flow sequence and spatial flow log according to the timestamp granularity and business entity identifier to construct a multimodal spatiotemporal transaction feature matrix. S2. Map the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space, extract graph structure feature vectors representing the flow of funds and materials, quantify the relative entropy of the graph structure feature vectors deviating from the historical baseline distribution to represent the data concept drift, analyze the spatial flow log to extract the physical link hysteresis damping coefficient, fuse the data concept drift and physical link hysteresis damping coefficient to generate the business anomaly exposure index, and extract high-frequency deviation samples from the multimodal spatiotemporal transaction feature matrix to construct a high-risk candidate sample pool; S3. Introduce the business anomaly exposure index as a distance metric weight, perform density-based feature space clustering analysis on the high-risk candidate sample pool to form multiple isomorphic risk clusters, calculate the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster, remove homogeneous redundant samples with low information gain, and extract representative transaction samples with centrality peak and marginal free distribution to generate a high-entropy refined sample subset. S4. Monitor the concurrent throughput of industrial transaction cycles and the real-time cognitive load threshold of the manual review resource pool. Sort the high-entropy refined sample subset in descending order according to the business anomaly exposure index to generate candidate sample slices. Establish a dynamic quota game mechanism between candidate sample slices and real-time cognitive load thresholds. Extract the optimal sample slice within the cognitive load carrying boundary and push it to the review terminal for manual identification to generate a labeled industrial transaction risk sample set.

2. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 1, characterized in that: Step S1 includes the following steps: The purchase order details and fund settlement records obtained from the supply chain enterprise resource planning system are integrated into a purchase transaction node sequence. The trajectory coordinates of the carrier vehicles and the scanning records of the warehouse throughput nodes uploaded by IoT devices are converted into spatial flow logs of physical warehousing and distribution nodes. Extract the globally unique order tracking code from the procurement transaction node flow sequence and spatial flow log as the business entity identifier, and perform time-dimensional slicing on the procurement transaction node flow sequence and spatial flow log according to the preset time window; Data segments with the same globally unique order tracking code and the same time dimension slice are concatenated into tensors to output a multimodal spatiotemporal transaction feature matrix.

3. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 1, characterized in that: The specific process of mapping the multimodal spatiotemporal transaction feature matrix to a high-dimensional topological space, extracting graph structure feature vectors representing the flow of funds and goods, and quantifying the relative entropy of the graph structure feature vectors deviating from the historical baseline distribution to represent the drift of data concepts is as follows: The transaction participants and warehousing operation areas in the multimodal spatiotemporal transaction feature matrix are analyzed as graph nodes, and the fund transfer path and logistics transportation route are analyzed as directed edges to construct a business entity relationship graph. The business entity relationship graph is mapped to a high-dimensional topological space, and the graph node centrality aggregation feature and directed edge weight feature are extracted and merged to generate a graph structure feature vector representing the flow of funds and materials. Extract the weighted graph structure baseline feature vectors within the historical normal trading cycle to generate a historical baseline distribution. Calculate the relative entropy between the probability density function of the current graph structure feature vector and the probability density function of the historical baseline distribution, which serves as the data concept drift degree.

4. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 3, characterized in that: The specific process of analyzing spatial flow logs to extract physical link hysteresis damping coefficients, fusing data concept drift and physical link hysteresis damping coefficients to generate a business anomaly exposure index, and extracting high-frequency deviation samples from the multimodal spatiotemporal transaction feature matrix to construct a high-risk candidate sample pool is as follows: Extract the standard estimated delivery timestamp and the actual node receipt timestamp from the spatial flow log, and calculate the time difference between the standard estimated delivery timestamp and the actual node receipt timestamp. The ratio of the time difference to the preset industry benchmark grace period is normalized to calculate the physical link hysteresis damping coefficient. Obtain the feature weights corresponding to the data concept drift degree and the physical link hysteresis damping coefficient, perform a linear combination and superposition on the weighted data concept drift degree and the physical link hysteresis damping coefficient, and output the business anomaly exposure index. By tracking the business anomaly exposure index of each transaction sample in the multimodal spatiotemporal transaction feature matrix using a time sliding window, transaction samples corresponding to abnormal fluctuation peaks that exceed the dynamic statistical confidence interval are extracted and marked as high-frequency deviation samples. High-frequency deviation samples are then merged to generate a high-risk candidate sample pool.

5. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 1, characterized in that: By introducing the business anomaly exposure index as a distance metric weight, density-based feature space clustering is performed on the high-risk candidate sample pool to form multiple isomorphic risk clusters. The specific process of calculating the industrial node influence centrality and intra-cluster information entropy of each isomorphic risk cluster is as follows: Extract the graph structure feature vector and the corresponding business anomaly exposure index of each sample in the high-risk candidate sample pool, convert the business anomaly exposure index into a spatial distance penalty coefficient, and calculate the weighted spatial distance between each sample by combining the graph structure feature vector. Based on weighted spatial distance execution density reachability clustering, the high-risk candidate sample pool is divided into multiple isomorphic risk clusters. The topological in-degree and out-degree of industrial circulation nodes contained in each isomorphic risk cluster are counted to generate the influence centrality of industrial nodes. Calculate the distribution dispersion of the feature vectors of the sample graph structure within each isomorphic risk cluster, and output the information entropy within the cluster.

6. The adaptive selection and labeling method for risk samples oriented towards the transaction process according to claim 5, characterized in that: The specific process of generating a high-entropy refined sample subset by stripping away homogeneous redundant samples with low information gain and extracting representative transaction samples with centrality peaks and marginal free distributions is as follows: Set a baseline entropy value, remove all samples in homogeneous risk clusters whose information entropy is lower than the baseline entropy value, complete the stripping of homogeneous redundant samples, and extract the samples in the remaining homogeneous risk clusters whose industrial node influence centrality reaches the highest value range as the centrality peak samples. Calculate the physical topological distance of each sample within the retained isomorphic risk cluster from the corresponding cluster center, extract the sample with the largest physical topological distance as the marginal free distribution sample, and perform tensor merging of the centrality peak sample and the marginal free distribution sample as the representative transaction sample, and output a high-entropy refined sample subset.

7. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 1, characterized in that: The specific process of monitoring the concurrent throughput of industrial transaction cycles and the real-time cognitive load threshold of the manual review resource pool, and generating candidate sample slices by sorting the high-entropy refined sample subset in descending order according to the business anomaly exposure index is as follows: Collect the order creation frequency and fund settlement frequency within the industrial transaction cycle, combine them to generate concurrent throughput, obtain the number of online review specialists in the manual review resource pool, the average historical review time, and the number of backlogged work orders, and generate a real-time cognitive load threshold through weighted combination calculation; Obtain the business anomaly exposure index corresponding to each sample in the high-entropy refined sample subset, sort the high-entropy refined sample subset in descending order according to the business anomaly exposure index, and perform a segmentation truncation operation on the high-entropy refined sample subset after the descending order according to a preset step size to generate multiple candidate sample slices.

8. The adaptive selection and labeling method for risk samples oriented towards transaction processes according to claim 7, characterized in that: The process of generating a labeled industrial transaction risk sample set by sorting a high-entropy refined sample subset in descending order according to the business anomaly exposure index, establishing a dynamic quota game mechanism between the candidate sample slices and the real-time cognitive load threshold, and selecting the optimal sample slice within the cognitive load carrying boundary to push to the review terminal for manual identification is as follows: The number of samples contained in the candidate sample slice is converted into the expected review time. The expected review time and the real-time cognitive load threshold are input into the supply and demand matching function to construct a dynamic quota game mechanism. In the dynamic quota game mechanism, candidate sample slices that are close to the expected review time and do not exceed the real-time cognitive load threshold are extracted, and the extracted candidate sample slices are marked as the optimal sample slices. The transaction records within the optimal sample slice and the corresponding spatial flow logs are packaged and pushed to the review terminal. The system receives risk qualitative feedback instructions returned by the audit terminal based on manual identification, converts the risk qualitative feedback instructions into classification labels and binds them to the corresponding samples in the optimal sample slice, and outputs a labeled industrial transaction risk sample set.

Citation Information

Patent Citations

  • Intelligent enterprise credit risk monitoring method and system based on machine learning

    CN120525628A

  • Abnormal transaction dynamic detection method and system fusing multi-scale analysis and information entropy

    CN121280143A