An inter-region identification method for a complex tax data system

By employing temporal coding embedding and adiabatic elimination principles to identify dominant factors in complex tax data systems within the tax control field, and combining this with frequent subgraph mining algorithms, the uninterpretability of deep learning models in the tax control field has been solved. This has enabled hierarchical partitioning and inter-region identification of complex systems, thereby improving the reliability and interpretability of the models.

CN115496570BActive Publication Date: 2025-11-25XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211311742.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-11-25
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing deep learning models suffer from uninterpretability issues in the tax control field, particularly in complex tax data systems where it is difficult to identify dominant factors and perform inter-regional segmentation, leading to insufficient interpretability and reliability of model predictions.

Method used

A nonlinear dynamic network model based on temporal coding embedding is adopted, which combines the adiabatic elimination principle and the frequent subgraph mining algorithm. By constructing a static snapshot and dynamic temporal embedding, the dominant factors of dynamic network evolution are identified, and a correlation matrix is ​​constructed based on binary vector coding to identify inter-regions.

Benefits of technology

It effectively improves the credibility and interpretability of the tax control system, enhances the interpretability of deep learning technology in the tax control field, and provides theoretical support for interpretable deep learning in other industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496570B_ABST
    Figure CN115496570B_ABST
Patent Text Reader

Abstract

The application discloses a kind of hierarchical division and meso-region identification method for complex dynamic network, comprising: first, by the method of "static snapshot construction-dynamic time sequence embedding" two stages, complex data system is transformed into semantic equivalent, including object, relationship, attribute and time sequence and the complex dynamic network of element;Second, based on the principle of adiabatic elimination in system science, the dominant factor of the subsystem concerned with dynamic network evolution is identified, and the boundary scale assumption space is constructed on this basis;Third, based on frequent subgraph mining algorithm, Motif in the subgraph instance of each boundary scale is mined;Finally, based on binary vector coding, the correlation matrix of each boundary scale is constructed, and based on conditional probability, the hierarchical coupling relationship of boundary scale is modeled, the hierarchical coupling of subgraph pattern is identified, and whether there is meso-region between two hypothesis spaces is determined by confidence threshold.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence and tax control technology, and particularly relates to a meso-region identification method for a complex tax data system. BACKGROUND

[0002] In recent years, deep learning technology has been successfully applied in the field of tax control, such as tax evasion detection, industry classification, and enterprise behavior anomaly detection. However, such end-to-end black-box intelligent learning models are not interpretable, which raises doubts about the credibility and fairness of the algorithm, and even causes legal problems, affecting the large-scale application of deep learning in the field of tax control. Deep learning models have a hierarchical structure and can automatically learn high-dimensional abstract features, but the acquired hierarchical features do not have a clear functional definition and cannot be associated with the problem to be solved, resulting in the uninterpretable nature of model prediction, i.e., knowing the result but not understanding the reason. The above phenomenon is an inherent shortcoming of existing deep learning models, which is difficult to solve by deep learning models themselves. The root cause lies in the non-decomposability of deep learning, which can only be solved with the help of external knowledge. Meso-science is a method for analyzing complex systems, which can discover meso-regions from data systems with spatiotemporal multi-scale dynamic structures, thereby reasonably decomposing complex systems. This idea is instructive for the construction of interpretable learning models. For example, invoice fraud detection in the tax control scene may exist in meso-regions such as "invoice-enterprise" and "enterprise-gang" at different levels. Therefore, reasonable hierarchical division and meso-region identification of complex tax data systems are the basis for enhancing the interpretability of deep learning technology in the field of tax control. Furthermore, hierarchical division and meso-region identification of complex data systems require the identification of dominant factors in the system, i.e., factors that play a major role in the evolution of the system. For example, the dominant factor in the evolution of the complex data system in the invoice fraud scene may be "invoice flow direction". Due to the difference, dynamics, and cross-temporal and spatial distribution characteristics of complex data systems, the identification of dominant factors is a difficult problem.

[0003] To solve this problem, the following documents provide solutions to the hierarchical division and meso-region identification of complex data systems:

[0004] Document 1: A method for improving the quality of deep learning data sets and the interpretability of models based on meso-science guidance (201910566328.2);

[0005] Document 2: An abnormal financial organization hierarchical division system based on neighborhood topological structure and its working method (202011009471.0);

[0006] Document 1 proposes a method for improving the quality of deep learning data sets and the interpretability of models using interdisciplinary guidance, which is based on the physical content of the object being processed, extracts the dominant mechanism, establishes a mesoscale model, and associates the mesoscale model with the deep learning process to improve the interpretability of the deep learning model.

[0007] Document 2 divides abnormal financial organizations into levels based on neighborhood topology, learns network representation of suspicious financial organization transaction records, represents financial organization accounts as corresponding low-dimensional dense vectors, and finally clusters the vectors through K-means to complete the hierarchical division of abnormal financial organizations, reducing manual intervention in such tasks. This method uses network representation learning to represent financial organization records and uses unsupervised clustering algorithms for classification and division. The data only includes financial organizations and their transaction records, and the data system is simple and can be divided. The dominant factor of the data system is known, and the interpretability of the representation learning and unsupervised clustering algorithm is weak, making it difficult to apply to tax control fields with high interpretability requirements and unknown dominant factors.

[0008] The above-mentioned traditional methods can solve specific mesoscale model problems and hierarchical division problems, but it is difficult to directly extend to the meso-region identification of complex dynamic systems in the invoice falsification scenario. The reason is that the research object of the above-mentioned method has the characteristics of small scale, not obvious dynamic evolution characteristics, or known dominant factors, which is very different from the complex data system required by this project, and it is difficult to match the characteristics of strong spatio-temporal evolution and high level coupling of the complex data system in the invoice falsification scenario. SUMMARY

[0009] The purpose of the present application is to provide a meso-region identification method for complex tax data systems. First, a nonlinear dynamic network model based on time series coding embedding is proposed to address the spatio-temporal multi-level and nonlinear characteristics of complex tax data systems. Through the "static snapshot construction-dynamic time series embedding" two-stage method, the complex data system is converted into a complex dynamic network containing objects, relationships, attributes, and time series elements. Second, based on the adiabatic elimination principle in system science, the dominant factor of the subsystem of interest is identified as the dynamic network evolves, and a boundary scale hypothesis space is constructed. Third, based on the frequent subgraph mining algorithm, the motif Motif in the subgraph instances of each boundary scale is mined. Finally, based on binary vector coding, a correlation matrix for each boundary scale is constructed, and based on conditional probability, the hierarchical coupling relationship of the boundary scale is modeled to identify the hierarchical coupling of the subgraph pattern. The confidence threshold is used to determine whether there is a meso-region between the two hypothesis spaces.

[0010] The present application is implemented by using the following technical solutions:

[0011] A method for inter-region identification of complex tax data system, comprising the following steps:

[0012] S101. Static snapshot construction-dynamic timing embedding; firstly, for the characteristics of complex tax data system in space-time multi-level and nonlinearity, based on the nonlinear dynamic network model of timing coding embedding, through the method of two-stage of "static snapshot construction-dynamic timing embedding", the complex data system is converted into a complex dynamic network containing objects, relations, attributes and timing elements which are semantically equivalent;

[0013] S102. Dynamic network evolution dominant factor identification based on adiabatic elimination; based on the adiabatic elimination principle in system science, the dominant factor of the subsystem concerned with the dynamic network evolution is identified, and the boundary scale hypothesis space is constructed on this basis;

[0014] S103. Boundary scale subgraph pattern mining; based on the frequent subgraph mining algorithm, the motif in the subgraph instance of each boundary scale is mined;

[0015] S104. Subgraph pattern hierarchical coupling identification; based on the binary vector coding, the correlation matrix of each boundary scale is constructed, and based on the conditional probability, the hierarchical coupling relationship of the boundary scale is modeled, the hierarchical coupling of the subgraph pattern is identified, and whether there is an inter-region between the two hypothesis spaces is determined by the confidence threshold.

[0016] The further improvement of the present application is that the method specifically comprises the following implementation steps:

[0017] 1) Static network snapshot construction

[0018] In order to convert the complex data system into a complex dynamic network containing objects, relations, attributes and timing elements which are semantically equivalent, firstly, the construction of static network snapshot is carried out, the ordered graph set of dynamic network in time is obtained, that is, the snapshot set of complex data system at different time;

[0019] 2) Dynamic timing embedding

[0020] In order to obtain the dynamic network evolution pattern while retaining the network structure information to the greatest extent, the dynamic network is subjected to timing coding embedding;

[0021] 3) Identification of dominant factor of dynamic network evolution

[0022] The process of a system changing from disorder to order or from low order to high order is called a "phase transition." The internal parameters of a system at the phase transition point are divided into slow relaxation variables and fast relaxation variables. Slow relaxation variables are few in number but decay slowly, and they are the fundamental variables that determine the phase transition of the system. Fast relaxation variables are relatively numerous but decay quickly. They are subordinate to slow relaxation variables and play an auxiliary role in the phase transition of the system. Slow relaxation variables change from nothing to something during the system evolution process, and through their dominance or enslavement of other fast relaxation variables, they dominate the overall evolution process of the system, indicate the formation of new structures, and reflect the degree of order of the new structures. These are the order parameters, that is, the dominant factors in the system evolution.

[0023] Based on the adiabatic elimination principle in systems science, we identify the dominant factor (DF) in the evolution of the subsystem of interest with the dynamic network, and construct a boundary scale hypothesis space under the guidance of DF.

[0024] 4) Boundary-scale subgraph pattern mining

[0025] Based on the frequent subgraph mining algorithm, the motifs of the ordering in subgraph instances at each boundary scale are mined. Based on graph pattern matching, the vector encoding of each node in the dynamic network under the motif is obtained. The vector encoding is used to indicate whether the corresponding motif is matched.

[0026] 5) Subgraph pattern hierarchical coupling identification

[0027] Based on the vector encoding in step 4), construct the correlation matrix W∈R between each boundary scale. n*n Based on conditional probability, the hierarchical coupling relationship between boundary scales is modeled to obtain a correlation matrix in probabilistic form. Finally, it is determined whether the hierarchical coupling relationship is satisfied between the two boundary scales, thereby identifying whether there is an inter-region between the two boundary scales.

[0028] A further improvement of this invention is that, in step 1), the construction of a static network snapshot specifically includes the following steps:

[0029] Step 1. Network Representation of Complex Tax Data Systems

[0030] To describe the structure of a complex tax data system, the system's network is defined based on the system's characteristics; G t =(V t E t V represents the network topology at time t. t With E t These represent the set of objects constituting the network and the set of relationships between those objects at that moment, respectively.

[0031] Step 2. Building Complex Tax Network Snapshots

[0032] Define a dynamic network model G = {G1, G2, ..., G...}T} is a temporally ordered atlas composed of static snapshots of a complex tax network at different times, with the network structure continuously adjusting as time progresses; where G t =(V t E t ) represents the network topology at time t, and the corresponding summary diagram is G. s =(V s E s ,L e W s Att s ), here, V s =V1∪V2∪…∪V T E s =E1∪E2∪…∪E T L e Represents the edge set E s The corresponding tag set is used to reflect the changes in network structure at different times, and its element l e It is a vector of dimension T containing only 0s or 1s, when l e When position t is 1, it means that edge e belongs to network G at time t. t =(V t E t Conversely, edge e does not belong to the network at that moment; W s Att represents the weight of the relationship between objects; s It is an attribute of the object, and each element's att e Each is a multidimensional matrix that stores the attribute characteristics of the object, such as Where n is the dimension of the attribute space.

[0033] A further improvement of the present invention is that, in Step 1), the complex tax data system is a whole formed by objects and relationships with different attributes under the action of temporal nonlinearity; the objects in the tax data system are heterogeneous data, including enterprises, legal persons, natural persons, tax authorities and industrial and commercial authorities.

[0034] A further improvement of this invention is that, in step 2), the dynamic timing embedding specifically includes the following steps:

[0035] Step 1. Data Preprocessing

[0036] Based on the network definition of the complex tax data system in step 1), the data is preprocessed into tuples with time-series labels:

[0037]

[0038] Where A represents the adjacency matrix of network nodes, X represents the node feature matrix, and Att represents the attribute feature matrix;

[0039] Step2. Network structure feature extraction within time step

[0040] The network structure features within the time step are obtained based on the GAT, and a static snapshot of the network at a time t is obtained, which is input to a two-layer GAT unit in the form of a tuple in Step 1, and the network structure feature vector H at the time t is obtained through weight learning, that is, H is:

[0041]

[0042] wherein, is the feature vector of the l layer, W t (l) is the projection matrix of the l layer, is the attention matrix of the l layer;

[0043] Step3. Network evolution time sequence embedding between time steps

[0044] The model weight information between time steps is updated to realize the time sequence embedding of network evolution, the model weight matrix W t (l) in Step 2 is regarded as the output of the dynamic system, and then based on the time sequence modeling capability of the Transformer neural network model, the output at the current time is obtained according to the input at the time t-1 and before, and the update formula is:

[0045]

[0046] wherein, W t (l) is the state information at the current time after time sequence embedding; Transformer() is a Transformer encoding unit; is the input at the current time and before, which is projected through the Transformer to obtain Then the current time t is obtained through the Transformer encoding unit, and the weight of 0 to t-1 time is obtained. The weighted sum is obtained to obtain the output at the current time. is initialized as the feature matrix of the node Wherein, n represents the number of input nodes, and d represents the feature vector dimension of the node.

[0047] The further improvement of the application is that in step 3), the dominant factor of dynamic network evolution is identified, which comprises the following steps:

[0048] Step1. Based on the adiabatic elimination principle in the synergetic theory, identify the dominant factors of the subsystem evolution with dynamic network, the dominant factors are the order parameters in the evolution process, the evolution direction of the order parameters determines the development direction of the system, as long as the order parameters are controlled, the development of the whole system can be grasped;

[0049] And identify the dominant factors of the subsystem evolution with dynamic network, the specific implementation steps are as follows:

[0050] Step101. Establish the reference system of the synergetic input state parameters of the data system

[0051] Suppose that the complex tax data system has a first-level state parameter:

[0052] X=[X1,X2,X3,…,X n ]

[0053] Where n is the dimension of the first-level state parameter;

[0054] At the same time, the first-level state parameter is further divided into second-level state parameters:

[0055] X i =[X i1 ,X i2 ,X i3 ,…,X ij ]

[0056] Where i is the serial number of the first-level state parameter, and j is the dimension of the current second-level state parameter;

[0057] In order to more comprehensively depict the external environment affecting the dynamic evolution of the complex tax data system, the network external state parameters are also established:

[0058] F=[F1,F2,F3,…,F n ]

[0059] F i =[F i1 ,F i2 ,F i3 ,…,F ij ]

[0060] Where n is the dimension of the first-level state parameter, i is the serial number of the first-level state parameter, and j is the dimension of the current second-level state parameter;

[0061] Step102. Establish the synergetic input-output constraint equation set of the data system

[0062] Taking the "input-output" model of synergetic network evolution as an example, let x i represent the synergetic input of the data system, q jThe cooperative output of the data system represents a nonlinear function f(x i , q j ) represents the constraint relationship between the cooperative input and the cooperative output of the data system, γ ij represents the constraint factor of the cooperative input state variable X on the cooperative output q, and therefore the following equation group exists:

[0063] q j =∑γ ij f(x i ,q j )

[0064] Wherein i and j represent the serial numbers of the cooperative input and output, i=1, 2, 3, …, n, and j=1, 2, 3, …, m.

[0065] Step 103. Adiabatic elimination is performed on the cooperative output of the data system

[0066] In order to perform linear stability analysis on the equation group pair in Step 102, the linear term is hidden, and it is assumed that the cooperative output is stable within a certain time, so that zero processing is performed, thereby reducing the dimension of the basic equation and reducing the degree of freedom of the equation. Let the cooperative output q j =0, and the equation group is obtained:

[0067] 0≈∑γ ij f(x i ,q j )

[0068] The state variable correlation table is obtained:

[0069]

[0070]

[0071] Through the state variable correlation table, the state variable with the smallest constraint effect is deleted, and the remaining state variables and the system output are adiabatically eliminated, and the sequence variable finally determining the evolution of the data system, i.e., the dominant factor of the subsystem evolution in the dynamic network, is obtained.

[0072] Step 2. Constructing an assumption space with boundary scales combined with domain knowledge

[0073] Under the guidance of the DF dominant factor, an assumption space with boundary scales ε={E1, E2, E3, …, E n} is constructed combined with the tax control field professional knowledge, m subgraph instances of each boundary scale in the dynamic network are extracted,

[0074] The further improvement of the present application is that in step 4), the boundary scale subgraph pattern mining specifically includes the following steps:

[0075] Step1. Based on the frequent subgraph mining algorithm, the subgraph instances of each boundary scale are mined

[0076] Based on the gSpan frequent subgraph mining algorithm, Motif in the graph set GD is mined, and is represented by Motif, and the specific steps are as follows:

[0077] Input: graph set GD, minimum support threshold min_sup

[0078] (1). The graph set GD is preprocessed as the basis of graph mining, and the necessary information such as the frequent edge set and the frequent node set is extracted from the graph set, in this part, GraphGen scans the graph set GD and obtains the frequent edge set, and the frequent edge set is arranged in descending order of frequency and increasing order of DFS code value;

[0079] (2). The frequent sub-tree is mined;

[0080] (3). For each frequent sub-tree, the inner edge that can be connected with the tree is found from the frequent edge set, and is added to the tree one by one, so that the frequent subgraph is formed, and the execution result of the algorithm is the frequent subgraph set FG;

[0081] Step2. Based on graph pattern matching, the binary vector code representation of each node j in the dynamic network V (|V| = N) in is obtained

[0082] Based on graph pattern matching, the binary vector code representation of each node j in the dynamic network V (|V| = N) in is obtained 1 represents that the corresponding Motif can be matched, and 0 represents that it cannot.

[0083] Further improvement of the application is that in step 5), the subgraph pattern hierarchical coupling identification specifically comprises the following steps:

[0084] Step1. Correlation matrix construction between boundary scales

[0085] Based on the binary vector coding, the correlation matrix of each boundary scale is constructed

[0086] W=∈R n*n

[0087] Wherein, W ij represents the number of times of Motif i occurrence; j

[0088] Step2. Hierarchical coupling relationship modeling between boundary scales

[0089] Based on the conditional probability, the hierarchical coupling relationship between the boundary scales is modeled, and first, the correlation matrix is converted into a probability form:​​​​

[0090]

[0091] Step3. Layered coupling relationship identification and inter-region identification

[0092] Based on the confidence threshold τ, it is determined whether E ij ≥τ, and whether there is a layered coupling relationship between E i and E j , and then it is identified that there is an inter-region between E i and E j .

[0093] The present application has at least the following beneficial technical effects:

[0094] The inter-region identification method for complex tax data system provided by the present application is used to solve the problem of uninterpretable intelligent learning model in complex tax data system. Through reasonable hierarchical division and inter-region identification of complex tax data system, the credibility and fairness of the subsequent tax control system results are effectively improved. Compared with the prior art, the present application has the following advantages:

[0095] (1) The present application divides the inter-region of complex tax data system in the scene of fake invoice, creatively introduces the adiabatic elimination principle in system science, divides the complex data system into multiple boundary scales and inter-regions, realizes the hierarchical division of complex system in the scene of fake invoice, and effectively enhances the explainability of deep learning technology in the field of tax control, and provides a basis for opening the black box model in the field of tax control.

[0096] (2) The present application comprehensively utilizes the sequence processing model and the graph convolution model, divides the tax scene into different processing tasks to achieve the effect of solving different problems separately, and provides a new solution for processing complex dynamic network structure problems.

[0097] (3) The present application has wide application, and is also applicable to other industries with high explainability requirements, is conducive to the popularization and application of explainable deep learning technology in the fields of medical treatment, automatic driving and finance, and provides theoretical support for building a new generation of explainable artificial intelligence. BRIEF DESCRIPTION OF DRAWINGS

[0098] Figure 1 It is a general flow chart of hierarchical division and inter-region identification based on complex dynamic network of the present application.

[0099] Figure 2 It is a dynamic network construction graph of the present application static snapshot construction-dynamic time sequence embedding two-stage.

[0100] Figure 3Flowchart for identifying inter-regional flow based on complex dynamic network of the present application. DETAILED DESCRIPTION

[0101] In order to more clearly illustrate the technical solutions of the present application, a kind of inter-regional identification method for complex tax data system of the present application is described in detail below in combination with the drawings and specific embodiments. Specifically, select the transaction network between 9 enterprises in the taxpayer information registered in a certain region from 2015 to 2017 in the state tax as a specific embodiment, which contains 5 industry categories.

[0102] In this example, an inter-regional identification method for complex tax data system provided by the present application includes the following steps:

[0103] S101. Static snapshot construction-dynamic time series embedding

[0104] Specifically, as Figure 2 shown is the dynamic network construction graph of the two-stage static snapshot construction-dynamic time series embedding in the present application, including the following steps:

[0105] Step1. Static network snapshot construction

[0106] In order to realize the conversion of complex data system into semantic equivalent, complex dynamic network containing objects, relationships, attributes and time series, etc. First, the construction of static network snapshot is carried out, and the ordered graph set of dynamic network in time is obtained, that is, the snapshot set of complex data system at different time; wherein G t =(V t , E t ) represents the network topology graph at time t, V t and E t represent the object set and the relationship set between objects at this time, respectively.

[0107] Specifically, in this embodiment, for simplicity, as Figure 2 shown, each black dot represents an enterprise, and the connection between the black dots represents a transaction between two enterprises at a certain time, wherein the enterprise has the attributes enterprise code (Qydm), enterprise type (Qylx), invoice code (Fpdm), invoice category (Fplb), industry code (Xydm), industry name (Xymc), registration type code (Djzclxdm) and invoice time (Kpsj). G t represents the transaction graph between all enterprises at time t.

[0108] Step2. Dynamic time series embedding

[0109] In order to obtain the dynamic network evolution mode while preserving the network structure information to the greatest extent, the dynamic network needs to be time series coding embedded, and the complex dynamic network is constructed by combining the "static snapshot construction-dynamic time series embedding" two stages, and the complex data system is converted into a semantically equivalent complex dynamic network.

[0110] Specifically, in the present embodiment, as shown in Figure 2 The combination of step 1) and step 2) is shown, where the black dot V t represents the enterprise object involved in the transaction, and the connection between the dots represents E t , which represents a transaction between objects, and G t =(V t ,E t ) represents the network topology at time t.

[0111] The overall system diagram is G s =(V s ,E s ,L e ,W s ,Att s ), where V s =V1∪V2∪…∪V T represents the set of all enterprises involved from time 0 to time T, E s =E1∪E2∪…∪E T represents the set of all transactions between enterprises from time 0 to time T, as shown in Figure 2 Take 9 enterprises as an example. L e represents the label set corresponding to the edge set E s , which is a T×E matrix, E represents the number of all edge sets from time 0 to time T, and each element l e of L e is a vector with only 0 or 1 in dimension T, when the position t is 1, it means that the edge e belongs to the network G t =(V t ,E t ) at time t, otherwise, the edge e does not belong to the network at that time; as l e =[1,0,1,…,0,1] T indicates that there is a transaction between the two enterprises connected at times t=1, 3, T. W s represents the weight of the object relationship, which is used here to indicate the size of the transaction between two enterprises, and is a T×q×q three-dimensional matrix, which is T×9×9 here, 9 is the number of enterprises involved from time 0 to time T, W s is used to represent the dynamic transaction between all enterprises from time 0 to time T. Att sis the attribute of the object, which is a matrix of size q x T x n, and its elements correspond to enterprises one by one, and is used to represent the attribute value of the enterprise at a certain time, and the element att e is a multi-dimensional matrix that stores the attribute characteristics of the enterprise, where n is the dimension of the attribute space. Here, n = 8, which respectively represents the enterprise code (Qydm), the enterprise type (Qylx), the invoice code (Fpdm), the invoice category (Fplb), the industry code (Xydm), the industry name (Xymc), the registration type code (Djzclxdm) and the invoice time (Kpsj), the invoice code of the party that does not issue an invoice in the transaction, and the invoice category attribute is empty, so here W s is a 3-dimensional matrix of size 9 x T x 8.

[0112] In this example, G t represents the static snapshot of the transaction of the system at time t, and is constantly adjusted as the time sequence changes, and Δ t→t+1 represents the transaction change between enterprises from time t to time t+1, and is used to indicate the update of the transaction network at the next time. In addition, the time sequence coding embedding is performed on the dynamic network, and the graph attention network GAT is used to extract the transaction network characteristics of the enterprise at time t, and the time sequence neural network (here, taking the Transformer as an example) is used to process the time sequence and combine the constructed static snapshot to convert into Figure 2 the complex dynamic network at the bottom of the middle, where each node V i is provided with a 3-dimensional matrix W s of size 9 x T x 3, which is used to represent its dynamics.

[0113] Specifically, in this embodiment, the graph attention network GAT t at time t is composed of multiple layers, and at time t, the adjacency matrix A t and the node embedding matrix are used as the input of the lth layer network, and the attention matrix and the weight matrix W t (l) The node embedding matrix is updated to obtain and output, and the update formula is:

[0114]

[0115] Here, is the normalized A t , which is defined as follows:

[0116]

[0117] where σ is the ReLU activation function, and the initialized embedding matrix is the node feature matrix, Thus which contains the high-level representation of features from the original graph nodes. Then the weight matrix As the output of the dynamic network, the Transformer encoding unit is used to model the input-output relationship, The update formula is:

[0118]

[0119] S102. Dynamic network evolution dominant factor identification based on adiabatic elimination; Specifically, the following steps are included:

[0120] Step1. Based on the adiabatic elimination principle in the theory of synergetics, identify the dominant factors of the subsystem of interest with the evolution of the dynamic network. The dominant factor is the order parameter that plays a dominant role in the evolution process. The evolution direction of the order parameter determines the development direction of the system. As long as the order parameter is controlled, the development of the entire system can be grasped.

[0121] To identify the dominant factors of the subsystem with the evolution of the dynamic network, the specific execution steps are as follows:

[0122] Step101. Establish the reference system of the cooperative input state parameter of the data system

[0123] Suppose there is a first-level state parameter in the complex tax data system:

[0124] X = [X1, X2, X3, …, X n ]

[0125] Where n is the dimension of the first-level state parameter.

[0126] Specifically, in this embodiment, the first-level state parameter is set as the capital investment of the enterprise, the transaction between enterprises, the business of the enterprise, the relationship capital of the enterprise, and the strategic goal of the enterprise. Thus, X = [X1, X2, X3, X4, X5].

[0127] At the same time, the first-level state parameter is further divided into second-level state parameters:

[0128] X i = [X i1 ,X i2 ,X i3 ,…,X ij ]

[0129] Where i is the serial number of the first-level state parameter, and j is the dimension of the current second-level state parameter.

[0130] Specifically, in this embodiment, it is set as:

[0131] X1 = [Employment, investment in equipment, expert consultation, investment in human resources, industrial layout investment]

[0132] X2 = [Inter-enterprise transaction time, transaction type, transaction amount, transaction frequency]

[0133] X3 = [number of business, business frequency, business type, business scope]

[0134] X4 = [value indicators of enterprises, transaction relationship networks of enterprises, invisible value of enterprises, tax level of enterprises]

[0135] X5 = [corporate culture, corporate reward and punishment mechanism, corporate development plan, corporate department coordination relationship]

[0136] The relaxation coefficient of the secondary state variable is much smaller than the relaxation coefficient of the primary state variable, and can be adiabatically eliminated. High-level state variables dominate low-level state variables, and order parameters exist in high-level state variables.

[0137] Further, in order to more comprehensively depict the external environment affecting the dynamic evolution of complex tax data systems, while establishing network external state parameters:

[0138] F = [F1, F2, F3, …, Fn] n ]

[0139] Where n is the dimension of the primary state variable, and i is the serial number of the state variable

[0140] Specifically, in this embodiment, it is assumed that:

[0141] F = [social development, industry status, market fluctuations, competitor status]

[0142] Step 102. Establish data system collaborative input-output constraint relationship equation set

[0143] Taking the “input-output” model of collaborative network evolution as an example, let x i (i = 1, 2, 3, …, n) represent the collaborative input of the data system, q i (i = 1, 2, 3, …, m) represent the collaborative output of the data system, there is a nonlinear function f(x i , q i ) represents the constraint relationship between the collaborative input and the collaborative output of the data system, γ ij represents the constraint factor of the collaborative input state variable X on the collaborative output q, so there is the following equation set:

[0144] q i = ∑γ ij f(xi , i )

[0145] where i represents the serial number of the collaborative input and output

[0146] Specifically, in the present embodiment, let q i (i = 1, 2, 3, 4, 5) each represent five collaborative outputs of enterprise risk, enterprise development trend, enterprise effectiveness, enterprise influence, and enterprise prospect. i (i = 1, 2, 3, 4, 5) respectively represent capital investment of the enterprise, transaction between enterprises, business operation of the enterprise, relationship capital of the enterprise, and strategic goal of the enterprise. Use γ ij f(x i , q i ) to represent the constraint relationship of the input and output state variables, where γ ij represents the constraint strength, then:

[0147] q1= γ 11 f(x1,q1)+ γ 21 f(x2,q1)+ γ 31 f(x3,q1)+ γ 41 f(x4,q1)+ γ 51 f(x5,q1)

[0148] q2= γ 12 f(x1,q2)+ γ 22 f(x2,q2)+ γ 32 f(x3,q2)+ γ 42 f(x4,q2)+ γ 52 f(x5,q2)

[0149] q3= γ 13 f(x1,q3)+ γ 23 f(x2,q3)+ γ 33 f(x3,q3)+ γ 43 f(x4,q3)+ γ 53 f(x5,q3)

[0150] q4= γ 14 f(x1,q4)+ γ 24 f(x2,q4)+ γ 34 f(x3,q4)+ γ 44 f(x4,q4)+ γ 54 f(x5,q4)

[0151] q5= γ 15 f(x1,q5)+ γ 25 f(x2,q5)+ γ 35f(x3, q5) + γ 45 f(x4, q5) + γ 55 f(x5, q5)

[0152] where γ ij represents the cooperative input state variable x i of the cooperative output q j .

[0153] Step 103. Adiabatic elimination of the cooperative output of the data system

[0154] To do linear stability analysis on the equation pair in Step 102, the linear term needs to be hidden, and assuming that the cooperative output is relatively stable within a certain time, zero processing is done, thereby reducing the dimension of the basic equation and reducing the degree of freedom of the equation, and setting the cooperative output q i = 0, the equation group is obtained:

[0155] 0 ≈ ∑ γ ij f(x i , q i )

[0156] Thus, the state variable correlation table is obtained:

[0157] X1 X2 X3 X4 X5 q1 True False True True True q2 True False False True True q3 True False False True False q4 False False True True False q5 False False True False True

[0158] Through the state variable correlation table, the state variable with the smallest constraint effect can be deleted, and then the remaining state variables and system outputs are adiabatically eliminated, and the order parameter that determines the evolution of the data system can be obtained, that is, the dominant factor of the subsystem evolution along with the dynamic network in the present application.

[0159] Specifically, in Figure 3 , each long strip represents a factor of system evolution, and through the principle of adiabatic elimination, the most critical factor of dimension i that determines the change of the system is found, that is, the dominant factor DF. From the above table, it can be seen that X2 and q i (i = 1, 2, 3, 4, 5) have no relationship, and X4 has the smallest change, which can be used as the dominant factor DF of the complex tax system.

[0160] Step 2. Constructing an assumption space with boundary scales combined with domain knowledge

[0161] Under the guidance of the DF dominant factor, an assumption space with boundary scales ε = {E1, E2, E3, …, E n} is constructed combined with tax control field professional knowledge, and m subgraph instances of each boundary scale in ε in the dynamic network are extracted,

[0162] S103. Boundary scale subgraph pattern mining, specifically comprising the following steps:

[0163] Step 1. Mining subgraph instances of each boundary scale based on a frequent subgraph mining algorithm

[0164] Mining Motifs in based on the gSpan frequent subgraph mining algorithm, denoted as, the specific steps are as follows: Input: graph set GD, minimum support threshold min_sup

[0165] (1) Preprocess the graph set GD as the basis for graph mining, extract the necessary information such as frequent edge sets and frequent node sets from the graph set, in this part, GraphGen scans the graph set GD and obtains the frequent edge set, arranges the frequent edge set in descending order of frequency and increasing order of DFS encoding value;

[0166] (1) Preprocess the graph set GD as the basis for graph mining, extract the necessary information such as frequent edge sets and frequent node sets from the graph set, in this part, GraphGen scans the graph set GD and obtains the frequent edge set, arranges the frequent edge set in descending order of frequency and increasing order of DFS encoding value;

[0167] Specifically, in this embodiment, input the complex tax network, traverse all static network snapshots at all times, calculate the frequency of all nodes and edges, compare the frequency with the minimum support threshold min_sup, and delete those nodes and edges that do not meet the threshold. Then reorder and number the remaining nodes and edges, calculate the frequency of the edges again, and then initialize the edge and perform the subtree mining process.

[0168] (2) Perform frequent subtree mining; specifically, in this embodiment, the current subgraph is recovered according to the graph encoding, and it is judged whether the current encoding is the minimum DFS encoding. If yes, it is added to the result set and the mining continues, if not, the mining ends.

[0169] (3) For each frequent subtree, find the inner edges that can be connected from the frequent edge set, and add them one by one to the tree, thereby forming a frequent subgraph, and the execution result of the algorithm is the frequent subgraph set FG.

[0170] Step 2. Obtain the binary vector code representation of each node j∈V(|V|=N) in the dynamic network under based on graph pattern matching

[0171] Specifically, based on graph pattern matching, obtain the binary vector code representation of each node j∈V(|V|=N) in the dynamic network under 1 represents that it can match the corresponding Motif, and 0 represents that it cannot.

[0172] S104. Subgraph pattern level coupling identification; specifically comprising the following steps:

[0173] Step 1. Correlation matrix construction between boundary scales

[0174] ​​Constructing correlation matrices for each boundary scale based on binary vector encoding.

[0175] W = ∈R n*n

[0176] Among them, W ij Motif i Motif in the following situations j Number of times it appears.

[0177] Step 2. Modeling the hierarchical coupling relationship between boundary scales

[0178] To model the hierarchical coupling relationship between boundary scales based on conditional probability, the correlation matrix is ​​first transformed into probabilistic form:

[0179]

[0180] Step 3. Identification of hierarchical coupling relationships and identification of inter-regions

[0181] Based on the confidence threshold τ, it is determined that P is satisfied. ij E ≥τ i and E j Whether there is a hierarchical coupling relationship between them, and thus identify E i and E j There is an intermediate region between them.

[0182] Specifically, in this embodiment, such as Figure 3 The right side shows the subgraph pattern hierarchical coupling identification module, which constructs the correlation matrix W∈R between each boundary scale based on the binary vector encoding in step 4. n*n W ij Motif i Motif in the following situations j The number of times it appears, that is, the number of times another subgraph exists given that a certain subgraph exists.

[0183] Furthermore, based on conditional probability mining of hierarchical coupling relationships between boundary scales, the correlation matrix is ​​first transformed into a probabilistic form P. ij =P(Motif) j |Motif i ) = W ij / N i N i =∑ j W ij Secondly, based on the confidence threshold τ, it is determined whether P is satisfied. ij E ≥τ i and E j The hierarchical coupling between them indicates a strong correlation, thus identifying E. i and Ej There is an intermediate region between the two edges, that is, there is an intermediate region between the two edges in this embodiment.

[0184] This embodiment uses a hierarchical division and intermediate region identification method of complex dynamic network, the specific process includes:

[0185] 1) Through the method of "static snapshot construction-dynamic time sequence embedding" two stages, the complex data system is transformed into a complex dynamic network containing objects, relationships, attributes and time sequence elements.

[0186] 2) Based on the adiabatic elimination principle in system science, the dominant factor of the subsystem concerned with the evolution of the dynamic network is identified, and the boundary scale assumption space is constructed on this basis.

[0187] 3) Based on the frequent subgraph mining algorithm, the Motif in the subgraph instance of each boundary scale is mined.

[0188] 4) Based on the binary vector coding, the correlation matrix of each boundary scale is constructed, and the boundary scale level coupling relationship is modeled based on the conditional probability, the subgraph pattern level coupling is identified, and whether there is an intermediate region between the two assumption spaces is determined by the confidence threshold.

[0189] Those skilled in the art will readily understand that the above description is only a method embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for inter-area identification for complex tax data systems, characterized in that, Comprising the following steps: S101. Static snapshot construction-dynamic timing embedding; first, for the characteristics of complex tax data system space-time multi-level, nonlinear, based on the nonlinear dynamic network model of timing coding embedding, through the method of "static snapshot construction-dynamic timing embedding" two stages, the complex data system is transformed into a semantic equivalent, containing object, relationship, attribute and timing element complex dynamic network; S102. Dynamic network evolution dominant factor identification based on adiabatic elimination; based on the principle of adiabatic elimination in system science, the dominant factor of the subsystem concerned with the evolution of dynamic network is identified, and the boundary scale hypothesis space is constructed; S103. Boundary scale subgraph pattern mining; based on the frequent subgraph mining algorithm, mining motifs in subgraph instances of each boundary scale ; S104. Subgraph pattern hierarchical coupling identification; Based on the binary vector coding, the correlation matrix of each boundary scale is constructed, and based on the conditional probability, the hierarchical coupling relationship of boundary scale is modeled, the hierarchical coupling of subgraph pattern is identified, and whether there is an intermediate area between the two hypothesis spaces is determined by the confidence threshold; The method specifically comprises the following implementation steps: 1) Static network snapshot construction In order to realize the transformation of complex data system into semantic equivalent, containing object, relationship, attribute and timing element complex dynamic network, first, the construction of static network snapshot is carried out, and the ordered graph set of dynamic network in time is obtained, that is, the snapshot set of complex data system at different time; 2) Dynamic timing embedding In order to obtain the dynamic network evolution pattern while preserving the network structure information to the greatest extent, the dynamic network is timing coding embedded; 3) Identify the dominant factor of dynamic network evolution The change process of system from disorder to order or from low order to high order is called "phase transition", the internal parameters of system at phase transition are divided into slow relaxation variable and fast relaxation variable, the number of slow relaxation variable is small but the decay speed is slow, which is the fundamental variable that determines the phase transition of system; the number of fast relaxation variable is relatively large but the decay speed is fast, which is subject to slow relaxation variable and plays an auxiliary role in the phase transition of system; the slow relaxation variable changes from nothing to something in the evolution process of system, and dominates the whole evolution process of system through the domination or service of other fast relaxation variables, indicates the formation of new structure, reflects the order degree of new structure, that is, the order parameter, that is, the dominant factor of system evolution; Based on the adiabatic elimination principle in system science, the dominant factors of the subsystem of interest evolving with the dynamic network are identified DF , under the guidance of DF , the boundary scale hypothesis space is constructed; 4) Boundary scale subgraph pattern mining Based on the frequent subgraph mining algorithm, the motifs in the boundary scale subgraph instances are mined , based on the graph pattern matching, the vector coding of each node in the dynamic network is obtained , which is used to indicate whether the corresponding is matched 5) Subgraph pattern hierarchical coupling identification Based on the vector coding in step 4), a correlation matrix between each boundary scale is constructed Based on the conditional probability, the hierarchical coupling relationship between boundary scales is modeled to obtain a probability form of the correlation matrix. Finally, it is determined whether the hierarchical coupling relationship between two boundary scales is satisfied, and whether there is an inter-region between the two boundary scales.

2. The method of claim 1, wherein, Step 1) of the static network snapshot construction specifically comprises the following steps: Step1. Network representation of complex tax data system In order to describe the complex tax data system structure, the network of the system is defined according to the characteristics of the tax data system; denotes the network topology at the moment, and denotes the object set and the relationship set between objects constituting the network at the moment, respectively; Step2. Complex tax network snapshot construction Define dynamic network model It is a time-ordered atlas, composed of static snapshots of a complex tax network at different moments, with the network structure continuously adjusting as time progresses; among which... express The network topology diagram at any given time, and the corresponding summary diagram are as follows: ,here, ; Representation of edge set The corresponding tag set is used to reflect the changes in network structure at different times, and its elements It is a dimension of A vector containing only 0s or 1s, when Middle position When it is 1, it represents an edge. Belongs to the The network of time Conversely, the side Not part of the network at that moment; Represents the weight of the relationship between objects; It is a property of the object, each element Each is a multidimensional matrix that stores the attribute characteristics of the object, such as ,in This represents the dimension of the attribute space.

3. The method of claim 2, wherein, In Step1 of Step1 of step 1), the complex tax data system is formed by objects and relationships with different attributes under the action of time sequence nonlinearity; the objects in the tax data system belong to heterogeneous data, including enterprises, legal persons, natural persons, tax authorities and business authorities.

4. The method of claim 2, wherein, In Step 2), dynamic timing embedding specifically comprises the following steps: Step1. Data preprocessing According to the network definition of complex tax data system in Step 1), the data is preprocessed to be a tuple with timing mark: Wherein, A represents the adjacency matrix of the network node, X represents the node feature matrix, Att represents the attribute feature matrix; Step2. Network structure feature extraction within time step Based on GAT, the network structure features within the time step are obtained. At a certain time t, a static snapshot of the network is obtained, which is represented as a tuple in Step1. The input is input into a two-layer GAT unit, and the network structure feature vector H at time t is obtained through weight learning, that is, H is: wherein, , , is a feature vector of the l-th layer, is a projection matrix of the l-th layer, is an attention matrix of the l-th layer; Step3. Network evolution time series embedding between time steps By updating the model weight information between time steps, the time sequence embedding of network evolution is realized. Based on the Transformer neural network model, the model weight matrix in Step 2 is regarded as the output of the dynamic system, and then based on the time sequence modeling capability of the Transformer, the output at the current time is obtained according to the input at t-1 and the previous time, and the update formula is: ​ wherein, is the state information of the current time after timing embedding; Transform() is a Transformer encoding unit; is the input at the current and previous time, which is obtained by Transformer projection Then, the current time t is obtained by the Transformer encoding unit for the time from 0 to t-1, and The weighted sum is calculated to obtain the output of the current time; is initialized as the feature matrix of the node wherein n represents the number of input nodes, and d represents the feature vector dimension of the node.

5. The method of claim 4, wherein, In Step 3), the dominant factors of dynamic network evolution are identified, which includes the following steps: Step1. Based on the adiabatic elimination principle in the theory of synergetics, the dominant factors of the subsystem concerned with the dynamic network evolution are identified. The dominant factor is the order parameter in the evolution process, which determines the development direction of the system. As long as the order parameter is controlled, the development of the whole system can be grasped; To identify the dominant factors of the subsystem with dynamic network evolution, the specific execution steps are as follows: Step101. Establish the reference system of the cooperative input state parameter of the data system Suppose that the complex tax data system has a first-level state parameter: Where n is the dimension of the first-level state parameter; At the same time, the first-level state parameter is further divided into a second-level state parameter: Where i is the serial number of the first-level state parameter, and j is the dimension of the current second-level state parameter; In order to more comprehensively depict the external environment that influences the dynamic evolution of the complex tax data system, a network external state parameter is also established: Where n is the dimension of the first-level state parameter, i is the serial number of the first-level state parameter, and j is the dimension of the current second-level state parameter; Step102. Establish the equation set of the cooperative input-output constraint relationship of the data system Taking the "input-output" model of the co-evolution of the network as an example, let represent the co-evolution input of the data system, represent the co-evolution output of the data system, and there is a nonlinear function represent the constraint relationship between the co-evolution input and the co-evolution output of the data system, represent the constraint factor of the co-evolution input state variable X on the co-evolution output q, so there is the following equation group: where i and j represent the sequence number of the cooperative input and output, , ; Step103. Adiabatic elimination of the cooperative output of the data system To do linear stability analysis for the system of equations in Step 102, first hide the linear terms and assume that the coordinated output is stable over a certain time. Then, do the zeroing process to reduce the dimension of the basic equations and the degree of freedom of the equations. Let the coordinated output = 0, and obtain the system of equations: The state parameter correlation table is obtained: Through the state parameter correlation table, the state parameter with the smallest constraint effect is deleted, and the remaining state parameters and system output are subjected to adiabatic elimination to obtain the order parameter that determines the evolution of the data system, i.e. the dominant factor of the subsystem with dynamic network evolution; Step2. Construct the hypothesis space of boundary scale combined with domain knowledge In DF Guided by the leading factors, the hypothesis space of boundary dimensions is constructed combined with the professional knowledge of tax control field , extract m subgraph instances of each boundary dimension in the dynamic network, .

6. The method of claim 5, wherein, Step 4) in which the boundary scale subgraph pattern mining includes the following steps: Step1. Based on the frequent subgraph mining algorithm, the subgraph instances of each boundary scale are mined Based on gSpan Frequent subgraph mining algorithm, mining In With Indicates, the specific steps are as follows: Input: atlas GD, Minimum support threshold min_sup (1). For the atlas GD Preprocessing is performed as the foundation for graph mining, extracting necessary information such as frequent edge sets and frequent node sets from the graph set. In this part, GraphGen Scanned Image Set GD And obtain the frequent edge set, then decrease the frequency of the frequent edge set and... DFS Arrange the encoded values ​​in ascending order; (2). Frequent sub-tree mining is performed; (3). For each frequent sub-tree, find the inner edges from the frequent edge set that can be connected to it, and add them one by one to the tree, thereby forming a frequent sub-graph, and the execution result of the algorithm is a frequent sub-graph set FG; Step2. Based on graph pattern matching, each node in the dynamic network is obtained Based on the graph pattern matching, each node in the dynamic network is obtained In The binary vector code representation , 1 indicates that it can be matched to the corresponding , 0 indicates that it cannot.

7. The method of claim 6, wherein, Step 5) in which the subgraph pattern hierarchical coupling identification includes the following steps: Step1. Correlation matrix construction between boundary scales Based on binary vector coding, the correlation matrix of each boundary scale is constructed wherein represents in the case where the number of times Step2. Hierarchical coupling relationship modeling between boundary scales Based on conditional probability, the hierarchical coupling relationship between boundary scales is modeled. First, the correlation matrix is converted to a probability form: Step3. Hierarchical coupling relationship identification and inter-region identification Based on a confidence threshold Determining whether satisfies have a hierarchical coupling relationship, and then identifying that There is an inter-region between

Citation Information

Patent Citations

  • Method for improving quality of deep learning data set and interpretability of model based on mesoscience guidance

    CN110533159A

  • A System and Method for Hierarchical Classification of Anomalies in Financial Organizations Based on Neighborhood Topology

    CN112150285B