Tax risk identification method and system based on multi-source data comparison
By constructing a risk transmission network for comparing multi-source data from pharmaceutical companies, the risk transmission path between pharmaceutical companies and related entities is dynamically identified, solving the complexity of tax risk identification for pharmaceutical companies in existing technologies and achieving efficient tax risk prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGZHONG PHARMA CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing tax risk identification methods are unable to effectively address the unique and complex supply chain networks of pharmaceutical companies. They lack network models that provide real-time information on the relationships between pharmaceutical companies and their affiliates, resulting in risk identification remaining at the "point-like" analysis level and failing to provide early warnings of chain reactions triggered by risks from affiliates.
By constructing a tax risk identification method based on multi-source data comparison, multi-source data of target entities and related entities are collected, dynamic risk feature vectors are generated, a risk transmission network is established, the temporal correlation between nodes is calculated, high-risk clusters are marked, and the tax risk level is assessed.
It achieves closed-loop management from data collection to strategy generation, improves the accuracy and timeliness of risk identification, and can proactively identify and prevent tax risks for pharmaceutical companies.
Smart Images

Figure CN121998781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tax risk identification technology, specifically to a tax risk identification method and system based on multi-source data comparison. Background Technology
[0002] As a strategic industry vital to national welfare and people's livelihood, the pharmaceutical industry's tax risk management is highly complex and unique due to its distinctive business model, regulatory policies, and supply chain structure. Pharmaceutical companies generally exhibit characteristics of "high investment, long cycles, and high risks" in their R&D, and are widely subject to policies such as additional deductions for R&D expenses and tax incentives for high-tech enterprises, leading to complex tax treatments. Simultaneously, the multi-stage distribution system from raw materials and finished products to commercial distribution, along with stringent regulatory policies such as the "two-invoice system" and centralized volume-based procurement, collectively constitute a unique tax risk environment for pharmaceutical companies. The complex network of relationships between pharmaceutical companies and CROs, CMOs, CSOs, and distributors further increases the concealment and cascading nature of risk transmission.
[0003] Existing tax risk identification methods are mostly general frameworks, which are difficult to effectively address the unique challenges of pharmaceutical companies. These methods are usually based on static ratio analysis or threshold comparison of financial and tax statement data at a single point in time, relying on general indicators such as tax burden rate and invoice matching degree. This "one-size-fits-all" analysis model cannot dynamically capture the risk evolution of pharmaceutical companies throughout their entire life cycle, nor can it design effective risk monitoring indicators for policy shocks (such as drastic changes in price and revenue structure caused by volume-based procurement). Furthermore, it lacks specialized algorithmic models for potential collaborative anomalies in the complex related-party transaction networks of pharmaceutical companies, resulting in serious deficiencies in their applicability and accuracy in the pharmaceutical company scenario.
[0004] The key bottleneck of existing technologies lies in the lack of dynamic topological awareness of the complex supply chain networks of pharmaceutical companies. Current analytical methods cannot construct network models that reflect the real-time relationships between pharmaceutical companies and their CROs, CMOs, CSOs, and distributors. This results in the system being unable to depict the transmission paths of risks between multiple nodes, and also makes it difficult to locate key risk hubs within the network. This lack of capability keeps risk identification at the "point-like" analysis level, failing to provide early warnings of chain reactions triggered by risks from related parties, constituting a major technical obstacle to the accurate prevention and control of tax risks in pharmaceutical companies. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a tax risk identification method and system based on multi-source data comparison, which aims to solve the above-mentioned problems described in the prior art.
[0006] A first aspect of the present invention is to provide a method for identifying tax risks based on multi-source data comparison, the method comprising: Data from multiple sources of the target entity and related entities are collected from tax, invoice and commercial credit systems according to a preset cycle, and corresponding datasets are constructed. Based on the dataset, time series data of preset key risk indicators within a preset observation period are extracted, and volatility and deviation are calculated based on the time series data of each indicator to generate a dynamic risk feature vector representing the risk status of each subject. Based on the relationship between the target entity and the associated entities, a risk transmission network is constructed; wherein network nodes represent the target entity and all associated entities, and directed edges represent risk transmission paths. In the risk transmission network, the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window is calculated; When the correlation coefficient exceeds a preset threshold, it is determined that there is risky collaboration and the associated entities corresponding to the multiple associated entity nodes are marked as a high-risk cluster; Assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and output quantitative risk level and response strategy data.
[0007] According to one aspect of the above technical solution, the step of constructing a risk transmission network based on the relationship between the target entity and the associated entity includes: Obtain the relationship data between the target entity and each related entity, including the equity investment ratio, the percentage of annual transaction volume, and the number of companies with overlapping shareholdings; The target entity and each associated entity are mapped to network nodes respectively, and directed edges are established between any network nodes based on the directionality of the association relationship to generate a preliminary network topology. Based on the association data, a weighted fusion algorithm is used to calculate the association strength value of each directed edge, and the association strength value is assigned as the weight of the corresponding directed edge to complete the construction of the risk transmission network.
[0008] According to one aspect of the above technical solution, based on the correlation data, a weighted fusion algorithm is used to calculate the correlation strength value of each directed edge, and the correlation strength value is assigned as the weight of the corresponding directed edge to complete the construction of the risk transmission network, including: The equity investment ratio, annual transaction volume ratio and number of companies with overlapping shareholdings in the aforementioned relationship data are normalized to obtain the standardized values corresponding to each indicator. Configure a preset weight coefficient for each type of association indicator, and linearly weight and fuse the standardized values of the multiple types of association indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the association strength value of the directed edge. The association strength value is mapped to the visual label of the corresponding directed edge, wherein the association strength value is positively correlated with the width and color depth of the edge, thereby generating a risk transmission network that includes node attributes and weighted edge attributes.
[0009] According to one aspect of the above technical solution, the step of configuring a preset weight coefficient for each type of correlation index, and linearly weighting and fusing the standardized values of the multiple correlation indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the correlation strength value of the directed edge includes: The information entropy of the equity investment ratio, the proportion of annual transaction volume, and the number of companies with overlapping shareholdings are calculated respectively. The degree of variation of each correlation indicator is determined based on the information entropy, so as to dynamically configure the preset weight coefficient of each correlation indicator. The standardized values corresponding to the equity investment ratio, annual transaction volume ratio, and number of companies with overlapping shareholdings are multiplied by the preset weight coefficient and summed to obtain the association strength value of the directed edge, which is then normalized.
[0010] According to one aspect of the above technical solution, in the risk transmission network, the step of calculating the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window includes: In the risk transmission network, the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node is calculated as the first correlation index. Calculate the dynamic time-normalized distance between the time series and convert it into a similarity score as a second correlation index.
[0011] According to one aspect of the above technical solution, the step of calculating the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node, respectively, as the first correlation indicator, includes: The time series of dynamic risk feature vectors corresponding to the target subject node and the associated subject node are respectively subjected to mean centering. The covariance and standard deviation are calculated based on the centered time series, and the Pearson correlation coefficient is obtained based on the ratio of the covariance to the standard deviation.
[0012] According to one aspect of the above technical solution, the step of calculating the dynamic time-normalized distance between the time series and converting it into a similarity score as a second correlation index includes: Construct the Euclidean distance matrix between the time series of the target subject node and the associated subject node, and find the optimal regular path using a dynamic programming algorithm; Calculate the cumulative distance from the first point to the second point on the optimal regularized path, and map it to a similarity score in the range of 0 to 1.
[0013] A second aspect of the present invention is to provide a tax risk identification system based on multi-source data comparison, applied to the method described in the above-mentioned technical solution, the system comprising: The data acquisition module is used to collect multi-source data of the target entity and related entities from the tax, invoice and commercial credit systems according to a preset period, and construct the corresponding dataset; The vector generation module is used to extract time-series data of preset key risk indicators within a preset observation period based on the dataset, and calculate volatility and deviation based on the time-series data of each indicator to generate dynamic risk feature vectors representing the risk status of each subject. The network construction module is used to construct a risk transmission network based on the relationship between the target entity and the associated entities; wherein network nodes represent the target entity and all associated entities, and directed edges represent risk transmission paths; The time-series calculation module calculates the time-series correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window in the risk transmission network. The cluster marking module is used to determine the existence of risky collaboration and mark the associated entities corresponding to multiple associated entity nodes as high-risk clusters when the correlation coefficient exceeds a preset threshold. The risk assessment module is used to assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and outputs quantitative risk level and response strategy data.
[0014] A third aspect of the present invention is to provide a readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method described in the above-described technical solution.
[0015] A fourth aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the above technical solutions.
[0016] Compared with existing technologies, the tax risk identification method and system based on multi-source data comparison shown in this invention have the following advantages: The tax risk identification method based on multi-source data comparison described in this invention realizes closed-loop management from data collection and risk quantification to strategy generation by constructing dynamic risk feature vectors and weighted transmission networks. This method not only improves the accuracy and timeliness of risk identification, but also effectively overcomes the limitations of traditional static analysis by introducing time-series correlation analysis and networked topology modeling, thus providing reliable technical support for enterprises to carry out proactive and intelligent tax risk prevention and control in complex business environments. Attached Figure Description
[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 A flowchart illustrating the tax risk identification method based on multi-source data comparison provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of a tax risk identification system based on multi-source data comparison provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be more thorough and complete.
[0019] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] Example 1 Please see Figure 1 The first aspect of the present invention is to provide a tax risk identification method based on multi-source data comparison, the method comprising steps S10-S60: Step S10: Collect multi-source data of the target entity and related entities from the tax, invoice and commercial credit systems according to the preset cycle, and construct the corresponding dataset.
[0022] In this embodiment, the method shown is a tax risk identification method based on multi-source data comparison. Its primary purpose is to automatically collect multi-source data on target entities and related entities from the tax system, invoice system, and commercial credit system according to a preset time period, such as weekly or monthly, and then construct a corresponding multi-source data dataset. The target entity refers to the enterprise or individual applying this method, also known as the target taxpayer, while the related entity refers to an enterprise or individual that has a relationship with the target entity, such as business dealings, also known as related taxpayers.
[0023] Specifically, the aforementioned multi-source data sources include, but are not limited to: 1) Tax system data, such as the enterprise's value-added tax and corporate income tax returns and their supplementary information. This data reflects the enterprise's core tax obligation fulfillment and declaration logic; 2) Invoice management system data, covering the entire lifecycle information of input and output invoices, including the status and timing of issuance, authentication, deduction, cancellation, and red-ink cancellation, which is key to tracking the authenticity of the transaction chain; 3) External business credit system data, such as business registration information, equity change records, judicial litigation information, and administrative penalty records, which are used to depict the enterprise's compliance profile and related networks.
[0024] Step S20: Based on the dataset, extract the time series data of preset key risk indicators within a preset observation period, and calculate the volatility and deviation based on the time series data of each indicator to generate a dynamic risk feature vector representing the risk status of each subject.
[0025] In this embodiment, the core objective of step S20 is to extract massive, heterogeneous time-series data into a digital profile that accurately and dynamically reflects the individual risk status of an enterprise, namely, a dynamic risk feature vector. Specifically, this includes constructing a key risk indicator system, quantifying and calculating time-series features, and generating the final feature vector.
[0026] Specifically, this involves constructing a pre-defined key risk indicator system and extracting time-series data, including building a multi-dimensional, quantifiable indicator analysis system based on tax risk driving factors. These indicators are pre-configured and have clear business implications, such as: tax burden volatility, discrepancies between input and output invoices, abnormal changes in the ratio of period expenses to revenue, and abnormal growth trends in VAT credit carryforwards. In the dataset output in step S10, according to a pre-defined observation period such as the past 24 months, the values of these indicators at different time points, such as the end of each month, are extracted, thereby generating a time-series data for each indicator. This aims to transform static data snapshots into dynamic behavioral trajectories to facilitate the identification of risk evolution patterns.
[0027] Furthermore, by introducing mathematical models, the abnormal performance of indicator time series can be precisely measured, namely, the quantification of volatility and deviation. Volatility measures the stability of indicator values, typically calculated using the standard deviation or coefficient of variation over the period. High volatility often indicates instability in business or financial behavior. Deviation measures the degree to which current or recent indicator values deviate from their historical normal range. By calculating the volatility and deviation of each key risk indicator separately, lengthy and difficult-to-compare time series can be quickly compressed into two statistically significant and comparable numerical features, thereby improving the density and comparability of risk information.
[0028] Furthermore, the volatility and deviation calculated based on each key risk indicator serve as two independent feature dimensions. Combining these feature dimensions of all indicators in a predetermined order generates a dynamic risk feature vector representing the overall risk status of the entity during the observation period. This vector fully maps the complex risk situation of a company within a specific period onto a structured mathematical space, allowing the risk similarity between different companies to be quantitatively assessed by calculating the distance or similarity between vectors. This provides standardized and computable input for subsequent risk transmission network analysis.
[0029] Step S30: Based on the relationship between the target entity and the associated entity, a risk transmission network is constructed.
[0030] In this network, nodes represent the target entity and all related entities, and directed edges represent risk transmission paths.
[0031] In this embodiment, the step of constructing a risk transmission network based on the association between the target entity and the associated entity includes: Obtain the relationship data between the target entity and each related entity, including the equity investment ratio, the percentage of annual transaction volume, and the number of companies with overlapping shareholdings; The target entity and each associated entity are mapped to network nodes respectively, and directed edges are established between any network nodes based on the directionality of the association relationship to generate a preliminary network topology. Based on the association data, a weighted fusion algorithm is used to calculate the association strength value of each directed edge, and the association strength value is assigned as the weight of the corresponding directed edge to complete the construction of the risk transmission network.
[0032] Specifically, the two basic elements of the risk transmission network are clearly defined: nodes and edges. Each network node uniquely represents a corporate entity, i.e., the target entity or a pre-identified related entity. Directed edges are used to abstract specific relationships between enterprises and clearly indicate the potential direction of risk transmission. For example, an edge from a holding company to a subsidiary represents a risk transmission path under a control relationship; an edge from a supplier to the core enterprise represents a risk transmission path under a supply chain dependency relationship. Through these precise mapping relationships, the originally scattered cluster of enterprises is transformed into a network diagram with a clear topological structure, thus facilitating the analysis and identification of tax risks.
[0033] The steps involved in constructing the risk transmission network include: calculating the association strength value of each directed edge using a weighted fusion algorithm based on the association data, and assigning the association strength value as the weight of the corresponding directed edge. The equity investment ratio, annual transaction volume ratio and number of companies with overlapping shareholdings in the aforementioned relationship data are normalized to obtain the standardized values corresponding to each indicator. Configure a preset weight coefficient for each type of association indicator, and linearly weight and fuse the standardized values of the multiple types of association indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the association strength value of the directed edge. The association strength value is mapped to the visual label of the corresponding directed edge, wherein the association strength value is positively correlated with the width and color depth of the edge, thereby generating a risk transmission network that includes node attributes and weighted edge attributes.
[0034] The steps involved in configuring preset weight coefficients for each type of association indicator, and linearly weighting and fusing the standardized values of multiple association indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the association strength value of the directed edge, include: The information entropy of the equity investment ratio, the proportion of annual transaction volume, and the number of companies with overlapping shareholdings are calculated respectively. The degree of variation of each correlation indicator is determined based on the information entropy, so as to dynamically configure the preset weight coefficient of each correlation indicator. The standardized values corresponding to the equity investment ratio, annual transaction volume ratio, and number of companies with overlapping shareholdings are multiplied by the preset weight coefficient and summed to obtain the association strength value of the directed edge, which is then normalized.
[0035] Specifically, in this embodiment, each directed edge is assigned a weight value representing the strength of the association. This weight value is calculated based on objective data collected in S10, such as equity investment ratio, annual transaction volume, and degree of overlap among senior executives, using a preset weighted fusion algorithm. For example, the higher the shareholding ratio and the larger the transaction volume, the greater the weight value of the directed edge, indicating a stronger probability and influence of risk transmission along that path. Through the construction of the aforementioned weighted network, the model can distinguish between strong and weak associations, thereby more accurately simulating the dynamic process of risk transmission. Graph algorithms can identify key hub nodes in the network, i.e., the center of risk diffusion, or simulate the transmission path and scope of impact after a risk erupts from a certain point.
[0036] Step S40: In the risk transmission network, calculate the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window.
[0037] Step S50: When the correlation coefficient exceeds a preset threshold, it is determined that there is risky collaboration and the associated entities corresponding to the multiple associated entity nodes are marked as high-risk clusters.
[0038] In this embodiment, a risk transmission network is constructed between the target entity and related entities. Based on the topological relationship in this risk transmission network, the temporal correlation of the dynamic risk feature vectors corresponding to the target entity node and the related entities corresponding to the related entity node within a specific time window, such as the past month, is calculated. By comparing the correlation coefficient, such as the Pearson correlation coefficient, with a preset threshold, if the Pearson correlation coefficient exceeds the preset threshold, it is determined that there is risk synergy. Multiple related entities corresponding to multiple related entity nodes are marked as high-risk clusters, that is, high-tax-risk enterprise clusters, and are subsequently subject to key monitoring and identification analysis.
[0039] Step S60: Assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and output quantified risk level and response strategy data.
[0040] In this embodiment, a differentiated assessment is performed based on the risk transmission network and its edge weights constructed in step S30, and the risk synergy strength calculated in step S40. The assessment model calculates direct risk and associated risk separately. Direct risk mainly targets the severity of the target entity's own risk characteristics; while associated risk is comprehensively calculated by analyzing the strength of the association path and the degree of risk synergy between the target entity and other nodes in the high-risk cluster. Finally, based on a preset risk matrix, the calculated risk value is mapped to multiple specific risk levels, thereby providing a clear priority judgment for decision-making and outputting corresponding response strategy data.
[0041] For example, for high-risk related risks, the strategy list in the response strategy data may include specific tax-related items such as initiating special due diligence on core related parties, prudently assessing and gradually reducing the scale of transactions with related parties, and preparing contemporaneous documentation for related party transactions for verification. The aim is to directly transform data analysis results into management actions, which greatly improves the efficiency and accuracy of tax risk management.
[0042] Compared with existing technologies, the tax risk identification method based on multi-source data comparison shown in this embodiment has the following advantages: The tax risk identification method based on multi-source data comparison described in this embodiment realizes closed-loop management from data collection and risk quantification to strategy generation by constructing dynamic risk feature vectors and weighted transmission networks. This method not only improves the accuracy and timeliness of risk identification, but also effectively overcomes the limitations of traditional static analysis by introducing time-series correlation analysis and networked topology modeling, thus providing reliable technical support for enterprises to carry out proactive and intelligent tax risk prevention and control in complex business environments.
[0043] Example 2 The second embodiment of the present invention also provides a tax risk identification method based on multi-source data comparison. The method shown in this embodiment is basically similar to the method shown in the first embodiment, except that: In the risk transmission network, the step of calculating the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject nodes within a specific time window includes: In the risk transmission network, the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node is calculated as the first correlation index. Calculate the dynamic time-normalized distance between the time series and convert it into a similarity score as a second correlation index.
[0044] The step of calculating the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node, respectively, as the first correlation indicator, includes: The time series of dynamic risk feature vectors corresponding to the target subject node and the associated subject node are respectively subjected to mean centering. The covariance and standard deviation are calculated based on the centered time series, and the Pearson correlation coefficient is obtained based on the ratio of the covariance to the standard deviation.
[0045] The step of calculating the dynamic time-normalized distance between the time series and converting it into a similarity score as a second relevance indicator includes: Construct the Euclidean distance matrix between the time series of the target subject node and the associated subject node, and find the optimal regular path using a dynamic programming algorithm; Calculate the cumulative distance from the first point to the second point on the optimal regularized path, and map it to a similarity score in the range of 0 to 1.
[0046] The second embodiment of this invention, while inheriting the core architecture of the first embodiment, deepens and expands the key technical aspect of time-series correlation analysis. Its core innovation lies in introducing a multi-algorithm fusion analysis framework. By combining the Pearson correlation coefficient with the Dynamic Time Warping (DTW) algorithm, it cross-validates and comprehensively assesses risk synergy effects from two dimensions: linear correlation and morphological similarity, significantly improving the robustness and accuracy of risk identification. This design effectively addresses the risk of misjudgment that may arise from a single algorithm. For example, two time series may have different overall trends but highly similar local patterns (in which case the Pearson coefficient may be low while the DTW similarity is high), or vice versa. Through dual-indicator verification, the system can capture more complex and hidden risk synergy patterns.
[0047] Specifically, when calculating the first correlation index, the Pearson correlation coefficient, the time series of dynamic risk characteristic vectors of the target and related subjects are mean-centered to eliminate baseline differences and focus the analysis on the volatility characteristics themselves. Then, the covariance and standard deviation of each series are calculated based on the centered series. Finally, the Pearson coefficient is obtained by dividing the covariance by the product of the standard deviations of the two series. This coefficient ranges from -1 to 1; the closer its absolute value is to 1, the stronger the linear correlation between the two series.
[0048] More importantly, to compensate for the Pearson coefficient's limitation of only measuring linear relationships, this embodiment also constructs an Euclidean distance matrix between corresponding points of the two sequences to quantify the local morphological differences between any two points in the sequences. Subsequently, a dynamic programming algorithm is used to traverse this matrix to find a path with the minimum cumulative distance from the starting point (first point) to the ending point (second point), i.e., the optimal regularized path. This aims to overcome the nonlinear scaling or translation phenomena that may occur in time series along the time axis. Finally, the calculated cumulative distance of the optimal path is transformed into a similarity score between 0 and 1 through mapping relationships such as the exponential decay function. The higher the score, the more similar the overall volatility patterns of the two risk sequences are, even if there is a phase difference in their occurrence times. Specifically, the analysis of the correlation between nodes can rely on the time series correlation calculation here to identify the risk transmission path.
[0049] Example 3 Please see Figure 2The third embodiment of the present invention provides a tax risk identification system based on multi-source data comparison, applied to the method described in any of the above embodiments, the system comprising: The data acquisition module is used to collect multi-source data of the target entity and related entities from the tax, invoice and commercial credit systems according to a preset period, and construct the corresponding dataset; The vector generation module is used to extract time-series data of preset key risk indicators within a preset observation period based on the dataset, and calculate volatility and deviation based on the time-series data of each indicator to generate dynamic risk feature vectors representing the risk status of each subject. The network construction module is used to construct a risk transmission network based on the relationship between the target entity and the associated entities; wherein network nodes represent the target entity and all associated entities, and directed edges represent risk transmission paths; The time-series calculation module calculates the time-series correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window in the risk transmission network. The cluster marking module is used to determine the existence of risky collaboration and mark the associated entities corresponding to multiple associated entity nodes as high-risk clusters when the correlation coefficient exceeds a preset threshold. The risk assessment module is used to assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and outputs quantitative risk level and response strategy data.
[0050] Example 4 A fourth embodiment of the present invention provides a readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any of the above embodiments.
[0051] Example 5 The fifth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the above technical solutions.
[0052] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0053] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A method for identifying tax risks based on multi-source data comparison, characterized in that, The method includes: Data from multiple sources of the target entity and related entities are collected from tax, invoice and commercial credit systems according to a preset cycle, and corresponding datasets are constructed. Based on the dataset, time series data of preset key risk indicators within a preset observation period are extracted, and volatility and deviation are calculated based on the time series data of each indicator to generate a dynamic risk feature vector representing the risk status of each subject. Based on the relationship between the target entity and the associated entities, a risk transmission network is constructed; wherein network nodes represent the target entity and all associated entities, and directed edges represent risk transmission paths. In the risk transmission network, the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window is calculated; When the correlation coefficient exceeds a preset threshold, it is determined that there is risky collaboration and the associated entities corresponding to the multiple associated entity nodes are marked as a high-risk cluster; Assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and output quantitative risk level and response strategy data.
2. The tax risk identification method based on multi-source data comparison according to claim 1, characterized in that, The steps for constructing a risk transmission network based on the relationship between the target entity and the associated entities include: Obtain the relationship data between the target entity and each related entity, including the equity investment ratio, the percentage of annual transaction volume, and the number of companies with overlapping shareholdings; The target entity and each associated entity are mapped to network nodes respectively, and directed edges are established between any network nodes based on the directionality of the association relationship to generate a preliminary network topology. Based on the association data, a weighted fusion algorithm is used to calculate the association strength value of each directed edge, and the association strength value is assigned as the weight of the corresponding directed edge to complete the construction of the risk transmission network.
3. The tax risk identification method based on multi-source data comparison according to claim 2, characterized in that, Based on the aforementioned correlation data, a weighted fusion algorithm is used to calculate the correlation strength value of each directed edge, and the correlation strength value is assigned as the weight of the corresponding directed edge to complete the construction of the risk transmission network. This includes the following steps: The equity investment ratio, annual transaction volume ratio and number of companies with overlapping shareholdings in the aforementioned relationship data are normalized to obtain the standardized values corresponding to each indicator. Configure a preset weight coefficient for each type of association indicator, and linearly weight and fuse the standardized values of the multiple types of association indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the association strength value of the directed edge. The association strength value is mapped to the visual label of the corresponding directed edge, wherein the association strength value is positively correlated with the width and color depth of the edge, thereby generating a risk transmission network that includes node attributes and weighted edge attributes.
4. The tax risk identification method based on multi-source data comparison according to claim 3, characterized in that, The steps of configuring preset weight coefficients for each type of association indicator, linearly weighting and fusing the standardized values of multiple association indicators corresponding to each directed edge with their respective preset weight coefficients to calculate the association strength value of the directed edge include: The information entropy of the equity investment ratio, the proportion of annual transaction volume, and the number of companies with overlapping shareholdings are calculated respectively. The degree of variation of each correlation indicator is determined based on the information entropy, so as to dynamically configure the preset weight coefficient of each correlation indicator. The standardized values corresponding to the equity investment ratio, annual transaction volume ratio, and number of companies with overlapping shareholdings are multiplied by the preset weight coefficient and summed to obtain the association strength value of the directed edge, which is then normalized.
5. The tax risk identification method based on multi-source data comparison according to any one of claims 1-4, characterized in that, In the risk transmission network, the step of calculating the temporal correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject nodes within a specific time window includes: In the risk transmission network, the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node is calculated as the first correlation index. Calculate the dynamic time-normalized distance between the time series and convert it into a similarity score as a second correlation index.
6. The tax risk identification method based on multi-source data comparison according to claim 5, characterized in that, The steps for calculating the Pearson correlation coefficient of the time series of the dynamic risk feature vectors corresponding to the target subject node and the associated subject nodes, respectively, as the first correlation indicator, include: The time series of dynamic risk feature vectors corresponding to the target subject node and the associated subject node are respectively subjected to mean centering. The covariance and standard deviation are calculated based on the centered time series, and the Pearson correlation coefficient is obtained based on the ratio of the covariance to the standard deviation.
7. The tax risk identification method based on multi-source data comparison according to claim 6, characterized in that, The steps of calculating the dynamic time-normalized distance between the time series and converting it into a similarity score as a second correlation indicator include: Construct the Euclidean distance matrix between the time series of the target subject node and the associated subject node, and find the optimal regular path using a dynamic programming algorithm; Calculate the cumulative distance from the first point to the second point on the optimal regularized path, and map it to a similarity score in the range of 0 to 1.
8. A tax risk identification system based on multi-source data comparison, characterized in that, The system, applicable to the method of any one of claims 1-7, comprises: The data acquisition module is used to collect multi-source data of the target entity and related entities from the tax, invoice and commercial credit systems according to a preset period, and construct the corresponding dataset; The vector generation module is used to extract time series data of preset key risk indicators within a preset observation period based on the dataset, and calculate volatility and deviation based on the time series data of each indicator to generate dynamic risk feature vectors representing the risk status of each subject. The network construction module is used to construct a risk transmission network based on the relationship between the target entity and the associated entities; wherein network nodes represent the target entity and all associated entities, and directed edges represent risk transmission paths; The time-series calculation module calculates the time-series correlation of the dynamic risk feature vectors corresponding to the target subject node and the associated subject node within a specific time window in the risk transmission network. The cluster marking module is used to determine the existence of risky collaboration and mark the associated entities corresponding to multiple associated entity nodes as high-risk clusters when the correlation coefficient exceeds a preset threshold. The risk assessment module is used to assess the direct and consequential tax risks posed by the related entities in the high-risk cluster to the target entity, and outputs quantitative risk level and response strategy data.
9. A readable storage medium having computer instructions stored thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-7.