Comprehensive Management Method and System for Financial Statement Data Based on Big Data Analysis
By building a data storage structure tree and a multi-dimensional cube structure, the problems of low efficiency and poor accuracy of traditional financial statement data management are solved, and efficient and accurate financial data management is achieved.
Patent Information
- Application Number
- CN202510387821.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing financial statement data management methods rely on traditional rules engines and manual annotations, resulting in low management efficiency and poor accuracy, and being unable to adapt to business changes and data growth rates.
Using a method based on big data analysis, a data storage structure tree and a multi-dimensional cube structure are built, financial statement data is mapped to multi-dimensional cubes, and data network management is carried out based on the association relationship to reduce manual intervention and rule adjustments.
Improve the management efficiency and accuracy of financial statement data, reduce human errors through automated processing, and adapt to data updates and business changes.
Smart Images

Figure CN119887422B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a comprehensive management method and system for financial statement data based on big data analysis. Background Art
[0002] In today's digital age, enterprises are faced with a vast amount of data such as financial statements and contract texts. It has become crucial to manage these data efficiently and retrieve them accurately. Currently, common financial statement data management methods mainly rely on traditional rule engines and manual annotation.
[0003] Traditional rule engines classify and label financial statement data based on predefined rules. For example, by setting specific keyword matching rules, when certain predefined keywords appear in a financial statement, it is classified into the corresponding category. However, with the development of the business and the diversification of the content of financial statements, new situations emerge continuously, and the rules need to be updated and adjusted constantly, which consumes a large amount of manpower and time. Manual annotation is to manually classify and label financial statement data by professional personnel. However, in the face of a large amount of financial statement data, the speed of manual annotation far lags behind the speed of data growth, and it is prone to human errors, resulting in inconsistencies in classification and labeling.
[0004] Therefore, in the process of managing financial statement data by existing methods, the management efficiency and management accuracy of financial statement data are low. Summary of the Invention
[0005] The present invention provides a comprehensive management method and system for financial statement data based on big data analysis to improve the management efficiency and management accuracy of financial statement data.
[0006] In a first aspect, the comprehensive management method for financial statement data based on big data analysis of the present invention includes:
[0007] Dividing the collected financial statement data, and constructing a data storage structure tree according to the data categories and hierarchical relationships of the data division results;
[0008] Mapping each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category;
[0009] Mapping each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network;
[0010] Associate the data elements corresponding to different nodes and different dimensions in the initial data network based on the association relationships between the financial statement data to obtain a final data network.
[0011] In a second aspect, the present invention further provides a comprehensive financial statement data management system based on big data analysis, which is applied to the comprehensive financial statement data management method based on big data analysis as described in the first aspect; the comprehensive financial statement data management system based on big data analysis includes:
[0012] A structure tree construction module, configured to perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data categories and hierarchical relationships of the data partitioning results;
[0013] A node mapping module, configured to map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category;
[0014] A data mapping module, configured to map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network;
[0015] A data association module, configured to associate the data elements corresponding to different nodes and different dimensions in the initial data network based on the association relationships between the financial statement data to obtain a final data network.
[0016] In a third aspect, the present invention further provides an electronic device, including: a memory, configured to store a computer software program; a processor, configured to read and execute the computer software program, and further implement the comprehensive financial statement data management method based on big data analysis as described in any one of the above.
[0017] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, it implements the comprehensive financial statement data management method based on big data analysis as described in any one of the above.
[0018] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the comprehensive financial statement data management method based on big data analysis as described in any one of the above.
[0019] The comprehensive management method for financial statement data based on big data analysis provided by the embodiments of the present invention manages financial statement data in the form of a data network through a data storage structure tree and a multi-dimensional cube structure. Therefore, during the management process of financial statement data, each data category in the financial statement data can be efficiently organized and stored through the data storage structure tree. Subsequently, when data is updated, only the relationship between nodes in the data storage structure tree needs to be adjusted, without continuously updating and adjusting rules, improving the management efficiency of financial statement data. At the same time, when new financial statement data is entered into the data network subsequently, the insertion position in the data storage structure tree can be quickly located according to the data category to which it belongs, and the data can be entered according to the association relationship between the parent and child nodes in the data storage structure tree, without manual participation, avoiding chaotic data storage and incorrect data warehousing, and improving the management accuracy of financial statement data. On the other hand, through the multi-dimensional cube structure, financial statement data can be managed from multiple dimensions, eliminating complex manual calculations, reducing human errors, and improving the management accuracy of financial statement data.
[0020] Therefore, the embodiments of the present invention manage financial statement data through a data network constructed by a data storage structure tree and a multi-dimensional cube structure, improving the management efficiency and management accuracy of financial statement data. Description of the Drawings
[0021] Figure 1 is a schematic flowchart of the comprehensive management method for financial statement data based on big data analysis provided by the embodiments of the present invention;
[0022] Figure 2 is a schematic structural diagram of the comprehensive management system for financial statement data based on big data analysis provided by the embodiments of the present invention;
[0023] Figure 3 is an embodiment diagram of the electronic device provided by the embodiments of the present invention;
[0024] Figure 4 is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. Detailed Embodiments
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0026] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0027] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0028] Optionally, refer to Figure 1 , Figure 1 which is a schematic flowchart of the comprehensive financial statement data management method based on big data analysis provided by the present invention. In the embodiments of the present invention, the execution subject of the comprehensive financial statement data management method based on big data analysis is a data management system. Therefore, the comprehensive financial statement data management method based on big data analysis includes:
[0029] Step 10: Perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data category and hierarchical relationship of the data partitioning result.
[0030] Optionally, the data management system collects the financial statement data of an enterprise and performs data partitioning on the financial statement data. Among them, the basis for classification can be various factors such as the nature, source, and use of financial data. For example, classified according to accounting subjects, the data is divided into large categories such as asset categories, liability categories, owner's equity categories, income categories, and expense categories. Under each large category, further subdivision can be carried out. For example, under the asset category, it can be divided into current assets and non-current assets, and current assets can be further divided into monetary funds, accounts receivable, etc.
[0031] Furthermore, the data management system constructs a data storage structure tree according to the data categories and their hierarchical relationships, specifically as described in steps 101 to 104. Among them, the root node of the tree can be the major category of data categories. Starting from the root node, each sub-node extends downward according to the hierarchical relationship. Each sub-node represents a specific data category, and the connection between nodes represents the hierarchical relationship.
[0032] In an embodiment, the annual financial statement data of an enterprise is collected, including the data of the balance sheet, income statement, and cash flow statement. Classified according to accounting subjects as:
[0033] 1. Asset category:
[0034] 1.1 Current assets: Monetary funds (including data such as cash and bank deposits); Accounts receivable; Inventories;
[0035] 1.2 Non-current assets: Fixed assets; Intangible assets;
[0036] 2. Liability category:
[0037] 2.1 Current liabilities: Short-term loans; Accounts payable;
[0038] 2.2 Non-current liabilities: Long-term loans;
[0039] 3. Owner's equity category: Paid-in capital; Capital reserve;
[0040] 4. Income category: Main business income; Other business income;
[0041] 5. Expense category: Main business cost; Administrative expenses; Selling expenses.
[0042] According to the above classification, a data storage structure tree is constructed. The root nodes can be "Asset category", "Liability category", "Owner's equity category", "Income category", and "Expense category". Starting from each root node, each root node is further divided into the next-level sub-nodes. For example, under the "Asset category", there are sub-nodes such as "Current assets" and "Non-current assets", and so on.
[0043] Step 20: Map each node in the data storage structure tree to a dimension of the multi-dimensional cube to construct a multi-dimensional cube structure.
[0044] Optionally, each node in the data storage structure tree represents a specific data category. The multidimensional cube is a structure used for data storage and analysis, and each dimension represents a feature or attribute of the data. Therefore, the data management system maps the nodes of the data storage structure tree to the dimensions, so that the data can be organized and stored in the multidimensional space according to the corresponding categories and hierarchical relationships, as described in steps 201 to 204. In the mapping process, it is necessary to ensure that the mapping of each node is unique and reasonable, and can accurately reflect the characteristics and relationships of the data.
[0045] Continuing with the above embodiment, the nodes in the data storage structure tree are mapped: "Asset Class" is mapped to one dimension of the multidimensional cube, for example, named "Asset Dimension". "Liability Class" is mapped to another dimension, named "Liability Dimension". "Owner's Equity Class" is mapped to "Owner's Equity Dimension". "Income Class" is mapped to "Income Dimension". "Expense Class" is mapped to "Expense Dimension".
[0046] For the sub-nodes under "Asset Class", such as "Current Assets" and "Non-Current Assets", the levels can be further subdivided under "Asset Dimension", and "Current Assets" and "Non-Current Assets" are used as different level elements under "Asset Dimension". Similarly, the sub-nodes under other categories are hierarchically mapped under the corresponding dimensions in this way. Finally, a multi-dimensional cube structure with multiple dimensions and hierarchical relationships is constructed to store and organize financial statement data.
[0047] Step 30, mapping each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network.
[0048] Optionally, each specific data element in the financial statement data, such as a specific monetary fund amount, a value of a certain account receivable, etc., is mapped to the corresponding node of the constructed data storage structure tree according to the data category to which it belongs by the data management system. At the same time, the data management system maps the data element to the corresponding position of the multidimensional cube according to the dimension of the multidimensional cube corresponding to the data category. Through the mapping operation, the data element has a clear position in the data storage structure tree and the multidimensional cube structure, and an initial data network is constructed, as described in steps 301 to 304, wherein the initial data network shows the relationship between the data element and the data category and dimension.
[0049] In one embodiment, in the balance sheet of an enterprise, the amount of cash and cash equivalents is 100,000 yuan. In the data storage structure tree, "cash and cash equivalents" belongs to the child node of "current assets" under "asset category", so this data element of 100,000 yuan is mapped to the node of "cash and cash equivalents". In the multi-dimensional cube structure, "asset category" corresponds to "asset dimension", "current assets" is a level under "asset dimension", and "cash and cash equivalents" is a specific element under "current assets", so 100,000 yuan is mapped to the position of "cash and cash equivalents" at the level of "current assets" under "asset dimension". Similarly, for other data elements in the financial statements, such as the accounts receivable amount of 50,000 yuan, the short-term borrowing amount of 80,000 yuan, etc., they are mapped to the data storage structure tree and the multi-dimensional cube structure respectively in a similar way, and finally an initial data network including all financial statement data elements is formed.
[0050] Step 40: Based on the association relationships between financial statement data, associate the data elements corresponding to different nodes and different dimensions in the initial data network to obtain the final data network.
[0051] Optionally, there are various internal association relationships between financial statement data. For example, the balance relationship between assets and liabilities (Assets = Liabilities + Owner's Equity), the calculation relationship between revenue, expenses and profit (Profit = Revenue - Expenses), etc. Therefore, according to the association relationships between financial statement data, the data management system reasonably associates the data elements located at different nodes (different data categories) and different dimensions in the initial data network to obtain the final data network, as specifically described in Steps 401 to 404. Therefore, through the association operation, the final data network is not just a simple collection of data elements, but an organic whole that can reflect the logical relationships between data.
[0052] In an embodiment of the present invention, the financial statement data is managed in the form of a data network through a data storage structure tree and a multi-dimensional cube structure. Therefore, in the process of managing the financial statement data, through the data storage structure tree, each data category in the financial statement data can be efficiently organized and stored. Subsequently, when the data is updated, only the relationship between the nodes of the data in the data storage structure tree needs to be adjusted, and there is no need to continuously update and adjust the rules, which improves the management efficiency of the financial statement data. At the same time, when new financial statement data is entered into the data network subsequently, the insertion position in the data storage structure tree can be quickly located according to the data category to which it belongs, and the data can be entered according to the association relationship between the parent and child nodes in the data storage structure tree, without manual participation, avoiding chaotic data storage and incorrect data warehousing, and improving the management accuracy of the financial statement data. On the other hand, through the multi-dimensional cube structure, the financial statement data can be managed from multiple dimensions, eliminating complex manual calculations, reducing human errors, and improving the management accuracy of the financial statement data.
[0053] In one embodiment, the descriptions of steps 101 to 104 are as follows:
[0054] Step 101, in the multi-dimensional financial data space, with each data point in the financial statement data as the center, determine the regional density of each data point according to the number of data points within a preset range for each data point, and cluster the data points with the same regional density to obtain an initial clustering.
[0055] Optionally, within the multi-dimensional financial data space, the data management system processes each data point in the financial statement. Taking a certain data point as the center of a circle, a specific preset range is set. Among them, the preset range can be a multi-dimensional space area determined based on factors such as the characteristic scale of the data and business requirements.
[0056] Furthermore, the data management system counts the number of data points falling within this preset range. Among them, the quantity represents the regional density of the data point. For example, if in a two-dimensional financial data space (the dimensions are income and cost), taking a certain data point (representing the income and cost values of a certain business unit) as the center, a circular range with a radius of r is set, and there are n other data points within this circular range, then the regional density of this data point is n.
[0057] Furthermore, the data management system groups the data points with the same regional density into one group to obtain an initial clustering. Therefore, through the above steps, the data points with similar density characteristics are initially aggregated together.
[0058] In one embodiment, there is a financial statement including the revenue and profit data of multiple business departments of an enterprise in different quarters. The data dimensions are quarter (4 values), business department (5), revenue value, and profit value, forming a four-dimensional financial data space. For a data point A, it represents that the revenue of business department X in the second quarter is 1 million yuan and the profit is 200,000 yuan. Set a multi-dimensional cube range centered on data point A. Within this range, after statistics, there are 10 other data points (these data points may come from the revenue and profit data of different business departments in similar quarters), so the regional density of data point A is 10. After calculating all data points, it is found that the regional densities of 30 data points are all 10, and the system clusters these 30 data points into an initial cluster. At the same time, there may be other data points with regional densities of 5, 15, etc., forming different initial clusters respectively.
[0059] Step 102, if the density change of the target data point at the boundary between the initial cluster and its adjacent cluster is less than or equal to the preset degree threshold, and the distance of the target data point in the feature space is less than or equal to the preset distance threshold, then fuse each initial cluster with its adjacent cluster to obtain the optimized cluster.
[0060] Furthermore, for each initial cluster, the data management system checks its situation at the boundary with the adjacent cluster, where the adjacent cluster refers to other clusters that are close in position to this initial cluster in the multi-dimensional financial data space. For the target data point at the boundary (i.e., the data point near the boundary of the initial cluster), the data management system calculates its density change, where the density change refers to the degree of difference between the density of the target data point in this cluster and its density in the adjacent cluster. At the same time, the data management system calculates the distance of the target data point in the feature space (i.e., the space composed of each dimension of the financial data), and the distance can use a suitable distance metric method such as the Euclidean distance.
[0061] Furthermore, if the density change is less than or equal to the preset degree threshold (predetermined according to business experience and data characteristics, used to judge whether the density change is small enough to indicate the similarity of the density characteristics of the two clusters), and the distance of the target data point in the feature space is less than or equal to the preset distance threshold (used to judge the proximity of the two clusters in space position), then it is determined that these two clusters have a strong correlation. Therefore, the data management system fuses the initial cluster with its adjacent cluster into a new cluster, that is, the optimized cluster, so that the clustering result better conforms to the internal distribution characteristics of the data.
[0062] Continuing with the above embodiment, there are initial clusters C1 and adjacent initial cluster C2. There is a data point B at the boundary of C1. The regional density of B in C1 is 12, and the regional density of the data point B' with a similar position to B in C2 is 10. Their density change is |12 - 10| = 2. The preset degree threshold is 3, satisfying the density change condition. In the feature space, the Euclidean distance between the data point B (whose coordinates are [Second quarter, Business unit X, 1 million yuan in revenue, 200,000 yuan in profit]) and the data point B' (coordinates are [Second quarter, Business unit Y, 950,000 yuan in revenue, 180,000 yuan in profit]) is calculated to be 5 (the preset distance threshold is 8), satisfying the distance condition. The initial clusters C1 and C2 are merged into an optimized cluster.
[0063] Step 103, for each optimized cluster, determine the sub - topic financial data category corresponding to each optimized cluster according to the distribution shape, density characteristics, and clustering characteristics of the data points within the optimized cluster.
[0064] Furthermore, for each optimized cluster, the data management system deeply analyzes various characteristics of the internal data points. Among them, various characteristics include the distribution shape, density characteristics, and clustering characteristics of the data points. The distribution shape refers to the distribution pattern of the data points in the multi - dimensional space, such as being linearly distributed, cluster - shaped distributed, or other specific shapes. The density characteristics may involve not only the regional density but also the change trend of the density, etc. The clustering characteristics include the size of the cluster (the number of data points), the compactness of the cluster, etc.
[0065] Furthermore, the data management system determines a corresponding sub - topic financial data category according to the distribution shape, density characteristics, and clustering characteristics of the data points within the optimized cluster. For example, if the data points in an optimized cluster mainly revolve around the data related to the R & D expenses of the enterprise, and the distribution shape shows a gradually increasing trend over time, and the density characteristics also conform to the characteristics of the R & D expense data, then it is determined as the "sub - topic financial data category of R & D expenses".
[0066] In an embodiment, there is an optimized cluster C, and the data points therein are mainly concentrated on the cost of sales data of different product lines of the enterprise. From the perspective of the distribution shape, it shows the characteristics of the cost of sales of different product lines fluctuating over time in the product dimension and the time dimension, and the data points are relatively concentrated. From the density characteristics, the regional density is relatively stable in different time periods and product lines. In terms of clustering characteristics, the size of the cluster is moderate and the compactness is relatively high. Considering these characteristics comprehensively, the data management system determines that the sub - topic financial data category corresponding to this optimized cluster is the "sub - topic financial data category of the cost of sales of product lines".
[0067] Step 104: Based on each sub - topic financial data category and the hierarchical relationship of each sub - topic financial data category in its topic financial data category, construct a data storage structure tree.
[0068] Further, for each sub - topic financial data category, the data management system determines the corresponding topic financial data category of each sub - topic financial data category.
[0069] Further, with the topic financial data category as the root node, the data management system connects the sub - topic financial data categories corresponding to the topic financial data category as child nodes under the root node, obtaining a hierarchical data storage structure tree, as specifically described in Steps 1041 to 1044. For example, the "R & D expense sub - topic financial data category" and the "Marketing expense sub - topic financial data category" may belong to the next level of the "Expense topic financial data category".
[0070] In one embodiment, there are the following relationships between topic financial data categories and sub - topic financial data categories:
[0071] 1. For the "Asset topic": It includes the "Current asset sub - topic" and the "Non - current asset sub - topic".
[0072] The "Current asset sub - topic": It includes the "Monetary fund sub - topic" and the "Accounts receivable sub - topic".
[0073] The "Non - current asset sub - topic": It includes the "Fixed asset sub - topic" and the "Intangible asset sub - topic".
[0074] 2. For the "Expense topic": It includes the "Operating expense sub - topic" and the "R & D expense sub - topic".
[0075] The "Operating expense sub - topic": It includes the "Sales expense sub - topic" and the "Administrative expense sub - topic".
[0076] Taking the "Asset topic" and the "Expense topic" as root nodes, there are child nodes of the "Current asset sub - topic" and the "Non - current asset sub - topic" under the "Asset topic", and child nodes of the "Monetary fund sub - topic" and the "Accounts receivable sub - topic" under the "Current asset sub - topic", and so on, to construct a data storage structure tree.
[0077] The embodiment of the present invention constructs a data storage structure tree, so it can store and manage financial statement data in a more reasonable way that conforms to the internal relationship of financial data and business requirements. Therefore, the management efficiency and management accuracy of financial statement data are improved through the data storage structure tree.
[0078] In one embodiment, the descriptions of Steps 1041 to 1044 are as follows:
[0079] Step 1041: For each sub-topic financial data category, determine the data level of each sub-topic financial data category within its sub-topic financial data category according to the amount of data in the sub-topic financial data category and the degree of association with the core business in its topic financial data category.
[0080] Optionally, the data management system conducts analysis for each determined sub-topic financial data category. On the one hand, consider the amount of data in the sub-topic financial data category. A sub-topic with a large amount of data may be in a relatively more important or fundamental position in the data level. On the other hand, evaluate the degree of association between the sub-topic and the core business in its topic financial data category. The higher the degree of association, the higher its position in the data level. Therefore, the data management system synthesizes these two factors to determine the relative data level of each sub-topic financial data category among all sub-topic financial data categories.
[0081] In one embodiment, under the "expense topic" financial data category, there are "R & D expense sub-topic", "sales expense sub-topic", "administrative expense sub-topic", etc. The amount of data in the "sales expense sub-topic" is the largest among all expense sub-topics, and the sales business is the core business of the enterprise, and the sales expense is closely related to the core business. While the amount of data in the "administrative expense sub-topic" is relatively small, and its direct association with the core business is not as strong as that of the sales expense. Therefore, the "sales expense sub-topic" can be determined to be at a higher data level, such as level one; the "administrative expense sub-topic" can be determined to be at a lower data level, such as level two.
[0082] Step 1042: For each sub-topic financial data category, divide the data points in the sub-topic financial data category into sub-categories according to the data characteristics of the data points in the sub-topic financial data category, and sort the sub-categories according to the data characteristic differences between the sub-categories to obtain the optimized data category of each sub-topic financial data category.
[0083] Furthermore, for each sub-topic financial data category, the data management system deeply analyzes the data characteristics of the data points therein. Among them, the data characteristics include the numerical range of the data, the time attribute of the data, the business processes involved in the data, etc., and divides the data points in the sub-topic financial data category into different sub-categories according to the data characteristics. For example, for the "accounts receivable sub-topic", the data points can be divided into sub-categories such as short-term accounts receivable (aging within 1 year), medium-term accounts receivable (aging from 1 to 3 years), long-term accounts receivable (aging over 3 years), etc.
[0084] Furthermore, the data management system sorts the sub-categories in each sub-topic financial data category according to the degree of difference in data characteristics between the sub-categories. Among them, the sorting can be based on the importance of features, chronological order, or business logic order, etc., to obtain the optimized data categories for each sub-topic financial data category.
[0085] In one embodiment, taking the "fixed assets sub-topic" as an example, the data characteristics of the data points include asset categories (such as buildings, machinery and equipment, transportation tools, etc.), acquisition time, usage status, etc. The data management system divides the data points into "building sub-categories", "machinery and equipment sub-categories", and "transportation tools sub-categories" according to the asset categories. Then, these sub-categories are sorted according to the degree of importance that the enterprise attaches to different asset categories and their importance in production and operation. Buildings are crucial for the production and operation of the enterprise, followed by machinery and equipment, and transportation tools are relatively less important. Then the sorting result is that the "building sub-category" ranks first, the "machinery and equipment sub-category" second, and the "transportation tools sub-category" last, forming the optimized data category of the "fixed assets sub-topic".
[0086] Step 1043: Taking each topic financial data category as the root node, taking the optimized data categories under each topic financial data category as the leaf nodes, and taking the data hierarchy of each sub-topic financial data category in its sub-topic financial data category as the node hierarchy of the leaf nodes, construct an initial storage structure tree.
[0087] Furthermore, the data management system takes each topic financial data category (such as "asset topic", "liability topic", "income topic", etc.) as the root node of the initial storage structure tree. For each sub-topic financial data category under each root node, connect its optimized data category as the leaf node to the corresponding root node. At the same time, according to the data hierarchy of each sub-topic financial data category in its sub-topic financial data category determined in Step 1041, assign the corresponding node hierarchy to these leaf nodes, and construct the initial storage structure tree, so that the constructed initial storage structure tree not only reflects the hierarchical relationship between the topic and the sub-topic, but also reflects the relative importance of the sub-topic and the hierarchical structure of data organization through the node hierarchy.
[0088] In one embodiment, for the "asset theme", there are "current asset sub - themes" and "non - current asset sub - themes" under it. After the processing in step 1042, the optimized data categories of the "current asset sub - theme" include "cash and cash equivalents sub - category", "accounts receivable sub - category", etc.; the optimized data categories of the "non - current asset sub - theme" include "fixed asset sub - category", "intangible asset sub - category", etc. In step 1041, the data level of the "current asset sub - theme" is determined to be level one, and the data level of the "non - current asset sub - theme" is level two. Then when constructing the initial storage structure tree, the "asset theme" is the root node, and the nodes corresponding to the "cash and cash equivalents sub - category", "accounts receivable sub - category", etc. are the leaf nodes corresponding to the "current asset sub - theme", and the node level is level one; the nodes corresponding to the "fixed asset sub - category", "intangible asset sub - category", etc. are the leaf nodes corresponding to the "non - current asset sub - theme", and the node level is level two.
[0089] Step 1044, based on the actual access frequencies and business logic associations among financial data categories, associate the corresponding nodes in the initial storage structure tree to obtain the data storage structure tree.
[0090] Furthermore, the data management system analyzes the access frequencies among financial data categories in actual business operations and their inherent business logic associations. For example, in financial analysis, the "main business income sub - theme" under the "income theme" and the "main business cost sub - theme" under the "expense theme" are often viewed simultaneously because they are directly related to the calculation of the enterprise's main business profit.
[0091] Furthermore, the data management system, according to the actual access frequencies and business logic relationships, connects the closely related nodes in the initial storage structure tree through corresponding methods (such as connection lines, establishing pointers, etc., which are conceptually related) to obtain the data storage structure tree.
[0092] In one embodiment, in actual business, an enterprise often conducts correlation analysis between the "inventory sub - theme" (belonging to the "current asset sub - theme" and under the "asset theme") and the "cost of goods sold sub - theme" (belonging to the "operating expense sub - theme" and under the "expense theme") because the consumption of inventory is directly related to the cost of goods sold. Therefore, in the initial storage structure tree, the nodes corresponding to the "inventory sub - theme" and the nodes corresponding to the "cost of goods sold sub - theme" are associated. Similarly, for the "short - term borrowing sub - theme" (under the "liability theme") and the "interest expense sub - theme" (under the "expense theme") that are often queried and analyzed together, node association is also performed to obtain the final data storage structure tree.
[0093] In an embodiment of the present invention, by considering the actual access frequency and business logic association, the initial storage structure tree is optimized, and relevant data nodes are connected, so that the data storage structure tree not only reflects the hierarchical relationship of the data, but also embodies the actual business needs, making the constructed financial data storage structure tree more scientific, reasonable and in line with the actual business of the enterprise. In terms of data storage, the storage space can be utilized more efficiently and the data can be organized reasonably. When querying and analyzing data, users can quickly find relevant data in the storage structure tree according to the business logic and common analysis paths, improving the management efficiency of financial data.
[0094] In one embodiment, the descriptions of steps 201 to 204 are as follows:
[0095] Step 201, identify a first target node related to time, a second target node related to financial subjects, and a third target node related to amounts in the data storage structure tree.
[0096] Optionally, the data management system traverses the data storage structure tree. For the first target node related to time, search for data nodes representing time attributes, such as nodes like "year", "quarter", "month", etc.
[0097] For the second target node related to financial subjects, the data management system identifies nodes such as "Assets - Current Assets - Monetary Funds", "Liabilities - Current Liabilities - Short - term Borrowings", etc. These nodes correspond to specific financial subject classifications and are the core classification identifiers of financial data.
[0098] For the third target node related to amounts, the data management system directly stores data nodes with specific amount values, such as the amount of a specific monetary fund, the amount of an accounts receivable, etc.
[0099] In one embodiment, there is a root node "Enterprise Annual Financial Statement Data" in the data storage structure tree, and there are multiple branches below it. One of the branches is the "Time Dimension Branch", which contains nodes such as "2024", "2025", etc. These are the first target nodes related to time. Under the "Financial Subject Dimension Branch", there are nodes such as "Assets - Current Assets - Accounts Receivable", "Liabilities - Non - Current Liabilities - Long - term Borrowings", etc. They belong to the second target nodes related to financial subjects. Among the leaf nodes of each financial subject node, such as "Accounts Receivable - The accounts receivable amount of Customer A is 50,000 yuan", the node where "50,000 yuan" is located is the third target node related to amounts.
[0100] Step 202: According to the node characteristics of each node in the data storage structure tree, map the first target node, the second target node, and the third target node to the first initial positions corresponding to the time dimension, subject dimension, and amount dimension of the multi-dimensional cube respectively.
[0101] Furthermore, the data management system determines the initial positions of each target node on the corresponding dimensions of the multi-dimensional cube based on the characteristics of each target node in the data storage structure tree. For the first target node (time-related), if it is a "year" node, it may be mapped to the time dimension in chronological order from early to late. For example, the "2020" node will be placed in a relatively forward position on the time dimension, and the "2025" node will be placed in a relatively backward position.
[0102] For the second target node (finance subject-related), it is mapped according to the hierarchical relationship and classification system of finance subjects. For example, on the subject dimension, "assets" as a major category will be in a higher-level position, its sub-category "current assets" is in a lower level under "assets", and "cash and cash equivalents" as a sub-node of "current assets" is in an even lower level.
[0103] For the third target node (amount-related), its position on the amount dimension is initially determined based on the finance subject and time node it belongs to. Generally, the amount nodes at different time points under the same finance subject will be arranged in chronological order on the amount dimension.
[0104] Continuing with the above embodiment, the "2024" node will be mapped to the time scale position corresponding to 2024 on the time dimension, and the "2025" node will be mapped to the time scale position corresponding to 2025. On the subject dimension, the "assets" node will be placed in a higher level, the "current assets" node is in the next lower level, and the "accounts receivable" node is in the next lower level under "current assets". For the third target node "accounts receivable - accounts receivable amount of customer A is 50,000 yuan", since it belongs to the data of the "accounts receivable" subject at a certain time point (assumed to be 2024), it will be initially mapped to the position corresponding to the "2024" time node and the "accounts receivable" subject node on the amount dimension.
[0105] Step 203: Based on the parent-child relationship and sibling relationship of each node in the data storage structure tree, optimize the first initial position of each node in the corresponding dimension to obtain the first target position of each node in the corresponding dimension.
[0106] Furthermore, the data management system determines the parent-child relationships and sibling relationships of the nodes in the data storage structure tree. For the parent-child relationships, for example, in the financial subject dimension, if "Asset Category" is the parent node of "Current Assets", then the position of the "Current Assets" node in the subject dimension should be below the "Asset Category" node with a certain logical hierarchical interval to reflect this parent-child hierarchical relationship. In the time dimension, if the "Quarter" node is the parent node of the "Month" node, the "Month" nodes should be arranged in sequence according to the month order within the time interval determined by the "Quarter" node to further optimize their positions.
[0107] For sibling relationships, for example, in the time dimension, different quarter nodes under the same year are sibling relationships and should be closely arranged in sequence according to the quarter order in the time dimension with a reasonable interval. In the financial subject dimension, "Cash and Cash Equivalents" and "Accounts Receivable" under "Current Assets" are sibling nodes, and they should be at the same level in the subject dimension. And according to factors such as business logic and data usage frequency, their relative positions are reasonably adjusted. By considering such parent-child relationships and sibling relationships, the first initial position of each node in its respective dimension is optimized to obtain a more reasonable first target position.
[0108] In one embodiment, in the subject dimension, "Asset Category" is the parent node, and "Current Assets" and "Non-Current Assets" under it are sibling nodes. "Cash and Cash Equivalents", "Accounts Receivable", etc. under "Current Assets" are also sibling nodes. According to the common analysis logic of financial data, the "Cash and Cash Equivalents" node with a higher usage frequency is placed in a relatively forward and easily accessible position among the child nodes of "Current Assets", and the "Accounts Receivable" node follows closely. In the time dimension, assuming there are four quarter nodes under "2024", namely "First Quarter", "Second Quarter", "Third Quarter", and "Fourth Quarter", which are sibling nodes. According to the quarter order, they are closely arranged in sequence within the time interval corresponding to the "2024" time node to obtain the optimized first target position.
[0109] Step 204, in the multi-dimensional cube, using the financial data category of each node as the attribute value of each node, combined with the first target position of each node in the dimension where it is located, construct a multi-dimensional cube structure.
[0110] Further, the data management system assigns an attribute value to each node of the multi-dimensional cube, and this attribute value is the financial data category represented by the node. For example, the attribute value of the node "Asset - Current Assets - Monetary Funds" is the "Monetary Funds Financial Data Category". Further, the data management system combines the first target positions of each node determined in the above steps in their respective dimensions (time dimension, subject dimension, amount dimension, etc.), places these nodes at the corresponding positions of the multi-dimensional cube, and constructs a complete multi-dimensional cube structure, so that in the multi-dimensional cube structure, the position and attributes of each node are clear, and the distribution and correlation of financial data in different dimensions can be intuitively displayed.
[0111] In one embodiment, when constructing the multi-dimensional cube structure, for the node "2024 - Asset - Current Assets - Monetary Funds - Amount: 100,000 yuan", its attribute value is the "Monetary Funds Financial Data Category". In the multi-dimensional cube, this node is located at the position corresponding to 2024 in the time dimension, at the hierarchical position corresponding to "Asset - Current Assets - Monetary Funds" in the subject dimension, and at the numerical position corresponding to 100,000 yuan in the amount dimension. By placing all similar nodes in the multi-dimensional cube in this way, the multi-dimensional cube structure is finally constructed. For example, the change in the amount of monetary funds in different years and quarters can be viewed from the time dimension, and the amount distribution under different financial subjects can be viewed from the subject dimension.
[0112] The embodiment of the present invention constructs an efficient, intuitive and financially logical multi-dimensional cube structure. Therefore, through the multi-dimensional cube structure, the financial statement data can be managed from multiple dimensions, eliminating complex manual calculations, reducing human errors, and improving the management accuracy of financial statement data.
[0113] In one embodiment, the descriptions of steps 301 to 304 are as follows:
[0114] Step 301: Extract features from each data element in the financial statement data to obtain the element features of each data element.
[0115] Optionally, the data management system extracts features from each data element in the financial statement data to obtain the element features of each data element, where the element features at least include the element name, element type, financial statement category to which it belongs, time attribute, financial subject information involved, and amount value involved.
[0116] For the element name, directly obtain the specific name represented by the data element, such as "monetary fund amount", "main business income", etc. The element type determines whether the data element is numerical, text-based, or of other types. For example, amount data is usually numerical, while financial account names are mostly text-based. The category of the financial statement to which it belongs clarifies whether the data element comes from the balance sheet, income statement, cash flow statement, etc. The time attribute determines the time point or time period corresponding to the data element, such as "the first quarter of 2024". The financial account information involved indicates the financial account associated with the data element, like "Asset - Current Asset - Accounts Receivable". For the amount value involved, extract the actual amount number in the data element if the data element itself is represented as an amount.
[0117] In one embodiment, there is a data element in the financial statement data, described as "In December 2024, in the balance sheet, the monetary fund amount is 80,000 yuan". The data management system extracts its element characteristics as follows: Element name: Monetary fund amount; Element type: Numerical; Category of the financial statement to which it belongs: Balance sheet; Time attribute: December 2024; Financial account information involved: Asset - Current Asset - Monetary fund; Amount value involved: 80,000 yuan.
[0118] Step 302, for each data element, perform path matching in the data storage structure tree according to the element characteristics to obtain the fourth target node in the data storage structure tree, and map each data element to the second initial position of the corresponding dimension in the multi-dimensional cube structure based on the element characteristics.
[0119] Furthermore, for each data element, the data management system performs path matching in the data storage structure tree. Taking the financial account information involved in the element as an example, starting from the root node of the data storage structure tree, gradually search downward along the hierarchy of the financial account classification until a completely matching node is found, and this node is the fourth target node. For example, for the financial account information of "Asset - Current Asset - Monetary fund", find the corresponding "Monetary fund" node in the data storage structure tree.
[0120] Meanwhile, for each data element, the data management system maps the data element to the second initial position of the corresponding dimension in the multi-dimensional cube structure according to the element characteristics. For example, according to the time attribute, map the data element to the corresponding time point position on the time dimension; according to the financial account information involved, map it to the corresponding account level position on the account dimension; according to the amount value involved, find the corresponding numerical position on the amount dimension. Through such mapping, the position of the data element in the multi-dimensional cube is initially determined.
[0121] In one embodiment, for the data element of "In December 2024, in the balance sheet, the amount of cash and cash equivalents is 80,000 yuan". In the data storage structure tree, through the path of "Asset - Current Assets - Cash and Cash Equivalents", the fourth target node of "Cash and Cash Equivalents" is found. In the multi - dimensional cube structure, in the time dimension, this data element is mapped to the time scale position corresponding to December 2024; in the account dimension, it is mapped to the corresponding hierarchical position of "Asset - Current Assets - Cash and Cash Equivalents"; in the amount dimension, it is mapped to the numerical position corresponding to 80,000 yuan, thus obtaining its second initial position.
[0122] Step 303, for each data element, optimize the second initial position according to the parent - child relationship and sibling relationship of the fourth target node in the data storage structure tree, to obtain the second target position of each data element in the corresponding dimension of the multi - dimensional cube structure.
[0123] Furthermore, the data management system determines the parent - child relationship and sibling relationship of the fourth target node of the data element in the data storage structure tree. In the account dimension, if the parent node of the fourth target node "Cash and Cash Equivalents" is "Current Assets", and the sibling nodes include "Accounts Receivable", etc. Considering the usage frequency and correlation relationship between cash and cash equivalents and accounts receivable in actual financial analysis, the position of this data element in the account dimension may be appropriately adjusted. If the analysis frequency of cash and cash equivalents data is high, its second initial position in the account dimension may be optimized in a more accessible direction. In the time dimension, if the time corresponding to the fourth target node is a month in a certain quarter, according to the order of months within the quarter and the logical relationship between data, optimize the position of this data element in the time dimension so that it maintains a reasonable order and interval with the data elements of other months in the same quarter. By considering such parent - child relationships and sibling relationships, adjust the second initial position of each data element in each dimension of the multi - dimensional cube to obtain a second target position that better conforms to the internal logic of financial data and business usage habits.
[0124] Continuing the above - mentioned embodiment, taking the cash and cash equivalents data element as an example. In the data storage structure tree, the "Cash and Cash Equivalents" node and the "Accounts Receivable" node are sibling relationships, and the enterprise often pays attention to the situation of cash and cash equivalents first in financial analysis. Then in the account dimension, the system will optimize the position of this cash and cash equivalents data element based on its second initial position to a more forward and accessible position to obtain the second target position. In the time dimension, if this data element is in December of the fourth quarter of 2024, the system will optimize its position in the time dimension to a position that is reasonably arranged with the data elements of October and November according to the time sequence within the quarter, so as to reflect the coherence and logic in time.
[0125] Step 304: Based on the fourth target node and the second target position, map each data element to the data storage structure tree and the multi-dimensional cube structure respectively, to obtain an initial data network.
[0126] Furthermore, the data management system maps each data element to the data storage structure tree and the multi-dimensional cube structure respectively according to the fourth target node and the second target position, to obtain an initial data network, as specifically described in Steps 3041 to 3044.
[0127] In the embodiment of the present invention, each data element in the financial statement data is mapped to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network. On the one hand, the data storage structure tree and the multi-dimensional cube structure can improve the management efficiency and management accuracy of the financial statement data. On the other hand, through the initial data network, the position of a specific data element in different structures can be quickly found, which is convenient for cross-structure data correlation analysis and improves the analysis efficiency of the financial statement data.
[0128] In one embodiment, the descriptions of Steps 3041 to 3044 are as follows:
[0129] Step 3041: Associate the financial data category of each data element as an attribute value with the second target positions of each dimension in the multi-dimensional cube structure.
[0130] Optionally, the data management system further refines the operations of each data element in the multi-dimensional cube structure. After determining the second target position of each data element in the corresponding dimension of the multi-dimensional cube structure in Step 303, use the financial data category of the data element as a key attribute value to establish a close connection with these second target positions. For example, for the financial data category of "cash and cash equivalents", there are second target positions of the corresponding data elements in the time dimension, subject dimension, and amount dimension of the multi-dimensional cube. Correlate the attribute value of "cash and cash equivalents" with these positions one by one. Therefore, in the multi-dimensional cube structure, the category of the data element represented by each position can be clearly identified, enabling quick positioning of the distribution of a certain type of financial data in each dimension in the multi-dimensional space.
[0131] In one embodiment, there is a data element whose financial data category is "main business income". In the multi-dimensional cube structure, the second target position in the time dimension corresponds to "the third quarter of 2024", the hierarchical position corresponding to "income category - main business income" in the subject dimension, and the numerical position corresponding to "250,000 yuan" in the amount dimension. The attribute value of "main business income" is associated with the second target positions in the three dimensions. That is, when the position combination of "the third quarter of 2024 - income category - main business income - 250,000 yuan" is queried in the multi-dimensional cube, the corresponding financial data category obtained is "main business income".
[0132] Step 3042, for each data element, if the position of the fourth target node in the data storage structure tree is consistent with the second target position of the corresponding dimension in the multi-dimensional cube structure, and the attribute of the fourth target node in the data storage structure tree is associated with the attribute of the corresponding dimension in the multi-dimensional cube structure, then a two-way index relationship is established between the position of the fourth target node and the second target position.
[0133] Furthermore, the data management system compares the relevant information of each data element in the data storage structure tree and the multi-dimensional cube structure. For a certain data element, first check whether the position of the fourth target node in the data storage structure tree is logically consistent with the second target position of the corresponding dimension in the multi-dimensional cube structure. For example, check whether the hierarchical structure where the "cash and cash equivalents" node is located in the data storage structure tree matches the hierarchical position corresponding to "cash and cash equivalents" in the subject dimension of the multi-dimensional cube. Furthermore, the data management system determines whether the attribute of the fourth target node (such as attributes like financial subject category, etc.) is associated with the attribute of the corresponding dimension in the multi-dimensional cube structure (also relevant attributes such as financial subject category, etc.). When both of these conditions are met, the data management system establishes a two-way index relationship between the position of the fourth target node and the second target position, where the two-way index indicates that from a certain node position in the data storage structure tree, its corresponding position in the multi-dimensional cube structure can be quickly found, and vice versa.
[0134] In one implementation example, take the data element of "Fixed assets - Buildings" as an example. In the data storage structure tree, the "Buildings" node is at the hierarchical position of "Asset category - Non-current assets - Fixed assets" (the fourth target node position). In the multi-dimensional cube structure, the hierarchical position (the second target position) corresponding to "Buildings" on the account dimension is logically consistent with the hierarchical position in the data storage structure tree. Moreover, the attributes of the "Buildings" node in the data storage structure tree (such as belonging to the fixed asset category) are associated with the attributes of the position of "Buildings" on the account dimension in the multi-dimensional cube (also indicating belonging to the fixed asset category). At this time, the data management system establishes a two-way index relationship between the two positions. When the user queries the "Buildings" node in the data storage structure tree, through this two-way index, it can quickly locate the corresponding position of "Buildings" on the account dimension in the multi-dimensional cube structure, as well as the data on the related time dimension and amount dimension; conversely, when querying the data at the position related to "Buildings" in the multi-dimensional cube structure, it can also quickly find the corresponding node in the data storage structure tree.
[0135] Step 3043: Map each data element to the fourth target node of the data storage structure tree and map each data element to the second target position of the corresponding dimension of the multi-dimensional cube structure.
[0136] Furthermore, the data management system accurately places each data element on the corresponding fourth target node in the data storage structure tree to obtain the mapped data storage structure tree, so that the data elements have clear attribution and position identification in the data storage structure tree. At the same time, the data management system places the data elements at the specific positions of the corresponding dimensions of the multi-dimensional cube structure according to the determined second target positions to obtain the mapped multi-dimensional cube structure.
[0137] Step 3044: Based on the two-way index relationship, the mapped data storage structure tree and the mapped multi-dimensional cube structure are fused to construct an initial data network.
[0138] Furthermore, the data management system connects the mapped data storage structure tree and the multi-dimensional cube structure through the two-way index relationship to construct an initial data network, where the two-way index relationship enables the data elements in the two structures to be mutually associated and accessed. Therefore, in the data storage structure tree, each node can find the corresponding data element position in the multi-dimensional cube structure through the two-way index, and vice versa.
[0139] In one embodiment, there are multiple financial subject nodes in the data storage structure tree, such as "Cash and Cash Equivalents", "Accounts Receivable", "Short-term Borrowings", etc. Each node is connected to the multi-dimensional cube structure through a bidirectional index. In the multi-dimensional cube structure, the positions of data elements on each dimension also correspond to the nodes in the data storage structure tree through the bidirectional index. When a user wants to query the change of Cash and Cash Equivalents throughout 2024, they can first find the "Cash and Cash Equivalents" node in the data storage structure tree, and quickly jump to the corresponding position of "2024" on the time dimension in the multi-dimensional cube structure through the bidirectional index to obtain all relevant data of Cash and Cash Equivalents in the amount dimension during this period. Or starting from the position of "2024" on the time dimension in the multi-dimensional cube structure, use the bidirectional index to find all financial subject nodes in the data storage structure tree related to this period, so as to comprehensively understand the financial data situation in 2024.
[0140] The embodiment of the present invention constructs a highly integrated, closely related, and easy-to-query and analyze financial data network. On the one hand, through the data storage structure tree and the multi-dimensional cube structure, it can improve the management efficiency and management accuracy of financial statement data. On the other hand, when an enterprise processes financial data, the initial data network can greatly improve the data processing efficiency.
[0141] In one embodiment, the descriptions of steps 401 to 404 are as follows:
[0142] Step 401: Analyze the financial statement data to obtain the first data elements with associated relationships in the same dimension and the second data elements with associated relationships in different dimensions.
[0143] Optionally, the associated relationships in the embodiment of the present invention include causal association, temporal sequence association, and subject association. Therefore, the data management system analyzes the financial statement data to obtain the first data elements with associated relationships in the same dimension and the second data elements with associated relationships in different dimensions.
[0144] For the data elements with associated relationships in the same dimension, that is, the first data elements, in terms of causal association, for example, in the amount dimension, there is a causal association between the cost data and the profit data of an enterprise. An increase in cost often leads to a decrease in profit. In terms of temporal sequence association, in the time dimension, the sales data for several consecutive quarters has a temporal sequence association. The sales volume in the subsequent quarter may be affected by factors such as the sales situation in the previous quarter and market trends. Subject association is reflected in the subject dimension. For example, "Main Business Income" and "Main Business Cost" are both financial subjects related to the main business and have a close subject association.
[0145] For the second data elements with correlation relationships in different dimensions, in terms of causal correlation, such as between the promotion activity time in the time dimension and the sales amount in the amount dimension, the launch of the promotion activity (time dimension) may trigger the growth of the sales amount (amount dimension). In terms of temporal correlation, the year-end closing time point in the time dimension will affect the data update and adjustment of some accounting subjects in the subject dimension, showing a chronological order correlation. Subject correlation can be manifested as the correlation between the "fixed assets" subject in the subject dimension and data elements such as the purchase amount and depreciation amount of fixed assets in the amount dimension, because different fixed asset subjects correspond to specific amount values. Through in-depth analysis of the financial statement data, the data management system identifies these data elements with different types and dimension correlations.
[0146] In one embodiment, in the financial statement data, two data elements, "cost of sales" and "gross profit", have a causal correlation in the amount dimension and belong to the first data elements. Because an increase in the cost of sales will directly lead to a decrease in the gross profit. At the same time, "sales in the first quarter of 2024" and "sales in the second quarter of 2024" have a temporal correlation in the time dimension and are also the first data elements. In the subject dimension, the "raw material procurement" subject and the "finished goods" subject have a subject correlation and belong to the category of the first data elements.
[0147] For the second data elements, there is a causal correlation between "promotion activity carried out in June 2024" (time dimension) and "substantial increase in sales amount in June 2024" (amount dimension). There is a temporal correlation between "year-end closing time" (time dimension) and "adjustment of some subject data in the financial statements" (subject dimension). There is a subject correlation between the "fixed assets - buildings" subject (subject dimension) and "purchase amount of buildings and annual depreciation amount" (amount dimension).
[0148] Step 402, for the first data elements, establish node associations between the corresponding nodes of the first data elements in the data storage structure tree and establish position associations between the corresponding positions of the first data elements in the multi-dimensional cube structure.
[0149] Furthermore, for the first data elements, in the data storage structure tree, the data management system finds their corresponding nodes and establishes associations. If the nodes in the data storage structure tree corresponding to the data elements "cost of sales" and "gross profit" are respectively the "cost of sales node" and the "gross profit node", establish a connection between the two nodes to indicate their causal correlation relationship. For the "sales in the first quarter of 2024 node" and the "sales in the second quarter of 2024 node" with a temporal correlation, establish a node association in chronological order to reflect the temporal relationship. In terms of subject correlation, if there is a subject correlation between the "raw material procurement node" and the "finished goods node", associate them.
[0150] In a multi-dimensional cube structure, associations are also established for the corresponding positions of the first data elements. In the amount dimension, a position association is established between the corresponding positions of "cost of goods sold" and "gross profit from sales" to show a causal association. In the time dimension, the corresponding positions of "sales volume in the first quarter of 2024" and "sales volume in the second quarter of 2024" are associated in chronological order. In the subject dimension, an association is established between the corresponding positions of "raw material procurement" and "merchandise inventory" to reflect the subject association.
[0151] Step 403: For the second data elements, cross-dimensional associations are established between the corresponding positions of the second data elements in the multi-dimensional cube structure.
[0152] Furthermore, for the second data elements with an association relationship in different dimensions, when there is a causal association, such as "promotion activities carried out in June 2024" (time dimension) and "substantial increase in sales volume in June 2024" (amount dimension), a cross-dimensional association is established between the position of "June 2024" in the time dimension and the corresponding position of "substantial increase in sales volume in June 2024" in the amount dimension. For chronological associations, like "year-end closing time" (time dimension) and "adjustment of some subject data in the financial statements" (subject dimension), an association is established between the year-end closing time position in the time dimension and the corresponding position of the relevant subject data adjustment in the subject dimension to reflect the chronological order. In terms of subject association, for the subject of "fixed assets - buildings" (subject dimension) and "purchase amount of buildings, annual depreciation amount" (amount dimension), a cross-dimensional association is established between the position of "fixed assets - buildings" in the subject dimension and the corresponding purchase amount and depreciation amount positions in the amount dimension.
[0153] In an embodiment, in the multi-dimensional cube structure, the time scale position of "June 2024" is found in the time dimension, and the corresponding amount value position of "substantial increase in sales volume in June 2024" is found in the amount dimension. An association line spanning the time dimension and the amount dimension is created to connect these two positions to show a causal association. For the position of "year-end closing time" in the time dimension and the corresponding position of the relevant subject data adjustment in the subject dimension, an association is established to indicate a chronological association. A cross-dimensional association is established between the hierarchical position of "fixed assets - buildings" in the subject dimension and the corresponding numerical positions of "purchase amount of buildings" and "annual depreciation amount" in the amount dimension to present a subject association.
[0154] Step 404: Based on the node association, position association, and cross-dimensional association, the first data elements and the second data elements in the initial data network are associated to obtain the final data network.
[0155] Furthermore, in the initial data network, the data management system closely links the first data elements through node association and location association, and associates and displays the second data elements in a multi-dimensional cube structure through cross-dimensional association. Finally, all these association relationships are integrated, enabling the first data elements and the second data elements to form an organic whole in the entire data network, thus constructing the final data network. In this final data network, all kinds of association relationships between different data elements are clearly presented. Whether it is the association within the same dimension or the association between different dimensions, they can be conveniently queried and analyzed, providing enterprises with more comprehensive and in-depth financial data insights.
[0156] In one embodiment, there are numerous first data elements and second data elements in the initial data network. The first data elements such as "cost of sales", "gross profit", "quarterly sales in 2024", "raw material procurement", "merchandise inventory", etc. have been linked through node association and location association in the data storage structure tree and the multi-dimensional cube structure. The second data elements such as "promotion activity time and sales volume change", "year-end closing time and account data adjustment", "fixed asset account and amount data", etc. are associated with each other in the multi-dimensional cube structure through cross-dimensional association. These different types of association relationships are unified and integrated to connect all relevant data elements. When financial personnel want to analyze the impact of the enterprise's sales strategy on profits, in the final data network, they can start from the promotion activity time node in the time dimension, find the data of sales volume change in the amount dimension through cross-dimensional association, then find the cost of sales data through location association, and further find the gross profit data through node association and location association, comprehensively understanding the data association relationships in the entire sales business chain, providing strong support for decision-making, and completing the construction and application of the final data network.
[0157] The embodiment of the present invention constructs a comprehensive, detailed and highly associated financial data network. The financial data network presents various association relationships between different data elements. Therefore, when an enterprise conducts financial analysis, whether it is the association within the same dimension or the association between different dimensions, it can utilize the final data network to deeply explore the internal connections between financial data from multiple perspectives, facilitating query and analysis, providing enterprises with more comprehensive and in-depth financial data insights, and improving the analysis efficiency and query efficiency.
[0158] In one embodiment, the descriptions of steps 4041 to 4042 are as follows:
[0159] Step 4041, optimize the initial association path of node association and the initial association method of location association respectively based on the association strength between the first data elements to obtain the optimized path and the optimized method.
[0160] Optionally, the data management system analyzes the association strength between the first data elements. The association strength can be determined based on multiple factors, such as the degree of change consistency between financial data, the closeness in business logic, etc. For the initial association path of node associations, the data management system examines the connection method between nodes in the data storage structure tree. If the nodes corresponding to the first data elements "cost of sales" and "gross profit" are initially associated through a relatively complex and indirect path, but it is found through analysis that the association strength between these two elements is extremely high, it will attempt to find a more direct and efficient association path. For example, in the hierarchical structure of the data storage structure tree, there may be a closer common parent node. By adjusting the association path, directly starting from this common parent node to connect the "cost of sales" and "gross profit" nodes, an optimized path is formed, enabling faster jumping between these two nodes during data query and analysis.
[0161] Furthermore, for the initial association method of position associations, in the multi-dimensional cube structure, if the initial association method between "sales in the first quarter of 2024" and "sales in the second quarter of 2024" in the time dimension is just a simple sequential connection, but considering that the association strength between these two data elements is high, and the enterprise often focuses on comparing the sales changes in these two quarters during analysis, the data management system optimizes the association method to connect them with a more prominent line that is different from other ordinary associations, or marks special prompt information on the association line, such as "focus on comparing quarterly sales", to obtain the optimized method.
[0162] In one embodiment, in the data storage structure tree, the initial association path between the "raw material procurement" node and the "production cost" node is connected through multiple intermediate-level nodes because they have undergone transformations in multiple links in the business process. However, from the analysis of the association strength of financial data, it is found that the change in raw material procurement volume is closely related to the fluctuation of production cost, and the association strength is high. After searching, it is found that "raw material and cost management" is the common parent node of these two nodes. Therefore, the association path is optimized to directly start from the "raw material and cost management" node and connect the "raw material procurement" and "production cost" nodes respectively, forming a more concise and efficient optimized path. In the multi-dimensional cube structure, the initial association method between "main business income" and "main business cost" in the amount dimension is an ordinary straight-line connection. Considering the core position of these two data elements in financial analysis and the extremely high association strength, the association method is optimized to connect them with a thick and red line, and mark "core association of main business" beside the line to obtain the optimized association method.
[0163] Step 4042, associate the first data elements in the initial data network based on the optimized path and the optimized method, and associate the second data elements in the initial data network based on cross-dimensional association to obtain the final data network.
[0164] Further, the data management system re-associates the first data element in the initial data network by using the optimized path and optimized method obtained in step 4041. Taking the data storage structure tree as an example, according to the optimized node association path, the connection relationship between the nodes corresponding to the first data element is re-established to ensure that the association between data elements in the tree structure is more in line with their actual strong association characteristics. In the multi-dimensional cube structure, according to the optimized position association method, the association display between the positions corresponding to the first data element is re-set to highlight the position relationship of strongly associated data elements.
[0165] Meanwhile, for the second data element, based on the cross-dimensional association established in step 403, the data management system maintains and improves the association between the corresponding positions of the second data element in different dimensions in the initial data network. By combining these two parts of work, the association relationship between the first data element and the second data element in the initial data network is integrated and optimized to obtain the final data network.
[0166] In one embodiment, in the initial data network, for the first data elements, such as "inventory goods" and "cost of goods sold", according to the optimized path, in the data storage structure tree, directly connect these two nodes from their common parent node "cost of goods operation" to strengthen their association in the tree structure. In the multi-dimensional cube structure, according to the optimized method, connect the positions corresponding to "inventory goods" and "cost of goods sold" in the amount dimension and the subject dimension with a specially designed association identifier (such as a colored line with an arrow) to highlight their strong association relationship. For the second data element, like "promotion activity time" (time dimension) and "amount of sales growth" (amount dimension), maintain the previously established cross-dimensional association. Finally, the data management system integrates these associations to construct the final data network.
[0167] The final data network generated by the embodiment of the present invention is more scientific, efficient and targeted. Therefore, when enterprise financial personnel conduct financial analysis, they can obtain strongly associated data elements and their relationships more quickly and accurately, improving the analysis efficiency and query efficiency.
[0168] Further, the financial statement data comprehensive management system based on big data analysis provided by the present invention is described below. The financial statement data comprehensive management system based on big data analysis described below can be mutually corresponding and referred to with the financial statement data comprehensive management method based on big data analysis described above.
[0169] Refer to Figure 2 , Figure 2It is a schematic structural diagram of a comprehensive financial statement data management system based on big data analysis provided by the present invention. The comprehensive financial statement data management system based on big data analysis includes.
[0170] A structure tree construction module 210, configured to perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data categories and hierarchical relationships of the data partitioning results;
[0171] A node mapping module 220, configured to map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category;
[0172] A data mapping module 230, configured to map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network;
[0173] A data association module 240, configured to associate the data elements corresponding to different nodes and different dimensions in the initial data network based on the association relationships between the financial statement data to obtain a final data network.
[0174] In an embodiment of the present invention, the financial statement data is managed through a data network constructed by a data storage structure tree and a multi-dimensional cube structure, improving the management efficiency and management accuracy of the financial statement data.
[0175] Please refer to Figure 3 , Figure 3 which is an embodiment diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:
[0176] Perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data categories and hierarchical relationships of the data partitioning results;
[0177] Map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category;
[0178] Map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network;
[0179] Based on the association relationships between the financial statement data, associate the data elements corresponding to different nodes and different dimensions in the initial data network to obtain a final data network.
[0180] Please refer to Figure 4 , Figure 4 which is the embodiment diagram of the computer-readable storage medium provided by the embodiment of the present invention. As Figure 4 shown, this embodiment provides a computer-readable storage medium 400, on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:
[0181] Perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data categories and hierarchical relationships of the data partitioning results;
[0182] Map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category;
[0183] Map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network;
[0184] Based on the association relationships between the financial statement data, associate the data elements corresponding to different nodes and different dimensions in the initial data network to obtain a final data network.
[0185] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the comprehensive management method of financial statement data based on big data analysis provided by the above-mentioned various methods.
[0186] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A comprehensive management method for financial statement data based on big data analysis, characterized in that, Including: Partition the collected financial statement data, and construct a data storage structure tree according to the data categories and hierarchical relationships of the data partition results; Map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; Each node in the data storage structure tree represents a data category; Map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network; Based on the association relationships between the financial statement data, associate the data elements corresponding to different nodes and different dimensions in the initial data network to obtain a final data network; Among them, the step of mapping each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure includes: Identify a first target node related to time, a second target node related to financial subjects, and a third target node related to amounts in the data storage structure tree; For the first target node related to time, map it to the first initial position corresponding to the time dimension of the multi-dimensional cube in chronological order from early to late; for the second target node related to financial subjects, map it to the first initial position corresponding to the subject dimension of the multi-dimensional cube according to the hierarchical relationship and classification system of financial subjects; for the third target node related to amounts, map it to the first initial position corresponding to the amount dimension of the multi-dimensional cube according to the financial subject and time node to which it belongs; the amount nodes at different time points under the same financial subject are arranged in chronological order on the amount dimension; Based on the parent-child relationship and sibling relationship of each node in the data storage structure tree, as well as business logic and data usage frequency, optimize the first initial position of each node in the dimension where the multi-dimensional cube is located to obtain the first target position of each node in the dimension; In the multi-dimensional cube, use the financial data category of each node as the attribute value of each node, and combine the first target position of each node in the dimension to construct the multi-dimensional cube structure; The step of partitioning the collected financial statement data and constructing a data storage structure tree according to the data categories and hierarchical relationships of the data partition results includes: In the multi-dimensional financial data space, with each data point in the financial statement data as the center, determine the regional density of each data point according to the number of data points within a preset range, and perform data clustering on the data points with the same regional density to obtain an initial clustering; For each initial clustering, if the density change of the target data points at the boundary between the initial clustering and its adjacent clustering is less than or equal to a preset degree threshold, and the distance of the target data points in the feature space is less than or equal to a preset distance threshold, then fuse each initial clustering with its adjacent clustering to obtain an optimized clustering; For each optimized clustering, determine the sub-topic financial data category corresponding to each optimized clustering according to the distribution shape, density characteristics, and clustering characteristics of the data points within the optimized clustering; Construct the data storage structure tree based on each sub - topic financial data category and the hierarchical relationship of each sub - topic financial data category in its topic financial data category.
2. The comprehensive management method for financial statement data based on big data analysis according to claim 1, wherein The constructing of the data storage structure tree based on each sub - topic financial data category and the hierarchical relationship of each sub - topic financial data category in its topic financial data category includes: For each sub - topic financial data category, determine the data level of each sub - topic financial data category in its sub - topic financial data category according to the data volume size of the sub - topic financial data category and its relevance to the core business in its topic financial data category. For each sub - topic financial data category, divide the data points in the sub - topic financial data category into sub - categories according to the data characteristics of the data points in the sub - topic financial data category, and sort the sub - categories according to the data characteristic differences between the sub - categories to obtain the optimized data category of each sub - topic financial data category. Using each topic financial data category as the root node, the optimized data category under each topic financial data category as the leaf node, and the data level of each sub - topic financial data category in its sub - topic financial data category as the node level of the leaf node, construct the initial storage structure tree. Based on the actual access frequency and business logic association between financial data categories, associate the corresponding nodes in the initial storage structure tree to obtain the data storage structure tree.
3. The comprehensive financial statement data management method based on big data analysis according to claim 1, characterized in that The mapping of each data element in the financial statement data to the data storage structure tree and the multi - dimensional cube structure to obtain the initial data network includes: Extract the features of each data element in the financial statement data to obtain the element features of each data element; the element features at least include the element name, element type, financial statement category to which it belongs, time attribute, financial account information involved, and the amount value involved. For each data element, perform path matching in the data storage structure tree according to the element features to obtain the fourth target node in the data storage structure tree, and map each data element to the second initial position of the corresponding dimension in the multi - dimensional cube structure based on the element features. For each data element, optimize the second initial position according to the parent - child relationship and sibling relationship of the fourth target node in the data storage structure tree to obtain the second target position of each data element in the corresponding dimension of the multi - dimensional cube structure. Based on the fourth target node and the second target position, map each data element to the data storage structure tree and the multi - dimensional cube structure respectively to obtain the initial data network.
4. The comprehensive management method for financial statement data based on big data analysis according to claim 3, characterized in that, The mapping of each data element to the data storage structure tree and the multi - dimensional cube structure respectively based on the fourth target node and the second target position to obtain the initial data network includes: Associate the financial data category of each data element as an attribute value with the second target position of each dimension in the multi - dimensional cube structure. For each data element, if the position of the fourth target node in the data storage structure tree is the same as the second target position in the corresponding dimension of the multi-dimensional cube structure, and the attribute of the fourth target node in the data storage structure tree is associated with the attribute in the corresponding dimension of the multi-dimensional cube structure, a two-way index relationship is established between the position of the fourth target node and the second target position; Map each data element to the fourth target node of the data storage structure tree, and map each data element to the second target position in the corresponding dimension of the multi-dimensional cube structure; Based on the two-way index relationship, the mapped data storage structure tree, and the mapped multi-dimensional cube structure, perform fusion to construct an initial data network.
5. The comprehensive management method for financial statement data based on big data analysis according to claim 1, characterized in that Associating the data elements corresponding to different nodes and different dimensions in the initial data network based on the association relationship between the financial statement data, to obtain the final data network, including: Analyze the financial statement data to obtain the first data elements with an association relationship in the same dimension and the second data elements with an association relationship in different dimensions; For the first data elements, establish node associations between the corresponding nodes in the data storage structure tree, and establish position associations between the corresponding positions in the multi-dimensional cube structure; For the second data elements, establish cross-dimensional associations between the corresponding positions in the multi-dimensional cube structure; Based on the node associations, position associations, and cross-dimensional associations, associate the first data elements and the second data elements in the initial data network to obtain the final data network.
6. The comprehensive management method for financial statement data based on big data analysis according to claim 5, characterized in that Associating the first data elements and the second data elements in the initial data network based on the node associations, position associations, and cross-dimensional associations, to obtain the final data network, including: Optimize the initial association path of the node association and the initial association method of the position association respectively based on the association strength between the first data elements to obtain the optimized path and the optimized method; Based on the optimized path and the optimized method, associate the first data elements in the initial data network, and based on the cross-dimensional association, associate the second data elements in the initial data network to obtain the final data network.
7. A comprehensive financial statement data management system based on big data analysis, characterized in that, Applied to the comprehensive management method of financial statement data based on big data analysis as described in any one of claims 1 to 6; The comprehensive management system of financial statement data based on big data analysis includes: A structure tree construction module, configured to perform data partitioning on the collected financial statement data, and construct a data storage structure tree according to the data category and hierarchical relationship of the data partitioning result; A node mapping module, configured to map each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure; each node in the data storage structure tree represents a data category; A data mapping module, configured to map each data element in the financial statement data to the data storage structure tree and the multi-dimensional cube structure to obtain an initial data network; A data association module, configured to associate data elements corresponding to different nodes and different dimensions in the initial data network based on the association relationships among the financial statement data, so as to obtain a final data network; Wherein, the mapping each node in the data storage structure tree to a dimension of a multi-dimensional cube to construct a multi-dimensional cube structure includes: Identifying a first target node related to time, a second target node related to financial subjects, and a third target node related to amounts in the data storage structure tree; For the first target node related to time, mapping it to the first initial position corresponding to the time dimension of the multi-dimensional cube in ascending order of years from early to late; for the second target node related to financial subjects, mapping it to the first initial position corresponding to the subject dimension of the multi-dimensional cube according to the hierarchical relationship and classification system of the financial subjects; for the third target node related to amounts, mapping it to the first initial position corresponding to the amount dimension of the multi-dimensional cube according to the financial subject and time node to which it belongs; the amount nodes at different time points under the same financial subject are arranged in sequence in the amount dimension according to the time order; Based on the parent-child relationship and sibling relationship of each node in the data storage structure tree, as well as business logic and data usage frequency, optimizing the first initial position of each node in the dimension where the multi-dimensional cube is located to obtain the first target position of each node in the dimension; In the multi-dimensional cube, using the financial data category of each node as the attribute value of each node, and combining the first target position of each node in the dimension, constructing the multi-dimensional cube structure; The data partitioning of the collected financial statement data, and constructing a data storage structure tree according to the data category and hierarchical relationship of the data partitioning result, includes: In the multi-dimensional financial data space, taking each data point in the financial statement data as the center, determining the regional density of each data point according to the number of data points within a preset range of each data point, and clustering the data points with the same regional density to obtain an initial clustering; For each initial clustering, if the density change of the target data points at the boundary between the initial clustering and its adjacent clustering is less than or equal to a preset degree threshold, and the distance of the target data points in the feature space is less than or equal to a preset distance threshold, then fusing each initial clustering with its adjacent clustering to obtain an optimized clustering; For each optimized clustering, determining the sub-topic financial data category corresponding to each optimized clustering according to the distribution shape, density feature, and clustering feature of the data points within the optimized clustering; Based on each sub-topic financial data category and the hierarchical relationship of each sub-topic financial data category in its topic financial data category, constructing the data storage structure tree.
8. A non-transitory computer-readable storage medium storing a computer software program, characterized in that, When the computer software program is executed by a processor, it implements the comprehensive management method for financial statement data based on big data analysis according to any one of claims 1 to 6.
Citation Information
Patent Citations
Power grid current hierarchical partitioned three-dimensional visualized display method
CN108763669A
Financial data accurate tracing method based on block chain MPT tree
CN114880331A