Enterprise data processing method and system based on structure tree

By using a structure tree-based enterprise data processing method and ant colony genetic algorithm optimization, the problem of low efficiency in enterprise revenue forecasting in existing technologies is solved, and efficient and accurate revenue forecasting is achieved.

CN120780888BActive Publication Date: 2026-05-12兵器装备集团财务有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
兵器装备集团财务有限责任公司
Filing Date
2025-06-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

现有技术无法通过大数据多维度分析方式有效降低人力,提高企业营收预测的数据统计效率,同时保障预测的准确性。

Method used

A structure tree-based enterprise data processing method is adopted. Enterprise operating data is acquired and analyzed through the server, a multi-level dimensional structure tree and extraction table are constructed, an operating function is generated, and revenue is predicted based on the enterprise similarity fusion comprehensive function. Ant colony algorithm and genetic algorithm are combined to optimize weights to improve prediction accuracy.

Benefits of technology

This approach achieves the goal of reducing manpower requirements, improving data statistics and forecasting efficiency, and enhancing the accuracy and credibility of enterprise revenue forecasts while ensuring the accuracy of revenue forecasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780888B_ABST
    Figure CN120780888B_ABST
Patent Text Reader

Abstract

The application provides an enterprise data processing method and system based on a structure tree, which can process enterprise data based on a structure tree, a genetic algorithm and an ant colony algorithm. The operation data of the enterprise can be quickly decomposed, classified, stored or extracted in the mode of the structure tree, and the corresponding algorithm is used for processing, so that the relative accuracy and rationality of the later calculation are effectively ensured. In the genetic algorithm, the population individuals constantly perform genetic operations, so that the individuals in the population continuously evolve towards the global optimal solution. The genetic algorithm has a relatively large search space coverage area and good parallel search capability. The ant colony algorithm has positive feedback characteristics due to the continuous accumulation of pheromone, which promotes the algorithm to converge to the optimal solution faster. The ant colony algorithm can perform multi-point search in the solution space, and has stronger robustness and optimization capability than other algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data processing technology, and more particularly to an enterprise data processing method and system based on a tree structure. Background Technology

[0002] Corporate revenue forecasting is an important tool for corporate strategic planning, budgeting, and investment decisions. It is usually estimated by combining historical data, market analysis, and statistical models.

[0003] Business operations often follow certain patterns and market cycles. Existing technologies cannot reduce the need for manual data collection through multi-dimensional big data analysis, thereby improving the efficiency of data collection for revenue forecasting while ensuring relatively accurate revenue figures. Summary of the Invention

[0004] This invention provides a structure tree-based enterprise data processing method and system that enables technical enterprise revenue forecasting. It reduces manpower and improves the efficiency of data statistics through multi-dimensional big data analysis, thereby improving forecasting efficiency while ensuring relatively accurate enterprise revenue forecasting.

[0005] A first aspect of this invention provides a method for processing enterprise data based on a structure tree, comprising:

[0006] The server obtains the first business data of the first enterprise, and classifies the first business data according to preset multi-level dimensions to obtain the first structure tree and the first extraction table;

[0007] The server receives the second business data of the second enterprise added by the analysis terminal, classifies it to obtain a second structure tree and a second extraction table, extracts the business specification information of the second structure tree, and partitions the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas;

[0008] The server constructs information points corresponding to each information area and generates corresponding first business functions according to the preset multi-level dimensions. The first business functions are then fused based on the similarity between the first enterprise and the second enterprise to obtain a comprehensive function.

[0009] Based on the comprehensive function, the server obtains the predicted revenue data and credibility information from the second operating data and feeds it back to the analysis end.

[0010] Optionally, in one possible implementation of the first aspect, the server obtains the first operating data of the first enterprise, and classifies the first operating data according to preset multi-level dimensions to obtain a first structure tree and a first extraction table, including:

[0011] The server filters companies in the database at preset time intervals, and adds a first label to the corresponding companies after determining that the first condition is met.

[0012] The server obtains the first operating data of the first enterprise, which includes at least financial information, personnel information, and shareholder information.

[0013] The first operating data is decomposed according to different categories and dimensions based on the first time series interval, resulting in a first structure tree and a first extraction table with multiple dimensions of time and category. The information in the first structure tree and the first extraction table is set accordingly.

[0014] Optionally, in one possible implementation of the first aspect, the server receives the second business data of the second enterprise added by the analysis terminal, categorizes it to obtain a second structure tree and a second extraction table, extracts the business specification information of the second structure tree, and partitions the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas, including:

[0015] The information of each dimension in the second operating data is decomposed according to the preset analysis points to obtain the operating specification information corresponding to each dimension. The operating specification information includes at least a time period.

[0016] A second structure tree and a second extraction table are generated based on the dimensions and time periods to categorize the second business data.

[0017] The first structure tree is processed and filtered based on the second structure tree to obtain multiple information areas. The first extraction table and the second extraction table are adjusted synchronously. Each information area has a corresponding category dimension and time dimension.

[0018] Optionally, in one possible implementation of the first aspect, the step of decomposing the information of each dimension in the second operating data according to preset analysis points to obtain the operating specification information corresponding to each dimension includes:

[0019] Obtain the maximum span time value of information for each dimension in the second operating data. If the maximum span time value is greater than or equal to a preset time value, then obtain the preset operating specification information.

[0020] If the maximum span time value is less than the preset time value, the maximum span time value is divided equally based on the number of preset analysis points to obtain the calculated business specification information.

[0021] Optionally, in one possible implementation of the first aspect, the first information region is obtained by processing and filtering the first information structure tree based on the second information structure tree, and the first extraction table and the second extraction table are adjusted synchronously. Each information region has a corresponding category dimension and time dimension, including:

[0022] The child nodes of the corresponding category dimension in the first tree structure are retained;

[0023] Determine the operational specifications of the grandchild nodes corresponding to the child nodes of each category dimension;

[0024] Based on the second structure tree, the grandchild nodes retained in the first structure tree are re-divided according to the business specification information, and non-corresponding nodes are removed to obtain the updated first structure tree, the first extraction table, and the corresponding information area.

[0025] Optionally, in one possible implementation of the first aspect, the server constructs information points corresponding to each information area and generates corresponding first business functions according to corresponding preset multi-level dimensions. A comprehensive function is obtained by fusing the first business functions based on the similarity between the first enterprise and the second enterprise, including:

[0026] The server uses the child nodes of the first tree structure as the initial information area, and retains the grandchild nodes in each initial information area based on the preset or calculated business specification information to obtain information points.

[0027] Based on the time dimension of the information points, all information points are statistically analyzed using functions to generate corresponding first business functions, with each information area corresponding to one first business function;

[0028] Calculate the corresponding similarity weights based on the similarity between the first company and the second company, and add them to the corresponding first structure tree.

[0029] A comprehensive function is obtained by fusing the similarity weights within the first structure tree and the first operating function.

[0030] Optionally, in one possible implementation of the first aspect, the calculation generates corresponding similarity weights based on the similarity between the first enterprise and the second enterprise, and a comprehensive function is obtained by fusing the similarity weights and the first operating function, including:

[0031] Extract the first and second attribute labels of the first and second enterprises, and calculate the attribute difference for each dimension by quantifying the first and second attribute labels.

[0032] Compare the original first and second structure trees to obtain the difference values ​​of child nodes and grandchild nodes;

[0033] The similarity between the first enterprise and the second enterprise is obtained based on the attribute difference, child node difference, and grandchild node difference. The similarity between the first enterprise and the second enterprise is then processed as a percentage to obtain the similarity weight. Based on the similarity weight, the first operating function is structured and fused to obtain the comprehensive function.

[0034] Optionally, in one possible implementation of the first aspect, the step of obtaining a comprehensive function by merging the first operating function based on similarity weights using a structure tree includes:

[0035] Each first enterprise's parent node is treated as a child node and then connected to the same parent node to obtain an analysis structure tree. The similarity weights are then mapped to the child nodes of the analysis structure tree.

[0036] Determine the dimensional weights of each grandchild node in the analysis tree structure;

[0037] Based on the analysis of the tree structure, the revenue forecast of the second operating data is obtained and the forecast data and credibility information are fed back to the analysis end.

[0038] A second aspect of the present invention is a method for processing enterprise data using the first aspect of the present invention, characterized in that it includes:

[0039] Step S1: During the enterprise revenue trend prediction calculation, initialize the ant colony-related parameters and the genetic algorithm-related parameters, including ant colony size, pheromone update rate, genetic algorithm population size, crossover and mutation probabilities, etc.

[0040] Step S2: Use the ant colony algorithm to generate a set of initial solutions and calculate the fitness of each solution;

[0041] A genetic algorithm is used to select a subset of individuals with high fitness from the initial solution to form a population.

[0042] Step S3: Perform crossover and mutation operations on the population to generate the next generation of individuals;

[0043] Step S4: Calculate the fitness of the next generation of individuals;

[0044] Step S5: If the specified number of iterations is reached or a preset satisfactory solution is found, output the result; otherwise, return to step S3 and continue optimization.

[0045] A third aspect of the present invention provides an enterprise data processing system based on a structure tree, comprising:

[0046] The acquisition module is used to enable the server to acquire the first business data of the first enterprise, and to classify the first business data according to preset multi-level dimensions to obtain the first structure tree and the first extraction table.

[0047] The extraction module is used to enable the server to receive and classify the second business data of the second enterprise added by the analysis terminal to obtain a second structure tree and a second extraction table, extract the business specification information of the second structure tree, and perform partitioning processing on the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas;

[0048] The generation module enables the server to construct information points corresponding to each information area and generate corresponding first business functions according to the corresponding preset multi-level dimensions. The first business functions are then fused based on the similarity between the first enterprise and the second enterprise to obtain a comprehensive function.

[0049] The prediction module is used to enable the server to obtain predicted data and credibility information based on the comprehensive function for the second operating data revenue prediction and feed it back to the analysis end. Attached Figure Description

[0050] Figure 1 A flowchart of a structure tree-based enterprise data processing method;

[0051] Figure 2 This is a structural diagram of an enterprise data processing system based on a tree structure. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0054] It should be understood that in the various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0055] It should be understood that in this invention, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0056] It should be understood that in this invention, "multiple" refers to two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B, and C", "Contains A, B, and C" means that all three A, B, and C are contained; "Contains A, B, or C" means that one of A, B, and C is contained; "Contains A, B, and / or C" means that any one, two, or three of A, B, and C are contained.

[0057] It should be understood that in this invention, "B corresponding to A", "B corresponding to A", "A and B correspond", or "B and A correspond" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Matching A and B is defined as a similarity between A and B that is greater than or equal to a preset threshold.

[0058] Depending on the context, "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection."

[0059] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0060] This invention provides a method for enterprise data processing based on a structure tree, such as... Figure 1 As shown, it includes:

[0061] The server obtains the first operating data of the first enterprise, and classifies the first operating data according to preset multi-level dimensions to obtain a first structure tree and a first extraction table. The technical solution provided by this invention first obtains the first operating data of the first enterprise, which can be considered as an enterprise pre-stored in a database. After obtaining the first operating data of the first enterprise, this invention decomposes it according to multiple dimensions to obtain the corresponding first structure tree and first extraction table. At this point, the first operating data has undergone initial structured processing.

[0062] Furthermore, the server acquires the first operating data of the first enterprise, and classifies the first operating data according to preset multi-level dimensions to obtain a first structure tree and a first extraction table, including:

[0063] The server filters companies in the database at preset time intervals. Once a first condition is met, a first label is added to the corresponding company. Since there may be various types of companies in the database, with different operating methods, operating times, and operating conditions, this invention filters companies to determine their data value. The first condition can be pre-set, such as operating time conditions, social security conditions, etc. Only when the condition is met will this invention add the first label to the corresponding company. Companies with the first label can be considered the first-tier companies.

[0064] The server retrieves the first company's initial operating data, which includes at least financial information, personnel information, and shareholder information. It should be noted that there are many types of operating data for a company; the above is only a partial list.

[0065] The invention decomposes the first operating data according to different categories and dimensions based on a first time series interval, resulting in a first structure tree and a first extraction table with multi-level dimensions of time and category. The information in the first structure tree and the first extraction table are set accordingly. The invention decomposes the first operating data according to different categories and dimensions based on a first time series interval, which can be a preset period, such as a week, a month, etc. The invention obtains a first structure tree and a first extraction table with multi-level dimensions of time and category based on the decomposition. The first structure tree includes at least a parent node, child nodes, and grandchild nodes. Child nodes can be seen as corresponding to category dimensions, such as income, expenditure, etc. Each child node connects to multiple grandchild nodes, and the grandchild nodes correspond to the income and expenditure of a period. The first structure tree and the first extraction table are corresponding. The invention stores the information of the first enterprise according to the first time series interval, resulting in faster data processing speeds in subsequent data processing.

[0066] The server receives and categorizes the second enterprise's second operating data added by the analysis end to obtain a second structure tree and a second extraction table. It then extracts the operating specification information from the second structure tree and partitions the first structure tree and the first extraction table based on this information to obtain multiple information areas. In the technical solution provided by this invention, the server categorizes the second enterprise's second operating data added by the analysis end to obtain a second structure tree and a second extraction table. The second enterprise can be considered as the enterprise requiring data processing and analysis. In this process, the second operating data is categorized to obtain a second structure tree and a second extraction table. The category dimension and time dimension within the second structure tree and the second extraction table are specific to the second enterprise and can be actively configured by the analysis end or set according to a preset logic. Both the category dimension and the time dimension can be preset.

[0067] Because the first and second tree structures may differ in terms of category and time, this invention extracts the business specifications information from the second tree structure and partitions the first tree structure and the first extraction table based on this information to obtain multiple information regions. For example, if the second tree structure does not have R&D expenses, while the company corresponding to the first tree structure does, then the first and second tree structures differ in the category dimension. As another example, if the company in the second tree structure was established in 2021, while the company corresponding to the first tree structure was established in 2011, then the first and second tree structures differ in the time dimension, as the second tree structure lacks the information from 2011 to 2021 found in the first tree structure.

[0068] Furthermore, the server receives and categorizes the second enterprise's second operating data added by the analysis terminal to obtain a second structure tree and a second extraction table. It extracts the operating specification information from the second structure tree and, based on the operating specification information, partitions the first structure tree and the first extraction table to obtain multiple information areas, including:

[0069] The invention decomposes the information of each dimension in the second operating data according to preset analysis points to obtain the corresponding operating specification information for each dimension. This operating specification information includes at least a time period. For example, the time period corresponding to R&D expenses is described in the following example.

[0070] Furthermore, the step of decomposing the information of each dimension in the second operating data according to preset analysis points to obtain the corresponding operating specification information for each dimension includes:

[0071] The maximum span time value of information in each dimension of the second operating data is obtained. If the maximum span time value is greater than or equal to a preset time value, the preset operating specification information is obtained. This invention obtains the maximum span time value of information in each dimension. Since the information length of each dimension may be different, this invention needs to determine the corresponding maximum span value. For example, a company may not have R&D expenses in 2021 but has R&D expenses in 2022. If the current time is 2023, then the maximum span time values ​​of financial income and R&D expenditure may be different. The preset time value can be 3 years, 2 years, etc. This invention obtains the operating specification information based on the maximum span time value of information in each dimension of the second operating data. This preset operating specification information can be 3 years or 2 years.

[0072] If the maximum span time value is less than the preset time value, the maximum span time value is evenly divided based on the preset number of analysis points to obtain the calculated operating specification information. In this case, it proves that the time of the corresponding dimension is short. At this time, the present invention will evenly divide the maximum span time value according to the preset number of analysis points to obtain the calculation. For example, if the preset number of analysis points is 5 and the maximum span time value is 5 months, then the calculated operating specification information is 1 month.

[0073] This invention generates a second structure tree and a second extraction table based on the second classification of operational data according to dimensions and time periods. The invention generates a second structure tree based on dimensions and time periods, where child nodes correspond to different dimensions, and each grandchild node corresponds to a specific time period under its child node. Simultaneously, the invention generates a second extraction table in tabular form.

[0074] Based on the second structure tree, the first structure tree is processed and filtered to obtain multiple information regions. The first extraction table and the second extraction table are adjusted synchronously. Each information region has a corresponding category dimension and time dimension. This invention processes and filters the first structure tree according to the second structure tree to obtain multiple information regions, so that information useful for the analysis of the second structure tree within the first structure tree is filtered to obtain the corresponding information regions.

[0075] Furthermore, the first information structure is processed and filtered based on the second structure structure to obtain multiple information regions. The first extraction table and the second extraction table are adjusted synchronously. Each information region has a corresponding category dimension and time dimension, including:

[0076] The invention first processes the child nodes, specifically saving the child nodes in the first and second structure trees that correspond to the category dimensions in the second structure tree. This ensures that the information in the first and second structure trees corresponds in the category dimension during subsequent data processing, avoiding excessive invalid information and improving the robustness of subsequent data processing.

[0077] The operational specifications of the grandchild nodes corresponding to the child nodes of each category dimension are determined. Then, this invention determines the operational specifications of the grandchild nodes of the child nodes for each category dimension, avoiding inconsistencies in the time dimension between the first and second structure trees.

[0078] Based on the second structure tree, the grandchild nodes retained in the first structure tree are re-divided according to business specification information, and non-corresponding nodes are removed, resulting in an updated first structure tree, a first extraction table, and corresponding information areas. This invention re-divides the grandchild nodes retained in the first structure tree according to business specification information based on the second structure tree, and removes non-corresponding nodes, specifically deleting grandchild nodes in the second structure tree that do not correspond to time periods in the first structure tree, thereby obtaining an updated first structure tree, a first extraction table, and corresponding information areas. It should be noted that the updated first structure tree, first extraction table, and corresponding information areas are the foundational data used for subsequent predictive calculations of the second business data.

[0079] Through the above methods, this invention can quickly align enterprise data of different multi-level dimensions based on the structure tree, and can quickly screen the first structure tree based on the second structure tree during data filtering and elimination, which greatly improves efficiency. Under the premise of ensuring the structure of data storage, it greatly improves the efficiency of subsequent data filtering.

[0080] The server constructs information points corresponding to each information area and generates corresponding first business functions according to preset multi-level dimensions. A comprehensive function is obtained by fusing the first business functions based on the similarity between the first enterprise and the second enterprise. In this invention, the server constructs information points corresponding to each information area and generates corresponding first business functions according to preset multi-level dimensions. These first business functions can be viewed as functions generated according to the time dimension. The parameters within the first business function include at least time, monetary values, personnel values, and equipment quantity values, such as R&D expense functions and R&D personnel quantity functions. In other words, this invention obtains information points corresponding to each information area and generates corresponding first business functions according to preset multi-level dimensions. It should be noted that there are multiple first enterprises, and this invention calculates the similarity between the first enterprise and the second enterprise and fuses the first business functions to obtain a comprehensive function.

[0081] Furthermore, the server constructs information points corresponding to each information area and generates corresponding first business functions according to the corresponding preset multi-level dimensions. Based on the similarity between the first enterprise and the second enterprise, the first business functions are fused to obtain a comprehensive function, including:

[0082] The server uses the child nodes of the first tree structure as the initial information area, and retains the grandchild nodes within each initial information area based on preset or calculated business specification information to obtain information points. In this invention, the child nodes of the first tree structure are used as the initial information area; one child node corresponds to one initial information area. The invention retains the grandchild nodes within each initial information area based on preset or calculated business specification information to obtain information points, which then have a corresponding time period.

[0083] This invention performs functionalized statistics on all information points according to their time dimension to generate corresponding first operating functions, with one first operating function for each information area.

[0084] The similarity weights generated based on the similarity between the first and second enterprises are calculated and added to the corresponding first structure tree. This invention further calculates the similarity weights between the first and second enterprises. It should be noted that there are multiple first enterprises, and this invention will obtain different similarity scores between the first and second enterprises. This allows the invention to determine the first enterprise whose business attributes are closer to those of the second enterprise and assign it a higher similarity weight. This method enables the invention to maximize the determination of similarity weights based on the differences between enterprises while considering many different samples of first enterprises.

[0085] A comprehensive function is obtained by fusing the similarity weights and first operating functions within the first structure tree. This invention, after obtaining the similarity weights, fuses the similarity weights and first operating functions to obtain the comprehensive function. This comprehensive function then includes multiple first operating functions and the similarity weights corresponding to each operating function.

[0086] Furthermore, the calculation generates corresponding similarity weights based on the similarity between the first enterprise and the second enterprise, and a comprehensive function is obtained by fusing the similarity weights and the first operating function, including:

[0087] This invention extracts the first and second attribute labels of the first and second enterprises, and calculates the attribute difference for each dimension by performing a quantitative difference calculation on the first and second attribute labels. The first attribute label can be the industry, the year of establishment, etc. This invention quantifies the industry and the year of establishment according to a preset quantification form, such as quantifying the software industry to 1, the design industry to 2, the agriculture industry to 5, etc. This quantification form is preset. This invention then comprehensively calculates the attribute difference for each dimension by performing a quantitative difference calculation on the first and second attribute labels.

[0088] The original first and second tree structures are compared to obtain the difference values ​​of child nodes and grandchild nodes. This invention compares the first and second tree structures and calculates the difference values. The more difference values ​​of child nodes and grandchild nodes obtained from the comparison of the first and second tree structures, the greater the difference between the two enterprises in the category and time dimensions of their operations, and the greater the corresponding difference.

[0089] The similarity between the first and second companies is obtained based on attribute differences, child node differences, and grandchild node differences. The similarity scores of all companies between the first and second companies are then weighted as percentages to obtain similarity weights. These similarity weights are then used to structure and fuse the first operating function into a comprehensive function. This invention obtains the similarity between the first and second companies based on attribute differences, child node differences, and grandchild node differences, and calculates the similarity using the following formula.

[0090]

[0091] Where, x i Let k be the similarity between the second company and the i-th first company. p Let the label weights of the second company and the i-th first company under the p-th attribute label be denoted as follows. This represents the quantified value of the first company under the p-th attribute label. Let be the quantified value of the second enterprise under the p-th attribute label, j be the upper limit value of the attribute label, α be the weight value of the child node, and s be the quantified value of the second enterprise under the p-th attribute label. Chi s represents the number of differences between child nodes, β represents the weight of the grandchild node, and s represents the number of differences between child nodes. Gra This represents the number of differences between the grandchild nodes. The overall similarity is obtained by combining the differences between the enterprise's own attributes and the business structure tree in the following way.

[0092] Furthermore, the comprehensive function obtained by fusion of the first operating function structure tree based on similarity weights includes:

[0093] Each parent node of the first enterprise is treated as a child node and connected to the same parent node to form an analysis structure tree. The similarity weights are then mapped to the child nodes of the analysis structure tree. It should be noted that each first enterprise and the second enterprise will have a similarity score. The higher the similarity score, the greater the corresponding similarity weight. The similarity score and the similarity weight can be the same or corresponding. This invention combines the parent node of each first enterprise as a child node with the same parent node to form an analysis structure tree, thereby combining all the first enterprises into a single overall analysis structure tree for analyzing the second enterprises.

[0094] Determine the dimensional weights of each grandchild node in the analysis structure tree. It should be noted that different types of dimensions will have different weights. This invention adds different weights to different types of dimensions, so that different management functions within the resulting comprehensive function have different weights.

[0095] In one possible implementation, the present invention generates functions for corresponding dimensions using the first operating function of each first enterprise, and substitutes the corresponding operating data of the second enterprise into the first operating function to obtain a prediction of the corresponding dimensions of the second enterprise from the perspective of the first enterprise's operation. Furthermore, a comprehensive function is obtained by statistically analyzing all the first operating functions. The prediction target of each first operating function is different and may correspond to a specific type of dimension. The present invention summarizes the first operating functions of different types of dimensions through the comprehensive function and calculates them by combining the weights of different dimensions to obtain the evaluation value of the corresponding enterprise. This evaluation value is used to indicate the degree of goodness of the enterprise's operation, which includes excellent, good, average, poor, etc.

[0096] Based on the analysis of the tree structure, the revenue forecast of the second operating data is obtained, and the predicted data and credibility information are fed back to the analysis end. This invention obtains predicted data for the revenue forecast of the second operating data, which includes multiple dimensions such as revenue, R&D expenditure, total expenditure, personnel information, etc. This invention also calculates the similarity between all second enterprises and the first enterprise, calculates the mean similarity, and obtains corresponding credibility information based on the mean range of the mean similarity. Each credibility information has a preset credibility level.

[0097] Based on the comprehensive function, the server obtains the predicted revenue data and credibility information from the second operating data and feeds it back to the analysis end.

[0098] Ant colony optimization (ACO) is a heuristic optimization algorithm that simulates the foraging behavior of ants. Ants leave pheromones along their paths while foraging; paths with higher pheromone concentrations are more likely to be chosen by other ants. Through this positive feedback mechanism, the ant colony can find the shortest path from the nest to the food source. In the optimization problem, the ants' path selection corresponds to a search in the solution space, and the updating of pheromones simulates the reinforcement of excellent solutions.

[0099] The technical solution provided by this invention introduces an ant colony algorithm to form a genetic ant colony algorithm. Chromosome encoding: Solutions in the ant colony algorithm (e.g., in weight optimization problems, a set of weight values ​​can be considered a solution) are encoded using chromosomes. Real-number encoding can be used, with each gene position corresponding to a weight value of a machine learning algorithm. Selection operation: Borrowing selection mechanisms from genetic algorithms, such as roulette wheel selection or tournament selection, individuals with higher fitness are selected for the next generation based on the fitness of each solution (individual ant) ​​(measured by metrics such as classification accuracy after combining multiple machine learning algorithms), increasing the proportion of excellent solutions in subsequent searches. Crossover operation: Crossover is performed on the selected individuals. For example, single-point crossover or multi-point crossover. During crossover, some genes (weight values) of two parent individuals are exchanged, thereby generating new offspring individuals and expanding the search range of the solution space. Mutation operation: The genes of individuals are mutated with a certain mutation probability. In weight optimization problems, some weight values ​​are randomly adjusted in small increments to avoid the algorithm getting trapped in local optima. Pheromone update: Combining the pheromone update rules of the ant colony algorithm, the pheromone is updated according to the quality of the solution after each iteration. The path traversed by a solution with high fitness (corresponding to a weight combination) has a higher pheromone concentration, guiding the search direction of subsequent ants.

[0100] The genetic ant colony algorithm is used to optimize the weights of multiple machine learning algorithms. Initialization: A certain number of ants (solutions) are randomly generated, each ant representing a weight combination of multiple machine learning algorithms. Parameters of the genetic ant colony algorithm are set, such as the number of ants, pheromone evaporation coefficient, crossover probability, and mutation probability. Fitness calculation: Each weight combination is applied to the combination of multiple machine learning algorithms. Training and prediction are performed using a training dataset. The fitness value of each ant is calculated using metrics such as classification accuracy and F1-score as fitness functions. Genetic operations include: Selection: Individuals are selected from the current population according to a selection strategy. Crossover: The selected individuals are crossovered to generate new offspring. Mutation: The offspring are mutated to obtain mutated individuals. Pheromone update: The pheromone concentration of the corresponding weight combination in the solution space is updated based on the ant's fitness. Higher fitness results in a greater increase in pheromone concentration. Iterative optimization: The steps of fitness calculation, genetic operations, and pheromone update are repeated until stopping conditions are met, such as reaching the maximum number of iterations or fitness convergence. Determining the optimal weights: At the end of the algorithm, the weight combination represented by the ant with the highest fitness is selected as the optimal weights for the multi-machine learning algorithm and applied to the actual classification task in order to achieve the best classification effect.

[0101] The present invention also provides an enterprise data processing method applying the first aspect of the embodiments of the present invention, characterized in that it includes:

[0102] Step S1: During the enterprise revenue trend prediction calculation, initialize the ant colony-related parameters and the genetic algorithm-related parameters, including ant colony size, pheromone update rate, genetic algorithm population size, crossover and mutation probabilities, etc.

[0103] Pheromones are initialized in the solution space (i.e., the space of all possible weight combinations), usually by setting the pheromone concentration at all locations to a small constant tau_0.

[0104] Step S2: Use the ant colony algorithm to generate a set of initial solutions and calculate the fitness of each solution.

[0105] Step S3: Perform crossover and mutation operations on the population to generate the next generation of individuals.

[0106] Step S4: Calculate the fitness of the next generation of individuals.

[0107] Step S5: If the specified number of iterations is reached or a preset satisfactory solution is found, output the result; otherwise, return to step S3 and continue optimization.

[0108] The technical solution provided by this invention involves ant colony generation: a certain number (let's say m) of ants are randomly generated, with each ant corresponding to a weight combination of multiple machine learning algorithms. Assuming there are n machine learning algorithms used for combined classification, the solution vector for each ant can be represented as \boldsymbol{w}_i=[w_{i 1},w_{i2},\cdots,w_{in}], where w_{ij} represents the weight assigned to the j-th machine learning algorithm by the i-th ant, and satisfies \sum_{j=1}^{n}w_{ij}=1 (ensuring the rationality of the weights).

[0109] Parameter settings: Set the key parameters of the genetic ant colony algorithm, such as the pheromone evaporation coefficient ρ_ho (usually between 0 and 1, used to simulate the natural evaporation of pheromones over time), crossover probability P_c (usually between 0.6 and 0.9, controlling the probability of crossover operation), mutation probability P_m (usually a small value, such as 0.01-0.1, determining the probability of mutation operation), and the maximum number of iterations T_{max}, etc.

[0110] Pheromones are initialized in the solution space (i.e., the space of all possible weight combinations), usually by setting the pheromone concentration at all locations to a small constant tau_0.

[0111] Algorithm Combination and Training: For each ant representing a weight combination \boldsymbol{w}_i, this combination is applied to a combination of n machine learning algorithms. Specifically, the prediction results of each algorithm are weighted and fused according to their respective weights to obtain the final prediction result. Then, this combined model is trained using the training dataset.

[0112] Fitness evaluation: The trained ensemble model is used to predict on the validation dataset, and the fitness value of each ant is calculated using appropriate metrics such as classification accuracy, F1 score, recall, and precision as the fitness function f(\boldsymbol{w}_i). The higher the fitness value, the better the classification performance of the multi-machine learning algorithm under that weight combination.

[0113] Selection: Individuals are selected from the current population to enter the next generation using methods such as roulette wheel selection and tournament selection. Taking roulette wheel selection as an example, the probability of each ant being selected is directly proportional to its fitness, i.e., P(\boldsymbol{w}_i)=\frac{f(\boldsymbol{w}_i)}{\sum_{j=1}^{m}f(\boldsymbol{w}_j)}, and the higher the fitness of the ant, the greater the probability of it being selected.

[0114] Crossover: Perform a crossover operation on the selected individuals (set as parent individuals). Taking single-point crossover as an example, randomly select a crossover point and exchange the genes (weight values) of the two parent individuals after the crossover point to generate two offspring individuals. For example, if the parent individuals \boldsymbol{w}_1=[w_{11},w_{12},w_{13},w_{14}] and \boldsymbol{w}_2=[w_{21},w_{22},w_{23},w_{24}], and the crossover point is 2, then the offspring individuals \boldsymbol{w}_{1}'=[w_{11},w_{12},w_{23},w_{24}] and \boldsymbol{w}_{2}'=[w_{21},w_{22},w_{13},w_{14}]. The purpose of crossover operations is to generate new weight combinations through gene recombination, thereby expanding the search space.

[0115] Mutation: Mutation is performed on the offspring individuals after crossover with a mutation probability P_m. For each gene locus, mutation is performed with a probability of P_m, which randomly changes the weight value corresponding to that gene locus (within a reasonable range, such as [0,1]). Mutation can increase the diversity of the population and prevent the algorithm from getting trapped in local optima too early.

[0116] Evaporation: Based on the pheromone evaporation coefficient ρrho, the pheromone concentration at all positions in the solution space is updated, i.e., ρ_{ij}(t+1)=(1-ρrho)ρ_{ij}(t), where ρ_{ij}(t) represents the pheromone concentration at position ij (corresponding to a certain weight combination) at the t-th iteration.

[0117] Deposition: Based on the ant's fitness, ants with higher fitness deposit more pheromones on the paths they traverse (corresponding to weighted combinations). Let the fitness of ant k be f(\boldsymbol{w}_k), then the amount of pheromones it deposits on path ij is \Delta\tau_{ij}^k=\frac{f(\boldsymbol{w}_k)}{\sum_{l=1}^{m}f(\boldsymbol{w}_l)}, and the updated pheromone concentration is \tau_{ij}(t+1)=(1-\rho)\tau_{ij}(t)+\sum_{k=1}^{m}\Delta\tau_{ij}^k.

[0118] Iterative optimization: Repeat the above steps of fitness calculation, genetic operation and pheromone update until the stopping condition is met, such as reaching the maximum number of iterations T_{max} or the fitness value converges (the change in fitness in adjacent iterations is less than a certain threshold).

[0119] Determining the optimal weights: At the end of the algorithm, the weight combination represented by the ant with the highest fitness is selected from the entire population as the optimal weights for the multi-machine learning algorithm, which is then applied to the actual classification task in the hope of achieving the best classification effect.

[0120] In genetic algorithms, individuals in the population continuously undergo genetic operations, causing them to evolve towards the global optimum. They have a large search space coverage and good parallel search capabilities. Ant colony algorithms, due to the continuous accumulation of pheromones, possess positive feedback characteristics, which promotes faster convergence towards the optimum. They can perform multi-point searches in the solution space, and their robustness and optimization capabilities are stronger than other algorithms.

[0121] Combining the advantages of the two algorithms, a strategy of fusion and improvement is proposed. The ant colony algorithm is used to initialize the population of the genetic algorithm to solve the problem of low initial population quality. The fusion and improvement algorithm is successfully applied to the data analysis and prediction model of financial enterprises.

[0122] To implement the enterprise data processing method based on a tree structure provided by this invention, this invention also provides an enterprise data processing system based on a tree structure, such as... Figure 2 As shown, it includes:

[0123] The acquisition module is used to enable the server to acquire the first business data of the first enterprise, and to classify the first business data according to preset multi-level dimensions to obtain the first structure tree and the first extraction table.

[0124] The extraction module is used to enable the server to receive and classify the second business data of the second enterprise added by the analysis terminal to obtain a second structure tree and a second extraction table, extract the business specification information of the second structure tree, and perform partitioning processing on the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas;

[0125] The generation module enables the server to construct information points corresponding to each information area and generate corresponding first business functions according to the corresponding preset multi-level dimensions. The first business functions are then fused based on the similarity between the first enterprise and the second enterprise to obtain a comprehensive function.

[0126] The prediction module is used to enable the server to obtain predicted data and credibility information based on the comprehensive function for the second operating data revenue prediction and feed it back to the analysis end.

[0127] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, is used to implement the methods provided in the various embodiments described above.

[0128] The storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, the storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be a component of the processor. The processor and storage medium can reside in an Application Specific Integrated Circuit (ASIC). This ASIC can also be located within a user device. Alternatively, the processor and storage medium can exist as discrete components in a communication device. Storage media can be read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0129] The present invention also provides a program product including execution instructions stored in a storage medium. At least one processor of the device can read the execution instructions from the storage medium, and the execution instructions by the at least one processor cause the device to implement the methods provided in the various embodiments described above.

[0130] In the above-described terminal or server embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing enterprise data based on a structure tree, characterized in that, include: The server obtains the first business data of the first enterprise, and classifies the first business data according to preset multi-level dimensions to obtain the first structure tree and the first extraction table; The server receives the second business data of the second enterprise added by the analysis terminal, classifies it to obtain a second structure tree and a second extraction table, extracts the business specification information of the second structure tree, and partitions the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas; The information of each dimension in the second operating data is decomposed according to the preset analysis points to obtain the operating specification information corresponding to each dimension. The operating specification information includes at least a time period. A second structure tree and a second extraction table are generated based on the dimensions and time periods to categorize the second business data. Based on the second structure tree, the first structure tree is processed and filtered to obtain multiple information areas. The first extraction table and the second extraction table are adjusted synchronously. Each information area has a corresponding category dimension and time dimension. The child nodes of the corresponding category dimension in the first tree structure are retained; Determine the operational specifications of the grandchild nodes corresponding to the child nodes of each category dimension; Based on the second structure tree, the grandchild nodes retained in the first structure tree are re-divided according to the business specification information, and non-corresponding nodes are removed to obtain the updated first structure tree, the first extraction table, and the corresponding information area. The server constructs information points corresponding to each information area and generates corresponding first business functions according to the preset multi-level dimensions. The first business functions are then fused based on the similarity between the first enterprise and the second enterprise to obtain a comprehensive function. Based on the comprehensive function, the server obtains the predicted revenue data and credibility information from the second operating data and feeds it back to the analysis end.

2. The enterprise data processing method based on a tree structure according to claim 1, characterized in that, The server acquires the first operating data of the first enterprise, and classifies the first operating data according to preset multi-level dimensions to obtain a first structure tree and a first extraction table, including: The server filters companies in the database at preset time intervals, and adds a first label to the corresponding companies after determining that the first condition is met. The server obtains the first operating data of the first enterprise, which includes at least financial information, personnel information, and shareholder information. The first operating data is decomposed according to different categories and dimensions based on the first time series interval, resulting in a first structure tree and a first extraction table with multiple dimensions of time and category. The information in the first structure tree and the first extraction table is set accordingly.

3. The enterprise data processing method based on a structure tree according to claim 1, characterized in that, The step of decomposing the information of each dimension in the second operating data according to preset analysis points to obtain the corresponding operating specification information for each dimension includes: Obtain the maximum span time value of information for each dimension in the second operating data. If the maximum span time value is greater than or equal to a preset time value, then obtain the preset operating specification information. If the maximum span time value is less than the preset time value, the maximum span time value is divided equally based on the preset number of analysis points to obtain the calculated business specification information.

4. The enterprise data processing method based on a structure tree according to claim 1, characterized in that, The server constructs information points corresponding to each information area and generates corresponding first business functions according to preset multi-level dimensions. Based on the similarity between the first enterprise and the second enterprise, the first business functions are fused to obtain a comprehensive function, including: The server uses the child nodes of the first tree structure as the initial information area, and retains the grandchild nodes in each initial information area based on the preset or calculated business specification information to obtain information points. Based on the time dimension of the information points, all information points are statistically analyzed using functions to generate corresponding first business functions, with each information area corresponding to one first business function; Calculate the corresponding similarity weights based on the similarity between the first company and the second company, and add them to the corresponding first structure tree. A comprehensive function is obtained by fusing the similarity weights within the first structure tree and the first operating function.

5. The enterprise data processing method based on a tree structure according to claim 4, characterized in that, The calculation generates corresponding similarity weights based on the similarity between the first enterprise and the second enterprise. A comprehensive function is obtained by fusing these similarity weights and the first business function, including: Extract the first and second attribute labels of the first and second enterprises, and calculate the attribute difference for each dimension by quantifying the first and second attribute labels. Compare the original first and second structure trees to obtain the difference values ​​of child nodes and grandchild nodes; The similarity between the first enterprise and the second enterprise is obtained based on the attribute difference, child node difference, and grandchild node difference. The similarity between the first enterprise and the second enterprise is then processed as a percentage to obtain the similarity weight. Based on the similarity weight, the first operating function is structured and fused to obtain the comprehensive function.

6. The enterprise data processing method based on a tree structure according to claim 5, characterized in that, The comprehensive function obtained by fusion of the first operating function structure tree based on similarity weights includes: Each first enterprise's parent node is treated as a child node and then connected to the same parent node to obtain an analysis structure tree. The similarity weights are then mapped to the child nodes of the analysis structure tree. Determine the dimensional weights of each grandchild node in the analysis tree structure; Based on the analysis of the tree structure, the revenue forecast of the second operating data is obtained and the forecast data and credibility information are fed back to the analysis end.

7. A method for processing enterprise data according to any one of claims 1 to 6, characterized in that, include: Step S1: During the enterprise revenue trend prediction calculation, initialize the ant colony-related parameters and the genetic algorithm-related parameters, including ant colony size, pheromone update rate, genetic algorithm population size, crossover and mutation probabilities; Step S2: Use the ant colony algorithm to generate a set of initial solutions and calculate the fitness of each solution; A genetic algorithm is used to select a subset of individuals with high fitness from the initial solution to form a population. Step S3: Perform crossover and mutation operations on the population to generate the next generation of individuals; Step S4: Calculate the fitness of the next generation of individuals; Step S5: If the specified number of iterations is reached or a preset satisfactory solution is found, output the result; Otherwise, return to step S3 and continue optimization.

8. A tree-structure-based enterprise data processing system, characterized in that, include: The acquisition module is used to enable the server to acquire the first business data of the first enterprise, and to classify the first business data according to preset multi-level dimensions to obtain the first structure tree and the first extraction table. The extraction module is used to enable the server to receive and classify the second business data of the second enterprise added by the analysis terminal to obtain a second structure tree and a second extraction table, extract the business specification information of the second structure tree, and perform partitioning processing on the first structure tree and the first extraction table based on the business specification information to obtain multiple information areas; The information of each dimension in the second operating data is decomposed according to the preset analysis points to obtain the operating specification information corresponding to each dimension. The operating specification information includes at least a time period. A second structure tree and a second extraction table are generated based on the dimensions and time periods to categorize the second business data. Based on the second structure tree, the first structure tree is processed and filtered to obtain multiple information areas. The first extraction table and the second extraction table are adjusted synchronously. Each information area has a corresponding category dimension and time dimension. The child nodes of the corresponding category dimension in the first tree structure are retained; Determine the operational specifications of the grandchild nodes corresponding to the child nodes of each category dimension; Based on the second structure tree, the grandchild nodes retained in the first structure tree are re-divided according to the business specification information, and non-corresponding nodes are removed to obtain the updated first structure tree, the first extraction table, and the corresponding information area. The generation module enables the server to construct information points corresponding to each information area and generate corresponding first business functions according to the corresponding preset multi-level dimensions. The first business functions are then fused based on the similarity between the first enterprise and the second enterprise to obtain a comprehensive function. The prediction module is used to enable the server to obtain predicted data and credibility information based on the comprehensive function for the second operating data revenue prediction and feed it back to the analysis end.