A regional industry development analysis and evaluation system

By constructing a multivariate mapping matrix and hierarchical analysis, the problems of one-sidedness and insufficient dynamic analysis in the assessment of regional industrial development in existing technologies are solved. This enables efficient data integration and refined assessment, improves early warning response speed and asset management efficiency, and supports decision-making and planning for governments and enterprises.

CN121032272BActive Publication Date: 2026-05-05LINGXI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LINGXI TECH CO LTD
Filing Date
2025-08-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies tend to overlook inter-industry correlations in regional industrial development assessments, resulting in insufficient dynamic analysis in spatial quantification and regional planning scenarios, making it difficult to identify early warning information. Furthermore, existing methods often rely on single indicators or data sources, leading to bias and limitations.

Method used

A regional industrial development analysis and evaluation system is adopted, including a data integration module, a sequence analysis module, a hierarchical analysis module, and a control verification module. By integrating industry index data with regional reference characteristics to construct a multivariate mapping matrix, time series analysis and hierarchical analysis are performed to generate a regional development model and output early warning information.

Benefits of technology

It has achieved the organic integration of different types of data, improved the efficiency and accuracy of data utilization, realized refined and hierarchical regional industry development assessment, improved early warning response speed and asset management efficiency, and provided strong support for government decision-making and corporate strategic planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032272B_ABST
    Figure CN121032272B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology, specifically a regional industrial development analysis and evaluation system. The system includes: determining industry index data and regional reference characteristics within a region to be identified; summarizing the industry index data and regional reference characteristics into a multivariate mapping matrix; applying time series analysis to the multivariate mapping matrix to extract the regional time distribution trend of attribute description indicators and determine the distribution proportion of each attribute description indicator under the regional time distribution trend; based on the distribution proportion of each attribute description indicator, sequentially obtaining hierarchical keywords at multiple levels, and assigning values ​​to regional hierarchical indicators according to each level of keywords; constructing a regional development model; solving for industry evaluation indicators based on the derivation mapping relationship between the regional development model and the multivariate mapping matrix; assessing the asset status of each industry according to the industry evaluation indicators; and outputting early warning information according to the asset assessment status of each region. This improves the accuracy of regional planning and the reliability of early warning prompts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically a regional industrial development analysis and evaluation system. Background Technology

[0002] Traditional assessment methods often rely on single indicators or data sources, such as GDP growth rate or industrial added value. While these indicators can reflect the development status of a region's industries to some extent, they are often one-sided and limited. Furthermore, with the rapid development of technologies such as big data and artificial intelligence, the types and sources of data are becoming increasingly diverse, making the effective integration and utilization of this data a pressing issue.

[0003] For example, Chinese Patent Publication No. CN117634952A discloses an industrial development assessment method, system, and electronic device. The method includes: collecting data on industrial chain enterprises and supply chain enterprises of the target industry within a target region; determining the supply and demand pairs between industrial chain enterprises and supply chain enterprises, as well as the branch relationships between headquarters enterprises and branch enterprises; constructing an industrial chain and supply chain coupling network, wherein the coupling network includes nodes and edges, with nodes representing individual enterprises and edges representing supply and demand pairs and branch relationships; establishing an assessment index system for the coupling network; calculating and obtaining multiple assessment indicators based on the coupling network; and assessing the development of the target industry within the target region based on the multiple assessment indicators.

[0004] For example, Chinese Patent Publication No. CN118917731A discloses a method, system, equipment, and medium for regional industrial development assessment. This invention includes: S1, obtaining statistical data on regional industries and enterprises within the same administrative region; S2, performing index assessment on regional industries based on the enterprise information statistics; and S3, collecting the index assessments, weighting and fusing them, and outputting the regional industrial development assessment result. This invention obtains the assessment result of regional industrial development by statistically analyzing the information on regional industries and enterprises within the same administrative region, performing index evaluation and weighted fusion. The current status of industrial development in the target region can be determined based on the regional industrial development assessment result.

[0005] Existing technologies describe how to assess industries under the consideration of supply chain structure and explain some assessment indices used. However, when assessing industries within a region, these methods tend to only consider the performance of some indices or industry characteristics, ignoring the correlation between these industries. In addition, there are significant differences in the industry and regional characteristics targeted by different regions, resulting in insufficient dynamic analysis in spatial quantification and regional planning scenarios, making it difficult to identify early warning information in regional planning. Summary of the Invention

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a regional industrial development analysis and evaluation system, including: a data integration module, used to determine the industry index data and regional reference characteristics in the region to be identified, and to summarize the industry index data and regional reference characteristics into a multivariate mapping matrix.

[0007] The sequence analysis module is used to apply time series analysis to the multivariate mapping matrix, extract the attribute description indicators corresponding to each data from the perspective of data regional distribution, as well as the regional time distribution trend of the attribute description indicators, and determine the distribution ratio of each attribute description indicator under the regional time distribution trend.

[0008] The hierarchical analysis module is used to perform hierarchical analysis on the multivariate mapping matrix based on the distribution ratio of each attribute description indicator, sequentially obtain hierarchical keywords of multiple levels, and assign regional hierarchical indicators according to the keywords of each level.

[0009] The model building module is used to construct a regional development model according to the assigned values ​​of regional level indicators, and to solve the industry evaluation indicators based on the derived mapping relationship between the regional development model and the multivariate mapping matrix.

[0010] The control and verification module is used to assess the asset status of various industries based on industry evaluation indicators, and output early warning information according to the asset assessment status of each region, and retrieve data from each region to the control and management platform from the early warning information.

[0011] The beneficial effects of this invention are as follows: First, by integrating industry index data and regional reference features within the region to be identified, this invention constructs a multivariate mapping matrix, enabling the organic combination of existing industry indices and corresponding regional reference features within the current region. Furthermore, it extracts attribute description indicators and their regional time distribution trends from the perspective of data regional distribution. The regional time distribution trends are constructed using each attribute description indicator as the core and regional reference features as an auxiliary, achieving the organic fusion of different types of data and improving the efficiency and accuracy of data utilization. Subsequently, identifying the corresponding distribution proportions of the combined data allows understanding the data proportions of different attribute description indicators in the time series, revealing the dynamic patterns and potential trends behind the data, and providing strong support for subsequent hierarchical analysis and the construction of regional development models.

[0012] Second, this invention employs hierarchical analysis to sequentially obtain hierarchical keywords at multiple levels, and assigns regional hierarchical indicators based on these keywords, thereby achieving a refined and hierarchical evaluation of the regional industry development index. For each level, adjustments are made according to its location, dividing the corresponding multiple indicators into top, middle, and bottom levels. The relative distribution within each level is identified based on its differences, semantics, average distance, and weights. Weights are used to describe the main objectives and analytical values ​​for the current regional analysis, enabling more layered analysis and completing a refined and hierarchical evaluation of the regional industry development index.

[0013] Third, this invention sets the data after hierarchical analysis as a regional development model, and then performs mapping analysis on the regional development model and the multivariate mapping matrix to obtain the distribution of endogenous latent variables and exogenous manifest variables in the current processing. The correlation between the current hierarchical analysis can be further evaluated to improve the accuracy of the numerical values ​​calculated by hierarchical analysis, thereby improving the accuracy and reliability of the regional industry development index assessment after multivariate data combination and hierarchical analysis, and providing strong data support for government decision-making, corporate strategic planning, etc.

[0014] Fourth, this invention performs multivariate mapping on the data of various industry indicators after time series and hierarchical analysis, and uses the industry indicators as the basis for judging each region to identify whether there are abnormal situations under the development of each region. Based on the obtained abnormal situations, early warning information is generated. After the early warning information is retrieved to the management and control platform, the decision-making and planning for the development of each region are adjusted according to the early warning information, thereby improving the asset management efficiency and early warning response speed under the development of each region. Attached Figure Description

[0015] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0016] Figure 1 This is a schematic diagram of a regional industrial development analysis and evaluation system.

[0017] Figure 2 This is a flowchart of the sequence analysis module of a regional industrial development analysis and evaluation system.

[0018] Figure 3 This is a flowchart illustrating the regional characteristic model within the sequence analysis module of a regional industrial development analysis and evaluation system.

[0019] Figure 4 This is a flowchart illustrating the distribution proportion of various attribute description indicators in the sequence analysis module of a regional industrial development analysis and evaluation system.

[0020] Figure 5This is a flowchart illustrating the hierarchical analysis module of a regional industrial development analysis and evaluation system.

[0021] Figure 6 This is a flowchart illustrating the model building module of a regional industrial development analysis and evaluation system. Detailed Implementation

[0022] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in the art or in accordance with the product manual.

[0023] See Figure 1 A regional industrial development analysis and evaluation system includes: a data integration module, a sequence analysis module, a hierarchical analysis module, a model building module, and a control and verification module; wherein, the output end of the data integration module is connected to the sequence analysis module, the output end of the sequence analysis module is connected to the hierarchical analysis module, the output end of the hierarchical analysis module is connected to the model building module, and the output end of the model building module is connected to the control and verification module.

[0024] The data integration module is used to determine the industry index data and regional reference features within the region to be identified, and to summarize the industry index data and regional reference features into a multivariate mapping matrix.

[0025] The sequence analysis module is used to apply time series analysis to the multivariate mapping matrix, extract the attribute description indicators corresponding to each data from the perspective of data regional distribution, as well as the regional time distribution trend of the attribute description indicators, and determine the distribution ratio of each attribute description indicator under the regional time distribution trend.

[0026] The hierarchical analysis module is used to perform hierarchical analysis on the multivariate mapping matrix based on the distribution ratio of each attribute description indicator, sequentially obtain hierarchical keywords of multiple levels, and assign regional hierarchical indicators according to the keywords of each level.

[0027] The model building module is used to construct a regional development model according to the assigned values ​​of regional level indicators, and to solve the industry evaluation indicators based on the derived mapping relationship between the regional development model and the multivariate mapping matrix.

[0028] The control and verification module is used to assess the asset status of various industries based on industry evaluation indicators, and output early warning information according to the asset assessment status of each region, and retrieve data from each region to the control and management platform from the early warning information.

[0029] The aforementioned multivariate mapping matrix combines regional reference features and industry indicator data to perform evaluation processing on the region that needs to be evaluated.

[0030] Industry index data includes not only the information represented by the aforementioned industry indices, such as output, sales, growth rate, profit, composition, and market share—data used in the calculation of industry indices using existing technologies, recorded monthly, quarterly, and annually—but also information on employment growth, including the number of new employees and the growth rate, as well as market, policy, and technological indicators corresponding to the industry. Regional reference features represent the geographical, economic, social, and industry characteristics of the region to be identified. For example, geographical features include the location of the target region, its topography, and natural resources, influencing the regional industry development. Economic features consider the region's total economic output, GDP per capita, industrial structure, population size, and age structure. Social features describe the region's support policies, tax policies, environmental regulations, and local culture. Industry features include the number of companies, types of companies, output value, and R&D investment within the industry chain. These mentioned geographical, economic, and social characteristics, to a certain extent, constrain and promote the distribution and development of industry characteristics. Using these four characteristics together as regional reference features is to identify the obstacles existing in the current region and the relative distribution of these characteristics among the identified multiple industry development indicators, in order to predict their development patterns. After multiple assessments, the comprehensive trend of different industries in the corresponding region can be obtained.

[0031] The multivariate mapping matrix establishes a mapping relationship between industry index data and regional reference features to combine various data from the region to be identified. This matrix can be viewed as a multi-dimensional dataset, where each row represents a data combination at a specific point in time or region, and the columns correspond to different regional reference features and industry indicators.

[0032] The process of determining the industry index data and regional reference features within the region to be identified also includes: acquiring text and image data within the region to be identified, extracting the industry index data and regional reference features of the region in sequence, matching the industry index data according to the geographical features, economic features, social features and industry features in the regional reference features, and concatenating the matched data to obtain a multivariate mapping matrix.

[0033] The matching method at this time is to set corresponding labels for the data collected from the industry index data according to geographical features, economic features, social features and industry features, and combine the relevant data into a category to determine how many industries exist in the current region, and what other features the industry has.

[0034] Preferably, in order to ensure the consistency of industry index data and regional reference features within the area to be identified, the implementation can also be expressed as follows: based on the data within the area to be identified, multiple industry index data and regional reference data are feature-matched to determine whether each industry index data and regional reference data belong to the same regional entity. This regional entity represents the data source corresponding to the industry index data and the regional reference data, and whether the region corresponding to the data source points to the same region under multivariate mapping, such as administrative division code, geographical name, etc., to ensure that they are consistent in the data. Then, it is verified whether the geographical coordinates corresponding to each region are consistent, thereby obtaining the industry index data and regional reference features of the area to be identified.

[0035] In one embodiment of the present invention, the attribute description index of the sequence analysis module is an index that describes industry indicator data and regional reference characteristics, such as competitiveness index, growth rate index, and percentage of profitable enterprises. The attribute description index can also be a keyword corresponding to the combination of industry indicator data and regional reference characteristics to summarize the indicators in the current industry. At the same time, the value of the attribute description index is recorded. The data in the multivariate mapping matrix is ​​divided into multiple intervals. According to the proportion of each interval to the total data, the distribution proportion of each attribute description index is obtained. The distribution proportion will further describe the distribution of different types of enterprises. The current distribution of enterprises and their industry indicators are combined to obtain a relative time trend.

[0036] When applying time series analysis to a multivariate mapping matrix, the mean, variance, trend slope, and periodic components of the attribute description indicators can be extracted based on their different distributions over time in the region to be identified. These parameters can then be used to describe the development of industry-related indicators within the region.

[0037] like Figure 2 As shown, the implementation of the sequence analysis module includes: extracting multiple attribute description indicators from the multivariate mapping matrix according to the historical dataset of the regional industry; generating regional division intervals corresponding to each attribute description indicator according to the industry economic growth rate, market share and R&D investment; the regional division intervals here are mainly set according to the different industry economic growth rate, market share and R&D investment, and multiple interval values ​​are set with each attribute description indicator.

[0038] Based on regional division intervals, a regional feature model is constructed with each attribute description indicator as the core and regional reference features as auxiliary. Clustering is performed using the values ​​of industry economic growth rate, market share, and R&D investment. The clustered data is used as the regional feature model to determine the relationship between each attribute description indicator and regional reference features. The regional feature model mainly describes the similarity of data within the regional division interval in terms of industry indicators and regional features, in order to discover industry groups with the same development trend, and whether the overall industry trend is different from the development trend of some industries.

[0039] The regional feature model is analyzed using time series data to generate a regional time distribution trend. The distribution of each attribute descriptor in the regional time distribution trend is output, yielding the distribution proportion of each attribute descriptor. When analyzing the regional time distribution trend, the focus is on identifying the central tendency and dispersion, described using the mean, slope, and other parameters of each attribute descriptor to determine the development status of the corresponding industry within the current region to be identified. For example, when analyzing the central tendency of the regional time distribution trend, the mean, median, and mode are typically used, while the standard deviation, variance, and interquartile range are used to display the distribution of each attribute descriptor.

[0040] When outputting the distribution proportion of each attribute description indicator, the obtained multiple attribute description indicators are compared with the total data to determine the distribution proportion of the attribute description indicators in the region to be identified. In some cases, the distribution proportion can also be represented by the completeness of the production chain and the proportion of the number of large enterprises in this region. That is, for the feature information corresponding to the regional reference features, after mapping this part of the feature information to the attribute description indicators, the proportion value of this part of the data is found to represent the distribution proportion of each attribute description indicator; that is, it expresses the proportion of the corresponding data of each attribute description indicator and other regional reference features in the total data when they co-occur.

[0041] To determine whether the interval feature model meets the needs of subsequent clustering and recognition, such as Figure 3 As shown, the implementation of the regional feature model also includes: initializing each regional division interval, and determining that the median value corresponding to the upper and lower limits of each regional division interval is the same as the average value of the regional division interval.

[0042] Based on the attribute description indicators corresponding to the regional division intervals, the data within the regional division intervals are traversed. If the data in adjacent regional division intervals overlap, the overlapping data is taken as the new regional division interval; this process continues until all regional division intervals no longer overlap.

[0043] Cluster analysis is performed on the traversed regional division intervals. Initial clustering is conducted using various attribute description indicators, followed by secondary clustering using regional reference features. The secondary clustering data is then used to construct a regional feature model. This secondary clustering is used to understand the relationships between these features and to perform subsequent correlation processing based on these relationships. For example, the correlation between features and the correlation between features and time are considered, and these two relationships are analyzed side-by-side to identify common and dissimilar trends.

[0044] The implementation method of constructing a regional feature model from the secondary clustering data during clustering also includes: obtaining multiple clusters based on the Euclidean distance between corresponding data in the regional division intervals; when the attribute description indicators on multiple clusters all reach their maximum values, the corresponding cluster is used as the data after the initial clustering; then, using the regional reference feature as the core, the data after the initial clustering is mapped to the regional reference feature, and multiple clusters are generated according to the mapping relationship; the mapped multiple clusters are superimposed, and the superimposed data is used as the data after the secondary clustering, thus completing the construction of the regional feature model. Then, these clustered data are combined with time series data to determine existing trend changes. Based on this trend, targeted development strategies can be formulated for different regions; for example, for regions with high economic growth rates, they can be encouraged to continue to increase R&D investment and expand market share; for regions with low economic growth rates, the reasons can be analyzed, and corresponding improvement measures can be proposed.

[0045] Then, regarding the generated regional time distribution trend, such as Figure 4As shown, the implementation of the distribution ratio of each attribute description indicator also includes: identifying the regional time distribution trend, and sequentially obtaining the degree of central tendency and dispersion in the regional time distribution trend. At this time, the degree of central tendency mainly represents the mean, median and mode in the regional time distribution trend, while the degree of dispersion mainly represents the standard deviation, variance and interquartile range in the regional time distribution trend. The degree of dispersion mainly quantifies the deviation between the data value and the mean, and uses the interquartile range to describe the relative difference in the middle 50% of the data range when there is deviation, in order to detect whether there are outliers. Suppose the analysis is on the growth of the number of enterprises in a certain region over the past ten years. If the average growth is 5% per year, the median is also 5%, and the mode is close to these values, it indicates that the annual growth rate of the region is relatively stable, and there are no obvious extreme values ​​affecting the overall trend. If the standard deviation is small, such as 1%, it means that the enterprise growth rate in most years fluctuates closely around the average, indicating that the enterprise growth in the region is relatively consistent and stable. Conversely, a large standard deviation, such as 5%, indicates significant differences in annual growth rates, potentially influenced by external factors like economic cycles or policy changes. For example, in the case of growth in the number of firms, a large interquartile range, such as an IQR of 10 percentage points per year, suggests that some years have very high growth rates while others are relatively low. This may reflect the instability of the region's economic development, influenced by external factors such as policy changes and market fluctuations. A large interquartile range may also indicate that different parts of the region are at different stages of development. For instance, some technology parks in certain areas may be experiencing rapid expansion, while other traditional industrial zones may be facing decline. This uneven development leads to increased dispersion in the overall data.

[0046] While the interquartile range primarily describes the middle 50% of the data range, a large interquartile range indirectly suggests the potential for more extreme or outlier values. These outliers may be caused by special events, such as the launch or closure of major projects, or natural disasters. A smaller interquartile range indicates that the middle 50% of the data points are concentrated within a narrower range, demonstrating higher stability and consistency. This typically means that if the interquartile range of the growth rate of the number of enterprises in a region is small, for example, only 2 percentage points, it implies that the growth rate in most years is very close to the average. This suggests that the region's economic development is relatively stable, without drastic fluctuations, policy implementation is effective, and the market environment is relatively stable. A small interquartile range may also reflect a more balanced development across different parts of the region. Whether it's emerging high-tech industries or traditional manufacturing, they maintain similar growth rates and development paces, demonstrating a good trend of coordinated regional development. From a risk management perspective, a smaller interquartile range implies a lower level of risk. Because the economic growth rate doesn't fluctuate much, investors and policymakers can more easily predict future trends and make more accurate plans accordingly.

[0047] Based on the degree of central tendency and dispersion in the regional time distribution trend, trend features are extracted. The trend features are then categorized by value type, and the trend distribution deviation and trend reference deviation are calculated for each value type. The trend features are then fitted according to the slope values ​​of the trend distribution deviation and trend reference deviation. The data proportions corresponding to the fitted trend features are used as the distribution proportions of the output attribute description indicators.

[0048] Trend characteristics here represent the degree of central tendency and dispersion based on the values ​​of mean, median, mode, standard deviation, variance, and interquartile range. Statistical analysis is performed according to these values ​​to identify multiple trend characteristics. Then, based on the values ​​of these trend characteristics, trend distribution bias and trend reference bias are calculated for different value types. Specifically, considering the data represented by each attribute indicator after generating the regional time distribution trend, within the intervals statistically analyzed for each trend characteristic, the trend distribution bias represents the difference between individual data points in these intervals and the overall mean. The overall mean represents the average of the corresponding total data within the currently analyzed region; this total data only represents the currently acquired data, not the total historical data stored in the database. The trend reference bias represents the difference between the current value and the historical standard value, which represents the average of historical data. The value type at this point indicates the central tendency and dispersion within the current regional time distribution trend. Multiple data segments are divided based on the values ​​of the mean, median, mode, standard deviation, variance, and interquartile range. For example, the data can be divided into three intervals—low, medium, and high—based on the values ​​of the mean, median, mode, standard deviation, variance, and interquartile range to set the value types of trend characteristics. Then, data with positive and negative deviations are further divided into value types corresponding to the trend characteristics. Positive deviations are data above the mean, and negative deviations are data below the mean. Next, for the data in each of the multiple value types, the slope values ​​of these data in the time series format are calculated sequentially. Taking simple linear regression as an example, a straight line is fitted using the least squares method, and the slope of this line is the required slope value. These slope values ​​are then fitted according to the trend characteristics. This fitted slope value can then represent which time periods in the entire time series are closer to the expected trend and which time periods deviate from the expectation.

[0049] This approach is illustrated using the following case: Suppose we want to analyze the changes in the innovation capability index of information technology companies in a province over the past ten years and formulate relevant policy recommendations based on this analysis.

[0050] Year: 2015 to 2024; Innovation Capability Index: The annual average of the innovation capability scores of each enterprise in each year; Average (central tendency): 70 points; Industry Excellence Level (Standard Value): 80 points.

[0051] For each year, calculate the average, standard deviation, and other statistical measures, and calculate the trend distribution deviation and trend reference deviation. For 2018, the innovation capability index is 85 points, then: trend distribution deviation = 85 - 70 = 15 points; trend reference deviation = 85 - 80 = 5 points. Perform linear regression analysis on the trend distribution deviation and trend reference deviation for all years. Assuming that the slope obtained is positive, it indicates that the overall trend is developing in a positive direction.

[0052] Suppose that analysis reveals that the trend distribution deviation and trend reference deviation are both positive for five years, with a positive slope, accounting for 50% of the total years. This means that during these five years, the innovation capabilities of the province's information technology companies have significantly improved and are approaching or even surpassing the industry's best level.

[0053] In the other five years, although the innovation capability index was below average in some years, the overall trend still showed a positive development trend.

[0054] The above analysis leads to the conclusion that the province's IT companies have shown a clear upward trend in innovation capabilities over the past decade, with significant breakthroughs achieved, particularly in certain key years. However, despite the overall upward trend, it is still necessary to pay attention to years where performance fell short of expectations, identify potential problems, and address them to ensure long-term, stable improvement in innovation capabilities.

[0055] In one embodiment of the present invention, during hierarchical analysis, hierarchical keywords are extracted from the data based on the distribution proportion calculated for each attribute description indicator. These keywords have different meanings depending on the level involved. For example, the target layer keywords, criterion layer keywords, and sub-criterion layer keywords in the hierarchical analysis method are considered as three levels: top, middle, and bottom. The bottom-level hierarchical keywords use the industry index in the corresponding data, and semantic analysis is performed sequentially to determine their index values ​​after semantic analysis, which are then used as the bottom-level hierarchical index values. The middle layer uses correlation analysis on the identified hierarchical keywords, performing correlation analysis on the hierarchical keywords under the trends of increasing and decreasing distribution proportions, and using the values ​​after correlation analysis as the hierarchical index values ​​of the middle layer. For example, words in the feature similarity and dissimilarity intervals are used for description, and the intersection of the hierarchical division is determined to judge whether the distribution of index values ​​is normal when the distribution proportion changes. The top layer uses the hierarchical keyword with the largest weight as the calculated index value. After mapping the index values ​​of the three levels to the industries in the region according to their weights, the comprehensive index is used as the assigned regional hierarchical index.

[0056] like Figure 5 As shown, the implementation of the hierarchical analysis module includes: extracting keywords from the distribution proportion of each attribute description indicator to obtain each benchmark keyword. The benchmark keywords are extracted from the data of the distribution proportion of each attribute description indicator. For example, when analyzing enterprise innovation capability, if the overall distribution proportion shows a positive growth, then the benchmark keyword will represent words related to the growth of innovation capability. Then, the key factors that can achieve the growth of innovation capability will be listed, such as the enterprise innovation capability index, the R&D investment ratio, etc., as well as the enterprise innovation capability scores for different years. The words represented by these values ​​will be considered as the benchmark keywords at this time.

[0057] Determine the hierarchical division of each benchmark keyword, and construct the hierarchical keywords for each level according to the multivariate mapping matrix under the corresponding level. Determining the hierarchical division of each benchmark keyword involves finding other benchmark keywords corresponding to each benchmark keyword, organizing these benchmark keywords according to the hierarchical structure, and dividing the data in this matrix into multiple levels according to the data corresponding to these benchmark keywords in the multivariate mapping matrix, thereby obtaining multiple hierarchical keywords. Hierarchical keywords are words that represent the benchmark keywords under the corresponding level.

[0058] Analyze the keywords at the top, middle, and bottom levels, determine the level indicator values ​​for each level, and assign the level indicator values ​​to the regional level indicators according to their weights.

[0059] The implementation method for analyzing the top, middle and bottom layer keywords is as follows: extract the industry index of each bottom layer as the bottom layer keyword, perform semantic analysis on the bottom layer keyword according to polysemous words, and set the bottom layer index value according to the semantic analysis results.

[0060] When performing semantic analysis, natural language processing libraries such as NLTK, spaCy, and transformers are used for word vector modeling and word frequency statistics. Contextual analysis is then used to identify the specific meanings of polysemous words. By comparing the underlying keywords with data in the dictionary as polysemous words, and calculating the cosine similarity of the keywords in different scenarios, we can determine how the same value of the underlying keywords can represent different meanings in different industries and scenarios. The cosine similarity of the underlying keywords in different scenarios is used as the semantic analysis result output. For example, consider the role of the price-to-earnings ratio (P / E ratio) in regional development in different industries and market stages. When the P / E ratio corresponds to a bull market or a bear market, the growth meaning represented by the same P / E ratio will be different. In this case, the underlying keywords will be set with hierarchical index values ​​based on the cosine similarity between these words and the data in the dictionary in the corresponding scenarios. This hierarchical index value will be biased towards describing the semantic similarity in the scenario, representing the specific meaning in the context of a specific industry, so that the subsequent hierarchical division can start from semantics, take relevance as the intermediate path, and take maximum weight as the goal to obtain a hierarchical analysis result. The value at each position in this result will represent the relevant elements in regional development, and how these relevant elements will change according to different goals.

[0061] Calculate the cosine similarity of the bottom-level keywords under the same distribution trend to obtain a similarity matrix. Merge the similarity matrices sequentially and set the intersections that exist when merging the similarity matrices as the middle-level keywords. Set the average Euclidean distance between the middle-level keywords and the corresponding bottom-level keywords as the middle-level index value.

[0062] At this stage, the underlying keywords are clustered, and then the intersections between these words are identified. For example, there might be an intersection between profitability and growth and risk level and liquidity, indicating a change in the position and correlation of these two sets of keywords in the hierarchical structure. Identifying this intersection allows the intermediate-level keywords to show significant changes in correlation when describing the distribution proportion of each attribute indicator. At the same time, when merging the similarity matrix, the similarity values ​​of corresponding keywords in the similarity matrix are compared. When the similarity value is greater than the similarity threshold, the data is merged. For example, if the similarity threshold is set to 0.7, the data with a similarity value greater than this are merged, the intersection of the merge is found, and words that can represent the connection relationship between different underlying-level keywords are selected according to the intersection to complete the construction of the relevant hierarchical analysis.

[0063] After obtaining the middle-level keywords, the top-level keywords need to be selected by filtering out the largest portion of the middle-level keywords and using it as the top-level keyword. That is, the data with the largest weight in the middle-level keywords is used as the top-level keyword, and the average value of the data corresponding to the top-level keyword is set as the top-level index value.

[0064] The hierarchical indicator values ​​of the top, middle and bottom layers are normalized, and the normalized data are then weighted and summed according to their weights to complete the assignment of regional hierarchical indicator values.

[0065] Throughout the hierarchical analysis process, a weight is assigned to each of the data representing the growth rate, sales revenue, output, and corresponding distribution proportion in the attribute description indicators to describe their relative importance in the regional development assessment. These weights are pre-set in the database, and the data set in the database is used in subsequent calculations to refine the assessment of the current region. Finally, these obtained hierarchical indicator values ​​are weighted and summed to obtain a regional hierarchical indicator value. This value represents the assessment value that the region can correspond to after the hierarchical analysis. This value can assist in making corresponding analysis and adjustments to the data structure analysis to fully understand the internal structure and internal relationships of various industries in the region, thereby completing the analysis and processing of the corresponding data.

[0066] The weights corresponding to the hierarchical index values ​​are set according to the proportion of the data corresponding to that hierarchical index value to the total number, so as to represent the weight of the data points under each level in the overall hierarchical analysis.

[0067] In one embodiment of the present invention, the model building module mainly assembles the data after hierarchical analysis into a regional development model, and performs mapping relationship processing on the multivariate mapping matrix according to the data existing in the regional development model to obtain the final industry evaluation indicators to be output.

[0068] After completing the assignment of values ​​at each level and setting the corresponding regional hierarchical indicators, we can determine the range and level of values ​​assigned to the multivariate mapping matrix for different industry characteristics and relative development in the corresponding regions. This allows the indicators and comprehensive values ​​at different levels to be displayed according to the industry indices that need to be identified, reflecting the situation of a region in terms of hierarchical analysis. Then, based on these hierarchical analysis contents, a corresponding regional development model is constructed. The regional development model represents the analysis model of the distribution proportion of indicators describing different attributes under this hierarchical analysis. To verify the accuracy and effectiveness of the model, the model is then used to find the derived mapping relationship with the multivariate mapping matrix, and parameter estimation is performed using linear structural equations. This determines whether the current multivariate mapping matrix and regional hierarchical indicators conform to the values ​​of the regional development model at the corresponding levels, thus judging whether the results of hierarchical analysis meet the judgment criteria. When the judgment criteria are met, the comprehensive indicator of the corresponding regional development model is used as the industry evaluation indicator at this time. In other words, the industry evaluation indicator is the value of the evaluated regional hierarchical indicator.

[0069] At this point, the existing derived mapping relationship is used to determine the correlation between the attribute description indicators targeted by the regional development model and the multivariate mapping matrix. For example, the indicator values ​​of the regional development model in multiple levels under the corresponding hierarchical analysis are divided into endogenous latent variables and exogenous manifest variables to verify whether there is consistency and correlation between the data at a single level of the regional development model and the multivariate mapping matrix itself. The industry index data and regional reference characteristics included in the multivariate mapping matrix are divided into exogenous manifest variables such as government subsidies, tax incentives, and GDP growth rate, and endogenous latent variables such as market share, innovation capability index, market size, growth rate index, profitability index, debt repayment capability index, and competitiveness index. Endogenous latent variables are those that cannot be directly observed but can be inferred or measured through other observable variables, i.e., manifest variables. Exogenous manifest variables are those that can be directly observed and measured. They are usually determined by external factors and affect endogenous latent variables as input variables.

[0070] like Figure 6 As shown, the model building module is implemented as follows: identify nodes at each level in the regional development model, map each node to a multivariate mapping matrix, and divide the endogenous latent variables and exogenous manifest variables corresponding to each node; each node represents the data corresponding to each level keyword, and the derivation mapping relationship at this time represents the mapping relationship between endogenous latent variables and exogenous manifest variables in each node.

[0071] Based on the mapping relationship between endogenous latent variables and exogenous manifest variables corresponding to nodes at each level, a linear structural equation for endogenous latent variables and exogenous manifest variables is constructed.

[0072] Based on the calculation results of the linear structural equations of endogenous latent variables and exogenous manifest variables, regional level indicators in the regional development model are selected and output as industry evaluation indicators.

[0073] For example, the linear structural equation for endogenous latent variables and exogenous manifest variables is expressed as: ;in, Indicates endogenous latent variables, The relation matrix represents the relationship between endogenous latent variables. The relation matrix expresses the relationship between each endogenous latent variable in the form of cosine similarity or Pearson correlation coefficient. The rows of the relation matrix represent an endogenous latent variable, and the columns represent other endogenous latent variables related to this endogenous latent variable. This represents the influence coefficient matrix from exogenous manifest variables to endogenous latent variables. Indicates exogenous manifest variables, This represents the residual term. The influence coefficient matrix from exogenous manifest variables to endogenous latent variables extracts data from the regional development model of the analytic hierarchy process (AHP). It identifies the sum of the shortest paths from the hierarchical nodes corresponding to the current exogenous manifest variables in the regional development model to the corresponding hierarchical nodes of the endogenous latent variables. This sum represents the influence coefficient of each exogenous manifest variable to the endogenous latent variable. For example, if the currently identified exogenous manifest variable is at a bottom-level node in the regional development model and the endogenous latent variable is at a middle-level node, then the path from the exogenous manifest variable to the endogenous latent variable in the corresponding data of the AHP is calculated. The sum of the paths is the sum of the weights of the hierarchical nodes traversed, which is used as the corresponding element value in the influence coefficient matrix. The rows of the influence coefficient matrix represent exogenous manifest variables, and the columns represent endogenous latent variables, to obtain the relative influence between these two values.

[0074] Linear structural equation modeling is used to analyze the relationships between endogenous latent variables and exogenous manifest variables. Data that conforms to normal relationships are output as industry evaluation indicators. The output of industry evaluation indicators includes regional level indicators, endogenous latent variables, exogenous manifest variables, as well as the corresponding relationship matrix, influence coefficient matrix, and residual terms of the endogenous latent variables and exogenous manifest variables to represent the corresponding relationship errors between the currently compared indicators. Finally, after receiving the corresponding calculation results and matrices, external staff can understand the evaluation process and results, thereby drawing conclusions about the development of the corresponding regional industry and improving the overall evaluation results.

[0075] In one embodiment of the present invention, the control and verification module mainly evaluates the data based on the solved industry evaluation indicators at multiple levels and the content of the industry evaluation indicators combined with multiple data in the multivariate mapping matrix, identifies whether there are corresponding abnormal situations in the currently described area, and after issuing warnings for these abnormal situations, retrieves the corresponding data to the control and management platform to realize regional management and control of each problem area.

[0076] The implementation of the control and verification module includes: dividing the image of the region to be identified according to the data corresponding to the industry evaluation indicators, and obtaining multiple sub-legend regions under the derived mapping relationship corresponding to the multivariate mapping matrix; each sub-legend region includes a legend label corresponding to the regional development model and the multivariate mapping matrix. The legend label represents the hierarchical keywords and other labels set under hierarchical analysis and time series analysis, and uses these hierarchical keywords, attribute description indicators and other features as the legend labels used at this time.

[0077] Edge extraction is performed on each sub-legend region to obtain a closed path containing a single legend label. This closed path indicates the region containing the same label after the sub-legend region is labeled. Connecting these regions is used as a closed path to evaluate the range covered by different legend labels in each region.

[0078] The process iterates through the closed paths of each legend label, identifies the boundary intersections of these paths, and overlays the data corresponding to these intersections to obtain the overlay amount for each sub-legend region. The overlaid data represents the combined situation of various indicators appearing at the same point, and this combined situation is used to further process the divided regions. For example, under a certain level keyword describing industry development, if multiple level keywords correspond to overlapping regions, the overlapping boundary points can be used to analyze the common problems or development conflicts existing in these overlapping areas. This helps identify whether the planned development content in each region is risky and whether the corresponding location requires early warning and control, thus completing the analysis and processing of the corresponding region. If the data of the overlapping legend labels at the boundary nodes can be directly overlaid and calculated, their weighted sum is directly calculated as the overlay amount for each sub-legend region. If it cannot be directly calculated, it is standardized to eliminate its dimensions before calculation, and the legend labels present at the boundary intersections are marked.

[0079] By using the overlay volume of each sub-legendary region and industry evaluation indicators, the sub-legendary regions are classified to obtain sub-legendary regions with different management priorities. Early warning information is then extracted according to these management priorities. Classifying the sub-legendary regions aims to identify areas with significant conflicts and areas with development risks. After determining the management priorities of these areas, the corresponding early warning information is output to complete the management of each region by the control platform.

[0080] Preferably, the method for classifying sub-legend regions includes: dividing the sub-legend regions into left and right adjacent regions based on the overlay amount and industry evaluation indicators; normalizing the distance between the left and right adjacent regions as the classification indicators for sub-legend regions; clustering the sub-legend regions using the classification indicators; setting management priorities for the clustered sub-legend regions; and then outputting the data corresponding to each category, which is the early warning information extracted at this time.

[0081] Preferably, the left adjacent region represents the region preceding the current sub-legend region, logically and spatially equivalent to the region preceding the current sub-legend region. The right adjacent region represents the following region, illustrating the data format of the region following the current sub-legend region. Then, Euclidean distance is calculated based on the positions of the left and right adjacent regions to determine their distance. Multiple clusters are generated using the overlay amount, industry evaluation indicators, and the distance between the left and right adjacent regions as clustering indicators. Management priorities are then assigned to the clustered data, with high-risk categories corresponding to high priority and low-risk categories to low priority. The management priorities are set according to the values ​​required by the three indicators for this clustering in the database. The corresponding data is then output according to the management priorities, resulting in multiple warning messages. These warning messages represent the management priorities and the data during clustering. Subsequent staff can review these data to identify problems during clustering, facilitating data control for the entire region.

[0082] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention, which are still covered within the protection scope of the present invention.

Claims

1. A regional industrial development analysis and evaluation system, characterized in that, include: The data integration module is used to determine the industry index data and regional reference features within the region to be identified, and to summarize the industry index data and regional reference features into a multivariate mapping matrix. The sequence analysis module is used to apply time series analysis to the multivariate mapping matrix, extract the attribute description indicators corresponding to each data from the perspective of data regional distribution, as well as the regional time distribution trend of the attribute description indicators, and determine the distribution ratio of each attribute description indicator under the regional time distribution trend. The hierarchical analysis module is used to perform hierarchical analysis on the multivariate mapping matrix based on the distribution ratio of each attribute description index, sequentially obtain hierarchical keywords of multiple levels, and assign regional hierarchical index values ​​according to the keywords of each level. The model building module is used to construct a regional development model according to the assigned values ​​of regional level indicators, and to solve the industry evaluation indicators based on the derived mapping relationship between the regional development model and the multivariate mapping matrix. The model building module is implemented as follows: Identify nodes at each level in the regional development model, map each node to a multivariate mapping matrix, and delineate the endogenous latent variables and exogenous manifest variables corresponding to each node. Based on the mapping relationship between endogenous latent variables and exogenous manifest variables corresponding to nodes at each level, a linear structural equation for endogenous latent variables and exogenous manifest variables is constructed. Based on the calculation results of the linear structural equations of endogenous latent variables and exogenous manifest variables, regional level indicators in the regional development model are selected and output as industry evaluation indicators. The control and verification module is used to assess the asset status of various industries based on industry evaluation indicators, and output early warning information according to the asset assessment status of each region, and retrieve data from each region to the control and management platform from the early warning information.

2. The regional industrial development analysis and evaluation system according to claim 1, characterized in that, Determining industry index data and regional reference characteristics within the region to be identified also includes: Obtain text and image data within the region to be identified, then extract industry index data and regional reference features from the region in sequence. Match the industry index data according to the geographical, economic, social, and industry features in the regional reference features, and then concatenate the matched data to obtain a multivariate mapping matrix.

3. The regional industrial development analysis and evaluation system according to claim 1, characterized in that, The implementation methods of the sequence analysis module include: Based on historical datasets of regional industries, multiple attribute description indicators are extracted from the multivariate mapping matrix. According to the industry economic growth rate, market share and R&D investment corresponding to the attribute description indicators, regional division intervals are generated for each attribute description indicator. Based on the regional division intervals, with each attribute description index as the core and regional reference features as the auxiliary, a regional feature model is constructed; clustering is performed using the values ​​of industry economic growth rate, market share and R&D investment, and the clustered data is used as the regional feature model to determine the relationship between each attribute description index and regional reference features. The regional feature model is analyzed using time series analysis to generate a regional time distribution trend; the distribution of each attribute descriptor in the regional time distribution trend is output to obtain the distribution ratio of each attribute descriptor.

4. The regional industrial development analysis and evaluation system according to claim 3, characterized in that, Other methods for implementing regional feature models include: Initialize each region into intervals, and determine that the median value corresponding to the upper and lower limits of each region interval is the same as the average value of the region interval; According to the attribute description indicators corresponding to the region division interval, the data within the region division interval is traversed. If the data in adjacent region division intervals overlap, the overlapping data is used as a new region division interval; until all region division intervals no longer overlap. Cluster analysis is performed on the traversed regional division intervals. Initial clustering is performed using each attribute description index, and secondary clustering is performed on the data after the initial clustering using regional reference features. Regional feature models are then constructed from the data after the secondary clustering.

5. The regional industrial development analysis and evaluation system according to claim 3, characterized in that, The implementation methods for the distribution proportion of each attribute description indicator also include: Identify the regional time distribution trend, and then obtain the degree of central tendency and the degree of dispersion in the regional time distribution trend. Based on the degree of central tendency and dispersion in the regional time distribution trend, trend features are extracted. The trend features are then categorized by value type, and the trend distribution deviation and trend reference deviation are calculated for each value type. The trend features are then fitted according to the slope values ​​of the trend distribution deviation and trend reference deviation. The data proportions corresponding to the fitted trend features are used as the distribution proportions of the output attribute description indicators.

6. The regional industrial development analysis and evaluation system according to claim 1, characterized in that, The implementation methods of the hierarchical analysis module include: Keyword extraction was performed on the distribution ratio of each attribute description indicator to obtain the benchmark keywords; Determine the hierarchical division of each benchmark keyword, and construct the hierarchical keywords for each level according to the multivariate mapping matrix under the corresponding level; Analyze the keywords at the top, middle, and bottom levels, determine the level indicator values ​​for each level, and assign the level indicator values ​​to the regional level indicators according to their weights.

7. A regional industrial development analysis and evaluation system according to claim 6, characterized in that, The implementation methods of top-level, middle-level, and bottom-level hierarchical keywords are analyzed as follows: Extract the industry indices of each underlying layer as the underlying layer keywords, perform semantic analysis on the underlying layer keywords according to polysemous words, and set the underlying layer index values ​​according to the semantic analysis results. Calculate the cosine similarity of the bottom-level keywords under the same distribution trend to obtain a similarity matrix. Merge the similarity matrices sequentially and set the intersections that exist when merging the similarity matrices as the middle-level keywords. Set the average Euclidean distance between the middle-level keywords and the corresponding bottom-level keywords as the middle-level index value. The data with the highest weight among the middle-level keywords is taken as the top-level keyword, and the average value of the data corresponding to the top-level keyword is set as the top-level index value. The hierarchical indicator values ​​of the top, middle and bottom layers are normalized, and the normalized data are then weighted and summed according to their weights to complete the assignment of regional hierarchical indicator values.

8. The regional industrial development analysis and evaluation system according to claim 1, characterized in that, The implementation methods of the control and verification module include: The image is segmented according to the data corresponding to the industry evaluation indicators, and multiple sub-legendary regions under the derived mapping relationship corresponding to the multivariate mapping matrix are obtained. Edge extraction is performed on each sub-legend region to obtain closed paths containing a single legend label. The closed paths of each legend label are traversed, the boundary intersections of the closed paths are identified, and the data corresponding to the boundary intersections are overlaid and analyzed to obtain the overlay amount of each sub-legend region. By using the overlay amount of each sub-legend region and industry evaluation indicators, the sub-legend regions are classified to obtain sub-legend regions with different management priorities, and early warning information is extracted according to the management priority.

9. A regional industrial development analysis and evaluation system according to claim 8, characterized in that, The method for classifying sub-legend regions includes: dividing the sub-legend regions into left and right adjacent regions based on the overlay amount and industry evaluation indicators; normalizing the distance between the left and right adjacent regions as the classification indicators for sub-legend regions; and clustering the sub-legend regions using the classification indicators.

Citation Information

Patent Citations

  • Industrial development evaluation method and system and electronic equipment

    CN117634952A

  • Regional industry development evaluation method, system, equipment and medium

    CN118917731A

  • Offshore wind power development planning and evaluation early warning method and system

    CN114429264A

  • Regional high and new technology industry evaluation method and system

    CN117436726A

  • A data analysis method and system for regional industry evaluation

    CN119761651A