Regional development element data processing method and system

By constructing a knowledge graph and using a hierarchical encoding method, the problem of strong macroscopic data decompression efficiency in regional economic data is solved, and rapid decoding and efficient analysis are achieved, and scientific and intelligent decision-making in regional development are supported.

CN120355107AActive Publication Date: 2025-07-22GUANGZHOU PLANNING DESIGN OFFICE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510847019.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the processing of regional economic data, the decompression efficiency of data with strong macroscopicity is inefficient, unable to meet the needs of rapid decision-making, and lacking an effective analysis mechanism for data macroscopy, making it difficult for macroscopic key information to be accurately identified and processed efficiently.

Method used

The hierarchical encoding method is adopted to set different encoding layers according to the macroscopic degree of the data. By constructing a knowledge graph, repetitive value nodes are removed, and the Hoffman coding algorithm is used to combine the hierarchical encoding idea to improve the decoding efficiency of data with a large macroscopic degree.

Benefits of technology

It improves data decoding efficiency, especially data with a large degree of macroscopicity, realizes rapid acquisition and efficient analysis, and supports the rapid response and precise formulation of regional development decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355107A_ABST
    Figure CN120355107A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to a regional development element data processing method and system. The method comprises the following steps: acquiring regional development element data and constructing a knowledge graph, calculating the macroscopic degree of each value according to the number of nodes after each value data, removing nodes with repeated values in the knowledge graph to obtain a simplified knowledge graph, performing continuous iterative cutting on the simplified knowledge graph according to the macroscopic degree, and obtaining the knowledge graph. And after each cutting, the obtained subsets are coded to obtain a multi-layer code. And the decoding efficiency of data with a large macroscopic degree is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method and system for processing regional development factor data. Background Art

[0002] In today's digital age, the research and decision-making on regional economic development highly rely on a vast amount of economic data. This data covers multi-dimensional information such as gross regional product, industrial structure, employment situation, trade transactions, fiscal revenue and expenditure, etc. They are not only huge in quantity, but also show an explosive growth in data scale over time and with the enrichment of monitoring means. Therefore, in the context of the rapid increase in regional economic data volume, in order to solve the problems of storage, transmission and analysis efficiency and achieve accurate economic decision-making, efficient compression of regional economic development data is the key to breaking through development obstacles and releasing data value.

[0003] Existing data processing solutions usually adopt a unified compression strategy without considering the importance differences of data at the macro decision-making level. In scenarios such as regional development planning and policy formulation, decision-makers often need to quickly obtain macro data such as GDP growth rate and industrial structure ratio in order to grasp the overall regional development trend. However, the unified compression method will lead to low efficiency in decompressing macro data, unable to meet the needs of rapid decision-making. For example, when there is a sudden economic policy adjustment, it is necessary to quickly analyze the macro trends of the regional industrial economy, but due to the long decompression time of macro data, the decision-making opportunity is delayed. In addition, the lack of an effective analysis mechanism for data macroscopity makes it difficult to accurately identify and efficiently process macro key information in the vast amount of data, seriously restricting the scientific and intelligent process of regional development. To meet the urgent need for rapid acquisition and efficient analysis of macro data in the field of regional development, a differential compression method based on the macro characteristics of data is urgently needed to ensure higher efficiency in decompressing macro data and assist in the rapid response and accurate formulation of regional development decisions. Summary of the Invention

[0004] In order to solve the problem of how to improve the decompression efficiency of data with high macroscopity, the present invention provides a method and system for processing regional development factor data.

[0005] In a first aspect, the present invention provides a method for processing regional development factor data, adopting the following technical solution: A method for processing regional development factor data includes steps: Obtain regional development factor data; Construct a knowledge graph using all regional development factor data, calculate the macro degree of each value according to the number of subsequent nodes of each value data, remove the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph, calculate the first segmentation times of the nodes corresponding to each value according to the macro degree, perform segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-layer subsets, use the first-layer subsets as encoding objects, encode the first-layer subsets according to the number of data of all values in the first-layer subsets, and use the encoding of the first-layer subsets as the first-layer encoding of each value therein; In response to the existence of a first-layer subset with a data quantity greater than 1, use the first-layer subset with a data quantity greater than 1 as an analysis subset; calculate the second segmentation times of the nodes corresponding to each value according to the macro degree, perform segmentation control on the analysis subset according to the second segmentation times to obtain the second-layer subsets, use the second-layer subsets as encoding objects, encode the second-layer subsets according to the number of data of all values in the second-layer subsets, and use the encoding of the second-layer subsets as the second-layer encoding of each value therein; In response to the non-existence of a second-layer subset with a data quantity greater than 1, store all-layer encodings of the regional development factor data.

[0006] The present invention reduces the number of matching times for each layer of encoding during decoding through a hierarchical encoding method, thereby improving the decoding efficiency; further, different encoding layers are set for the data according to the macro degree of the data, so that the encoding layer of the data with a large macro degree is smaller, thereby improving the decoding efficiency of the data with a large macro degree; further, during the encoding process, each layer of subset is encoded according to the number of data of all values in each layer of subset, so that the encoding with a large data usage amount is set shorter, and the encoding with a small data usage amount is set longer, effectively improving the compression effect.

[0007] Preferably, the calculation of the macro degree of each value according to the number of subsequent nodes of each value data includes: ; wherein, represents the number of nodes on the i-th path after each node, N represents the number of paths after each node, H represents the macro degree of each node, and norm() represents linear normalization processing; Take the mean value of the macro degrees of all nodes with the same value as the macro degree of this value.

[0008] The present invention accurately reflects the content of information covering various attributes after the node through the number of subsequent nodes of each node, and further accurately reflects the macro degree situation of the node.

[0009] Preferably, the removal of the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph includes: For any value, all nodes with this value are retrieved from the knowledge graph. Among all nodes with this value, the node with the smallest difference in macroscopic degree from this value is denoted as the reserved node, and other nodes with this value except the reserved node are removed; During the removal process, if there are nodes before and after other nodes with this value, after removing the nodes with this value, the nodes before and after it are connected; if other nodes with this value only exist before or after nodes, the nodes with this value are directly removed.

[0010] The present invention improves the encoding quality by removing nodes with duplicate values in the knowledge graph, so as to set different encodings for data with the same value during the subsequent encoding process.

[0011] Preferably, calculating the first splitting times of nodes corresponding to each value according to the macroscopic degree includes: Retrieving the number of connection edges of each node in the simplified knowledge graph, taking the maximum value of the number of connection edges of all nodes in the simplified knowledge graph as the first reference value, and performing normalization processing on the product of the macroscopic degree and the first reference value to obtain the first splitting times of nodes corresponding to each value.

[0012] The present invention introduces the macroscopic degree when calculating the splitting times, so that nodes with a larger macroscopic degree have a larger splitting times, and then enables nodes with a larger macroscopic degree to be split into independent data faster, providing a basis for reducing the encoding layers of data with a larger macroscopic degree.

[0013] Preferably, controlling the splitting of the simplified knowledge graph according to the first splitting times to obtain the first-layer subset includes: Retrieving the connected nodes of any node in the simplified knowledge graph, obtaining the absolute value of the difference in macroscopic degree between the value corresponding to this node and the value corresponding to the connected nodes, arranging all connected nodes in descending order according to the absolute value of the difference in macroscopic degree to obtain a connected node sequence, and taking the first K1 connected nodes in the connected node sequence as splitting nodes, where K1 represents the first splitting times; Cutting off the connection edges between this node and the splitting nodes, classifying nodes with a connection relationship with each other into a subset, and denoting the obtained subset as the first-layer subset.

[0014] Preferably, taking the first-layer subset as the encoding object, encoding the first-layer subset according to the number of data of all values in the first-layer subset, and taking the encoding of the first-layer subset as the first-layer encoding of each value in it includes: Obtain the quantity of regional development factor data for each value in the first-level subset, take the sum of the quantities of regional development factor data for all values in the first-level subset as the occurrence frequency of the first-level subset, use the first-level subset as the coding object, and perform coding processing on all first-level subsets based on the occurrence frequency using a coding compression algorithm to obtain the code for each first-level subset, and use the code of the first-level subset as the first-level code for the value of each node within it.

[0015] The present invention takes into account that the Huffman coding algorithm will set shorter codes for data with a large occurrence frequency, so the sum of the data of the regional development factor data for all values in the subset is used to set the occurrence frequency, providing a basis for effectively reducing the length of the codes with a large amount of data used and increasing the compression ratio.

[0016] Preferably, calculating the second splitting times of the corresponding nodes for each value according to the macroscopic degree includes: Obtain the quantity of remaining connection edges of each node, take the maximum value of the quantities of remaining connection edges of all nodes in the analysis subset as the second reference value, and perform normalization processing on the product of the macroscopic degree and the second reference value to obtain the second splitting times of the corresponding nodes for each value.

[0017] Preferably, calculating the second splitting times of the corresponding nodes for each value according to the macroscopic degree, and performing splitting control on the analysis subset according to the second splitting times to obtain the second-level subset includes: Obtain the connected nodes of any node in the analysis subset, obtain the difference in macroscopic degree between this node and the connected nodes, arrange all the connected nodes in descending order according to the difference in macroscopic degree to obtain a connected node sequence, and take the first K2 connected nodes in the connected node sequence as the splitting nodes, where K2 represents the second splitting times; Cut off the connection edges between this node and the splitting nodes, and classify the nodes that have a connection relationship with each other in the analysis subset into a subset, and denote the newly obtained subset as the second-level subset.

[0018] Preferably, coding the second-level subset according to the quantity of data for all values in the second-level subset, and using the code of the second-level subset as the second-level code for each value within it includes: Obtain the quantity of regional development factor data for each value in the second-level subset, take the sum of the quantities of regional development factor data for all values in the second-level subset as the occurrence frequency of the second-level subset, use the second-level subset as the coding object, and perform coding processing on all second-level subsets in the analysis subset based on the occurrence frequency using a coding compression algorithm to obtain the code for each second-level subset, and use the code of the second-level subset as the second-level code for each value within it.

[0019] In a second aspect, the present invention provides a system for processing regional development factor data, adopting the following technical solutions: A system for processing regional development factor data includes: a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned method for processing regional development factor data is implemented.

[0020] By adopting the above technical solutions, the above-mentioned method for processing regional development factor data is generated into a computer program and stored in the memory, so as to be loaded and executed by the processor. Thus, a terminal device is manufactured based on the memory and the processor, which is convenient to use.

[0021] The present invention has the following technical effects: The present invention reduces the number of matching times for each layer of coding during decoding by means of hierarchical coding, thereby improving the decoding efficiency; Furthermore, different coding layers are set for the data according to the macroscopic degree of the data, so that the coding layer of the data with a larger macroscopic degree is smaller, and thus the decoding efficiency of the data with a larger macroscopic degree is improved; Furthermore, during the coding process, each layer of subsets is coded according to the number of data of all kinds of values in each layer of subsets, so that the coding with a larger data usage amount is set shorter, and the coding with a smaller data usage amount is set longer, effectively improving the compression effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the method in a method for processing regional development factor data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] An embodiment of the present invention discloses a method for processing regional development factor data, referring to Figure 1 , including steps S1 - S5: S1: Obtain regional development factor data.

[0024] Specifically, obtain regional development factor data. The types of regional development factor data include but are not limited to the following aspects: industrial data, enterprise data, and market data.

[0025] Among them, industrial data includes but is not limited to the following aspects: industrial structure data, industrial development scale data, industrial growth data, and industrial employment data of the entire region; industrial structure data, industrial development scale data, industrial growth data, and industrial employment data of a local area. Enterprise data includes but is not limited to the following aspects: enterprise operation status data, innovation ability data, and market share data of the entire region; enterprise operation status data, innovation ability data, and market share data of a local area. Market data includes but is not limited to the following aspects: market supply and demand data, price index data, and trade data of the entire region; market supply and demand data, price index data, and trade data of a local area.

[0026] S2: Construct a knowledge graph using all regional development factor data, calculate the macro degree of each value according to the number of subsequent nodes of each value data, remove the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph, calculate the first segmentation times of the nodes corresponding to each value according to the macro degree, perform segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-layer subsets, use the first-layer subsets as the coding objects, encode the first-layer subsets according to the number of data of all values in the first-layer subsets, and use the encoding of the first-layer subsets as the first-layer encoding of each value in it.

[0027] It should be noted that the Huffman coding compression algorithm, as a commonly used compression algorithm, has a good compression effect. Therefore, the Huffman coding compression algorithm can be used for data compression. The coding obtained by the Huffman coding compression algorithm is a variable-length coding, so the segmentation points of the codes of different data cannot be identified. Therefore, this algorithm needs to be decoded sequentially according to the order of the data during decoding. Using the traditional Huffman coding compression algorithm for compression cannot achieve the priority decoding of data with a large macro degree. In order to improve the decoding efficiency of data with a high macro degree, it is considered to add the idea of hierarchical coding to the traditional Huffman coding algorithm, that is, set different coding layers for different data according to the macro situation of the data, set fewer coding layers for data with a large macro degree, and set more coding layers for data with a small macro degree, so that data with a large macro degree can be decoded quickly.

[0028] S20: Construct a knowledge graph using all regional development factor data.

[0029] It should be noted that to set the coding layer of data according to the macro degree of the data, the macro degree of the data needs to be obtained first.

[0030] It should be further noted that the knowledge graph can better reflect the association relationship and coverage relationship between data. Therefore, the knowledge graph can better reflect the macro situation of the data. First, a knowledge graph needs to be constructed.

[0031] Preferably, as an example, a knowledge graph is constructed using all regional development factor data, including: Based on all regional development factor data, a knowledge graph is constructed using existing knowledge graph construction methods. The existing knowledge graph construction methods can be the Neo4j method or the Apache Jena method, or other knowledge graph construction methods, which are not specifically limited in this embodiment.

[0032] S21: Calculate the macro degree of each value according to the number of subsequent nodes of each value data.

[0033] Preferably, as an example, calculate the macro degree of each node according to the position of each node in the knowledge graph, including:

[0034] Among them, represents the number of nodes on the i-th path after each node, N represents the number of paths after each node, H represents the macro degree of each node, norm() represents linear normalization processing. In this embodiment, the maximum-minimum normalization method is adopted, and other embodiments can adopt other normalization methods, which are not specifically limited in this embodiment; Take the mean of the macro degrees of all nodes with the same value as the macro degree of this value.

[0035] It can be understood that the more nodes there are on each path after this node, the more information of sub-attributes this node contains, and thus the greater the macro degree of this node. The larger it is, the more sub-attribute information there is after this node, and thus the greater the macro degree of this node.

[0036] S22: Remove the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph.

[0037] It should be noted that in order to make data with different macro degrees have different coding layers, consider using the method of dividing subsets for coding. That is, continuously iterate and divide the large subset until the subset contains a single data. Each time it is divided, coding is performed once. Therefore, the more times it is divided, the more coding layers there are.

[0038] It should be further noted that when using the Huffman coding algorithm for coding, one type of value data corresponds to one type of coding. When using the method of subset division for coding, it is easy to divide data with the same value into different subsets, resulting in different codings for data in different subsets. To prevent this phenomenon from occurring, only one node of each value needs to be retained in the knowledge graph.

[0039] Preferably, as an example, obtaining a simplified knowledge graph by removing nodes with duplicate values in the knowledge graph, including: For any one value, obtain all nodes with this value in the knowledge graph, and among all nodes with this value, obtain the node with the smallest difference in macroscopic degree from this value and denote it as the reserved node, and remove other nodes with this value except the reserved node; During the removal process, if there are nodes before and after other nodes with this value, then after removing the nodes with this value, connect the nodes before and after it; if other nodes with this value only exist before or after nodes, then directly remove the nodes with this value.

[0040] It can be understood that through the above processing, only one node with a certain value can be retained in the knowledge graph, providing a basis for subsequent hierarchical coding.

[0041] S23: Calculate the first splitting times of the nodes corresponding to each value according to the macroscopic degree.

[0042] It should be noted that when performing subset splitting, consider implementing subset splitting by splitting the simplified knowledge graph. In order to make data with a larger macroscopic degree have fewer splitting times, it is necessary to set the splitting times of the nodes corresponding to the data for each split according to the macroscopic degree of the data, so that the nodes corresponding to the data with a larger macroscopic degree have more splitting times for each split, and then make the nodes corresponding to the data with a larger macroscopic degree split from other nodes faster, reducing the splitting times of this data, and further providing a basis for reducing the number of coding layers subsequently.

[0043] Preferably, as an example, calculating the first splitting times of the nodes corresponding to each value according to the macroscopic degree, including: Obtain the number of connection edges of each node in the simplified knowledge graph, take the maximum value of the number of connection edges of all nodes in the simplified knowledge graph as the first reference value. If the product of the macroscopic degree of the nodes corresponding to any one value in the simplified knowledge graph and the first reference value is less than the number of connection edges, then take the floor value of the product of the macroscopic degree and the first reference value as the first splitting times of the nodes corresponding to this value; if the product of the macroscopic degree of the nodes corresponding to this value and the first reference value is not less than the number of connection edges, then take the number of connection edges of this node as the first splitting times.

[0044] It can be understood that the splitting times obtained by this method can ensure that data with a larger macroscopic degree has more splitting times, so that data with a larger macroscopic degree splits from other nodes faster, reducing the number of times this data participates in splitting, and further providing a basis for reducing the number of coding layers subsequently.

[0045] S24: Perform segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-level subsets.

[0046] Preferably, as an example, performing segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-level subsets includes: Obtain the connected nodes of any node in the simplified knowledge graph, obtain the absolute value of the difference in the macroscopic degree between the corresponding value of this node and the corresponding values of the connected nodes, sort all the connected nodes in descending order according to the absolute value of the difference in the macroscopic degree to obtain a connected node sequence, and use the first K1 connected nodes in the connected node sequence as the segmentation nodes, where K1 represents the first segmentation times; Cut off the connection edges between this node and the segmentation nodes, classify the nodes that have a connection relationship with each other into a subset, and denote the obtained subset as the first-level subset.

[0047] It can be understood that through the segmentation control by the first segmentation times, the data with a larger macroscopic degree can be separated from other nodes faster, reducing the number of times this data participates in the segmentation, and thus reducing the number of encoding layers for the subsequent process.

[0048] S25: Use the first-level subsets as the encoding objects, encode the first-level subsets according to the number of data with all kinds of values in the first-level subsets, and use the encoding of the first-level subsets as the first-level encoding of each kind of value therein.

[0049] It should be noted that in order to improve the compression ratio, the length of the encoding for the data with a higher usage frequency should be set shorter, and the length of the encoding for the data with a lower usage frequency should be set longer. Therefore, encoding processing needs to be carried out based on this.

[0050] Preferably, as an example, using the first-level subsets as the encoding objects, encoding the first-level subsets according to the number of data with all kinds of values in the first-level subsets, and using the encoding of the first-level subsets as the first-level encoding of each kind of value therein includes: Obtain the number of data of the regional development factor with each kind of value in the first-level subsets, use the sum of the number of data of the regional development factor with all kinds of values in the first-level subsets as the occurrence frequency of the first-level subsets, use the first-level subsets as the encoding objects, and perform encoding processing on all the first-level subsets based on the occurrence frequency using an encoding compression algorithm to obtain the encoding of each first-level subset, and use the encoding of the first-level subsets as the first-level encoding of the value of each node therein.

[0051] It can be understood that each value in the first - layer subset will use the encoding of this layer subset. Therefore, the more data of all values in the first - layer subset, the more data using the encoding of this layer subset. Thus, the encoding of the first - layer subset is set shorter to save the compression amount. The Huffman coding algorithm controls the encoding length based on the occurrence frequency. Therefore, the occurrence frequency is set according to the number of data of all values in the first - layer subset to adjust the encoding length of the data with a high occurrence frequency.

[0052] S3: In response to the number of data in the first - layer subset being greater than 1, take the first - layer subset with the number of data greater than 1 as the analysis subset; calculate the second splitting times of the corresponding nodes of each value according to the macroscopic degree, perform splitting control on the analysis subset according to the second splitting times to obtain the second - layer subset, take the second - layer subset as the encoding object, encode the second - layer subset according to the number of data of all values in the second - layer subset, and use the encoding of the second - layer subset as the second - layer encoding of each value therein.

[0053] It should be noted that the number of data in the first - layer subset being greater than 1 indicates that it has not been split into independent data. Therefore, it is necessary to continue splitting and continue encoding.

[0054] S30: Calculate the second splitting times of the corresponding nodes of each value according to the macroscopic degree.

[0055] Preferably, as an example, calculating the second splitting times of the corresponding nodes of each value according to the macroscopic degree includes: Obtain the number of remaining connection edges of each node, take the maximum value of the number of remaining connection edges of all nodes in the analysis subset as the second reference value. If the product of the macroscopic degree of the node corresponding to any value in the analysis subset and the second reference value is less than the number of remaining connection edges, take the floor value of the product of the macroscopic degree and the second reference value as the second splitting times of the node corresponding to this value; if the product of the macroscopic degree of the node corresponding to this value and the second reference value is less than the number of remaining connection edges, take the number of remaining connection edges of this node as the second splitting times.

[0056] S31: Perform splitting control on the analysis subset according to the second splitting times to obtain the second - layer subset.

[0057] Preferably, as an example, performing splitting control on the analysis subset according to the second splitting times to obtain the second - layer subset includes: Obtain the connected nodes of any node in the analysis subset, obtain the difference in macroscopic degree between this node and the connected nodes, sort all the connected nodes in descending order according to the difference in macroscopic degree to obtain a connected - node sequence, and take the first K2 connected nodes in the connected - node sequence as the splitting nodes, where K2 represents the second splitting times; Cut off the connection edge between this node and the splitting node, classify the nodes with connection relationships in the analysis subset into one subset, and denote the newly obtained subset as the second-layer subset.

[0058] S32: Take the second-layer subset as the coding object, encode the second-layer subset according to the number of data of all types of values in the second-layer subset, and use the encoding of the second-layer subset as the second-layer encoding of each value therein.

[0059] Preferably, as an example, taking the second-layer subset as the coding object, encoding the second-layer subset according to the number of data of all types of values in the second-layer subset, and using the encoding of the second-layer subset as the second-layer encoding of each value therein includes: Obtain the number of regional development factor data of each value in the second-layer subset, take the sum of the number of regional development factor data of all types of values in the second-layer subset as the occurrence frequency of the second-layer subset, take the second-layer subset as the coding object, and use the coding compression algorithm to encode all the second-layer subsets in the analysis subset based on the occurrence frequency to obtain the encoding of each second-layer subset, and use the encoding of the second-layer subset as the second-layer encoding of each value therein.

[0060] S4: In response to the non-existence of the number of data in the second-layer subset being greater than 1, store all-layer encodings of the regional development factor data.

[0061] It should be noted that when the number of data in the second-layer subset is less than or equal to 1, it means that there is only a single data in each second-layer subset. At this time, no splitting is required and the hierarchical encoding is completed.

[0062] Preferably, as an example, storing all-layer encodings of the regional development factor data includes: For any layer, if there is no encoding of any regional development factor data in this layer, use the empty encoding as the encoding of this data in this layer, and connect the encodings of all regional development factor data in this layer together for storage; in the same way, store the encodings of all regional development factor data in each layer. Record the correspondence between all types of values of all regional development factor data and each layer subset, and store the Huffman tree obtained for each layer subset.

[0063] Specifically, if there is only one type of value data in the subset, then take this value as the coding object, take the number of occurrences of this value in all regional development factor data as the occurrence frequency, and perform encoding based on the occurrence frequency.

[0064] It should be noted that, for the sake of convenience of explanation, the method for setting the empty code will be exemplified below. Suppose there are 4 codes in this layer, namely 00, 11, 100, and 101. By observation, it can be found that the code 01 does not exist among these 4 codes, so the empty code can be set to 01.

[0065] S5: Perform decoding processing.

[0066] Preferably, as an example, the decoding processing includes: Perform subset partitioning according to the correspondence between all kinds of values and the subsets of each layer to obtain the subsets of each layer.

[0067] Starting from the first layer, obtain the codes of the regional development element data in each layer. Based on the mapping relationship between the codes and the data sets in the Huffman tree of each layer subset, decode the subsets to which each data belongs. And so on, until the corresponding code is the empty code, the decoding ends. Obtain the data corresponding to the code of the upper layer of the empty code to implement data decoding.

[0068] It can be understood that compared with all the regional development element data, the amount of data in each layer subset is less, and the number of code types in each layer subset is relatively small. Therefore, there are fewer objects to be matched during decoding, so the decoding efficiency of each layer subset is relatively high. Therefore, the data with fewer coding layers can be decoded quickly.

[0069] An embodiment of the present invention also discloses a regional development element data processing system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a method for processing regional development element data according to the present invention is implemented.

[0070] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.

[0071] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as, for example, a resistive random access memory, a dynamic random access memory, a static random access memory, an enhanced dynamic random access memory, a high-bandwidth memory, a hybrid storage cube, etc., or any other medium that can be used to store the required information and can be accessed by an application program, a module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device.

Claims

1. A method for processing regional development factor data, characterized in that, Including the steps: Obtain regional development factor data; Construct a knowledge graph using all the regional development factor data, calculate the macro degree of each value according to the number of subsequent nodes of each value data, remove the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph, calculate the first segmentation times of the nodes corresponding to each value according to the macro degree, perform segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-layer subsets, use the first-layer subsets as the coding objects, encode the first-layer subsets according to the number of data of all values in the first-layer subsets, and use the encoding of the first-layer subsets as the first-layer encoding of each value therein; In response to the existence of a first-layer subset with a data quantity greater than 1, use the first-layer subset with a data quantity greater than 1 as the analysis subset; calculate the second segmentation times of the nodes corresponding to each value according to the macro degree, perform segmentation control on the analysis subset according to the second segmentation times to obtain the second-layer subsets, use the second-layer subsets as the coding objects, encode the second-layer subsets according to the number of data of all values in the second-layer subsets, and use the encoding of the second-layer subsets as the second-layer encoding of each value therein; In response to the non-existence of a second-layer subset with a data quantity greater than 1, store all the layer encodings of the regional development factor data.

2. The method for processing regional development element data according to claim 1, characterized in that The calculating the macro degree of each value according to the number of subsequent nodes of each value data includes: ; Among them, represents the number of nodes on the i-th path after each node, N represents the number of paths after each node, H represents the macroscopic degree of each node, and norm() represents linear normalization processing; Taking the mean value of the macro degrees of all nodes with the same value as the macro degree of this value.

3. A method for processing regional development element data according to claim 1, characterized in that The removing the nodes with duplicate values in the knowledge graph to obtain a simplified knowledge graph includes: For any value, obtain all the nodes of this value in the knowledge graph, obtain the node with the smallest difference in macro degree from the macro degree of this value among all the nodes of this value and denote it as the reserved node, and remove the other nodes of this value except the reserved node; During the removing process, if there are nodes before and after the other nodes of this value, after removing the nodes of this value, connect the nodes before and after it; if the other nodes of this value only exist before or after, directly remove the nodes of this value.

4. The method for processing regional development element data according to claim 1, wherein, The calculating the first segmentation times of the nodes corresponding to each value according to the macro degree includes: Obtain the number of connection edges of each node in the simplified knowledge graph, take the maximum value of the number of connection edges of all nodes in the simplified knowledge graph as the first reference value, and perform regularization processing on the product of the macro degree and the first reference value to obtain the first segmentation times of the nodes corresponding to each value.

5. A method for processing regional development element data according to claim 1, characterized in that, The performing segmentation control on the simplified knowledge graph according to the first segmentation times to obtain the first-layer subsets includes: Obtain the connected nodes of any node in the simplified knowledge graph, obtain the absolute value of the difference in macro degree between the value corresponding to this node and the value corresponding to the connected nodes, sort all the connected nodes in descending order according to the absolute value of the difference in macro degree to obtain a connected node sequence, and use the first K1 connected nodes in the connected node sequence as the segmentation nodes, where K1 represents the first segmentation times; Cut the connection edge between this node and the splitting node, classify the nodes with connection relationships into a subset, and denote the obtained subset as the first-layer subset.

6. The method for processing regional development element data according to claim 1, wherein, Taking the first-layer subset as the coding object, coding the first-layer subset according to the number of data of each type of value in the first-layer subset, and taking the coding of the first-layer subset as the first-layer coding of each type of value therein, including: Obtain the number of regional development element data of each type of value in the first-layer subset, take the sum of the number of regional development element data of all types of values in the first-layer subset as the occurrence frequency of the first-layer subset, take the first-layer subset as the coding object, and use the coding compression algorithm to code all the first-layer subsets based on the occurrence frequency to obtain the coding of each first-layer subset, and take the coding of the first-layer subset as the first-layer coding of the value of each node therein.

7. A method for processing regional development element data according to claim 1, characterized in that Calculating the second splitting times of the corresponding nodes of each type of value according to the macroscopic degree, including: Obtain the number of remaining connection edges of each node, take the maximum value of the number of remaining connection edges of all nodes in the analysis subset as the second reference value, and regularize the product of the macroscopic degree and the second reference value to obtain the second splitting times of the corresponding nodes of each value.

8. A method for processing regional development factor data according to claim 1, characterized in that Calculating the second splitting times of the corresponding nodes of each type of value according to the macroscopic degree, and controlling the splitting of the analysis subset according to the second splitting times to obtain the second-layer subset, including: Obtain the connected nodes of any node in the analysis subset, obtain the difference in macroscopic degree between this node and the connected nodes, sort all the connected nodes in descending order according to the difference in macroscopic degree to obtain the connected node sequence, and take the first K2 connected nodes in the connected node sequence as the splitting nodes, where K2 represents the second splitting times; Cut the connection edge between this node and the splitting nodes, classify the nodes with connection relationships in the analysis subset into a subset, and denote the newly obtained subset as the second-layer subset.

9. A method for processing regional development element data according to claim 1, characterized in that, Coding the second-layer subset according to the number of data of each type of value in the second-layer subset, and taking the coding of the second-layer subset as the second-layer coding of each type of value therein, including: Obtain the number of regional development element data of each type of value in the second-layer subset, take the sum of the number of regional development element data of all types of values in the second-layer subset as the occurrence frequency of the second-layer subset, take the second-layer subset as the coding object, and use the coding compression algorithm to code all the second-layer subsets in the analysis subset based on the occurrence frequency to obtain the coding of each second-layer subset, and take the coding of the second-layer subset as the second-layer coding of each type of value therein.

10. A regional development element data processing system, characterized in that, Including: A processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a method for processing regional development element data according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Risk prediction method and device, equipment and storage medium

    CN113822494A

  • Intelligent financial query method and device, equipment and medium

    CN118796984A

  • Knowledge graph construction and intelligent question and answer method and device based on deep learning

    CN119691135A

  • Knowledge extraction method, apparatus, electronic device, and storage medium

    WO2021212682A1

  • Risk prediction method and apparatus, and device and storage medium

    WO2023065545A1