A data optimization acquisition method and system applied to urban planning
By analyzing the type of urban planning data and the geographical location, storage space and priority of nodes, the storage relationship between data and nodes is determined, and the problem of low storage efficiency in urban planning data acquisition is solved, and efficient data collection and storage is achieved.
Patent Information
- Application Number
- CN202510352653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Due to the large amount of data, urban planning data cannot be effectively stored during the collection process, which affects the efficiency of data calling. The distributed storage ignores geographical location relationships in the allocation of storage space and access relationships, which cannot effectively improve storage efficiency.
Through the urban planning management platform, the data type and node geographical location, storage space and priority are obtained, the access relationship and storage space relationship between nodes and data are analyzed, the storage relationship between data and nodes are determined, and distributed storage optimization is performed.
It improves the efficiency of storage after the acquisition of urban planning data, reduces transmission delay, optimizes the user's reading efficiency of data, and realizes efficient data collection and storage.
Smart Images

Figure CN119862241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a data optimization acquisition method and system applied to urban planning. Background Art
[0002] It has become an important means for urban construction planning to collect data related to urban planning for modern urban development planning. The data related to urban planning includes, but is not limited to, basic geographic information data, such as remote sensing image data, terrain elevation data, etc., ecological planning data, business audit data, government affairs data, network open-source data, etc.; by performing fusion analysis on the urban planning data, more comprehensive and accurate information support can be provided for urban development planning.
[0003] Since the data related to urban planning includes a large variety of data types and complex data styles, including but not limited to big data, image data, three-dimensional image data, document reports, and even code data, multiple types of data in the same urban planning usually exist in the same urban planning data, and the data packet size will be relatively large. In the process of collecting urban planning data, its storage process is subject to certain tests. If the data related to urban planning is collected and uniformly stored in the storage center, during the user's call process, affected by transmission delay and storage space, the collection process will reduce the storage efficiency due to the sequential storage of data packets, and at the same time seriously reduce the user's reading efficiency of urban planning data. Therefore, distributed storage can be adopted for the data related to urban planning to achieve data storage; however, during the distributed storage process, the nodes have storage space limitations, and at the same time, due to the different degrees of demand for the data related to each urban planning by different nodes, only based on the collection process and access frequency for storage allocation will ignore the geographical location relationship and access relationship between the urban planning data and each node, and cannot effectively improve the storage efficiency after collection for each node user. Summary of the Invention
[0004] The present invention provides a data optimization acquisition method and system applied to urban planning to solve the problem that the existing urban planning data cannot be effectively stored during the acquisition process due to the huge data volume, which affects the call of relevant data. The specific technical solutions adopted are as follows:
[0005] The present invention proposes a data optimization acquisition method applied to urban planning. The method includes the following steps:
[0006] Obtain the data related to urban planning through the urban planning management platform, and record the type of each relevant data, as well as the geographical location distribution, storage space, and priority of each node in the platform, and obtain the access relationship between each node and each relevant data;
[0007] Analyze the access relationship between each node and each relevant data, and combine the types of each relevant data to determine the access correlation factor of each node to each type of relevant data; based on the acquisition process of each relevant data, analyze the size relationship between each relevant data and the storage space of each node to obtain the storage adaptation factor of each relevant data to each node;
[0008] Analyze the priority of each node and its access correlation factor to each type of relevant data, and based on the geographical location distribution of each node and the storage adaptation factor, determine the storage relationship between each relevant data and each node;
[0009] Based on the storage relationship between each relevant data and each node, perform distributed storage after collecting each relevant data.
[0010] Optionally, the specific method included in analyzing the access relationship between each node and each relevant data, combining the types of each relevant data, and determining the access correlation factor of each node to each type of relevant data is as follows:
[0011] Based on the access frequency of each node to each relevant data per week and the type of each relevant data, obtain the total access frequency of each node to each type of relevant data per week;
[0012] Take the ratio of the total access frequency of any node to any type of relevant data in any week to the sum of the access frequencies of all relevant data of this node in this week as the relative access frequency of this node to this type of relevant data in this week;
[0013] Based on the ratio of the mean value to the standard deviation of the relative access frequency of this node to this type of relevant data per week, and combining the time series change similarity relationship between the total access frequency of this node to this type of relevant data per week and the sum of the access frequencies of all relevant data of this node per week, obtain the access correlation factor of this node to this type of relevant data.
[0014] Optionally, the specific method included in analyzing the size relationship between each relevant data and the storage space of each node based on the acquisition process of each relevant data to obtain the storage adaptation factor of each relevant data to each node is as follows:
[0015] Record the most recently collected relevant data as the current relevant data, and record the type corresponding to the current relevant data as the current type of relevant data; record the ratio of the mean value of the access frequencies of the th node to each relevant data under the current type of relevant data to the mean value of the total access frequencies of all nodes to each relevant data under the current type of relevant data as the access frequency of the
[0016] th node to the current type of relevant data; The ratio of the access frequency of a node to any relevant data under the relevant data of the current type to the total access frequency of all nodes to the relevant data under the relevant data of the current type is used as the access frequency of the node to the relevant data under the relevant data of the current type;
[0017] According to the access frequencies of each node to the relevant data of each type or the relevant data thereunder, combined with the geographical location distribution and storage space of the nodes, obtain the storage adaptation factors of each relevant data for each node.
[0018] Optionally, the specific method included in obtaining the storage adaptation factors of each relevant data for each node according to the access frequencies of each node to the relevant data of each type or the relevant data thereunder, combined with the geographical location distribution and storage space of the nodes is:
[0019] The storage adaptation factor of the current relevant data for the th node is calculated as:
[0020]
[0021] where represents the access frequency of the th node to the relevant data of the current type, represents the geographical location distance between the th node and the node where the current relevant data is collected, represents the data packet size of the current relevant data, represents the th node's storage space size, represents the exponential function with the natural constant as the base;
[0022] The storage adaptation factor of the th relevant data for the th node is calculated as:
[0023]
[0024] where represents the access frequency of the th node to the th relevant data, represents the geographical location distance between the th node and the node where the th relevant data is collected, represents the th relevant data's data packet size.
[0025] Optionally, for analyzing the priorities of each node and the access-related factors of each node to various types of relevant data, based on the geographical location distribution of each node and the storage adaptation factor, to determine the storage relationship between each relevant data and each node, the specific method included is as follows:
[0026] Based on the access-related factors of each node to various types of relevant data, combined with the priorities of each node, obtain the storage priorities of each node for various types of relevant data;
[0027] Based on the storage priorities of each node for various types of relevant data, analyze the delay impact of different nodes on the transmission of relevant data, adjust the storage adaptation factor, and obtain the storage non-essential factors of each relevant data for each node;
[0028] Based on the storage non-essential factors, combined with the size relationship between the relevant data and the storage space of the node, construct the storage model of each node;
[0029] By optimizing the storage model of each node, determine the storage relationship between each relevant data and each node.
[0030] Optionally, for analyzing the delay impact of different nodes on the transmission of relevant data based on the storage priorities of each node for various types of relevant data, adjusting the storage adaptation factor, and obtaining the storage non-essential factors of each relevant data for each node, the specific method included is as follows:
[0031] For any one node and other nodes, based on the storage priorities for the current type of relevant data, obtain the difference in the storage priorities of this node and each node for the current type of relevant data;
[0032] Based on the geographical location distance between different nodes and the packet size of the current relevant data, through the data transmission speed and the signal transmission speed, combined with the difference in the storage priorities of different nodes for the current type of relevant data, obtain the transmission interference factor of the current relevant data between different nodes, and the transmission interference factor of the current relevant data between different nodes has a negative correlation with the difference in the storage priorities;
[0033] Based on the transmission interference factor of the current relevant data between different nodes, adjust the storage adaptation factor of the current relevant data for each node, and obtain the storage non-essential factors of the current relevant data for each node.
[0034] Optionally, for obtaining the storage non-essential factors of the current relevant data for each node, the specific method included is as follows:
[0035] Subtract the storage adaptation factor of the current relevant data for the th node from 1 to get the difference value, and the storage adaptation factor of the current relevant data for the The product of the transmission interference factors of the nodes is used as the storage non-essential factor of the current relevant data for the th node.
[0036] Optionally, based on the storage non-essential factor, combining the size relationship between the relevant data and the storage space of the nodes, constructing the storage model of each node, the specific method included is:
[0037] For the th node, arrange each relevant data in ascending order according to its storage non-essential factor for the th node to form the data call sequence of the th node; the storage model of the th node is as follows:
[0038]
[0039]
[0040] Among them, is the output value of the objective function of the storage model of the th node, represents the storage space occupied when traversing to the th relevant data in the data call sequence of the th node, represents the number of relevant data traversed when traversing to the th relevant data in the data call sequence of the th node, represents the parameter indicating whether the th relevant data in the data call sequence of the th node is stored in the th node, represents the data packet size of the th relevant data in the data call sequence of the th node, represents the storage non-essential factor of the th relevant data for the th node in the data call sequence of the th node, represents the storage space size of the th node; represents the preset storage ratio, is a hyperparameter.
[0041] Optionally, by optimizing the storage model of each node, determining the storage relationship between each relevant data and each node, the specific method included is:
[0042] Arrange the nodes in ascending order according to their own priorities to form a node traversal sequence, and mark the first node in the node traversal sequence as the current node;
[0043] Optimize the objective function of the storage model of the current node, with the goal of maximizing the function output value, to obtain several relevant data that need to be stored in the current node;
[0044] Before the last relevant data stored in the current node in the data call sequence of the current node, all relevant data not stored in the current node is denoted as the un-stored relevant data of the current node;
[0045] During the process of calculating the transmission interference factor of each un-stored relevant data for the current node, the node corresponding to the transmission interference factor determined by each un-stored relevant data is used as the neighboring broadcast node of the current node, and the corresponding un-stored relevant data is stored in the neighboring broadcast node.
[0046] The present invention also proposes a data optimization acquisition system applied to urban planning. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0047] The beneficial effects of the present invention are as follows: By performing distributed storage on the relevant data of urban planning after collection, and considering the broadcast effect between adjacent nodes in the management platform, relevant data is selectively stored according to the access relationship during the storage process of each node, so as to effectively utilize the storage space of each node, improve the calling and reading efficiency of users for the relevant data of urban planning through the nodes, and then ensure the high efficiency of the relevant data for urban planning during the collection process, thereby realizing the optimized collection of urban planning data; among them, by analyzing the temporal changes of the access relationship of nodes to the same type of relevant data, quantifying the access-related factors of nodes to various types of relevant data, and reflecting the importance of nodes in the access relationship of the corresponding type of relevant data, it provides a basis for correcting the priority of storing the corresponding type of relevant data by nodes after collection; at the same time, according to the access relationship of nodes to relevant data, combined with the collection process of relevant data and the size relationship between it and the storage space of nodes, the impact of transmission delay on the relevant data with high-frequency access is reduced, so as to obtain the access-related factors; through the access-related factors, further optimize the storage priority of nodes for various types of relevant data, and analyze the broadcast effect between adjacent nodes, so as to adjust the storage adaptation factor of relevant data for nodes, thereby effectively utilizing the broadcast effect between nodes and the storage space of each node, and then by constructing the storage model of nodes, while ensuring the effective utilization of storage space and storing more necessary relevant data, improve the storage efficiency of relevant data for urban planning after collection, and at the same time reduce the impact of transmission delay on the calling and reading of relevant data in centralized storage or direct distributed storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a schematic flow chart of a data optimization collection method applied to urban planning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0051] Please refer to Figure 1, which shows a flowchart of a data optimization acquisition method applied to urban planning provided by an embodiment of the present invention. The method includes the following steps:
[0052] Step S001: Obtain relevant data for urban planning through the urban planning management platform, record the types of each relevant data, as well as the geographical location distribution, storage space, and priority of each node in the platform, and obtain the access relationship between each node and each relevant data.
[0053] The purpose of this embodiment is to comprehensively consider the geographical location distribution, storage space, and priority of nodes in the platform during the process of collecting and storing relevant data for urban planning, so as to make full use of the broadcast effect between adjacent nodes in the platform, that is, the rapid call between data packets stored in adjacent nodes; and combined with the access relationship between nodes and relevant data, effectively utilize the storage space of each node, and finally achieve the high efficiency of storing urban planning data after collection.
[0054] Specifically, obtain various types of relevant data for urban planning through the urban planning management platform, including basic geographic information data, ecological planning data, business audit data, government affairs data, and network open-source data, etc. Each relevant data is collected through each node in the platform, and at the same time, record the type to which each relevant data belongs and the size of its data packet.
[0055] Furthermore, the urban planning management platform is constructed based on a digital twin model. The storage center is used as the central node in the model to store all relevant data for urban planning. The central node and other each node are mapped into the digital twin model according to the real geographical location distribution, then the geographical distance between each node and the central node can be obtained. At the same time, during the model construction process, each node corresponds to the government functional departments of different cities, and the priority is assigned according to the government affairs level of each node (the higher the administrative level of the city, the higher the priority), and at the same time, record the size of the storage space of the server corresponding to each node.
[0056] Furthermore, since the urban planning management platform is established, each time a node accesses each relevant data, a record is made, and the access relationship between each node and each relevant data is obtained at one time.
[0057] Step S002: Analyze the access relationship between each node and each relevant data, combine the types of each relevant data, and determine the access correlation factor of each node to each type of relevant data; according to the collection process of each relevant data, analyze the size relationship between each relevant data and the storage space of each node, and obtain the storage adaptation factor of each relevant data for each node.
[0058] Preferably, in an embodiment of the present invention, by analyzing the access relationships between each node and each relevant data, and combining the types of each relevant data, access correlation factors of each node for each type of relevant data are determined. The specific method includes:
[0059] It should be noted that there are differences in the access relationships of different nodes to relevant data, which are specifically reflected in the differences in the variation of access frequencies in terms of time sequence for different types of relevant data. By performing time sequence analysis on the access of each node to historical relevant data of various types, the access correlation relationship between the node and the relevant data type is quantified. Thus, in the subsequent process of selective storage according to priorities, further limitations are imposed on the basis of priorities to ensure that the relevant data for urban planning stored in the subsequent distributed storage process is more compatible with the access and reading of nodes.
[0060] Specifically, for any node, obtain the access frequency of the node to each relevant data per week since the establishment of the urban planning management platform; it should be especially noted that if it is less than one week as of the current time, the remaining nearest few days after screening each week are used as one week for subsequent processing; the sum of the access frequencies of the node to the same type of relevant data within the same week is used as the total access frequency of the node to this type of relevant data in this week; the ratio of the total access frequency to the sum of the access frequencies of the node to all relevant data in this week is used as the relative access frequency of the node to this type of relevant data in this week; then the th node's access correlation factor for the th type of relevant data is calculated as follows:
[0061]
[0062] wherein, represents the mean value of the relative access frequencies of the th node to the th type of relevant data for all weeks, represents the standard deviation of the relative access frequencies of the th node to the th type of relevant data for all weeks, represents the sequence formed by the total access frequencies of the th node to the th type of relevant data per week in time sequence, represents the sequence formed by the sum of the access frequencies of the th node to all relevant data per week in time sequence, represents the cosine similarity obtained by calculating the two sequences, To avoid the hyperparameter of the denominator being zero, in this embodiment, is used for description.
[0063] It should be noted that the access relative frequency reflects the proportion of the access frequency of each week of a node to the relevant data of the corresponding type in the overall situation. The larger its mean value indicates that the proportion of each week is relatively large. At the same time, the smaller the standard deviation, the more similar the proportion of the access frequency between each week is, and overall, it tends to be that the node accesses the relevant data of this type more frequently; at the same time, the closer the change relationship between the total access frequency and the sum of the access frequencies of each week is, that is, the larger the cosine similarity is, the more easily the change of the access relationship of this node to the relevant data is affected by the relevant data of this type, and the larger the access correlation factor between the two is.
[0064] Preferably, in an embodiment of the present invention, according to the acquisition process of each relevant data, analyze the size relationship between each relevant data and the storage space of each node, and obtain the storage adaptation factor of each relevant data for each node. The specific method included is as follows:
[0065] It should be noted that the storage space in the servers of each node is limited. At the same time, affected by the geographical location distribution of the nodes, the relevant data of urban planning is collected through the nodes. Then, when storing after collection, it is necessary to give priority to considering the acquisition process of the relevant data by the nodes to avoid a large amount of transmission time generated by overall transmission and then storage in the storage center, which will affect the acquisition efficiency; at the same time, since the relevant data is collected through the corresponding nodes, the access frequency of the corresponding node to this relevant data will be relatively large. If it is stored in the storage center, it will cause the subsequent high-frequency call of the relevant data to be affected by the transmission delay, seriously reducing the call and reading efficiency of the relevant data.
[0066] Furthermore, it should be noted that based on the acquisition process, the larger the proportion of the access frequency of a node to the relevant data of the same type in the total access frequency of the relevant data of the corresponding type, the greater the possibility that the corresponding relevant data is frequently accessed by the corresponding node; at the same time, consider the size relationship between the data packet size of the relevant data and the storage space of the node server to ensure that the high-frequency call is less affected by the transmission delay and effectively utilize the storage space for distributed storage.
[0067] Specifically, record the recently acquired relevant data as the current relevant data, and record the type corresponding to the current relevant data as the current type of relevant data; take the ratio of the mean value of the access frequency of the th node to each relevant data under the current type of relevant data to the mean value of the total access frequency of all nodes to each relevant data under the current type of relevant data as the access frequency of the th node to the current type of relevant data; take the ratio of the access frequency of the th node to any one relevant data under the current type of relevant data to the total access frequency of all nodes to this relevant data under the current type of relevant data as the access frequency of the th node to this relevant data under the current type of relevant data.
[0068] It should be noted that for the current relevant data, which has just been collected and has not been accessed by any node, it is necessary to analyze based on the access frequency of the relevant data of the corresponding type; while for the relevant data obtained from the platform, the nodes where the data is collected and the access relationship of each node to it have been recorded, then during the calculation of the storage adaptation factor, it is directly quantified through its access frequency.
[0069] Specifically, for the storage adaptation factor of the current relevant data for the th node the calculation method is as follows:
[0070]
[0071] Among them, represents the access frequency of the th node to the relevant data of the current type, represents the geographical location distance between the th node and the node where the current relevant data is collected, represents the data packet size of the current relevant data, represents the th node's storage space size, represents the exponential function with the natural constant as the base; it should be noted that if the node where the current relevant data is collected is the th node, then the geographical location distance is actually 0, but it is not calculated with 0, and the part is directly replaced with 1, and the following formula takes the same operation.
[0072] Similarly, for the storage adaptation factor of the th relevant data for the th node the calculation method is as follows:
[0073]
[0074] Among them, represents the access frequency of the th node to the th relevant data. It should be noted that the th relevant data is not necessarily the relevant data of the current type, it can be relevant data of any type, and the access frequency of the node to it is obtained according to the above method; represents the geographical location distance between the th node and the node where the th relevant data is collected, represents the th relevant data's data packet size.
[0075] It should be noted that if there is no access relationship between the newly collected relevant data and the nodes, then ratio analysis is carried out based on the access frequency of the relevant data of the same type at the nodes, while other relevant data is directly analyzed based on the overall ratio of the access frequency of the nodes to it. The greater the ratio of the access frequency of the nodes to the overall access frequency of the relevant data, the more necessary it is to store the relevant data in the corresponding nodes for subsequent call and reading. At the same time, the greater the geographical distribution distance between the collection node of the relevant data and the node, the more it should be stored in the node in the case of frequent calls to reduce the impact of transmission delay. If the node is the collection node, the necessity of storing the relevant data reaches the maximum, that is, the corresponding parameter is adjusted to 1. And the smaller the storage space occupied by the relevant data packet at the node, the smaller the burden on the node server to store the relevant data, and storing it under relatively frequent access will not interfere with the normal operation of the node server.
[0076] So far, by analyzing the temporal changes in the access relationship of the nodes to the relevant data of the same type, quantifying the access-related factors of the nodes to various types of relevant data, and reflecting the importance of the nodes in the access relationship of the corresponding types of relevant data, it provides a basis for correcting the priority of storing the corresponding types of relevant data after node collection. At the same time, based on the access relationship of the nodes to the relevant data, combined with the collection process of the relevant data and the size relationship between it and the node storage space, the impact of transmission delay on the relevant data with high-frequency access is reduced, thereby obtaining the access-related factors.
[0077] Step S003: Analyze the priority of each node and its access-related factors for each type of relevant data, and based on the geographical location distribution of each node and the storage adaptation factor, determine the storage relationship between each relevant data and each node.
[0078] Preferably, in an embodiment of the present invention, the specific method included in this step is:
[0079] Based on the access-related factors of each node to each type of relevant data, combined with the priority of each node, obtain the storage priority of each node for each type of relevant data;
[0080] Based on the storage priority of each node for each type of relevant data, analyze the impact of different nodes on the transmission delay of the relevant data, adjust the storage adaptation factor, and obtain the storage non-necessary factor of each relevant data for each node;
[0081] Based on the storage non-necessary factor, combined with the size relationship between the relevant data and the storage space of the node, construct the storage model of each node;
[0082] By optimizing the storage model of each node, determine the storage relationship between each relevant data and each node.
[0083] It should be noted that the priority of a node itself is the priority determined by administrative, geographical and other factors for the city where the node is located in the urban planning management platform. For example, the priority of a provincial administrative unit is higher than that of a municipal administrative unit, and the relevant data it stores and calls will include more urban planning-related data of cities compared to the actual situation. On this basis, there will be differences in the types of relevant data collected by each node itself, so it is necessary to access relevant factors to correct and obtain the storage priority of the node for type-related data.
[0084] As an example, based on the access-related factors of each node to each type of relevant data and combining the priorities of each node, the storage priorities of each node for each type of relevant data are obtained. The specific methods include:
[0085] Specifically, taking the th node and the th type of relevant data as an example, the higher the node priority, the larger the priority value. Then, multiply the priority of the th node by the access-related factor of the th node for the th type of relevant data, and use the result as the storage priority of the th node for the th type of relevant data.
[0086] It should be noted that after determining the storage priority, it is necessary to adjust the storage adaptation factor based on the storage priority and the geographical location distribution among the nodes. That is, if there is a node with a higher storage priority for the same type within a relatively short distance of a node, the corresponding relevant data can be stored in the node with a higher storage priority, effectively utilizing the storage space of the node.
[0087] As an example, based on the storage priorities of each node for each type of relevant data, analyze the time delay impact of different nodes on relevant data transmission, and adjust the storage adaptation factor to obtain the storage non-essential factors of each relevant data for each node. The specific methods include:
[0088] Specifically, taking the current relevant data and its corresponding current type of relevant data as an example, obtain the storage priority of each node for the current type of relevant data, calculate the several differences obtained by subtracting the storage priority of the th node for the current type of relevant data from the storage priorities of other each node for the current type of relevant data, and perform linear normalization on all the differences. The result obtained is used as the difference in the storage priority of the th node and the corresponding node for the current type of relevant data.
[0089] Furthermore, the data transmission speed between each node in the urban planning management platform is known, and the signal transmission speed is known. Then, for the current relevant data, for the The calculation method of the transmission interference factor of a node is as follows:
[0090]
[0091] Among them, represents the transmission interference factor between the current relevant data between the th node and the th node, represents the storage priority difference between the th node and the th node for the current type of relevant data, represents the data transmission speed, represents the signal transmission speed, represents the packet size of the current relevant data, represents the th node and the th node's geographical location distance, represents the maximum value of the packet sizes of all relevant data, represents the maximum value of the geographical location distances between all nodes; adding 1 to the denominator aims to avoid the denominator being 0 and affecting the calculation result; according to the above method, obtain the transmission interference factor between the current relevant data between the th node and other each node, and take the minimum value among them as the transmission interference factor of the current relevant data for the th node.
[0092] Furthermore, multiply the difference obtained by subtracting the storage adaptation factor of the current relevant data for the th node from 1 by the transmission interference factor of the current relevant data for the th node, and take it as the storage non-necessary factor of the current relevant data for the th node.
[0093] It should be noted that the larger the storage adaptation factor, the more necessary it is to store at the node, and the smaller the storage non-necessary factor; and if it is not stored at the corresponding node, the broadcast effect between nodes needs to be considered, that is, the data transmission between neighboring nodes. Then, the greater the storage priority difference, the higher the storage priority of the corresponding node for the current type of relevant data, and the stronger the broadcast effect when the transmission delay is smaller, and the corresponding transmission interference factor is smaller. During the transmission process, it is necessary to send a signal first to determine the transmission and then perform the transmission of the data packet. The transmission delay is obtained comprehensively, and the smaller the transmission interference factor, the more the corresponding node can rely on the broadcast effect of neighboring nodes to call and read relevant data, thereby reducing the necessity of storage and increasing the storage non-necessary factor.
[0094] It should be noted that the storage non-essential factor reflects the necessity of storing relevant data at the corresponding node. Then, the relevant data is traversed one by one in ascending order of the storage non-essential factor, and at the same time, the proportion of the already stored relevant data in the node storage space is judged to construct a storage model.
[0095] As an example, based on the storage non-essential factor and combining the size relationship between the relevant data and the storage space of the node, a storage model for each node is constructed. The specific method includes:
[0096] Specifically, for the th node, the relevant data is arranged in ascending order of its storage non-essential factor for the th node to form the data call sequence of the th node; the storage model of the th node is as follows:
[0097]
[0098]
[0099] Among them, is the output value of the objective function of the storage model of the th node, represents the storage space occupied when traversing to the th relevant data in the data call sequence of the th node, represents the number of relevant data traversed when traversing to the th relevant data in the data call sequence of the th node, represents a parameter indicating whether the th relevant data in the data call sequence of the th node is stored in the th node. If so, , otherwise ; represents the data packet size of the th relevant data in the data call sequence of the th node, represents the storage non-essential factor of the th relevant data in the data call sequence of the th node for the th node, represents the storage space size of the th node; represents a preset storage ratio, which is described using in this embodiment; To avoid hyperparameters with a denominator of 0, this embodiment uses for description.
[0100] Furthermore, according to the above method, the objective function of the storage model is obtained for each node.
[0101] It should be noted that under the objective function of the node, it is necessary to ensure that the storage necessity of the relevant data stored in the corresponding node is large enough, and at the same time make full use of the storage space. Therefore, to ensure the normal operation of the node server, a preset storage ratio is set to construct the objective function of the storage model and perform subsequent optimization.
[0102] It should be further noted that after constructing the storage model, it is necessary to traverse according to the priority of the node itself, and optimize the storage model one by one, so as to finally obtain the storage relationship between all relevant data and each node.
[0103] As an example, by optimizing the storage model of each node, the storage relationship between each relevant data and each node is determined. The specific method includes:
[0104] Specifically, the nodes are arranged in ascending order according to their own priorities to form a node traversal sequence. The first node in the node traversal sequence is denoted as the current node. Then, the objective function of the storage model of the current node is optimized. The optimization objective is to maximize the function output value, so as to obtain several relevant data that need to be stored in the current node. At the same time, all relevant data that are not stored in the current node before the last relevant data stored in the data call sequence of the current node are denoted as the un-stored relevant data of the current node. During the calculation of the transmission interference factor of each un-stored relevant data for the current node, the node corresponding to the transmission interference factor determined by each un-stored relevant data (that is, the transmission interference factor is determined by the minimum value of the transmission interference factors of other nodes, and the minimum value corresponds to a node) is used as the neighboring broadcast node of the current node, and the corresponding un-stored relevant data is stored in the neighboring broadcast node.
[0105] Furthermore, the objective function of the storage model of the next node in the node traversal sequence is optimized. If the node is the neighboring broadcast node of the previous node in the sequence, then in the objective function of the storage model and its corresponding data call sequence, the un-stored relevant data of the previous node stored in the neighboring broadcast node is no longer calculated, that is, the corresponding un-stored relevant data is fixedly stored in this node, and the situation of not storing this un-stored relevant data is not considered during the optimization of the storage model, and the calculation of the non-storage necessary factor needs to be involved.
[0106] And so on, finally obtaining the relevant data that need to be stored in each node, and thus obtaining the storage relationship between each relevant data and each node.
[0107] Furthermore, after new relevant data is collected, the storage non-essential factors of the new relevant data for each node are obtained according to the above method. Starting from the node with the lowest storage priority for the relevant data of the type to which the new relevant data belongs, on the premise that the storage model objective function of this node has reached optimization, it is judged whether adding the new relevant data will increase the objective function. If it increases, the data is stored; otherwise, it is no longer stored. And so on, the storage relationship of the new relevant data is judged for each node to obtain the storage relationship.
[0108] So far, by accessing relevant factors, the storage priorities of each type of relevant data for nodes are further optimized, and the broadcast effect between neighboring nodes is analyzed, thereby adjusting the storage adaptation factors of relevant data for nodes, so as to effectively utilize the broadcast effect between nodes and the storage space of each node. Furthermore, by constructing the storage model of nodes, while ensuring the effective utilization of storage space and storing more necessary relevant data, the storage efficiency after the collection of relevant data for urban planning is improved, and at the same time, the transmission delay impact during the call and reading of relevant data by centralized storage or direct distributed storage is reduced.
[0109] Step S004: Based on the storage relationship between each relevant data and each node, after collecting each relevant data, perform distributed storage to achieve the optimization of data collection for urban planning.
[0110] It should be noted that each relevant data is collected through different nodes during the collection process. Through the analysis of the access relationship between each type of relevant data and each node, the storage relationship between each relevant data and each node is obtained, which can comprehensively consider the occupation of the storage space of relevant data for nodes and the transmission delay impact during the access process of relevant data by nodes after collection, ensuring the storage efficiency after collection and improving the call and reading efficiency of each node for relevant data at the same time.
[0111] Specifically, after any node in the platform collects any relevant data for urban planning, the storage relationship between it and each node in the platform is quantified, that is, whether it needs to be stored in the corresponding node. If it needs to be stored, the relevant data is transmitted and stored in the corresponding node server; if it does not need to be stored, the corresponding node does not store the relevant data, so as to achieve distributed storage of each relevant data after collection, and then optimize the collection process of relevant data for urban planning.
[0112] So far, by performing distributed storage on the relevant data of urban planning after collection, and considering the broadcast effect between adjacent nodes in the management platform, relevant data is selectively stored according to the access relationship during the storage process of each node, so as to effectively utilize the storage space of each node, improve the calling and reading efficiency of users for the relevant data of urban planning through the nodes, and then ensure the high efficiency of the relevant data for urban planning during the collection process, thereby realizing the optimized collection of urban planning data.
[0113] It should be noted that this embodiment uses a model to present the inverse proportional relationship and normalization processing. is an exponential function with the natural constant as the base. is the input of the model, and the implementer can set the inverse proportional function and the normalization function according to the actual situation.
[0114] Another embodiment of the present invention provides a data optimization collection system applied to urban planning. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the above method steps S001 to step S004 are implemented.
[0115] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data optimization collection method for urban planning, characterized in that: The method comprises the following steps: Obtain relevant data for urban planning through the urban planning management platform, record the type of each relevant data, the geographical location distribution, storage space and priority of each node in the platform, and obtain the access relationship between each node and each relevant data; Analyze the access relationship between each node and each related data, and determine the access correlation factor of each node to each type of related data in combination with the type of each related data; analyze the size relationship between each related data and the storage space of each node according to the collection process of each related data, and obtain the storage adaptation factor of each related data for each node; Analyze the priority of each node and its access-related factors to each type of related data, and determine the storage relationship between each related data and each node based on the geographical location distribution of each node and the storage adaptability factor; Based on the storage relationship between each relevant data and each node, each relevant data is collected and distributedly stored; The analysis of the priority of each node and its access-related factors to each type of related data determines the storage relationship between each related data and each node based on the geographical location distribution of each node and the storage adaptability factor, including the specific method of: Based on the access-related factors of each node to each type of related data and the priority of each node, the storage priority of each node to each type of related data is obtained; According to the storage priority of each node for each type of related data, the delay influence of different nodes on the transmission of related data is analyzed, and the storage adaptation factor is adjusted to obtain the storage non-essential factor of each related data for each node; Based on the storage non-essential factors and in combination with the size relationship between the relevant data and the storage space of the node, a storage model of each node is constructed; By optimizing the storage model of each node, the storage relationship between each related data and each node is determined. The specific method is as follows: Arrange the nodes in ascending order according to their priority to form a node traversal sequence. The first node in the node traversal sequence is recorded as the current node. Optimize the objective function of the storage model of the current node, the optimization goal is to maximize the function output value, and obtain a number of related data that need to be stored in the current node; All relevant data not stored in the current node before the last relevant data stored in the current node in the data call sequence of the current node are recorded as unstored relevant data of the current node; In the process of calculating the transmission interference factor of each unstored related data for the current node, the node corresponding to the transmission interference factor determined by each unstored related data is used as the neighboring broadcast node of the current node, and the corresponding unstored related data is stored in the neighboring broadcast node.
2. The data optimization collection method for urban planning according to claim 1, characterized in that: The specific method of analyzing the access relationship between each node and each related data and determining the access correlation factor of each node to each type of related data in combination with the type of each related data is as follows: According to the access frequency of each node to each relevant data in each week and the type of each relevant data, the total access frequency of each node to each type of relevant data in each week is obtained; The ratio of the total access frequency of any node to any type of related data in any week to the sum of the access frequencies of the node to all related data in that week is taken as the relative frequency of the node's access to the related data of that type in that week; Based on the ratio of the mean to the standard deviation of the relative frequency of the node's access to this type of related data each week, combined with the similarity relationship between the total frequency of the node's access to this type of related data each week and the sum of the node's access frequencies to all related data each week, the node's access correlation factor for this type of related data is obtained.
3. The data optimization collection method for urban planning according to claim 2 is characterized in that: According to the collection process of each relevant data, the size relationship between each relevant data and the storage space of each node is analyzed to obtain the storage adaptation factor of each relevant data for each node, including the specific method of: The most recently collected relevant data is recorded as the current relevant data, and the type corresponding to the current relevant data is recorded as the current type relevant data; The ratio of the mean access frequency of each node to each related data under the current type of related data to the mean total access frequency of all nodes to each related data under the current type of related data is taken as the first The access frequency of each node to the data related to the current type; The first The ratio of the access frequency of a node to any related data under the current type of related data to the total access frequency of all nodes to the related data under the current type of related data is taken as the first The access frequency of each node to the related data under the current type of related data; According to the access frequency of each node to each type of related data or its related data, combined with the geographical location distribution and storage space of the nodes, the storage adaptation factor of each related data for each node is obtained.
4. The data optimization collection method for urban planning according to claim 3 is characterized in that: The method of obtaining the storage adaptation factor of each node for each related data according to the access frequency of each node to each type of related data or the related data thereunder, combined with the geographical location distribution and storage space of the node, includes the following specific methods: The current relevant data for Storage Adaptability Factor of Nodes The calculation method is: in, Indicates The access frequency of nodes to data related to the current type, Indicates The geographical distance between each node and the node where the current relevant data is collected, Indicates the current data packet size of the relevant data, Indicates The storage space size of each node, represents an exponential function with a natural constant as base; No. The relevant data for Storage Adaptability Factor of Nodes The calculation method is: in, Indicates Node pair The access frequency of relevant data, Indicates The node and The geographical location distance of the nodes where the relevant data is collected, Indicates The data packet size of the relevant data.
5. The data optimization collection method for urban planning according to claim 1 is characterized in that: The method of analyzing the delay influence of different nodes on the transmission of relevant data according to the storage priority of each node for each type of relevant data, adjusting the storage adaptation factor, and obtaining the storage non-essential factor of each relevant data for each node includes the following specific methods: For any node and other nodes, based on the storage priority of the current type of related data, the storage priority difference between the node and other nodes in the current type of related data is obtained; Based on the geographical location distance between different nodes and the data packet size of the current relevant data, the transmission interference factor of the current relevant data between different nodes is obtained by data transmission speed and signal transmission speed, combined with the storage priority difference of the current type of relevant data of different nodes, and the transmission interference factor of the current relevant data between different nodes is negatively correlated with the storage priority difference; Based on the transmission interference factor of the current relevant data between different nodes, the storage adaptation factor of the current relevant data for each node is adjusted to obtain the storage non-essential factor of the current relevant data for each node.
6. The data optimization collection method for urban planning according to claim 5, characterized in that: The specific method of obtaining the non-essential factors of the storage of the current relevant data for each node includes: Subtract 1 from the current relevant data for the The difference between the storage adaptation factor of the node and the current relevant data for the The product of the transmission interference factors of the nodes is used as the current relevant data for the The storage of each node is not a necessary factor.
7. The data optimization collection method for urban planning according to claim 1, characterized in that: The storage model of each node is constructed based on the non-essential storage factor and in combination with the size relationship between the relevant data and the storage space of the node, and the specific method includes: For nodes, and sort the relevant data according to their The storage non-essential factors of the nodes are arranged from small to large to form the The data call sequence of nodes; The storage model of each node is as follows: in, For the The output value of the storage model objective function of each node, Indicates The data call sequence of nodes is traversed to the The storage space occupied by the relevant data is Indicates The data call sequence of nodes is traversed to the The number of related data traversed when traversing related data, Indicates The first node in the data call sequence Is the relevant data stored in The parameters in the node, Indicates The first node in the data call sequence The data packet size of the relevant data, Indicates The first node in the data call sequence The relevant data for The storage of nodes is not a necessary factor. Indicates The storage space size of each node; Indicates the preset storage ratio. is a hyperparameter.
8. A data optimization collection system for urban planning, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of a data optimization collection method applied to urban planning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Distributed file system storage optimization energy saving method based on greedy firefly algorithm
CN106547854A
Distributed storage method and system for clinical test data
CN119323046A