Distributed Receiving and Storing Method and System for Data Records of Fund Assets
By calculating the thermal strategy coefficient of the label and identifying the thermal system tags, and dynamically adjusting the data storage strategy, the thermal sharding problem under high load situations in distributed systems is solved, and system performance and efficiency are improved.
Patent Information
- Application Number
- CN202510083577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art is difficult to effectively deal with the thermal sharding phenomenon under high load situations in distributed systems, resulting in some nodes becoming performance bottlenecks and affecting the overall performance of the system.
By configuring server nodes and data tags, calculating the label thermal strategy coefficient, identifying the thermal tags, and recording and storing data in combination with the thermal tag judgment results, dynamically adjusting the data storage strategy to reduce the risk of thermal sharding.
Effectively quantify and reduce the risk of hot sharding, improve the critical distinction ability of server nodes to store content and data content, reduce the risk of performance inefficiency, and improve the service performance and efficiency of fund asset data.
Smart Images

Figure CN119520549B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed storage, and particularly relates to a method and system for distributed receiving and storage of data records of fund assets. Background Art
[0002] In traditional financial markets, the data records of fund assets often rely on centralized data processing systems, which often have the risks of single-point failures in data storage and data leakage caused by cyberattacks. Therefore, DLT distributed ledger technology and blockchain technology have emerged. At the same time, these technologies have gradually been confronted with the conflict between effectively expanding nodes in a distributed system to cope with the sudden increase in the amount of fund asset data while maintaining the system complexity at a relatively stable state. Existing technologies often use data sharding mechanisms to solve this trade-off competition. By horizontally splitting data into different nodes, the scalability of the distributed system is improved, which is an important means to cope with massive data processing. However, sharding is usually carried out according to certain rules, such as hashing or range sharding according to fund IDs. This static partitioning does not consider the dynamic changes in data access. The data access in the fund market is highly dynamic and uncertain. The trading volume of some funds may surge during a specific period, resulting in a rapid increase in the load of these shards, forming the phenomenon of hot shards. That is, in a high-load situation in a distributed system, due to the uneven load distribution of data shards, some shards bear significantly higher access requests than other shards, leading to the problem of prominent performance bottlenecks in related nodes, especially at present, high-frequency trading accounts for an increasing proportion of the daily trading volume, which will further lead to an increase in request processing latency and even node downtime, directly affecting the overall performance of the system, especially the response speed, which is quite important in the processing of fund asset data.
[0003] However, currently, the phenomenon of hot shards in a high-load situation in a distributed system is generally reduced by server load balancing algorithms, ignoring the phenomenon of hot shards caused by the cumulative imbalance of the key content of each node itself. The key imbalance effect occurs because the data carried by different data tags or shards has different importance in business. For example, the asset data of high-frequency trading and popular funds may be read and written more frequently or have more sudden characteristics than other data, resulting in the load generated by them being balanced overall in time observation. However, due to the high business impact of these contents, some nodes are still more likely to become the bottleneck of the distributed storage system; in addition, traditional hashing or range sharding strategies usually only consider the distribution of data and cannot dynamically respond to changes in market demand. Therefore, when the statically partitioned data has a sudden increase in load, it is impossible to consider the sudden high demand of the business, resulting in overloading of the shards of specific data. Summary of the Invention
[0004] The object of the present invention is to provide a method and system for distributed reception and storage of data records of fund assets, so as to solve one or more technical problems existing in the prior art, and at least provide a beneficial option or create conditions.
[0005] To achieve the above object, according to one aspect of the present invention, there is provided a method for distributed reception and storage of data records of fund assets, the method comprising the following steps:
[0006] Configure server nodes and data tags in the data storage environment; obtain the tag status of the data tags from the server nodes; calculate the tag heat policy coefficient according to the tag status, and use the tag heat policy coefficient to identify the hot system tags from each server node; combine the determination result of the hot system tags to perform data record storage;
[0007] Further, the method for identifying each server node and data tag in the data storage environment is: the data storage environment includes several server nodes, simply referred to as nodes;
[0008] The data records of fund assets are stored in each node through a data sharding mechanism. The data records are composed of several data tags. The several data tags of the data records are stored in different nodes respectively, and each node stores the fund asset data of each data tag.
[0009] Among them, the data types of fund asset data tags are each shard processed by a regularized data sharding algorithm; or the data tags are fund types, fund names, fund managers, fund unit net values, dividend information, net asset values, investment groups, etc.
[0010] Further, the method for obtaining the tag status of the data tags from the server nodes is: preset a monitoring interval rg, rg ∈ [0.5, 24] hours, and its default value is 1 hour; obtain the tag status of the data tags every rg. Among them, the tag status of the data tag is the product of the operation order value and the tag pressure. The calculation method of the operation order value is: the average value of the CPU occupancy rate of the server node within the monitoring interval is recorded as the occupancy amount, and the minimum value among the occupancy amounts of all nodes is recorded as the occupancy base value. Then, the difference between the occupancy amount of any node and the occupancy base value is used as the operation order value; the calculation method of the tag pressure is: in a node, the number of read and write operations on the same data tag is used as the unit read and write amount, and the numerical set composed of the corresponding read and write amounts of each data tag under the node is subjected to minmax normalization processing. The normalized value of the unit read and write amount of the data tag is defined as its tag pressure.
[0011] Further, the method for calculating the tag heat policy coefficient according to the tag status is: use any data tag as the determination tag and a preset time interval as the monitoring interval;
[0012] Alternatively, for the determination tag, the moment when the tag state value is zero is taken as the zero point position; starting from the current moment, any node searches in the reverse time direction to obtain the first zero point position as the current zero point position of the node, and the number of moments between the current moment and the current zero point position is taken as the zero point span of the node; in the set composed of the zero point spans of each node, the median value Zps.mid, the interquartile range Zps.ir, and the maximum value Zps.mx of the set are respectively obtained, and a zero point interval Zps_rg is constructed, Zps_rg = min{(Zps.mid + Zps.ir), Zps.mx}; the time period from the current moment to Zps_rg in its reverse time direction is taken as the monitoring interval.
[0013] Identify the thermal system tag according to the tag state of the determination tag within the monitoring interval.
[0014] Furthermore, the method for identifying the thermal system tag according to the tag state of the determination tag within the monitoring interval is: within the monitoring interval, the minimum values of the tag states at different moments of the same node and at different nodes at the same moment are respectively obtained, and are respectively denoted as the first state quasi-value and the second state quasi-value; define the difference between the tag state and the first state quasi-value as the quasi-node differential degree SLbv.
[0015] For any node, if the first state quasi-value and the second state quasi-value corresponding to any non-zero tag state are both zero, then the moment corresponding to the tag state is denoted as the state valley point.
[0016] Among them, the role of the state valley point can be regarded as a key feature on the tag state curve. By analyzing these state valley points, the local fluctuations of the load change in the time series can be directly captured, especially those sudden sharp drops. Therefore, it can play a role similar to a warning line in the time series, marking the inflection points of the low-value data, and the inflection points, as the basis of the relative steady state, provide a more scientific measurement standard for quantifying the change trajectory of the system load.
[0017] Construct all the state valley points at all moments into a set Fn, the number of elements in the set Fn is denoted as Sfn, and the difference between the tag state of any state valley point and the minimum tag state in the set Fn is denoted as the node differential degree of the state valley point.
[0018] The principle of constructing the state valley points into a set to calculate the quasi-node differential degree is that combining the difference degree of the time series can enable the algorithm to quickly analyze the key fluctuations of the load with relatively low computational complexity, so that the system is more sensitive to sudden changes in the load.
[0019] Take the label state at any moment as the valley point threshold. The ratio of the number of elements in the set Fn with label states less than the valley point threshold to the number of elements with label states greater than the valley point threshold is denoted as the state evolution degree Eff at this moment. If the label states of all elements in the set Fn are greater than the valley point threshold or all less than the valley point threshold, then the state evolution degree at this moment is 1 / Sfn; The label thermal strategy coefficient Ltsc calculated according to the state valley point is calculated as follows:
[0020] ;
[0021] where j1 is the cumulative variable, SLbv j1 and Lbv j1 are the quasi-node differential degree and node differential degree of the j1-th state valley point respectively. exp() represents the exponential function with the natural constant e as the base, max() is the function to obtain the maximum value, and lg() is the logarithm function with base 10;
[0022] Since the label thermal strategy coefficient calculated using the node differential degree focuses on the identification of basic state differences, it can quickly quantify the thermal fragmentation risk caused by the cumulative imbalance of the key content of each node in real time. However, because this method is highly dependent on the original data, it will lead to the overfitting problem of quantitative analysis, especially when there are a large number of rapidly rising label states that occur frequently. However, the existing technologies cannot solve the overfitting problem induced by the strong dependence of the real-time quantification process on the original data. To better solve this problem and eliminate the phenomenon of data distortion caused by overfitting, the present invention proposes a more preferable solution as follows:
[0023] Preferably, the method of identifying the thermal system label according to the label state of the determined label within the monitoring interval can be replaced by: within the monitoring interval, if the label state value at a moment is greater than that of its previous moment and the label state value of its previous moment is not zero, then define this moment as the state increment position; The average value of each state increment position is denoted as the state phase value. If the label state at any moment is greater than the state phase value, it is determined that it satisfies the first phase value condition;
[0024] Take any state increment position and its corresponding label state to form an increment binary group. Use the plrep and splev functions to perform cubic spline interpolation on the set constructed by the increment binary group to obtain a smooth curve. The difference between the curve value of any non-state increment position and the label state is the increment estimation amount. When the increment estimation amount is less than zero, the non-state increment position has an inefficient estimation, otherwise it has an effective estimation;
[0025] In the fund asset management system, high-frequency trading may lead to uneven data acquisition density. Therefore, a fitting curve constructed by effective state values can be used to complement the missing density data points. Especially in the scenario where the load suddenly increases, the fitting curve can accurately estimate the load peak through the fluctuation trend of the effective states of the front and back tags.
[0026] Search in the reverse time direction from any effective state increment position to obtain the first non-consecutive effective state increment position as its reference position, and use the time period between this effective state increment position and its reference position as an estimation reference interval; exclude the estimation reference interval containing the zero position.
[0027] Calculate the strategy sensitivity value Psenv of the estimation reference interval according to the increment estimation amount: Psenv = Len × lg(1 + eff.rp × eff.dsv); where eff.rp represents the proportion of the moments with effective estimations among all moments between the estimation reference interval and the first zero position in its reverse time direction, eff.dsv is the maximum value of the effective state of the increment estimation amount label at each moment with effective estimations within the estimation reference interval; Len is the proportion of the number of moments in the estimation reference interval to the total number of moments in the monitoring interval, and lg is the logarithm symbol with base 10.
[0028] If the number of times of effective estimations in the estimation reference interval is greater than the number of times of inefficient estimations, it is considered that this estimation reference interval meets the second effective state condition; both the first effective state condition and the second effective state condition are used as effective state conditions, and the number of current moments that meet the effective state conditions is recorded as the effective state condition achievement value Glv.
[0029] Calculate the label heat strategy coefficient Ltsc according to each estimation reference interval:
[0030] ;
[0031] where e is the natural constant, i1 is the serial number of the estimation reference interval, MPs represents the average value of the strategy sensitivity values corresponding to all estimation reference intervals, noz is the number of estimation reference intervals, avg{} is the average value function, and Psenv i1 is the strategy sensitivity value of the i1-th estimation reference interval, exp() is the exponential function with base e as the natural constant, rk() is the sorting coefficient function, and the return value obtained through the sorting coefficient function rk(i1) is: obtain the maximum value of the label effective state within the i1-th estimation reference interval, and use the percentile value of the obtained maximum value among all label effective states in the monitoring interval as the return value of the sorting coefficient function.
[0032] Beneficial effects: Since the label heat policy coefficient is calculated based on the estimated reference interval, it can accurately mark and downgrade the positions where load pressure collapse occurs, thereby forming a phased defect analysis of the thermal risk accumulation effect. Therefore, it can effectively quantify the thermal sharding risk caused by the cumulative imbalance of content criticality under the same data label for each server node, improve the criticality discrimination ability of the server node for its stored content and data content, and reduce the performance inefficiency risk that an individual server node becomes a functional bottleneck in the distributed storage system under the high business impact of fund data.
[0033] Further, the method for identifying heat-related labels from each server node using the label heat policy coefficient is as follows: For any data label, define the median value in the label heat policy coefficients corresponding to each node as the policy division variable. If the label heat policy coefficient of a node is greater than the policy division variable, then it is considered that the data label belongs to the heat-related label in this node; otherwise, it is a non-heat-related label.
[0034] Further, the method for storing data records in combination with the determination result of heat-related labels is as follows: Any data record to be written contains several data labels, and heat risk exclusion storage is performed for each of its data labels. Specifically: Obtain the storage nodes of the current data label. If the current data label is a heat-related label in any storage node, then define this node as a risk node, and obtain the heat risk exclusion return value: When all storage nodes are risk nodes, then it is considered that the return value is FALSE, and a new storage node is reselected for the current data label; otherwise, the return value is TRUE; If the heat risk exclusion return values of all data labels are TRUE through heat risk exclusion storage, then the data record to be written is completed for storage.
[0035] Preferably, all variables not defined in the present invention can be manually set thresholds if there is no clear definition.
[0036] The present invention also provides a distributed receiving and storing system for data records of fund assets. The distributed receiving and storing system for data records of fund assets includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method for distributed receiving and storing data records of fund assets. The distributed receiving and storing system for data records of fund assets can run on computing devices such as desktop computers, laptop computers, palmtop computers, and cloud data centers. The operable system can include, but is not limited to, a processor, a memory, and a server cluster. The processor executes the computer program and runs in the following units of the system:
[0037] A distributed storage environment layout unit, configured to identify and obtain each server node and data label in the data storage environment;
[0038] A real-time monitoring unit for obtaining the tag status of data tags from server nodes;
[0039] A thermal tag dynamic discrimination unit for calculating a tag thermal policy coefficient based on the tag status and identifying thermal tags from each server node using the tag thermal policy coefficient;
[0040] A new data storage unit for recording and storing data in combination with the determination result of thermal tags.
[0041] The beneficial effects of the present invention are as follows: The present invention provides a method and system for distributed reception and storage of data records of fund assets, analyzes the phased defects in the formation of the thermal risk accumulation effect, effectively quantifies the thermal sharding risk caused by the accumulation of the content criticality imbalance phenomenon of each server node under the same data tag, thereby improving the ability of the server node to distinguish the criticality of its stored content and data content, reducing the performance inefficiency risk that an individual server node becomes a functional bottleneck in the distributed storage system under the high business impact of fund data, providing a more stable distributed storage strategy for fund assets in a data environment with high-frequency reading and writing, and enhancing the service performance and efficiency of fund asset data. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present invention will become more obvious. The same reference numerals in the drawings of the present invention represent the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0043] Figure 1 It shows a flowchart of the method for distributed reception and storage of data records of fund assets;
[0044] Figure 2 It shows a structural diagram of the system for distributed reception and storage of data records of fund assets. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with the embodiments and the drawings to fully understand the purpose, solution and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0046] As Figure 1 shown is a flowchart of the method for distributed reception and storage of data records of fund assets. The following will describe the method for distributed reception and storage of data records of fund assets according to the embodiments of the present invention in combination with Figure 1 to elaborate. The method includes the following steps:
[0047] Configure server nodes and data tags in a data storage environment; obtain the tag status of data tags from the server nodes; calculate the tag heat policy coefficient according to the tag status, and use the tag heat policy coefficient to identify hot system tags from each server node; combine the determination results of the hot system tags to perform data record storage;
[0048] Further, the method for identifying each server node and data tag in the data storage environment is: the data storage environment includes several server nodes, referred to as nodes;
[0049] The data records of the fund assets are stored in each node through a data sharding mechanism. The data records are composed of several data tags. The several data tags of the data records are stored in different nodes respectively, and each node stores the fund asset data of each data tag.
[0050] Among them, the data tags are fund type, fund name, fund manager, net asset value per fund unit, dividend information, net asset value, and investment grouping.
[0051] Further, the method for obtaining the tag status of data tags from the server nodes is: preset a monitoring interval rg, and rg is 1 hour When; Obtain the tag status of the data tag once every rg. Among them, the tag status of the data tag is the product of the operation order value and the tag pressure. The calculation method of the operation order value is: the average value of the CPU occupancy rate of the server node within the monitoring interval is recorded as the occupancy amount, and the minimum value among the occupancy amounts of all nodes is recorded as the occupancy base value. Then, the difference between the occupancy amount of any node and the occupancy base value is used as the ratio of the difference to the occupancy base value as its operation order value; the calculation method of the tag pressure is: in a node, the number of read and write operations on the same data tag is used as the unit read and write amount, and the numerical set composed of the read and write amounts corresponding to each data tag under the node is subjected to minmax normalization processing, and the normalized value of the unit read and write amount of the data tag is defined as its tag pressure.
[0052] Further, the method for calculating the tag heat policy coefficient according to the tag status is: use any data tag as the determination tag;
[0053] Alternatively, for the determination tag, the moment when the tag state value is zero is taken as the zero point position; starting from the current moment, any node searches in the reverse time direction to obtain the first zero point position as the current zero point position of the node, and the number of moments between the current moment and the current zero point position is taken as the zero point span of the node; in the set composed of the zero point spans of each node, the median value Zps.mid, the interquartile range Zps.ir, and the maximum value Zps.mx of the set are respectively obtained, and the zero point interval Zps_rg is constructed, Zps_rg = min{(Zps.mid + Zps.ir), Zps.mx}; the time period from the current moment to its reverse time direction of Zps_rg is taken as the monitoring interval;
[0054] where the interquartile range is the difference between the upper quartile and the lower quartile, and min{} is the minimum value function.
[0055] Identify the thermal system tag according to the tag state of the determination tag within the monitoring interval.
[0056] Furthermore, the method for identifying the thermal system tag according to the tag state of the determination tag within the monitoring interval is: within the monitoring interval, the minimum values of the tag states at different moments of the same node and at different nodes at the same moment are respectively obtained, and are respectively denoted as the first state quasi-value and the second state quasi-value; the difference between the tag state and the first state quasi-value is defined as the quasi-node differential degree SLbv; the tag state only adopts the corresponding value under the determination tag;
[0057] For any node, if the first state quasi-value and the second state quasi-value corresponding to any non-zero tag state are both zero, then the moment corresponding to the tag state is denoted as the state valley point;
[0058] All the state valley points at all moments are constructed into a set Fn, the number of elements in the set Fn is denoted as Sfn, and the difference between the tag state of any state valley point and the minimum tag state in the set Fn is denoted as the node differential degree of the state valley point;
[0059] Taking the tag state of any moment as the valley point threshold, the ratio of the number of elements in the set Fn with a tag state less than the valley point threshold to the number of elements with a tag state greater than the valley point threshold is denoted as the state evolution degree Eff of that moment. If the tag states of all elements in the set Fn are greater than the valley point threshold, or all are less than the valley point threshold, then the state evolution degree of that moment is 1 / Sfn; the tag thermal strategy coefficient Ltsc calculated according to the state valley point, its calculation method is:
[0060] ;
[0061] where j1 is the cumulative variable, SLbv j1 and Lbv j1They are respectively the quasi-node differential degree and the node differential degree of the j1-th effective state valley point. exp() represents the exponential function with the natural constant e as the base, max() is the function to obtain the maximum value, and lg() is the logarithm function with base 10;
[0062] Preferably, the method for identifying the thermal system label according to the label effective state of the judgment label in the monitoring interval can be replaced by: in the monitoring interval, if the label effective state value at a certain moment is greater than that at the previous moment and the label effective state value at the previous moment is not zero, then define this moment as the effective state increment position, otherwise it is the non-effective state increment position; the average value of each effective state increment position is recorded as the effective state phase value. If the label effective state at any moment is greater than the effective state phase value, it is determined that it satisfies the first phase value condition;
[0063] The label effective state only adopts the value corresponding to the judgment label;
[0064] Taking any effective state increment position and its corresponding label effective state to form an increment binary group, using the plrep and splev functions to perform cubic spline interpolation on the set constructed by the increment binary group to obtain a smooth curve, and taking the fitting value of each moment on the obtained smooth curve as the curve value of the corresponding moment; the difference between the curve value of any non-effective state increment position and the label effective state is the increment estimation amount. When the increment estimation amount is less than zero, the non-effective state increment position has an inefficient estimation, otherwise it has an effective estimation;
[0065] Searching backward in time from any effective state increment position to obtain the first non-consecutively occurring effective state increment position as its reference position, and taking the time period between this effective state increment position and its reference position as an estimation reference interval; excluding the estimation reference interval containing the zero position; where the continuously occurring effective state increment positions refer to several effective state increment positions that are temporally continuous with the effective state increment position at the search starting position. Therefore, each continuous effective state increment position belongs to the estimation reference interval formed by the effective state increment position that is closest to the current moment in time;
[0066] Calculating the strategy sensitivity value Psenv of the estimation reference interval according to the increment estimation amount: Psenv = Len × lg(1 + eff.rp × eff.dsv); where eff.rp represents the proportion of the moments with effective estimation among all moments between the estimation reference interval and the first zero position in the reverse time direction, eff.dsv is the maximum value of the increment estimation amount label effective state at each moment with effective estimation in the estimation reference interval; Len is the proportion of the number of moments in the estimation reference interval to the total number of moments in the monitoring interval, and lg is the logarithm symbol with base 10;
[0067] When zs.rp cannot obtain the first zero position in the reverse time direction, the moment at the farthest end in the monitoring interval is used as the first zero position obtained by the search;
[0068] If the number of effective estimations in the estimated reference interval is greater than the number of inefficient estimations, then it is considered that the estimated reference interval meets the second phase value condition; both the first phase value condition and the second phase value condition are used as phase value conditions, and the number of phase value conditions satisfied at the current moment is recorded as the phase value condition achievement value Glv;
[0069] If any moment cannot be included in any estimated reference interval, then inherit the second phase value condition corresponding to the first estimated reference interval searched in the reverse time direction; the number of inefficient estimations refers to the number of non-effective state increments of inefficient estimations, and the number of effective estimations is the same; calculate the label heat strategy coefficient Ltsc according to each estimated reference interval:
[0070] ;
[0071] where e is the natural constant, i1 is the serial number of the estimated reference interval, MPs represents the average value of the strategy sensitivity values corresponding to all estimated reference intervals, noz is the number of estimated reference intervals, avg{} is the average value function, Psenv i1 is the strategy sensitivity value of the i1-th estimated reference interval, exp() is the exponential function with the natural constant e as the base, rk() is the sorting coefficient function, and the return value obtained through the sorting coefficient function rk(i1) is: obtain the maximum value of the label state in the i1-th estimated reference interval, and use the percentile value of the obtained maximum value among all label states in the monitoring interval as the return value of the sorting coefficient function.
[0072] Furthermore, the method for identifying hot system labels from each server node using the label heat strategy coefficient is: for any data label, define the median value of the label heat strategy coefficients corresponding to each node as the strategy division variable. If the label heat strategy coefficient of a node is greater than the strategy division variable, then it is considered that the data label belongs to the hot system label in this node, otherwise it is a non-hot system label.
[0073] Furthermore, the method for data record storage in combination with the determination result of hot system labels is: any data record to be written contains several data labels, and perform hot risk exclusion storage for each of its data labels. Specifically: obtain each storage node of the current data label. If the current data label is a hot system label in any storage node, then define this node as a risk node, and obtain the hot risk exclusion return value: when all storage nodes are risk nodes, then it is considered that the return value is FALSE, and reselect the storage node for the current data label, otherwise the return value is TRUE; through hot risk exclusion storage, if the hot risk exclusion return values of all data labels are TRUE, then the data record to be written is completed for storage.
[0074] If the current data tag of a data record has a thermal risk, it is considered that the storage state of the data tag is prone to accumulating thermal node risks. If there are too many data tags with such thermal node risks, it will affect the load balancing between server nodes, and there is a risk of uneven load in the content heat or content criticality direction in the entire fund asset data storage system.
[0075] The distributed receiving and storage system for data records of fund assets provided by the embodiments of the present invention, as Figure 2 shown in the structural diagram of the distributed receiving and storage system for data records of fund assets of the present invention. The distributed receiving and storage system for data records of fund assets in this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the embodiment of the above-mentioned distributed receiving and storage method for data records of fund assets.
[0076] The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it runs in the following units of the system:
[0077] The distributed storage environment layout unit is used to identify and obtain each server node and data tag in the data storage environment;
[0078] The real-time monitoring unit is used to obtain the tag status of the data tag from the server node;
[0079] The thermal system tag dynamic discrimination unit is used to calculate the tag thermal policy coefficient according to the tag status, and use the tag thermal policy coefficient to identify the thermal system tags from each server node;
[0080] The new data storage unit is used to perform data record storage in combination with the determination result of the thermal system tag.
[0081] The distributed receiving and storage system for data records of fund assets can run on computing devices such as desktop computers, laptop computers, palmtop computers, and cloud servers. The distributed receiving and storage system for data records of fund assets, the system that can run can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above examples are only examples of the distributed receiving and storage system for data records of fund assets, and do not constitute a limitation on the distributed receiving and storage system for data records of fund assets. It can include more or fewer components than the examples, or combine some components, or different components. For example, the distributed receiving and storage system for data records of fund assets can also include input and output devices, network access devices, buses, etc.
[0082] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the operating system of the data recording distributed receiving and storing system of the fund assets, and connects various parts of the operable system of the data recording distributed receiving and storing system of the fund assets through various interfaces and lines.
[0083] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of the data recording distributed receiving and storing system of the fund assets. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0084] Although the description of the present invention has been quite detailed and several of the described embodiments have been described in particular, it is not intended to be limited to any of these details or embodiments or any particular embodiment, so as to effectively cover the intended scope of the present invention. In addition, the present invention is described above in terms of embodiments foreseeable by the inventor for the purpose of providing a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.
Claims
1. A distributed receiving and storing method for fund asset data records, characterized in that: The method comprises the following steps: configuring server nodes and data tags in a data storage environment; obtaining the tag effectiveness state of the data tags from the server nodes; calculating the tag thermal strategy coefficient according to the tag effectiveness state, and using the tag thermal strategy coefficient to identify the thermal system tags from each server node; and recording and storing data in combination with the thermal system tag determination result; The method for obtaining the label validity state of the data label from the server node is: calculating the operation order value according to the CPU occupancy rate between different servers, calculating the label pressure according to the number of read and write times of the data label, and recording the product of the operation order value and the label pressure as the label validity state; The method for calculating the label thermal strategy coefficient according to the label effective state is: within a preset period, the label effective state time series is compared horizontally to obtain the first effective state quasi-value, the label effective state node is compared horizontally to obtain the second effective state quasi-value, the quasi-node difference dimension is calculated through the first effective state quasi-value, the effective state valley point is identified according to the first effective state quasi-value and the second effective state quasi-value, and the effective state derivative is obtained from the occurrence frequency of the effective state valley point, and the label thermal strategy coefficient is calculated using the quasi-node difference dimension, the node difference dimension and the effective state derivative; The method for storing data records in combination with the hot tag determination result is that, for each data tag of the data record to be written, when all the allocated storage nodes are not hot tags, the writing and storage of the data record is completed.
2. The distributed receiving and storing method for fund asset data records according to claim 1, characterized in that: The method for identifying and obtaining each server node and data label in a data storage environment is as follows: the data storage environment includes a plurality of server nodes, referred to as nodes; The data records of fund assets are sharded and stored in each node through the data sharding mechanism. The data records are composed of several data tags. The several data tags of the data records are stored in different nodes respectively. Each node stores the fund asset data of each data tag.
3. The distributed receiving and storing method for fund asset data records according to claim 1, characterized in that: The method for obtaining the label validity state of the data label from the server node is: preset a monitoring interval rg, rg∈[0.5, 24] hours, and obtain the label validity state of the data label every rg, wherein the label validity state of the data label is the product of the operation order value and the label pressure, and the calculation method of the operation order value is: the average value of the CPU occupancy rate of the server node within the monitoring interval is recorded as the occupancy, and the minimum value of the occupancy of all nodes is recorded as the occupancy base value, then the occupancy of any node is subtracted from the occupancy base value, and the ratio of the difference to the occupancy base value is used as its operation order value; the calculation method of the label pressure is: in a node, the number of read and write times of the same data label is taken as the unit read and write amount, and the value set consisting of the corresponding read and write amounts of each data label under the node is minmax normalized, and the normalized value of the unit read and write amount of the data label is defined as its label pressure.
4. The distributed receiving and storing method for fund asset data records according to claim 1, characterized in that: The method for calculating the label thermal strategy coefficient according to the label effectiveness state is: taking any data label as the judgment label and taking a preset time interval as the monitoring interval; and identifying the thermal system label according to the label effectiveness state of the judgment label within the monitoring interval.
5. The distributed receiving and storing method for fund asset data records according to claim 4 is characterized in that: The method of identifying the thermal system label according to the label validity state of the label determined in the monitoring interval is: respectively obtain the minimum value of the label validity state at different times of the same node and at different nodes at the same time, and record them as the first validity state quasi-value and the second validity state quasi-value respectively; define the difference between the label validity state and the first validity state quasi-value as the quasi-node difference dimension SLbv; for any node, if the first validity state quasi-value and the second validity state quasi-value corresponding to any non-zero label validity state are both zero, then the moment corresponding to the label validity state is recorded as the validity state valley point; construct each validity state valley point at all times into a set Fn, the number of elements in the set Fn is recorded as Sfn, and the difference between the label validity state of any validity state valley point and the minimum value of the label validity state in the set Fn is recorded as the node difference dimension of the validity state valley point; The label effectiveness state at any moment is taken as the valley threshold. The ratio of the number of elements whose label effectiveness state is less than the valley threshold to the number of elements whose label effectiveness state is greater than the valley threshold in the set Fn is recorded as the effectiveness state derivative at that moment. The label thermal strategy coefficient is calculated based on the effectiveness state derivative, effective state valley, quasi-node difference dimension and node difference dimension.
6. The distributed receiving and storing method for fund asset data records according to claim 4, characterized in that: The method of identifying thermal system tags according to the tag effectiveness state of the tag determined within the monitoring interval can be replaced as follows: if the tag effectiveness state value at a moment is greater than that at the previous moment and the tag effectiveness state value at the previous moment is not zero, then the moment is defined as an effectiveness state increase; the average value of each effectiveness state increase is recorded as the effectiveness state phase value, and if the tag effectiveness state at any moment is greater than the effectiveness state phase value, then it is determined that it meets the first phase value condition; An increase-position binary is formed by any effective increase-position and its corresponding label effective state. The set constructed by the increase-position binary is interpolated by cubic spline using the plrep and splev functions to obtain a smooth curve. The difference between the curve value of any non-effective increase-position and the label effective state is the increase-position estimation value. When the increase-position estimation value is less than zero, the non-effective increase-position is inefficiently estimated, otherwise it is effectively estimated. From any effective state increase, search in reverse time direction to obtain the first non-continuous effective state increase as its reference position, and use the time period between the effective state increase and its reference position as an estimated reference interval; eliminate the estimated reference interval containing the zero point; calculate the strategy sensitivity value of the estimated reference interval based on the increase estimate; if the number of valid estimates in the estimated reference interval is greater than the number of inefficient estimates, the estimated reference interval is considered to meet the second phase value condition; both the first phase value condition and the second phase value condition are used as phase value conditions, and the number of phase value conditions that meet the current moment is recorded as the phase value condition achievement value; calculate the label thermal strategy coefficient based on the phase value condition achievement value and each estimated reference interval.
7. The distributed receiving and storing method for fund asset data records according to claim 1, characterized in that: The method of using the label thermal strategy coefficient to identify the hot label from each server node is: for any data label, define the median value of the label thermal strategy coefficient corresponding to each node as the strategy partition variable. If the label thermal strategy coefficient of the node is greater than the strategy partition variable, then the data label is considered to be a hot label at the node, otherwise it is a non-hot label.
8. The distributed receiving and storing method for fund asset data records according to claim 1, characterized in that: The method for storing data records in combination with the result of thermal label determination is: any data record to be written contains several data labels, and thermal risk exclusion storage is performed for each data label, specifically: each storage node of the current data label is obtained, if the current data label in any storage node is a thermal label, the node is defined as a risk node, and the thermal risk exclusion return value is obtained: when all storage nodes are risk nodes, the return value is FALSE, and the storage node is reselected for the new current data label, otherwise the return value is TRUE; By using thermal risk exclusion storage, the thermal risk exclusion return values of all data tags are TRUE, and the data records to be written are stored.
9. The distributed receiving and storage system for fund asset data records is characterized by: The distributed receiving and storage system for data records of the fund assets comprises: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the distributed receiving and storage method for data records of the fund assets described in any one of claims 1 to 8 are implemented. The distributed receiving and storage system for data records of the fund assets runs on computing devices such as desktop computers, laptop computers, PDAs, and cloud data centers.
Citation Information
Patent Citations
Data storage method and device based on node access popularity
CN112749004A
Data storage method and device, computer equipment and storage medium
CN115686385A