Intelligent data analysis system and method based on multi-terminal synchronization
By analyzing and sorting the priority of target data in the database and performing adaptive index reorganization, the problem that traditional index reorganization does not consider time locality and data correlation is solved, and data query speed and multi-end synchronization efficiency are improved.
Patent Information
- Application Number
- CN202510217498.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When indexing and reorganizing B+ trees under the traditional method, time locality and data correlation are not considered, resulting in poor retrieval performance of index search data, and increasing the frequency and resource utilization of index reorganization.
By obtaining the index values in the leaf nodes of the B+ tree that need to be indexed and reorganized in the database on each device, analyzing the historical data change timestamps, the number of associated data, the frequency of data modification and the stability of the index page, calculating the data relationship priority and data position priority, sorting the target data and performing adaptive index reorganization.
Reduces the frequency of index reorganization, improves data query speed, improves the efficiency of multi-terminal synchronization, and reduces the use of CPU and disk I/O resources.
Smart Images

Figure CN120123345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an intelligent data analysis system and method based on multi-terminal synchronization. Background Art
[0002] With the rapid development of information technology, data has become an important basis for enterprise operation decision-making. In practical applications, enterprises often need to process and analyze data on multi-terminal devices (such as computers, mobile phones, tablets, etc.), which increases the frequency of data synchronization. During the data synchronization process, data is usually searched according to the index of the data, and synchronization operations such as addition, deletion, and modification are performed in the database to improve the efficiency of data synchronization. However, frequent data synchronization (especially addition and deletion) will cause index fragmentation, thus reducing the query performance. Therefore, when the index fragmentation rate reaches 10% to 30%, it is necessary to reorganize the index, that is, to reduce fragmentation by rearranging the data in the index page and releasing unused space, thereby improving the query performance during the data synchronization process.
[0003] When reorganizing the index of the B+ tree in the traditional way, only the index structure is adjusted by means of merging, splitting, etc. on the basis of the original data order remaining unchanged. According to the principle of temporal locality and data correlation, if a certain data is operated on, after a period of time, this data may be accessed and operated on again, and the data associated with this data may also be accessed and operated on. After the index reorganization of this data in the traditional way, when this data is operated on next time, the traversal time of its corresponding index value is relatively long. For example, when traversing the index values in the leaf nodes of the B+ tree in the order from left to right, the corresponding index value or the index value corresponding to the data associated with it is very likely to appear in the rightmost leaf node of the B+ tree, resulting in a relatively long traversal time when operating on this data or the data associated with it in the order from left to right next time. Therefore, the index reorganization in the traditional way does not consider the temporal locality and data correlation, and thus does not optimize the retrieval performance of searching data according to the index of the data. At the same time, the frequency of index reorganization may increase, thereby occupying a large amount of CPU and disk I / O resources, resulting in a decline in the performance of other queries or transactions and increasing the operation and maintenance complexity.
[0004] Therefore, how to perform adaptive index reorganization on the B+ tree to reduce the frequency of index reorganization and improve the efficiency of multi-terminal synchronization has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide an intelligent data analysis system and method based on multi-terminal synchronization to solve the problem of how to perform adaptive index reorganization according to the temporal locality to reduce the frequency of index reorganization and improve the efficiency of multi-terminal synchronization.
[0006] In a first aspect, an intelligent data analysis method based on multi-terminal synchronization is provided in an embodiment of the present invention. The method includes the following steps:
[0007] After data synchronization is performed on each device side, in the database of any one device side, obtain the index values in each leaf node of the B+ tree that needs to be index-reorganized, and obtain the data corresponding to each of the index values, denoted as target data;
[0008] For any one of the target data, in the database log, obtain the data change timestamp of the data under the index value of the any one of the target data within a historical period. According to the data change timestamp of the any one of the target data, obtain the number of associated data of the any one of the target data. According to the data change timestamp and the number of associated data, obtain the data relationship priority of the any one of the target data;
[0009] According to the data modification frequency feature and modification interval feature of the data under the index value of the any one of the target data, obtain the modification frequency of the any one of the target data. According to the stability of the index page under the index value of the any one of the target data and the modification frequency, obtain the data location priority of the any one of the target data;
[0010] Obtain the data relationship priority and data location priority of each of the target data. According to the data relationship priority and data location priority of each of the target data, sort all the target data to obtain a corresponding sorting result. According to the sorting result, perform adaptive index reorganization on the B+ tree.
[0011] Preferably, the step of obtaining the number of associated data of the any one of the target data according to the data change timestamp of the any one of the target data includes:
[0012] Obtain the data change timestamps of the data under the index values of all the target data within the historical period. Divide the historical period into at least two sub-periods. According to all the data change timestamps, form a time series sequence with the index values corresponding to the data change timestamps within the same sub-period, and obtain the time series sequence corresponding to each of the sub-periods;
[0013] Denote the index value of the any one of the target data as the target index value, and denote the index values of the other target data except the any one of the target data as other index values. For any one of the other index values, denote the time series sequence that simultaneously includes the target index value and the any one of the other index values as the first target time series sequence. If the number of the first target time series sequences is greater than a preset number, then regard the target data corresponding to the any one of the other index values as the associated data of the any one of the target data;
[0014] Denote the number of all the associated data of the any one of the target data as the number of associated data of the any one of the target data.
[0015] Preferably, obtaining the data relationship priority of any target data according to the data change timestamp and the associated data quantity includes:
[0016] If the associated data quantity of any target data is 0, then record the data relationship priority of any target data as 0;
[0017] If the associated data quantity of any target data is not 0, then record the time series sequence including the index value of any target data as the second target time series sequence. For any second target time series sequence, obtain the data change timestamp under the first index value in any second target time series sequence. In any second target time series sequence, calculate the time interval between the data change timestamp under the index value of any target data and the data change timestamp under the first index value to obtain the influence coefficient of any target data in the sub-period corresponding to any second target time series sequence;
[0018] Calculate the average value of the influence coefficients of any target data in the sub-period corresponding to each second target time series sequence to obtain the average influence coefficient of any target data, and calculate the reciprocal of the sum of the average influence coefficient and the constant 1 to obtain the first influence degree of any target data;
[0019] Subtract the reciprocal of the associated data quantity of any target data from the constant 1 to obtain the second influence degree of any target data;
[0020] Perform weighted summation on the first influence degree and the second influence degree to obtain the data relationship priority of any target data.
[0021] Preferably, obtaining the modification frequency of any target data according to the data modification frequency feature and the modification interval feature under the index value of any target data includes:
[0022] Obtain the data modification frequency under the index value of any target data according to the quantity of data change timestamps under the index value of any target data within the historical period;
[0023] Calculate the time interval between every two adjacent data change timestamps under the index value of any target data within the historical period to obtain the average time interval. Subtract the reciprocal of the time interval from the constant 1 to obtain the data modification density under the index value of any target data;
[0024] Perform weighted summation on the data modification frequency and the data modification density to obtain the modification frequency of any target data.
[0025] Preferably, obtaining the data location priority of any target data according to the stability of the index page under the index value of any target data and the modification frequency includes:
[0026] Denote the index page under the index value of any target data as the target index page, and within the historical period, obtain the stability rate of the target index page according to the change in the number of index values corresponding to the target index page between every two adjacent historical moments;
[0027] Calculate the product of the stability rate of the target index page and the modification frequency of any target data to obtain the data location priority of any target data.
[0028] Preferably, obtaining the stability rate of the target index page according to the change in the number of index values corresponding to the target index page between every two adjacent historical moments includes:
[0029] For any historical moment, form a first set with the index values of the target index page at the any historical moment, form a second set with the index values of the target index page at the previous historical moment of the any historical moment. In the first set, obtain the number of index values that are the same as the index values in the second set, and calculate the proportion of the number of index values in all elements of the first set to obtain the consistency rate of the target index page at the any historical moment;
[0030] Obtain the consistency rate of the target index page at each historical moment to obtain the mean value and variance of the consistency rate. Calculate the reciprocal of the sum of the constant 1 and the variance of the consistency rate, and take the product of the reciprocal and the mean value of the consistency rate as the first variable;
[0031] Form a consistency rate sequence with the consistency rates of the target index page at each historical moment, calculate the absolute value of the difference between the first data and the last data in the consistency rate sequence, and take the reciprocal of the sum of the constant 1 and the absolute value of the difference as the second variable;
[0032] Calculate the mean value between the first variable and the second variable to obtain the stability rate of the target index page.
[0033] Preferably, sorting all target data according to the data relationship priority and data location priority of each target data to obtain the corresponding sorting result includes:
[0034] Arrange all target data in descending order according to the data relationship priority, and arrange the target data with the same data relationship priority in descending order according to the data location priority to obtain a target data sequence.
[0035] Preferably, the adaptive index reorganization of the B+ tree according to the sorting result includes:
[0036] Updating the target data under each index value in accordance with a preset traversal order of the index values in the leaf nodes of the B+ tree for the target data sequence, and adaptively reorganizing the index of the B+ tree according to the updated result.
[0037] In a second aspect, an embodiment of the present invention further provides an intelligent data analysis system based on multi-terminal synchronization, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements an intelligent data analysis method based on multi-terminal synchronization as described in the first aspect.
[0038] The beneficial effects of the embodiments of the present invention compared with the prior art are:
[0039] After data synchronization is performed on each device end in the present invention, in the database of any device end, the index values in each leaf node of the B+ tree that needs to be reorganized by index are obtained, and the data corresponding to each index value is obtained, which is recorded as target data; for any target data, in the database log, the data change timestamp under the index value of the any target data in the historical period is obtained, and according to the data change timestamp of the any target data, the number of associated data of the any target data is obtained, and according to the data change timestamp and the number of associated data, the data relationship priority of the any target data is obtained; according to the data modification frequency feature and the modification interval feature under the index value of the any target data, the modification frequency of the any target data is obtained, and according to the stability of the index page under the index value of the any target data and the modification frequency, the data position priority of the any target data is obtained; the data relationship priority and the data position priority of each target data are obtained, and all target data are sorted according to the data relationship priority and the data position priority of each target data to obtain a corresponding sorting result, and the B+ tree is adaptively reorganized by index according to the sorting result. Among them, considering temporal locality and data correlation, the B+ tree is adaptively reorganized by index according to the priority of the target data (data relationship priority), the modification frequency of the target data, and the change situation (stability) of the index page, optimizing the result of reorganizing the index of the B+ tree in the traditional way (when reorganizing the index of the B+ tree in the traditional way, only the index structure is adjusted by means of merging, splitting, etc. on the basis of the original data order remaining unchanged). While reducing index fragmentation, the query speed of data is improved, and the efficiency of data synchronization is enhanced. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 is the flowchart of a method for intelligent data analysis based on multi - terminal synchronization provided in the first embodiment of the present invention;
[0042] Figure 2 is a schematic diagram for updating the target data under each index value in the leaf nodes of a B + tree provided in the first embodiment of the present invention. Detailed implementation manners
[0043] The following details the embodiments of the present disclosure. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation to the present disclosure.
[0044] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0045] To illustrate the technical solutions of the present invention, the following will be described through specific embodiments.
[0046] See Figure 1 , which is the flowchart of a method for intelligent data analysis based on multi - terminal synchronization provided in the first embodiment of the present invention. As Figure 1 shown, the method may include:
[0047] Step S101, after data synchronization is performed on each device end, in the database of any device end, obtain the index values in each leaf node of the B + tree that needs to be index - reorganized, and obtain the data corresponding to each of the index values, denoted as target data.
[0048] A multi - terminal structure refers to a multi - server or multi - client structure. Taking the multi - client structure as an example, there are multiple clients and one main server. The clients and the main server have the same initial database. When the data in any client changes, data synchronization is required: the database generates a log and uploads the changed data to the server at the same time. After receiving the message from the client, the server updates its own database. At the same time, to maintain data consistency, after the server's database is updated, it needs to notify other clients and transmit the corresponding data to other clients. After receiving the message from the server, other clients receive the data and update their own databases. When the server changes, the database content of all clients is modified in the same way.
[0049] During the process of multi - terminal data synchronization, data is usually searched according to its index, and operations such as addition, deletion, and modification are performed on the index in the database. Among them, the B + tree is a widely used index structure. When data is deleted or inserted, the corresponding index values also need to be added or deleted. These operations will cause the index pages of the B + tree to split or merge, resulting in index fragmentation. When the index fragmentation rate reaches 10% to 30%, the index needs to be reorganized, that is, by rearranging the data in the index pages and releasing the unused space, the structure of the B + tree is adjusted, thereby reducing index fragmentation and improving the query performance during data synchronization.
[0050] In the traditional method of reorganizing the B + tree index, the index structure is adjusted only by means of merging, splitting, etc. on the basis of keeping the original data order unchanged. According to the principle of temporal locality and data correlation, if a piece of data has been operated on, it may be accessed and operated on again after a period of time, and the data related to this data may also be accessed and operated on. After the index reorganization of this data in the traditional method, when this data is operated on next time, the traversal time of its corresponding index value is relatively long. For example, when traversing the index values in the leaf nodes of the B + tree in the order from left to right, the corresponding index value or the index value corresponding to the data related to it is very likely to appear in the right - most leaf node of the B + tree. This results in a relatively long traversal time when operating on this data or the data related to it in the order from left to right next time. Therefore, the traditional index reorganization does not consider temporal locality and data correlation, and thus does not optimize the retrieval performance of searching for data according to the data index. At the same time, the frequency of index reorganization may increase, thereby occupying a large amount of CPU and disk I / O resources, resulting in a decline in the performance of other queries or transactions and an increase in the complexity of operation and maintenance.
[0051] Therefore, in the embodiments of the present invention, according to temporal locality and data correlation, the B+ tree is adaptively reorganized to improve the efficiency of multi-terminal synchronization. When it is detected that the index fragmentation rate in the database of any device terminal is greater than 10% and less than 30%, this time is the current time. At the same time, the index values in each leaf node of the B+ tree at the current time are obtained, and the data corresponding to each index value is recorded as the target data for subsequent analysis. The obtaining of the index fragmentation rate is a prior art and will not be elaborated here.
[0052] Step S102, for any target data, in the database log, obtain the data change timestamp of the data under the index value of the any target data within the historical period. According to the data change timestamp of the any target data, obtain the number of associated data of the any target data. According to the data change timestamp and the number of associated data, obtain the data relationship priority of the any target data.
[0053] According to the database log, obtain the data change timestamp of the data under the index value of each target data within 10 minutes before the current time. There is no limitation here. The implementer can set the historical period according to the specific scenario for analyzing each target data according to temporal locality and data correlation, so as to sort all target data. Then, according to the sorting result, the B+ tree is adaptively reindexed and reorganized, so that when traversing the index values in the leaf nodes of the B+ tree, the query speed of the data that may need to be frequently accessed and operated is improved, and thus the efficiency of multi-terminal synchronization is improved.
[0054] When a data modification operation occurs, according to data correlation, that is, there is a certain connection or dependency between two or more data, so that a change in one data will cause a corresponding change in another or more data. For example, in the membership mechanism of the personal account in a video APP, there is a correlation between the membership rights and the membership level: when the membership level is upgraded, the membership rights will change accordingly; there is also a correlation between the membership level and the membership experience: when the membership experience is upgraded to a certain value, the membership level will also be upgraded, etc. Therefore, in the embodiments of the present invention, according to the data change timestamp of the data under the index value of each target data, it is judged which data under the index values change and are repetitive within the same short time. There is likely an association between the target data under these index values, and a higher priority should be assigned to them. When sorting each target data, the higher the priority of the target data, the more forward its position. When traversing the index values in the leaf nodes of the B+ tree, the search time for the data with a higher priority is shorter, the query efficiency is improved, and thus the data synchronization efficiency is improved.
[0055] In an embodiment of the present invention, taking the i-th target data as an example, first, obtain the data change timestamps corresponding to the index values of each target data within 10 minutes before the current moment, and divide the 10 minutes before the current moment into 120 sub-periods evenly, with each sub-period being 5 seconds. There is no limit here, and the implementer can set it according to the specific scenario. Then, according to all the data change timestamps, form a time series sequence with the index values corresponding to the data change timestamps within the same sub-period, and obtain the time series sequence corresponding to each of the sub-periods. It should be noted that the number of index values in each time series sequence may be different. For example, the index values corresponding to the data change timestamps from 5 to 10 seconds are 27, 38, and 56, and these three index values form a time series sequence (27, 38, 56); the index values corresponding to the data change timestamps from 10 to 15 seconds are 27, 38, 20, and 56, and these four index values form a time series sequence (27, 38, 20, 56). Determine which index values in all the time series sequences appear and are repetitive within the same sub-period as the index value of the i-th target data. The data under these index values is very likely to be related to the i-th target data, and record the number of these index values as the associated data quantity of the i-th target data. Finally, according to the associated data quantity and the data change timestamp corresponding to the index value of the i-th target data, obtain the data relationship priority of the i-th target data.
[0056] Among them, the specific process of obtaining the associated data quantity of the i-th target data is as follows:
[0057] Record the index value of the i-th target data as the target index value, and record the index values of other target data except the i-th target data as other index values. For any other index value, record the time series sequence that contains both the target index value and the any other index value as the first target time series sequence. If the number of the first target time series sequences is greater than 3 (there is no limit here, and the implementer can set it according to the specific scenario), then regard the target data corresponding to the any other index value as the associated data of the i-th target data;
[0058] Record the number of all the associated data of the i-th target data as the associated data quantity of the i-th target data.
[0059] If the index value of the i-th target data is 27, the first time series is (27, 58, 36, 34), the second time series is (27, 58, 20, 36), the third time series is (27, 36, 58, 34), and the fourth time series is (27, 58, 36, 23), then the number of time series that contain both the index value 58 and the index value 27 of the i-th target data is 4, the number of time series that contain both the index value 36 and the index value 27 of the i-th target data is 4, the number of time series that contain both the index value 34 and the index value 27 of the i-th target data is 2, the number of time series that contain both the index value 20 and the index value 27 of the i-th target data is 1, and the number of time series that contain both the index value 23 and the index value 27 of the i-th target data is 1. Therefore, it is confirmed that the target data corresponding to the index values 58 and 36 is the associated data of the i-th target data.
[0060] If the number of associated data of the i-th target data is 0, it means that there is no other data associated with the i-th target data. Therefore, the data relationship priority of the i-th target data is recorded as 0; if the number of associated data of the i-th target data is not 0, then according to the number of associated data and the data change timestamp under the index value corresponding to the i-th target data, the data relationship priority of the i-th target data is obtained. Specifically:
[0061] The time series containing the index value of the i-th target data is recorded as the second target time series. For any second target time series, obtain the data change timestamp under the first index value in the any second target time series. In the any second target time series, calculate the time interval between the data change timestamp under the index value of the i-th target data and the data change timestamp under the first index value to obtain the influence coefficient of the i-th target data in the sub-period corresponding to the any second target time series;
[0062] Calculate the average value of the influence coefficients of the i-th target data in the sub-periods corresponding to each second target time series to obtain the average influence coefficient of the i-th target data, and calculate the reciprocal of the sum of the average influence coefficient and the constant 1 to obtain the first influence degree of the i-th target data;
[0063] Subtract the reciprocal of the number of associated data of the i-th target data from the constant 1 to obtain the second influence degree of the i-th target data;
[0064] Perform weighted summation on the first influence degree and the second influence degree to obtain the data relationship priority of the i-th target data.
[0065] In an embodiment, the calculation formula for the data relationship priority of the i-th target data is:
[0066]
[0067] Among them, α i represents the data relationship priority of the i-th target data, s represents the number of associated data of the i-th target data, t represents the average influence coefficient of the i-th target data, w 1 represents the first weight, w 2 represents the second weight, and 1 represents a constant.
[0068] It should be noted that is the first degree of influence; the smaller t is, it indicates that the time when the data under the target index value (the index value corresponding to the i-th target data) is changed in each sub-period is earlier, and the i-th target data is more likely to be the data that affects the change of other target data, and then the larger, α i the larger, the higher the priority of the i-th target data; is the second degree of influence. The larger s is, it indicates that there are more data associated with the i-th target data, and the greater the influence generated when the i-th target data is changed, and then the larger, α i the larger, the higher the priority of the i-th target data; set w 1 = 0.4, w 2 = 0.6. There is no limit here, and the implementer can set it according to the specific scenario.
[0069] Thus, the data relationship priority of the i-th target data is obtained.
[0070] Step S103, according to the data modification frequency characteristics and modification interval characteristics of the data under the index value of any one of the target data, obtain the modification frequency of any one of the target data, and according to the stability of the index page under the index value of any one of the target data and the modification frequency, obtain the data position priority of any one of the target data.
[0071] According to temporal locality: If a certain piece of data has been operated on, it may be accessed and operated on again after a period of time. If a certain piece of data has been accessed and operated on multiple times within a past period of time, then within a future time period, it is very likely that the data will also be accessed and operated on multiple times. Therefore, according to the data change timestamp within 10 minutes before the current moment of the index value of the i-th target data, there is no restriction here, and the implementer can set the time according to the specific scenario to obtain the modification frequency of the index value of the i-th target data. The greater the modification frequency, the more likely it is that the i-th target data will be modified again within a future time period. When sorting each target data, the position of the i-th target data is more forward, so that when traversing the index values in the leaf nodes of the B+ tree, the search time for the i-th target data is shorter, improving the query efficiency.
[0072] Considering that data modification within a period of time is divided into two cases: one is that the data is accessed and operated on multiple times within a short period of time, and the second is that the data is accessed and operated on within each time period. From the perspective of long-term use, the second case is more likely to increase the I / O overhead. Therefore, on the basis of the modification frequency of the index value of the i-th target data, it is also necessary to combine the time interval between every two adjacent data change timestamps under the i-th target index value to obtain the modification frequency of the i-th target data. The greater the modification frequency of the i-th target data, the more likely it is that the i-th target data will be modified again within a future time period. When sorting each target data, the position of the i-th target data is more forward, so that when traversing the index values in the leaf nodes of the B+ tree, the search time for the i-th target data is shorter, improving the query efficiency. Then the specific method for obtaining the modification frequency of the i-th target data is as follows:
[0073] According to the number of data change timestamps under the index value of the i-th target data within 10 minutes before the current moment, obtain the data modification frequency under the index value of the i-th target data. There is no restriction here, and the implementer can set the historical time period according to the specific scenario;
[0074] Calculate the time interval between every two adjacent data change timestamps under the index value of the i-th target data within the historical period to obtain the average time interval. Subtract the reciprocal of the time interval from the constant 1 to obtain the data modification density under the index value of the i-th target data;
[0075] Perform a weighted sum of the data modification frequency and the data modification density to obtain the modification frequency of the i-th target data.
[0076] In an implementation manner, the calculation formula for the modification frequency of the i-th target data is:
[0077]
[0078] Among them, β i represents the modification frequency of the i-th target data, p represents the data modification frequency under the index value of the i-th target data, ρ represents the data modification density under the index value of the i-th target data, and w 3 represents the first weight, and w 4 represents the second weight, and 1 represents a constant.
[0079] It should be noted that the larger p is, the greater the frequency of data modification under the index value of the i-th target data within a period of time, and thus the larger β i is, and the greater the modification frequency of the i-th target data; That is, it is the data modification density under the index value of the i-th target data. The larger ρ is, the longer the time interval between every two adjacent modifications of the data under the index value of the i-th target data within a period of time, and the greater the long-term I / O overhead. Thus is larger, β i is larger, and the greater the modification frequency of the i-th target data; Set w 3 = 0.7, w 4 = 0.3. There is no restriction here, and the implementer can set it according to the specific scenario.
[0080] Also considering that data deletion and data insertion will cause the leaf nodes of the B+ tree to merge and split. Among them, the most affected is data insertion. After data insertion, the leaf nodes of the B+ tree may split, which will increase the depth of the B+ tree. The greater the depth, the more node splitting and merging operations may be required for insertion and deletion operations, affecting performance.
[0081] Therefore, in the embodiments of the present invention, the index page corresponding to the index value of the i-th target data is denoted as the target index page. First, according to the change in the number of index values corresponding to the target index page between every two adjacent historical moments, the stability rate of the target index page is obtained. The smaller the stability rate, the more likely it is to have insert and delete operations in the index value range of the target index page. When sorting each target data, the position of the target data corresponding to the index value in the target index page should be placed at the back. At this time, the positions of the target data corresponding to the index values in the index pages after the target index page will move forward, preventing the index of the data from moving backward due to data insertion when querying a certain data in the future, thereby increasing the query time for the data. For example, the index values in a certain index page of a B+ tree are (27, 34, 36, 58), and the index value of the data to be queried is 36. During the query process, it is only necessary to traverse 27, 34, 36 in sequence to find it. If data insertion occurs at this time and the index values in this index page become (27, 30, 34, 36, 58), then during the query process, it is necessary to traverse 27, 30, 34, 36 in sequence, increasing the number of queries; if the target data under the index value 27 is updated to the data to be queried, the number of queries is the same.
[0082] Among them, the specific method for obtaining the stability rate of the target index page according to the change in the number of index values corresponding to the target index page between every two adjacent historical moments within 10 minutes before the current moment is as follows:
[0083] For any historical moment, the index values of the target index page at the any historical moment are formed into a first set, and the index values of the target index page at the previous historical moment of the any historical moment are formed into a second set. In the first set, the number of index values that are the same as the index values in the second set is obtained, and the ratio of the number of index values in all elements in the first set is calculated to obtain the consistency rate of the target index page at the any historical moment;
[0084] The consistency rates of the target index page at each historical moment are obtained to obtain the mean value and variance of the consistency rates. The reciprocal of the sum of the constant 1 and the variance of the consistency rates is calculated, and the product of the reciprocal and the mean value of the consistency rates is used as the first variable;
[0085] The consistency rates of the target index page at each historical moment are formed into a consistency rate sequence, and the absolute value of the difference between the first data and the last data in the consistency rate sequence is calculated. The reciprocal of the sum of the constant 1 and the absolute value of the difference is used as the second variable;
[0086] The mean value between the first variable and the second variable is calculated to obtain the stability rate of the target index page.
[0087] In one embodiment, the formula for calculating the stability rate of the target index page is:
[0088]
[0089] where γ represents the stability rate of the target index page, represents the mean value of the consistency rate, S 2 represents the variance of the consistency rate, v 1 represents the first data in the consistency rate sequence, v t represents the last data in the consistency rate sequence, 1 represents a constant, and | | represents the absolute value symbol.
[0090] It should be noted that the larger, the less the index value in the target index page changes over a period of time, and thus the larger γ, indicating that the stability of the target index page is stronger; the smaller S 2 the smaller, the more stable the change of the index value in the target index page at each historical moment, and the more reliable the mean value of the consistency rate the more reliable, and thus the larger γ, indicating that the stability of the target index page is stronger; the smaller |v 1 -v t |, the longer the index value in the target index page remains unchanged when it changes, and thus the larger γ, indicating that the stability of the target index page is stronger.
[0091] Furthermore, according to the stability rate of the target index page and the modification frequency, the data position priority of the i-th target data is obtained. Specifically:
[0092] Calculate the product of the stability rate of the target index page and the modification frequency of the i-th target data to obtain the data position priority of the i-th target data.
[0093] In one embodiment, the formula for calculating the data position priority of the i-th target data is:
[0094] δ i =γ×β i
[0095] where δ i represents the data position priority of the i-th target data, β i represents the modification frequency of the i-th target data, γ represents the stability rate of the target index page (the index page where the index value of the i-th target data is located), and 1 represents a constant.
[0096] It should be noted that β iThe larger it is, it indicates that the data under the index value of the $i$-th target data has been modified more frequently within a period of time, and the modification time is relatively fixed. When sorting all target data, the position of the $i$-th target data is more forward, and then $\delta$ i The larger it is, when sorting all target data, the position of the $i$-th target data is more forward, and the data position priority of the $i$-th target data is higher; the larger $\gamma$ is, it means that the index page of the $i$-th target data has fewer split and merge situations in the historical period, and the index page (target index page) where the index value of the $i$-th target data is located is more stable, and then $\beta$ i The larger it is, when sorting all target data, the position of the $i$-th target data is more forward, and the data position priority of the $i$-th target data is higher.
[0097] Thus, the data position priority of the $i$-th target data is obtained.
[0098] Step S104, obtain the data relationship priority and data position priority of each of the target data, sort all the target data according to the data relationship priority and data position priority of each of the target data to obtain the corresponding sorting result, and perform an adaptive index reorganization on the B+ tree according to the sorting result.
[0099] According to Step S102 and Step S103, obtain the data relationship priority and data position priority of each target data, and then sort all target data according to the data relationship priority and data position priority of each target data: first sort all target data in descending order according to the data relationship priority, and for target data with the same data relationship priority, sort them in descending order according to the data position priority. Finally, obtain the target data sequence. The larger the data relationship priority and the higher the modification frequency of the target data, the more stable the index page where its index value is located, and the more forward the position of the target data, so that when traversing the index values in the leaf nodes of the B+ tree, the search time for the target data is shorter and the query efficiency is higher.
[0100] Furthermore, update the target data under the index value of each leaf node of the B+ tree according to the target data sequence. For example: in the order of traversing the index values in the leaf nodes of the B+ tree, assume that the order of the target data obtained under each index value is $(a, b, c, d, e, f, g, h, j, k, y,)$. After sorting all target data according to the data relationship priority and data position priority of each target data, the obtained target data sequence is $(c, e, a, b, h, f, e, j, y, d, k,)$. Refer to Figure 2 , which is a schematic diagram for updating the target data under each index value in the leaf nodes of the B+ tree. Figure 2Among them, the leaf nodes correspond to the numbers (3, 14, 17, 18, 20, 23, 27, 34, 36, 45, 58) in a row of white squares, which are the index values of each target data. The thick black arrows indicate the order of traversing each index value in the leaf nodes of the B+ tree. The letters pointed by the dashed black arrows are the target data corresponding to each index value, and the letters in the gray squares are the data after updating the target data under each index value according to the target data sequence.
[0101] After the update is completed, according to the number of elements in each node of the B+ tree, operations such as splitting and merging the nodes of the B+ tree are performed to adjust the overall structure of the B+ tree and complete the adaptive index reorganization. Among them, performing operations such as splitting and merging the nodes of the B+ tree according to the number of elements in each node of the B+ tree to adjust the overall structure of the B+ tree is a prior art and will not be elaborated here.
[0102] In summary, after data synchronization is performed on each device side, in the database of any device side, the index values in each leaf node of the B+ tree that needs to be reorganized are obtained, and the data corresponding to each of the index values is obtained, which is recorded as the target data; for any target data, in the database log, the data change timestamp under the index value of the any target data within the historical period is obtained, and according to the data change timestamp of the any target data, the number of associated data of the any target data is obtained, and according to the data change timestamp and the number of associated data, the data relationship priority of the any target data is obtained; according to the data modification frequency characteristics and modification interval characteristics under the index value of the any target data, the modification frequency of the any target data is obtained, and according to the stability of the index page under the index value of the any target data and the modification frequency, the data position priority of the any target data is obtained; the data relationship priority and data position priority of each target data are obtained, and according to the data relationship priority and data position priority of each target data, all target data are sorted to obtain the corresponding sorting result, and the B+ tree is adaptively reorganized according to the sorting result. Among them, considering temporal locality and data correlation, the B+ tree is adaptively reorganized according to the priority of the target data (data relationship priority), the modification frequency of the target data, and the change situation (stability) of the index page, which optimizes the result of reorganizing the B+ tree in the traditional way (when reorganizing the B+ tree in the traditional way, only the index structure is adjusted by means of merging, splitting, etc. on the basis of the original data order remaining unchanged). While reducing index fragmentation, it improves the query speed of data and enhances the efficiency of data synchronization.
[0103] Based on the same inventive concept as the above method, an embodiment of the present invention further provides an intelligent data analysis system based on multi-terminal synchronization, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods of an intelligent data analysis method based on multi-terminal synchronization.
[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. An intelligent data analysis method based on multi-terminal synchronization, applicable to the field of data synchronization of multi-terminal devices, characterized in that: The intelligent data analysis method based on multi-terminal synchronization includes: After data synchronization is performed on each device, in the database of any device, the index value of each leaf node of the B+ tree that needs to be reorganized is obtained, and the data corresponding to each index value is obtained and recorded as the target data; For any target data, in the database log, obtain the data change timestamp under the index value of any target data in the historical period, obtain the number of associated data of any target data according to the data change timestamp of any target data, and obtain the data relationship priority of any target data according to the data change timestamp and the number of associated data; According to the data modification frequency characteristics and modification interval characteristics under the index value of any target data, the modification frequency of any target data is obtained, and according to the stability of the index page under the index value of any target data and the modification frequency, the data location priority of any target data is obtained; The data relationship priority and data location priority of each target data are obtained, and all target data are sorted according to the data relationship priority and data location priority of each target data to obtain corresponding sorting results, and the B+ tree is adaptively indexed and reorganized according to the sorting results.
2. According to claim 1, the intelligent data analysis method based on multi-terminal synchronization is characterized in that: The acquiring the quantity of associated data of any target data according to the data change timestamp of any target data comprises: Obtain the data change timestamps under the index values of all target data in the historical period, divide the historical period into at least two sub-periods, and form a time series sequence with the index values corresponding to the data change timestamps in the same sub-period according to all data change timestamps, to obtain the time series sequence corresponding to each sub-period; Recording the index value of any target data as the target index value, recording the index values of other target data except the any target data as other index values, and for any other index value, recording a time series sequence including both the target index value and the any other index value as a first target time series sequence, and if the number of the first target time series sequences is greater than a preset number, recording the target data corresponding to the any other index value as associated data of the any target data; The number of all associated data of the any target data is recorded as the number of associated data of the any target data.
3. According to claim 2, the intelligent data analysis method based on multi-terminal synchronization is characterized in that: The obtaining the data relationship priority of any target data according to the data change timestamp and the number of associated data includes: If the number of associated data of any target data is 0, the data relationship priority of any target data is recorded as 0; If the number of associated data of any target data is not 0, the time series sequence containing the index value of any target data is recorded as the second target time series sequence, and for any second target time series sequence, the data change timestamp under the first index value in any second target time series sequence is obtained, and in any second target time series sequence, the time interval between the data change timestamp under the index value of any target data and the data change timestamp under the first index value is calculated to obtain the influence coefficient of any target data in the sub-period corresponding to any second target time series sequence; Calculate the average value of the influence coefficient of any target data in the sub-period corresponding to each second target time series sequence to obtain the average influence coefficient of any target data, calculate the inverse of the sum of the average influence coefficient and a constant 1, and obtain the first influence degree of any target data; Subtract the reciprocal of the number of associated data of any target data from a constant 1 to obtain a second influence degree of any target data; A weighted sum of the first influence degree and the second influence degree is taken to obtain a data relationship priority of any target data.
4. The intelligent data analysis method based on multi-terminal synchronization according to claim 1 is characterized in that: The obtaining the modification frequency of any target data according to the data modification frequency characteristic and the modification interval characteristic under the index value of any target data comprises: Obtaining a data modification frequency under the index value of any target data according to the number of data change timestamps under the index value of any target data in the historical period; Calculate the time interval between each two adjacent data change timestamps under the index value of any target data in the historical period to obtain the mean of the time interval, and subtract the reciprocal of the time interval from a constant 1 to obtain the data modification density under the index value of any target data; The data modification frequency and the data modification density are weightedly summed to obtain the modification frequency of any target data.
5. The intelligent data analysis method based on multi-terminal synchronization according to claim 1 is characterized in that: The obtaining the data location priority of any target data according to the stability of the index page under the index value of any target data and the modification frequency includes: Recording an index page under the index value of any target data as a target index page, and obtaining a stability rate of the target index page according to a change in the number of index values corresponding to the target index page between each two adjacent historical moments during the historical period; The product of the stability rate of the target index page and the modification frequency of any target data is calculated to obtain the data location priority of any target data.
6. The intelligent data analysis method based on multi-terminal synchronization according to claim 5 is characterized in that: The obtaining the stability rate of the target index page according to the change in the number of index values corresponding to the target index page between each two adjacent historical moments includes: For any historical moment, the index value of the target index page at any historical moment is grouped into a first set, and the index value of the target index page at the previous historical moment is grouped into a second set. In the first set, the number of index values that are the same as the index values in the second set is obtained, and the proportion of the number of index values in the number of all elements in the first set is calculated to obtain the consistency rate of the target index page at any historical moment. Obtain the consistency rate of the target index page at each of the historical moments, obtain a consistency rate mean and a consistency rate variance, calculate the inverse of the sum of a constant 1 and the consistency rate variance, and use the product of the inverse and the consistency rate mean as the first variable; The consistency rate of the target index page at each of the historical moments is used to form a consistency rate sequence, the absolute value of the difference between the first data and the last data in the consistency rate sequence is calculated, and the reciprocal of the sum of the constant 1 and the absolute value of the difference is used as the second variable; The mean value between the first variable and the second variable is calculated to obtain the stability rate of the target index page.
7. The intelligent data analysis method based on multi-terminal synchronization according to claim 1 is characterized in that: The step of sorting all target data according to the data relationship priority and the data location priority of each target data to obtain corresponding sorting results includes: All target data are arranged in descending order according to the data relationship priority, and target data with the same data relationship priority are sorted in descending order according to the data position priority to obtain a target data sequence.
8. The intelligent data analysis method based on multi-terminal synchronization according to claim 7 is characterized in that: The step of adaptively reorganizing the index of the B+ tree according to the sorting result includes: The target data sequence is updated according to the preset traversal order of the index values in the leaf nodes of the B+ tree, and the B+ tree is adaptively indexed and reorganized according to the updated result.
9. An intelligent data analysis system based on multi-terminal synchronization, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of an intelligent data analysis method based on multi-terminal synchronization as described in any one of claims 1-8 are implemented.