Data processing method and apparatus, device, and storage medium

By obtaining the matching index expression rules in the global index expression rules, writing data to the corresponding storage partition without storing the mapping relationship between data and storage partitions, the problem of wasting storage space and increasing database maintenance resources is solved, and the effect of saving storage space and reducing database maintenance resources is achieved.

WO2025119400A1PCT designated stage expired Publication Date: 2025-06-12CHINA UNIONPAY

Patent Information

Application Number
PCT/CN2024/138208
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-12-10
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the prior art, the mapping relationship storing indexes and data wastes storage space and increases the resources for database maintenance mapping relationships.

Method used

By obtaining the global index expression rules matching the first data in the M global index expression rules, the first data is written to the corresponding memory partition, without storing the relationship data of the first data and the memory partition.

Benefits of technology

When multiple data corresponds to the same global index expression rule, it is implemented without storing the relational data mapped between the data that complies with the global index expression rule and its storage partition, saving a lot of storage space and reducing database maintenance resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138208_12062025_PF_FP_ABST
    Figure CN2024138208_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method and apparatus, a device, and a storage medium. The method comprises: by means of a data writing request which is sent by a requester and used for requesting to write first data into a storage space, acquiring a result of whether M global index expression rules for data division for storage partitions comprise a first global index expression rule matched with the first data, and when the first global index expression rule matched with the first data is obtained from among the M global index expression rules, writing the first data into a first storage partition corresponding to the first global index expression rule among N storage partitions, and not storing data of the relationship between the first data and the first storage partition.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device, equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311648679.0, filed on December 4, 2023, entitled “Data processing method, device, equipment and storage medium,” and the entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present application belongs to the field of data processing technology, and in particular relates to a data identification method, device, equipment and storage medium. Background Art

[0004] With the rapid development of big data technology, the amount of data has exploded. To improve data query speed, databases typically store and maintain a one-to-one mapping relationship between indexes and data. This allows for quick query of the data corresponding to the index through the mapping relationship.

[0005] However, storing the mapping relationship requires wasting a large amount of storage space, and when data in the database is rewritten, the corresponding index also needs to be adaptively adjusted, which increases the resources of the database for maintaining the mapping relationship. Summary of the Invention

[0006] The embodiments of the present application provide a data processing method, apparatus, device and storage medium, which can solve the problem in related technologies of wasting storage space and increasing database resources for maintaining mapping relationships by storing mapping relationships between indexes and data.

[0007] In a first aspect, an embodiment of the present application provides a data processing method, which may include:

[0008] Receive a data write request sent by a requester, where the data write request is used to request writing first data into a storage space, where the storage space includes N storage partitions, where N is an integer greater than 1;

[0009] When a first global index expression rule matching the first data is obtained from among the M global index expression rules, the first data is written to the first storage partition, and the relationship data between the first data and the first storage partition is not stored. The global index expression rule is a rule for dividing the storage partition of the data, and the first storage partition is the storage partition corresponding to the first global index expression rule among the N storage partitions.

[0010] In a second aspect, an embodiment of the present application provides a data processing device, which may include:

[0011] a receiving module, configured to receive a data write request sent by a requesting party, wherein the data write request is used to request writing first data into a storage space, where the storage space includes N storage partitions, where N is an integer greater than 1;

[0012] A processing module is used to write the first data into the first storage partition when a first global index expression rule matching the first data is obtained from M global index expression rules, and not store the relationship data between the first data and the first storage partition. The global index expression rule is a rule for dividing the storage partition of data, and the first storage partition is a storage partition corresponding to the first global index expression rule among the N storage partitions.

[0013] In a third aspect, an embodiment of the present application provides a computer device, the computer device comprising: a processor and a memory storing computer program instructions;

[0014] When the processor executes the computer program instructions, the data processing method shown in the first aspect is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon, which, when executed by a processor, implements the data processing method shown in the first aspect.

[0016] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the data processing method shown in the first aspect.

[0017] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the data processing method shown in the first aspect.

[0018] The data processing method, apparatus, device and storage medium of the embodiments of the present application obtain the result of whether the global index expression rules of M storage partitions for dividing data include a first global index expression rule that matches the first data through a data write request sent by the requesting party for requesting to write the first data into the storage space. When the first global index expression rule that matches the first data is obtained from the M global index expression rules, the first data is written into the first storage partition corresponding to the first global index expression rule among the N storage partitions, and the relationship data between the first data and the first storage partition is not stored. In this way, the correlation of the storage partitions of the divided data is described by the global index expression rules, and then it is determined whether the first data to be written in the data write request satisfies the global index expression rules. If the first data satisfies the first global index expression rule among the M global index expression rules, the first data is directly written into the first storage partition corresponding to the first global index expression rule, and there is no need to store the relationship data mapped between the first data and the first storage partition. This achieves that when multiple data correspond to the same global index expression rule, there is no need to store the relationship data mapped between the data that meets the global index expression rule and its storage partition, saving a lot of storage space. In addition, there is no need for the database to maintain the associated data between the data that meets the global index expression rule and its storage partition, reducing the maintenance resources of the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] FIG1 is a schematic diagram of the structure of a data processing system provided in an embodiment of the present application;

[0021] FIG2 is a flow chart of a data processing method provided in an embodiment of the present application;

[0022] FIG3 is a schematic diagram of a binary tree model storing association relationships in a data processing method provided in an embodiment of the present application;

[0023] FIG4 is a schematic diagram of a data query process in a data processing method provided in an embodiment of the present application;

[0024] FIG5 is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0025] FIG6 is a schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0026] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0027] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0028] The acquisition, storage, use, and processing of data (including but not limited to the features and information herein) in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0029] An index is a data structure used by storage engines to quickly locate records. For example, databases can create indexes to quickly access data in database tables or views. Common indexes are implemented by maintaining an accurate mapping between the index and the data. This generally requires a significant amount of storage space and requires dynamic maintenance when adding, deleting, or modifying data in a table. This slows down data write operations and, when data is frequently written and updated, can degrade database performance.

[0030] With the further development of digitalization and intelligentization of information technology, the computing and storage scenarios of large amounts of data can no longer be supported by single-machine storage alone, which has led to a further increase in the demand and usage of distributed databases. In related technologies, massive amounts of data can be split horizontally and the split data can be stored in different physical nodes in a distributed storage manner, so that the same table data is maintained in multiple partitions. Since the local index key values ​​of a distributed database may exist in all partitions, the local index can no longer meet the query requirements of distributed storage. Therefore, in order to avoid data omissions, it is necessary to perform a scan based on the local index in all partitions. In this way, the scanning method that relies solely on the local index is inefficient and will cause a waste of processor and input / output resources on a global scale, reducing database performance.

[0031] In order to solve the above technical problems, the embodiments of the present application provide a data processing method, apparatus, computer equipment and storage medium.

[0032] Based on this, the data processing method, apparatus, computer equipment and storage medium of the embodiments of the present application will be described in detail below with reference to Figures 1 to 6. It should be noted that these embodiments are not intended to limit the scope of disclosure of the present application.

[0033] FIG1 is a schematic diagram of a data processing system provided in an embodiment of the present application.

[0034] As shown in Figure 1, the data processing system 10 provided in the embodiment of the present application can be set in a database or a retrieval engine. The database can be a distributed database, such as a column storage database based on Hadoop or Cassandra, a MongoDB document database, and a Redis key-value storage database. Among them, the data processing system 10 may include a request receiving module 101, an expression rule verification module 102, an index mapping relationship module 103, and a storage space 104. Furthermore, the storage space 104 includes N storage partitions, such as a first storage partition, a second storage partition, a third storage partition, ..., an Nth storage partition, where N is an integer greater than 1.

[0035] The data processing system 10 is described in detail below, as shown below.

[0036] The request receiving module 101 performs data interaction with the expression rule verification 102 and the index mapping relationship module 103 respectively. Among them, the request receiving module 101 is used to receive various types of requests sent by the requester, such as data write requests, data query requests and data modification requests. And, according to the type of request, it is determined to perform data interaction with the expression rule verification 102 or the index mapping relationship module 103. Specifically, if the request receiving module 101 receives a data write request sent by the requester, it sends its data write request to the expression rule verification 102. If the request receiving module 101 receives a data query request sent by the requester, it sends its data query request to the index mapping relationship module 103.

[0037] The expression rule verification module 102 stores M global index expression rules and exchanges data with the request receiving module 101, the index mapping relationship module 103, and the storage space 104. The expression rule verification module 102 is configured to, upon receiving a data write request from the request receiving module 101, determine whether a first global index expression rule matching the first data can be obtained from the M global index expression rules; if a first global index expression rule matching the first data is obtained from the M global index expression rules, write the first data into the first storage partition in the storage space 104, and not store the relationship data between the first data and the first storage partition. In addition, the expression rule verification 102 can also be used to, when a data query request sent by the request receiving module 101 is received and the index mapping relationship module 103 does not include the second correlation global index corresponding to the third data that the data query request wants to query, receive the data query request sent by the index mapping relationship module 103; determine whether a second global index expression rule matching the third data can be obtained from the M global index expression rules; when a second global index expression rule matching the third data is obtained from the M global index expression rules, read the third data from the fifth storage partition corresponding to the second global index expression rule, and send the third data to the requesting party.

[0038] The index mapping module 103 stores association information between a relevance global index and a partition key of a storage partition. The index mapping module 103 can interact with the expression rule validation 102, the request receiving module 101, and the storage space 104 for data exchange. Specifically, upon receiving a data write request from the request receiving module 101 and failing to obtain a target global index expression rule matching the first data in the data write request from the expression rule validation 102, the index mapping module 103 is configured to construct a first relevance global index, which represents the correspondence between the first data and the second storage partition, and write the first data to the second storage partition in the storage space 104. The index mapping module 103 can also be configured to, upon receiving a data read request from the request receiving module 101, traverse a binary tree model; if the binary tree model includes a second relevance global index corresponding to the third data to be queried by the data read request, read third data from a fourth storage partition corresponding to the partition key of the fourth storage partition based on the association information between the second relevance global index stored in the binary tree model and the partition key of the fourth storage partition; and send the third data to the requesting party.

[0039] The storage space 104 can interact with the expression rule verification 102 and the index mapping relationship module 103 respectively to store data.

[0040] It should be noted that the data processing method provided in the embodiments of the present application is applicable to scenarios of distributed data retrieval, as well as scenarios in which global indexes are created for fields with specific relevance to partition keys to store data.

[0041] Therefore, the data processing system in the embodiment of the present application can describe the storage partition rules of most data through the global index expression rules in the expression rule verification 102, and specify the storage partitions of some special data that do not meet the global index expression rules through the association information between the correlation global index and the partition key of the storage partition in the index mapping relationship module 103. When writing data, if it is determined that the data does not match the global index expression rule, the data is maintained in the accurate correlation global index; when querying data, the storage partition of the data is first queried through the association information between the accurate correlation global index and the partition key of the storage partition. If the association information between the accurate correlation global index and the partition key of the storage partition cannot be fully matched to the specific partition, the storage partition where the data to be queried is located is calculated according to the global index expression rule.

[0042] In this way, the correlation of the storage partitions for dividing the data is described by the global index expression rules, and then it is determined whether the first data to be written in the data write request satisfies the global index expression rules. If the first data satisfies the first global index expression rule among the M global index expression rules, the first data is directly written to the first storage partition corresponding to the first global index expression rule, and there is no need to store the relationship data mapped between the first data and the first storage partition. This achieves that when multiple data correspond to the same global index expression rule, there is no need to store the relationship data mapped between the data that meets the global index expression rule and its storage partition, saving a lot of storage space. In addition, when writing data that meets the global index expression rule, there is no need for the database to maintain the associated data between the data that meets the global index expression rule and its storage partition, reducing the maintenance resources of the database. In addition, in the application embodiment, there is no need to add or modify the accurate index mapping relationship, which improves the writing speed. The accurate index mapping relationship only needs to maintain a small amount of special data, has a faster retrieval and update speed, can improve the processing performance of the write operation, avoid resource waste and ensure memory performance.

[0043] Based on this, in order to better illustrate the contents of the embodiment of the present application, a data processing method provided in the embodiment of the present application is described below in combination with Figure 2, as shown below.

[0044] FIG2 is a flow chart of a data processing method provided in an embodiment of the present application.

[0045] As shown in FIG2 , the data processing method can be applied to a data processing system of a data application party, and the data processing system can be set in a computer device. The data processing method can specifically include the following steps:

[0046] Step 210: Receive a data write request sent by the requesting party, where the data write request is used to request writing the first data into the storage space, where the storage space includes N storage partitions, where N is an integer greater than 1; Step 230: When a first global index expression rule matching the first data is obtained from among the M global index expression rules, write the first data into the first storage partition, and do not store the relationship data between the first data and the first storage partition, where the global index expression rule is a rule for dividing the storage partitions for data, and the first storage partition is the storage partition corresponding to the first global index expression rule among the N storage partitions.

[0047] In this way, the correlation of the storage partitions of the divided data is described by the global index expression rules, and then it is determined whether the first data to be written in the data write request satisfies the global index expression rules. If the first data satisfies the first global index expression rule among the M global index expression rules, the first data is directly written into the first storage partition corresponding to the first global index expression rule, and there is no need to store the relationship data mapped between the first data and the first storage partition. This achieves that when multiple data correspond to the same global index expression rule, there is no need to store the relationship data mapped between the data that meets the global index expression rule and its storage partition, saving a lot of storage space. In addition, there is no need for the database to maintain the associated data between the data that meets the global index expression rule and its storage partition, reducing the maintenance resources of the database.

[0048] The above steps are described in detail below.

[0049] First, regarding step 220, the embodiment of the present application provides the following four methods of obtaining a first global index expression rule that matches the first data from M global index expression rules, as shown below.

[0050] In some embodiments of the present application, the first data includes a first key field. Based on this, before step 220, the data processing method may further include steps 2301 and 2302.

[0051] Step 2301: Match the first key field with the key field of each global index expression rule in the M global index expression rules to obtain a first matching result.

[0052] Step 2302: When the first matching result indicates that there is a target key field matching the first key field among the M global index expression rules, the global index expression rule including the target key field is determined as the first global index expression rule.

[0053] Exemplarily, the key field may be a flow identifier (Flow id), which may be customized by the data processing system. In this case, the first key field, such as Flow id1, may be matched with the key field of a pre-stored global index expression rule, and the global index expression rule including the first key field Flow id1 among the M global index expression rules may be determined as the first global index expression rule. In this way, the function rule can be used in business to reversely calculate the partition key (create_time) of the storage partition.

[0054] In some other embodiments of the present application, before step 220 , the data processing method may further include steps 2401 and 2402 .

[0055] Step 2401: Calculate the similarity values ​​between the data content of the first data and the data content divided by each of the M global index expression rules through a similarity algorithm.

[0056] For example, if the data content of the first data relates to business 1, then the similarity value between the data content divided by each global index expression in the M global index expression rules and business 1 can be calculated.

[0057] Step 2402: When the similarity value is greater than or equal to a preset threshold, the global index expression rule corresponding to the similarity value is determined as the first global index expression rule.

[0058] Exemplarily, based on the above step 2401, M similarity values ​​can be obtained. At this time, the global index expression rule corresponding to the largest similarity value among the M similarity values ​​can be determined as the first global index expression rule to filter out the global index expression rule used to divide the data of business 1 from the M global index expression rules.

[0059] In some other embodiments of the present application, before step 220 , the data processing method may further include steps 2501 and 2502 .

[0060] Step 2501: Obtain a first data type of first data, where the first data type is a numeric type or a non-numeric type.

[0061] Step 2502: Determine the global index expression rule used to divide data of the first data type among the M global index expression rules as the first global index expression rule.

[0062] Exemplarily, the M global index expression rules may include global index expression rule 1 for dividing numeric data and global index expression rule 2 for dividing non-numeric data. In this way, if the first data type of the first data is, for example, numeric, global index expression rule 1 may be determined as the first global index expression rule.

[0063] In some other embodiments of the present application, before step 220 , the data processing method may further include steps 2601 and 2602 .

[0064] Step 2601: Obtain the application service scenario to which the first data belongs.

[0065] For example, the service scenario to which the acquired first data belongs is an offline payment scenario or a transfer scenario.

[0066] Step 2602: Determine the global index expression rule used to divide data related to the application service scenario among the M global index expression rules as the first global index expression rule.

[0067] Exemplarily, the global index expression rule for dividing data for offline payment scenarios among the M global index expression rules is determined as the first global index expression rule; or, the global index expression rule for data for account transfer scenarios among the M global index expression rules is determined as the first global index expression rule.

[0068] In addition, in an embodiment of the present application, there is also provided a data writing process that is executed if the first global index expression rule that matches the first data is not obtained in the M global index expression rules, that is, after step 210, the N storage partitions include the second storage partition. Based on this, the data processing method may also include steps 2701 and 2702. Here, it should be noted that the application principle of the global index expression rule in the embodiment of the present application can be illustrated by the following example. For example, if the global index expression rule indicates that the data from D1 to D100 should be stored in the storage space A1, then if the first data is D2, then D2 will be matched with the M global index expression rules. If there is a global index expression rule that divides the first data into the storage space A1, it can be determined as the first global index expression rule that matches the first data. Conversely, if the first data is divided into the storage space A2, it can be represented that there is no global index expression rule that matches the first data in the M global index expression rules.

[0069] Based on this, in step 2701, if no target global index expression rule matching the first data is obtained from the M global index expression rules, a first relevance global index is constructed, where the first relevance global index is used to represent the correspondence between the first data and the second storage partition. For example, the first data D1 is stored in storage partition A2.

[0070] Therefore, the correlation global index in the embodiment of the present application directly associates the corresponding relationship between the data and the storage partition, so as to facilitate users to directly query the data.

[0071] Furthermore, the embodiments of the present application provide at least four methods for constructing the first relevance global index, as shown below.

[0072] In some embodiments of the present application, step 2701 may specifically include:

[0073] The first data is associated with a first storage partition among the N storage partitions by using a preset partitioning rule corresponding to the N storage partitions to obtain a first correlation global index.

[0074] Exemplarily, the preset partition rule may include a partition_rule partition rule.

[0075] In other embodiments of the present application, the relationship between data and storage partitions can be customized by flexibly defining expression rules. For example, due to the correlation between transaction variables, a correlation relationship can be used to represent the uncertain correlation between variables. In most business scenarios, the creation time and update time of the records, and the start time and end time of the business can have a linear correlation. For example, there can also be a correlation between some indicator quantities in the data processing system and some characteristic values ​​in data analysis.

[0076] Therefore, embodiments of the present application can associate data and storage partitions based on the content of the data and the attributes of the data itself to determine a first relevance global index, such as calculating time relevance. Using creation time as the partition key, the creation time and update time of online business scenarios will be very close. Using both as inputs to the range partitioning rule, the calculated values ​​are generally the same partition. The update time of batch business scenarios is generally a fixed number of days different from the creation time. A first relevance global index can be established based on the creation time, create_time; if it is an online scenario, the first relevance global index can be established based on the update time, update_time; and if it is a T+1 batch scenario, the first relevance global index can be established based on the update time, update_time.

[0077] In some other embodiments of the present application, if the applicant specifies a function for partition calculation, the data and storage partitions may be associated according to the function specified by the applicant. Based on this, the first data carries a first function name, and the function corresponding to the first function name is the associated function specified by the requester. Based on this, step 2701 may specifically include:

[0078] The first data is associated with the first storage partition of the N storage partitions through an association function corresponding to the first function name to obtain a first correlation global index.

[0079] For example, in the embodiment of the present application, the first function name may be proname, pronamespace, or proowner. Based on this, if the first function name is proname, the associated function corresponding to proname may be called, and the first data may be associated with the first storage partition among the N storage partitions through the associated function to obtain a first correlation global index.

[0080] In some other embodiments of the present application, a correlation global index can be created through a database adaptive fitting method, and the M global index expression rules include a first linear fitting relationship expression rule. Based on this, step 2701 can specifically include steps 27011 to 27013.

[0081] Step 27011: convert the data stored in each of the N storage partitions within the first time window to obtain N data segments, where the N data segments are continuous data segments, and the data in the N data segments are linearly positively correlated.

[0082] Step 27012: When the target data segment is included in the N data segments, the target data in the target data segment is obtained by using a clustering algorithm, and the amount of data in the target data segment is greater than or equal to a preset threshold.

[0083] For example, a clustering algorithm may be used to remove some discrete data to obtain a data segment with a larger data volume among the N data segments.

[0084] Step 27013: Perform linear fitting on the target data to obtain a first linear fitting relationship expression rule.

[0085] Specifically, here, the target data may be linearly fitted using the least squares method to obtain a first linear fitting relationship expression rule.

[0086] Based on this, steps 27011 to 27013 are explained below through an example.

[0087] For example, in a distributed retrieval scenario, index creation using the adaptive fitting method only supports the establishment of a global index based on correlation for time and numeric type fields. The segmented adaptive linear fitting implementation algorithm can use cluster analysis to obtain a fittable interval, and use the least squares method to perform linear fitting within the interval, ultimately obtaining a segmented linear fitting expression. When the specific adaptive fitting algorithm is implemented, all partition ranges are converted into linearly growing continuous data segments. For example, [0,100) belongs to partition 1, [100,200) belongs to partition 2, and [200,300) belongs to partition 3.

[0088] When the amount of data in a partition reaches the cluster analysis threshold, for example, 10,000, the K-means algorithm cluster analysis can be used to obtain the index segmentation range corresponding to the partition and eliminate outlier data. Afterwards, linear fitting is performed within the valid data range. Generally, the least squares method can be used to implement the fitting to obtain the linear fitting relationship expression rule. According to the linear fitting relationship expression rule, the data partition range that the partition can adapt to is calculated. It should be noted that segmented linear fitting relationship expression rules can be obtained based on different partitions, and the index segmentation ranges of the linear fitting relationship expression rules are not allowed to overlap. For example, data from D1 to D100 should be assigned to storage partition A1, and data from D101 to D220 should be assigned to storage partition A2. At this time, the data from D99-D110 should not be assigned to storage partition A3.

[0089] In addition, an embodiment of the present application also provides a method for updating the first linear fitting relationship expression rule, that is, the N storage partitions include a second storage partition and a third storage partition, and the second storage partition is used to store data that does not match the global index expression rule. Based on this, after step 27013, the data processing method in the embodiment of the present application can also include steps 27014 to 27016.

[0090] Step 27014: Obtain a second linear fitting relationship expression rule for the data stored in each storage partition of the N storage partitions within the second time window.

[0091] Step 27015: traverse the second storage partition based on the second linear fitting relationship expression rule.

[0092] Step 27016: When second data matching the second linear fitting relationship expression rule is obtained from the second storage partition, the second data is moved from the second storage partition to a third storage partition corresponding to the second linear fitting relationship expression rule.

[0093] In addition, it should be noted that the association information between the second storage partition and the first correlation global index is stored in a binary tree model. The M global index expression rules and the binary tree model can be updated in the following manner. Based on this, after step 27016, the data processing method in the embodiment of the present application can also include steps 27017 and 27018.

[0094] Step 27017: Remove the association information between the first correlation global index and the partition key of the second storage partition from the binary tree model.

[0095] Step 27018: Replace the first linear fitting relationship expression rule in the M global index expression rules with the second linear fitting relationship expression rule.

[0096] Exemplarily, online adjustment of the association information between the global index of correlation and the partition key of the storage partition can be achieved through the following process, wherein the database can be selected to adaptively fit the correlation between the index and the storage partition. Since the adaptive fitting based on partial data may deviate significantly from the final actual data distribution, it is desirable to be able to adaptively balance the correlation after the data distribution changes. Here, before the actual online adjustment is performed, a set of second linear fitting relationship expression rules will be recalculated based on the current data distribution. During the transition process of performing the online adjustment, the current first linear fitting relationship expression rules and the adjusted second linear fitting relationship expression rules are simultaneously effective, and a partition key mapping relationship binary tree, such as a B+ tree, is rebuilt for the data that satisfies the association information between the global index of correlation and the partition key of the storage partition. When the partition key mapping relationship B+ tree is rebuilt, data that meets the adjusted expression is removed from the B+ tree, and data that does not meet the adjusted expression is added to the B+ tree. After the partition key mapping relationship B+ tree is rebuilt, the adjusted expression is used as the matching rule, and the original expression is no longer effective.

[0097] In this way, online adjustment of the global index of correlation can be achieved. When the proportion of correlation changes, seamless adaptive transition conversion is supported to achieve online adjustment.

[0098] Step 2702: Write the first data into the second storage partition.

[0099] In which, the embodiment of the present application also provides a method for storing association information between the first correlation global index and the partition key of the second storage partition, that is, after step 2702, the data processing method can also include step 2703.

[0100] Step 2703: Store the association information between the first relevance global index and the partition key of the second storage partition in a binary tree model. The binary tree model includes leaf nodes and non-leaf nodes. The leaf nodes are used to store the partition key of the second storage partition, and the non-leaf nodes are used to store the first relevance global index and primary key information. The primary key information is used to identify the association between the first relevance global index and the partition key of the second storage partition.

[0101] For example, as shown in FIG3 , an example diagram of the mapping relationship between indexes and partition keys is shown. Specifically, the binary tree model in the embodiment of the present application may be a B+ tree. Therefore, for special data, it is necessary to maintain an accurate relationship between the relevance global index and the partition key. In this embodiment, a typical B+ tree is used, with the partition key stored in the leaf nodes, and non-leaf nodes only storing the relevance global index and primary key information.

[0102] Based on the above-mentioned data writing process, an embodiment of the present application also provides a data query process, that is, the N storage partitions also include a fourth storage partition. Based on this, after step 220, the data processing method can also include steps 2801 to 2804.

[0103] Step 2801: Receive a data query request sent by a requesting party, where the data query request is used to request to query third data from a storage space.

[0104] Exemplarily, the requester wants to query third data such as D3.

[0105] Step 2802: traverse the binary tree model according to the data query request.

[0106] For example, since the amount of association information between the correlation global index stored in the binary tree model and the partition key of the storage partition is small, during the data query process, the data is first matched with the association information between the correlation global index and the partition key of the second storage partition.

[0107] Step 2803, in the case where the binary tree model includes a second correlation global index corresponding to the third data, the third data is read from the fourth storage partition corresponding to the partition key of the fourth storage partition based on the association information between the second correlation global index stored in the binary tree model and the partition key of the fourth storage partition.

[0108] For example, if the storage partition A3 of D3 is matched through the association information between the correlation global index and the partition key of the second storage partition, the storage partition A3 can be directly searched without the need for other data matching.

[0109] In this way, matching time can be reduced and query efficiency can be improved.

[0110] Step 2804: Send the third data to the requesting party.

[0111] Alternatively, when the second correlation global index corresponding to the third data is not included in the binary tree model, the N storage partitions also include a fifth storage partition. Based on this, after step 2802, the data processing method may also include steps 2805 to 2807.

[0112] Step 2805: When the binary tree model does not include the second correlation global index corresponding to the third data, the M global index expression rules are matched with the third data to obtain a second matching result.

[0113] Step 2806, when the second matching result indicates that the M global index expression rules include the third data that matches the second global index expression rule, read the third data from the fifth storage partition corresponding to the second global index expression rule.

[0114] Step 2807: Send the third data to the requesting party.

[0115] For example, if the association information between the correlation global index and the partition key of the second storage partition does not match the storage partition A3 of D3, then M global index expression rules are used to match it. In this way, when the second global index expression rule that matches D3 is obtained among the M global index expression rules, the storage partition A5 associated with the second global index expression rule is determined as the storage partition where D3 is located. In this way, D3 can be directly extracted from A5. In this way, the operation of traversing all storage partitions can be reduced, thereby improving query efficiency. Based on the above-mentioned data writing process, the embodiment of the present application also provides a data rewriting process. Based on this, after step 220, the data processing method can also include steps 2901 to 2902.

[0116] Step 2901: Receive a data modification request sent by a requesting party, where the data modification request is used to request that first data stored in a first storage partition be modified into fourth data.

[0117] Exemplarily, take the modification of the first data D1 of the first storage partition A1 into the fourth data D4 as an example.

[0118] Step 2902: When the fourth data matches the first global index expression rule, the fourth data is written into the first storage partition, and the relationship data between the fourth data and the first storage partition is not stored.

[0119] For example, since the first global index expression rule matching D1 limits the data responses of D1-D100 to the first storage partition A1, the modified D4 does not exceed the range of D1-D100. Therefore, there is no need to adjust the storage partition of D4, nor is there any need to store the relationship data between the fourth data and the first storage partition.

[0120] In this way, during the data rewriting process, if the rewritten data still matches the global index expression rules associated with the data before rewriting, the database does not need to maintain the associated data between the data that meets the global index expression rules and its storage partition, which reduces the database maintenance resources and improves the performance of database rewriting and writing data.

[0121] Alternatively, after step 2901 , the N storage partitions further include a fifth storage partition. Based on this, the data processing method may further include steps 2903 to 2904 .

[0122] Step 2903: When the fourth data does not match the first global index expression rule, each of the M global index expression rules is matched with the fourth data to obtain a third matching result.

[0123] Step 2904, when the third matching result indicates that the M global index expression rules include the third global index expression rule that matches the fourth data, the fourth data is written to the fifth storage partition corresponding to the third global index expression rule, and the relationship data between the fourth data and the fifth storage partition is not stored.

[0124] For example, still taking the above example as an example, if the fourth data is D150, then according to the global index expression rules, the storage partition of the data D101-D200 should be A2, then, at this time, the modified fourth data D150 will be adjusted to A2.

[0125] Therefore, the storage partition rules of most data can be described by the global index expression rules in the expression rule verification 102, and the storage partitions of some special data that do not meet the global index expression rules can be specified by the association information between the correlation global index and the partition key of the storage partition in the index mapping relationship module 103. When writing data, if it is determined that the data does not match the global index expression rule, the data is maintained in the accurate correlation global index; when querying data, the storage partition of the data is first queried through the association information between the accurate correlation global index and the partition key of the storage partition. If the association information between the accurate correlation global index and the partition key of the storage partition cannot be completely matched to the specific partition, the storage partition where the data to be queried is located is calculated according to the global index expression rule.

[0126] In this way, the correlation of the storage partitions for dividing the data is described by the global index expression rules, and then it is determined whether the first data to be written in the data write request satisfies the global index expression rules. If the first data satisfies the first global index expression rule among the M global index expression rules, the first data is directly written to the first storage partition corresponding to the first global index expression rule, and there is no need to store the relationship data mapped between the first data and the first storage partition. This achieves that when multiple data correspond to the same global index expression rule, there is no need to store the relationship data mapped between the data that meets the global index expression rule and its storage partition, saving a lot of storage space. In addition, when writing data that meets the global index expression rule, there is no need for the database to maintain the associated data between the data that meets the global index expression rule and its storage partition, reducing the maintenance resources of the database. In addition, in the application embodiment, there is no need to add or modify the accurate index mapping relationship, which improves the writing speed. The accurate index mapping relationship only needs to maintain a small amount of special data, has a faster retrieval and update speed, can improve the processing performance of the write operation, avoid resource waste and ensure memory performance.

[0127] The present application also provides a data processing device, which is described in detail in conjunction with FIG5 .

[0128] FIG5 is a schematic structural diagram of a data processing device provided in one embodiment of the present application.

[0129] In some embodiments of the present application, the data processing device shown in FIG. 5 may be provided in the data processing system shown in FIG. 1 .

[0130] As shown in FIG5 , the data processing device 50 may specifically include:

[0131] A receiving module 501 is configured to receive a data write request sent by a requesting party, where the data write request is used to request writing first data into a storage space, where the storage space includes N storage partitions, where N is an integer greater than 1;

[0132] Processing module 502 is used to write the first data into the first storage partition when a first global index expression rule matching the first data is obtained from M global index expression rules, and not store the relationship data between the first data and the first storage partition. The global index expression rule is a rule for dividing the storage partition of data, and the first storage partition is the storage partition corresponding to the first global index expression rule among the N storage partitions.

[0133] The data processing device 50 in the embodiment of the present application is described in detail below.

[0134] In some embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a construction module and a first writing module; wherein,

[0135] A construction module is used to construct a first correlation global index when N storage partitions include a second storage partition and no target global index expression rule matching the first data is obtained from M global index expression rules. The first correlation global index is used to characterize the correspondence between the first data and the second storage partition.

[0136] The first writing module is used to write the first data into the second storage partition.

[0137] In some other embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a storage module; wherein,

[0138] A storage module, configured to store association information between the first correlation global index and the partition key of the second storage partition in a binary tree model;

[0139] Among them, the binary tree model includes leaf nodes and non-leaf nodes, the leaf nodes are used to store the partition key of the second storage partition, and the non-leaf nodes are used to store the first correlation global index and primary key information, and the primary key information is used to identify the association relationship between the first correlation global index and the partition key of the second storage partition.

[0140] In some other embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a first association module; wherein,

[0141] The first associating module is configured to associate the first data with a first storage partition among the N storage partitions by using a preset partitioning rule corresponding to the N storage partitions to obtain a first correlation global index.

[0142] In some further embodiments of the present application, the data processing device 50 in the embodiment of the application may further include a second association module; wherein,

[0143] The second association module is used to associate the first data with the first storage partition among N storage partitions through the association function corresponding to the first function name, so as to obtain a first correlation global index when the first data carries a first function name and the function corresponding to the first function name is the association function specified by the requester.

[0144] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a conversion module, a first acquisition module and a fitting module; wherein,

[0145] a conversion module, configured to, when the M global index expression rules include a first linear fitting relationship expression rule, convert data stored in each of the N storage partitions within a first time window to obtain N data segments, where the N data segments are continuous data segments and data in the N data segments are linearly positively correlated;

[0146] A first acquisition module is configured to acquire target data in the target data segment by using a clustering algorithm when the target data segment is included in the N data segments, and the amount of data in the target data segment is greater than or equal to a preset threshold;

[0147] The fitting module is used to perform linear fitting on the target data to obtain a first linear fitting relationship expression rule.

[0148] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a second acquisition module, a first traversal module and a movement module; wherein,

[0149] A second acquisition module is configured to, when the N storage partitions include a second storage partition and a third storage partition, and the second storage partition is configured to store data that does not match the global index expression rule, acquire a second linear fitting relationship expression rule for data stored in each storage partition in the N storage partitions within a second time window;

[0150] A first traversal module, configured to traverse the second storage partition based on a second linear fitting relationship expression rule;

[0151] The moving module is used to move the second data from the second storage partition to a third storage partition corresponding to the second linear fitting relationship expression rule when the second data matching the second linear fitting relationship expression rule is obtained from the second storage partition.

[0152] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a removal module and a replacement module; wherein,

[0153] a removal module, configured to remove the association information between the first correlation global index and the partition key of the second storage partition from the binary tree model when the association information between the second storage partition and the first correlation global index is stored in the binary tree model;

[0154] The replacement module is used to replace the first linear fitting relationship expression rule in the M global index expression rules with the second linear fitting relationship expression rule.

[0155] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a first matching module and a first determining module; wherein,

[0156] a first matching module, configured to, when the first data includes a first key field, match the first key field with the key field of each of the M global index expression rules to obtain a first matching result;

[0157] The first determining module is configured to determine the global index expression rule including the target key field as the first global index expression rule when the first matching result indicates that there is a target key field matching the first key field among the M global index expression rules.

[0158] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a calculation module and a second determination module; wherein,

[0159] a calculation module, configured to calculate, by a similarity algorithm, a similarity value between the data content of the first data and the data content divided by each of the M global index expression rules;

[0160] The second determining module is configured to determine, when the similarity value is greater than or equal to a preset threshold, the global index expression rule corresponding to the similarity value as the first global index expression rule.

[0161] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a third acquisition module and a third determination module; wherein,

[0162] A third acquisition module is used to acquire a first data type of the first data, where the first data type is a numerical type or a non-numerical type;

[0163] The third determining module is configured to determine the global index expression rule for dividing data of the first data type among the M global index expression rules as the first global index expression rule.

[0164] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a fourth acquisition module and a fourth determination module; wherein,

[0165] A fourth acquisition module is used to acquire the application service scenario to which the first data belongs;

[0166] The fourth determining module is used to determine the global index expression rule used for dividing data of the application service scenario among the M global index expression rules as the first global index expression rule.

[0167] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a second traversal module, a first reading module and a first sending module; wherein,

[0168] The receiving module 501 may also be configured to, when the N storage partitions further include a fourth storage partition, receive a data query request sent by a requester, where the data query request is used to request querying third data from the storage space;

[0169] The second traversal module is used to traverse the binary tree model according to the data query request;

[0170] a first reading module, configured to, when the binary tree model includes a second relevance global index corresponding to the third data, read the third data from a fourth storage partition corresponding to the partition key of the fourth storage partition based on association information between the second relevance global index stored in the binary tree model and the partition key of the fourth storage partition;

[0171] The first sending module is configured to send third data to the requesting party.

[0172] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a second matching module, a second reading module and a second sending module; wherein,

[0173] a second matching module, configured to, when the N storage partitions also include a fifth storage partition and the binary tree model does not include a second correlation global index corresponding to the third data, match the M global index expression rules with the third data to obtain a second matching result;

[0174] A second reading module is configured to read the third data from a fifth storage partition corresponding to the second global index expression rule when the second matching result indicates that the M global index expression rules include a second global index expression rule that matches the third data;

[0175] The second sending module is used to send third data to the requesting party.

[0176] In some further embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a second writing module; wherein,

[0177] The receiving module 501 may also be configured to receive a data modification request sent by a requesting party, where the data modification request is configured to request that the first data stored in the first storage partition be modified into fourth data.

[0178] The second writing module is used to write the fourth data into the first storage partition when the fourth data matches the first global index expression rule, and not store the relationship data between the fourth data and the first storage partition.

[0179] In some further embodiments of the present application, the data processing device 50 in the embodiment of the application may further include a third matching module; wherein,

[0180] a third matching module, configured to, when the N storage partitions also include a fifth storage partition and the fourth data does not match the first global index expression rule, respectively match each of the M global index expression rules with the fourth data to obtain a third matching result;

[0181] The processing module 502 can also be used to write the fourth data into the fifth storage partition corresponding to the third global index expression rule when the third matching result indicates that the M global index expression rules include the fourth data that matches the third global index expression rule, and not store the relationship data between the fourth data and the fifth storage partition.

[0182] Thus, the data processing device of the embodiment of the present application obtains the result of whether the global index expression rules of the M storage partitions for dividing the data include the first global index expression rule that matches the first data through the data write request sent by the requester for requesting to write the first data into the storage space. When the first global index expression rule that matches the first data is obtained from the M global index expression rules, the first data is written to the first storage partition corresponding to the first global index expression rule among the N storage partitions, and the relationship data between the first data and the first storage partition is not stored. In this way, the correlation of the storage partitions for dividing the data is described by the global index expression rule, and then it is determined whether the first data to be written in the data write request satisfies the global index expression rule. If the first data satisfies the first global index expression rule among the M global index expression rules, the first data is directly written to the first storage partition corresponding to the first global index expression rule, and there is no need to store the relationship data mapped between the first data and the first storage partition. This achieves that when multiple data correspond to the same global index expression rule, there is no need to store the relationship data mapped between the data that meets the global index expression rule and its storage partition, saving a lot of storage space, and eliminating the need for the database to maintain the associated data between the data that meets the global index expression rule and its storage partition, reducing the maintenance resources of the database.

[0183] Based on the same inventive concept, the present application also provides a computer device, which will be described in detail with reference to FIG6 .

[0184] FIG6 is a schematic diagram of the structure of a computer device provided in one embodiment of the present application.

[0185] As shown in Figure 6, the computer device may include at least one of the following involved in the embodiments of the present application: an electronic device, a server. The computer device may include a processor 601 and a memory 602 storing computer program instructions.

[0186] Specifically, the processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASTC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0187] The memory 602 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 602 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid-state memory. In a specific embodiment, the memory 602 includes a solid-state storage (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0188] The processor 601 implements any one of the data processing methods in the above embodiments by reading and executing computer program instructions stored in the memory 602 .

[0189] In one example, the computer device may further include a communication interface 603 and a bus 610. As shown in FIG6, the processor 601, the memory 602, and the communication interface 603 are connected via the bus 610 and communicate with each other.

[0190] The communication interface 603 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0191] Bus 610 comprises hardware, software or both, and the parts of flow control device are coupled to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard system (ETSA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard system (TSA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel system (MCA) bus, peripheral component interconnection (PCT) bus, PCT-Express (PCT-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more above these combinations.In suitable case, bus 610 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0192] The data processing device can execute the data processing method in the embodiment of the present application, thereby realizing the data processing method and apparatus described in conjunction with Figures 1 to 6.

[0193] In addition, in conjunction with the data processing methods in the above embodiments, embodiments of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the data processing methods in the above embodiments is implemented.

[0194] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0195] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0196] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0197] The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application.

Claims

1. A data processing method, comprising: Receive a data write request sent by a requester, wherein the data write request is used to request to write first data into a storage space, wherein the storage space includes N storage partitions, where N is an integer greater than 1; When a first global index expression rule matching the first data is obtained from among the M global index expression rules, the first data is written into a first storage partition, and relationship data between the first data and the first storage partition is not stored. The global index expression rule is a rule for dividing storage partitions for data, and the first storage partition is a storage partition among the N storage partitions corresponding to the first global index expression rule.

2. The method according to claim 1, wherein: The N storage partitions include a second storage partition; and the method further includes: In the case that a target global index expression rule matching the first data is not obtained from the M global index expression rules, a first correlation global index is constructed, where the first correlation global index is used to characterize the corresponding relationship between the first data and the second storage partition; The first data is written into the second storage partition.

3. The method according to claim 2, wherein: The method further comprises: Storing association information between the first correlation global index and the partition key of the second storage partition in a binary tree model; Among them, the binary tree model includes leaf nodes and non-leaf nodes, the leaf nodes are used to store the partition key of the second storage partition, and the non-leaf nodes are used to store the first correlation global index and primary key information, and the primary key information is used to identify the association relationship between the first correlation global index and the partition key of the second storage partition.

4. The method according to claim 2, wherein: The step of constructing a first correlation global index comprises: The first data is associated with a first storage partition among the N storage partitions by using a preset partitioning rule corresponding to the N storage partitions to obtain the first correlation global index.

5. The method according to claim 2, wherein: The first data carries a first function name, and the function corresponding to the first function name is a correlation function specified by the requesting party; and the constructing of a first correlation global index includes: The first data is associated with a first storage partition among the N storage partitions through an association function corresponding to the first function name to obtain the first correlation global index.

6. The method according to claim 2, wherein: The M global index expression rules include a first linear fitting relationship expression rule; and constructing a first correlation global index includes: Converting the data stored in each storage partition of the N storage partitions within the first time window to obtain N data segments, wherein the N data segments are continuous data segments, and the data in the N data segments are linearly positively correlated; In the case where the N data segments include a target data segment, obtaining target data in the target data segment by using a clustering algorithm, and the amount of data in the target data segment is greater than or equal to a preset threshold; Performing linear fitting on the target data to obtain the first linear fitting relationship expression rule.

7. The method according to claim 6, wherein: The N storage partitions include a second storage partition and a third storage partition, the second storage partition being used to store data that does not match a global index expression rule; The method further comprises: Obtain a second linear fitting relationship expression rule for data stored in each storage partition of the N storage partitions within a second time window; Based on the second linear fitting relationship expression rule, traverse the second storage partition; When second data matching the second linear fitting relationship expression rule is obtained from the second storage partition, the second data is moved from the second storage partition to a third storage partition corresponding to the second linear fitting relationship expression rule.

8. The method according to claim 7, wherein: The association information between the second storage partition and the first correlation global index is stored in a binary tree model; the method further includes: Remove association information between the first correlation global index and the partition key of the second storage partition from the binary tree model; The first linear fitting relationship expression rule among the M global index expression rules is replaced by the second linear fitting relationship expression rule.

9. The method according to claim 1, wherein: The first data includes a first key field; the method further includes: Matching the first key field with the key field of each global index expression rule in the M global index expression rules respectively to obtain a first matching result; When the first matching result indicates that there is a target key field matching the first key field among the M global index expression rules, the global index expression rule including the target key field is determined as the first global index expression rule.

10. The method according to claim 1, wherein: The method further comprises: By using a similarity algorithm, respectively calculate the similarity value between the data content of the first data and the data content divided by each global index expression rule in the M global index expression rules; When the similarity value is greater than or equal to a preset threshold, the global index expression rule corresponding to the similarity value is determined as the first global index expression rule.

11. The method according to claim 1, wherein: The method further comprises: Acquire a first data type of the first data, where the first data type is a numeric type or a non-numeric type; The global index expression rule used to divide data of the first data type among the M global index expression rules is determined as the first global index expression rule.

12. The method according to claim 1, wherein: The method further comprises: Acquire the application service scenario to which the first data belongs; The global index expression rule used to divide the data related to the application service scenario among the M global index expression rules is determined as the first global index expression rule.

13. The method according to claim 3, wherein: The N storage partitions also include a fourth storage partition; and the method further includes: receiving a data query request sent by the requesting party, wherein the data query request is used to request to query third data from the storage space; According to the data query request, traverse the binary tree model; In a case where the binary tree model includes a second relevance global index corresponding to the third data, reading the third data from a fourth storage partition corresponding to the partition key of the fourth storage partition according to association information between the second relevance global index stored in the binary tree model and the partition key of the fourth storage partition; The third data is sent to the requesting party.

14. The method according to claim 13, wherein: The N storage partitions also include a fifth storage partition; and the method further includes: In a case where the binary tree model does not include a second correlation global index corresponding to the third data, matching the M global index expression rules with the third data to obtain a second matching result; In a case where the second matching result indicates that the M global index expression rules include a second global index expression rule that matches the third data, reading the third data from the fifth storage partition corresponding to the second global index expression rule; The third data is sent to the requesting party.

15. The method according to claim 1, characterized in that The method further comprises: receiving a data modification request sent by the requesting party, wherein the data modification request is used to request that the first data stored in the first storage partition be modified into fourth data; In the case where the fourth data matches the first global index expression rule, the fourth data is written into the first storage partition, and the relationship data between the fourth data and the first storage partition is not stored.

16. The method according to claim 15, wherein: The N storage partitions also include a fifth storage partition; and the method further includes: In the case that the fourth data does not match the first global index expression rule, each of the M global index expression rules is matched with the fourth data to obtain a third matching result; When the third matching result indicates that the M global index expression rules include a third global index expression rule that matches the fourth data, the fourth data is written to the fifth storage partition corresponding to the third global index expression rule, and the relationship data between the fourth data and the fifth storage partition is not stored.

17. A data processing device, comprising: A receiving module, used for receiving a data write request sent by a requester, wherein the data write request is used for requesting to write first data into a storage space, wherein the storage space includes N storage partitions, where N is an integer greater than 1; A processing module is used to write the first data into a first storage partition when a first global index expression rule matching the first data is obtained from among M global index expression rules, and not store relationship data between the first data and the first storage partition, wherein the global index expression rule is a rule for dividing storage partitions for data, and the first storage partition is a storage partition among the N storage partitions corresponding to the first global index expression rule.

18. A computer device, comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the data processing method according to any one of claims 1 to 16 is implemented.

19. A storage medium storing computer program instructions, wherein the computer program instructions, when executed by a processor, implement the data processing method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method for storing data, memory and computer system

    CN106201350A

  • Data reading method, data writing method and object file system

    CN111399759A

  • Flash-based data access method, device and apparatus

    CN112463020A

  • Storage processing method, electronic equipment and readable storage medium

    CN116954511A

  • Data processing method and device, equipment and storage medium

    CN117762927A

Cited By

  • Data processing method and device, equipment and storage medium

    CN121117049A

  • Data processing method and device based on global index, computer equipment and readable storage medium

    CN121542271A