Data Storage Method, Device, Equipment and Medium under Multimodal Big Data System

By adjusting the data storage location in the multimodal big data system, and building an adaptive storage optimization system based on query requests and access constraints, the problem of low query efficiency in unified multimodal data queries is solved, and the query efficiency is significantly improved.

CN113946597BActive Publication Date: 2025-07-18CETC JINCANG (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111228924.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-21
Publication Date
2025-07-18
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

In the unified query of multimodal data in the prior art, the performance impact of the same data is greatly affected by different models, and the query efficiency of periodic query is low.

Method used

Obtain the data source, data table name and data access constraints according to the query request, determine the initial and target data storage locations, and adjust the storage location of the query data according to the relationship between the two, and build an adaptive storage optimization system.

Benefits of technology

The data query efficiency under multimodal big data systems has been improved, especially queries with certain periodicity, achieving significant acceleration effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113946597B_ABST
    Figure CN113946597B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data storage method, apparatus, device, and medium under a multimodal big data system, including: obtaining the data source, data table name, and data access constraint conditions of query data corresponding to a first query request according to the first query request, where one query request includes at least one piece of query data, determining the initial data storage locations of the respective query data according to the data source and the data table name, and determining the target data storage locations of the respective query data according to the data access constraint conditions; adjusting the storage locations of the query data according to the relationship between the initial data storage locations and the target data storage locations to ensure the efficiency of periodic query data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data storage, and in particular, to a data storage method, apparatus, device, and medium under a multi-modal big data system. Background Art

[0002] In recent years, the big data-related industries and technologies have developed rapidly, and data is diverse. For example, according to different data sources and usage requirements, data is presented in various forms such as structured, semi-structured, and unstructured, and correspondingly, various data models for each type have emerged, such as relational, key-value, document, graph, and other data models. Due to the existence of various data types, processing these data requires relying on various database systems. How users can achieve unified query for different types of data in complex query scenarios is a problem called unified query of multi-modal data.

[0003] To solve the problem of unified query of multi-modal data, many multi-modal big data unified query systems have emerged. Since the relational model has been used for the longest time and the query language is the most well-known, many relational database vendors have developed query support functions for other data models on the basis of supporting the management of the original relational data. At the same time, with the development and maturity of new data types such as documents, graphs, and key-values, there is a possibility of mutual conversion between their logical structures and relational data, which becomes the theoretical basis for realizing cross-model unified query. According to this theory, many unified engines for query processing of multiple data models have emerged on top of the relational model. These engines use SQL expression, SQL parsing and model mapping, and SQL query processing and optimization to make full use of the conversion between the relational model and non-relational models, and access multiple relational and non-relational model data through SQL and its extended languages to achieve unified access to multi-modal data.

[0004] In the prior art, when performing unified query of multi-modal data, since the same data stored in different models will have a great impact on performance, the query efficiency of queries with a certain period is relatively low. Summary of the Invention

[0005] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a data storage method, apparatus, device, and medium under a multi-modal big data system to improve data query efficiency.

[0006] In a first aspect, an embodiment of the present disclosure provides a data storage method under a multi-modal big data system, including:

[0007] According to a first query request, obtain the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request, where one query request includes at least one piece of the query data;

[0008] Determine the initial data storage locations of the query data according to the data source and the data table name, and determine the target data storage locations of the query data according to the data access constraint conditions;

[0009] Adjust the storage locations of the query data according to the relationship between the initial data storage location and the target data storage location.

[0010] Optionally, the determining to adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes:

[0011] When it is determined that the target data storage location is the same as the initial data storage location, keep the storage locations of the query data unchanged;

[0012] When it is determined that the target data storage location is different from the initial data storage location, obtain the first query time of the query data at the initial data storage location and the second query time of the query data at the target data storage location;

[0013] Adjust the storage location of the query data according to the relationship between the first query time and the second query time.

[0014] Optionally, the adjusting the storage location of the query data according to the relationship between the first query time and the second query time includes:

[0015] When the first query time is equal to the second query time, copy the query data at the initial data storage location to the target data storage location;

[0016] When the first query time is greater than the second query time, cut and paste the query data at the initial data storage location to the target data storage location.

[0017] Optionally, before adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location, it further includes:

[0018] Determine the target data table names of the query data according to the target data storage location;

[0019] The adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes:

[0020] Store the query data in the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location.

[0021] Optionally, the method further includes:

[0022] Obtaining the number of target data table names in the second query request according to the second query request;

[0023] In the same second query request, when the number of target data table names meets a preset number, obtaining the workload of the storage engine corresponding to the target data table name;

[0024] Adjusting the storage location of the query data in the second query request according to the relationship between the workload and the preset load.

[0025] Optionally, the method further includes:

[0026] Obtaining a first query request in a query cycle;

[0027] The adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes:

[0028] Adjusting the storage location of the query data in a query cycle according to the relationship between the initial data storage location and the target data storage location.

[0029] Optionally, the obtaining, according to the first query request, the data source, the data table name, and the data access constraint conditions of the query data corresponding to the first query request includes:

[0030] Determining the data source, the data table name, and the data access constraint conditions of the query data according to the query conditions in the first query request.

[0031] In a second aspect, an embodiment of the present disclosure provides a data storage device under a multimodal big data system, including:

[0032] A data acquisition module, configured to obtain the data source, the data table name, and the data access constraint conditions of the query data corresponding to the first query request according to the first query request, where one query request includes at least one piece of the query data;

[0033] A data storage location determination module, configured to determine the initial data storage location of each piece of the query data according to the data source and the data table name, and determine the target data storage location of each piece of the query data according to the data access constraint conditions;

[0034] A storage location adjustment module, configured to adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location.

[0035] In a third aspect, embodiments of the present disclosure provide an electronic device, including:

[0036] One or more processors;

[0037] A storage device for storing one or more programs,

[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement the data storage method as described in any one of the first aspects.

[0039] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data storage method as described in any one of the first aspects.

[0040] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:

[0041] The data storage method, apparatus, device, and medium under the multi-modal big data system provided by the embodiments of the present disclosure obtain the data source, data table name, and data access constraint conditions corresponding to the first query request according to the first query request, where one query request includes at least one query data; determine the initial data storage location of each query data according to the data source and the data table name, and determine the target data storage location of each query data according to the data access constraint conditions; adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location. The present invention is based on a multi-modal data unified query architecture, constructs an adaptive storage optimization system for user queries, and has an obvious acceleration effect on queries with a certain periodicity. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0044] Figure 1 is a flowchart of a data storage method provided by an embodiment of the present disclosure;

[0045] Figure 2 is a flowchart of another data storage method provided by an embodiment of the present disclosure;

[0046] Figure 3 It is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure;

[0047] Figure 4 It is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure;

[0048] Figure 5 It is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure;

[0049] Figure 6 It is a schematic structural diagram of a data storage device provided by an embodiment of the present disclosure;

[0050] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0051] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0052] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0053] The technical solutions of the present disclosure can be applied to electronic devices, where the electronic devices can be computers, tablets, mobile phones or other intelligent terminal devices, etc. The electronic device has a display screen, where the display screen can be a touch screen or a non-touch screen. For an electronic device with a touch screen, the user can implement interactive operations with the electronic device through gestures, fingers or touch tools (for example, a stylus). For a non-touch screen electronic device, interactive operations with the electronic device can be achieved through external devices (for example, a mouse, a keyboard or a camera, etc.) or voice recognition or expression recognition, etc.

[0054] Among them, the present disclosure does not limit the type of the operating system of the electronic device. For example, Android system, Linux system, Windows system, iOS system, etc.

[0055] Figure 1It is a schematic flowchart of a data storage method provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of processing data. The method of this embodiment can be executed by a data processing device, which can be implemented in a hardware / or software manner and can be configured in an electronic device, and can implement the data processing method described in any embodiment of the present application.

[0056] In actual application scenarios, the mixed storage of multimodal data is mainly reflected in the fact that there are a large number of different data that need to be analyzed uniformly. During the analysis process, one of the problems that need to be solved is the data storage problem, and the data storage mode will have an obvious impact on the query performance. The problem that adaptive optimization wants to solve is to adjust the data storage mode to maximize the user's query efficiency. Adaptive optimization is applicable to the following scenarios: the queries submitted by users are periodic, and data operations avoid overly frequent transaction operations such as addition, deletion, and modification. In the prior art, the storage method of multimodal data is freely determined by users. Therefore, when querying the multimodal data stored by users, the storage location of the data will affect the query performance. To ensure the data query efficiency, an embodiment of the present disclosure provides a data storage method.

[0057] As Figure 1 shown, the method specifically includes the following:

[0058] S10. According to the first query request, obtain the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request.

[0059] Among them, one query request includes at least one query data.

[0060] Specifically, for each first query request, the first query request obtains the logical plan tree of each first query request through the interface provided by the computing engine. In this logical plan tree, all leaf nodes are scan nodes, and the meaning of this node is that the computing engine obtains data from the database it is connected to, while the remaining nodes are computing nodes responsible for performing join and aggregation operations in relational calculations. The scan node provides the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request. The process of collecting the information of the leaf nodes of the logical plan tree essentially completes the decomposition and screening of the query statement in the first query request.

[0061] Specifically, determine the data source, data table name, and data access constraint conditions of the query data according to the query conditions in the first query request.

[0062] (1) When the query condition is an indefinite value - including the use of inequality signs, between, in, and is_null, the relational engine is most suitable. In most cases, the amount of data returned by this type of scan is relatively large, and it takes a lot of time and computing resources when converting cross-model types;

[0063] (2) When the query condition is a fixed value and involves multiple attributes, since document storage has the advantage of dynamic query and the result return value is relatively small through fixed-value constraints, its performance will be better than relational storage when choosing document storage;

[0064] (3) In the fixed-value condition of a single attribute, the conditions suitable for the key-value engine are relatively harsh. This is because in multi-table queries, the role that key-value storage can play is limited, and at the same time, the data types supported by key-value data are limited. For example, decimal can only be stored in the form of double in the key-value engine, and when executing decimal-related conditions, all data can only be delivered to the computing engine for screening, which instead reduces performance. Therefore, currently, only the single-attribute constraint of integer selection is handed over to the key-value engine for processing.

[0065] S20. Determine the initial data storage location of each query data according to the data source and the data table name, and determine the target data storage location of each query data according to the data access constraint conditions.

[0066] After obtaining the data source, the data table name, and the data access constraint conditions of the query data corresponding to each first query request through the interface provided by the computing engine, the data source and the data table name of the query data corresponding to the first query request form a set of mapping relationships. This relationship provides the distribution pattern of each query data before optimization, and the data access constraint conditions corresponding to the first query request jointly predict the target data storage location after adaptive optimization. And the location information corresponding to the initial data storage location and the location corresponding to the target data storage location will determine the final storage adjustment method of the query data.

[0067] S30. Adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location.

[0068] In this application, since the initial data storage location is the storage location set by the user, when querying data according to the storage location set by the user, in a multi-modal big data system, the storage location set by the user may consume some resources in the query request, reducing the data query efficiency. Therefore, in this application, after determining the target data storage location, the storage location of the query data is adjusted according to the relationship between the initial data storage location and the target data storage location, so as to provide different storage schemes according to different query requirements of the user and ensure the query efficiency.

[0069] The data storage method under the multi-modal big data system provided by the embodiments of the present disclosure obtains the data source, data table name, and data access constraint conditions corresponding to the first query request according to the first query request, where one query request includes at least one query data; determines the initial data storage location of each query data according to the data source and the data table name, and determines the target data storage location of each query data according to the data access constraint conditions; adjusts the storage location of the query data according to the relationship between the initial data storage location and the target data storage location. Based on the unified query architecture of multi-modal data, the present invention constructs an adaptive storage optimization system for user queries, and has an obvious acceleration effect on queries with a certain periodicity.

[0070] Figure 2 It is a schematic flowchart of another data storage method provided by the embodiments of the present disclosure. The embodiments of the present disclosure are based on the above embodiments, as Figure 2 shown, one realizable manner of step S30 includes:

[0071] S31. Determine whether the target data storage location is the same as the initial data storage location. When the target data storage location is the same as the initial data storage location, execute step S32. When the target data storage location is different from the initial data storage location, execute step S33 and step S34.

[0072] S32. Keep the storage locations of all query data unchanged.

[0073] S33. Obtain the first query time when the query data is at the initial data storage location, and the second query time when the query data is at the target data storage location.

[0074] S34. Adjust the storage location of the query data according to the relationship between the first query time and the second query time.

[0075] To ensure that the provided data storage method improves the query efficiency, after determining the initial data storage location of each query data according to the data source and the data table name, and determining the target data storage location of each query data according to the data access constraint conditions, determine whether the determined target data storage location is the same as the initial data storage location. When the determined target data storage location is the same as the initial data storage location, keep the storage locations of all query data unchanged. When the determined target data storage location is different from the initial data storage location, by obtaining the first query time when the user executes the query request to obtain this query data from the initial data storage location, and the second query time when obtaining this query data from the target data storage location, adjust the storage location of the query data according to the relationship between the first query time and the second query time.

[0076] The specific process of adjusting the storage location of query data according to the relationship between the first query time and the second query time is as follows: when the first query time is equal to the second query time, copy the query data at the initial data storage location to the target data storage location; when the first query time is greater than the second query time, cut and paste the query data at the initial data storage location to the target data storage location.

[0077] Determine the adjustment method of the storage location of query data by judging the first query time of the query data at the initial data storage location and the second query time of the query data at the target data storage location. When the first query time is equal to the second query time, copy the query data at the initial data storage location to the target data storage location; when the first query time is greater than the second query time, cut and paste the query data at the initial data storage location to the target data storage location, ensuring that the adaptive storage optimization system built for user queries has an obvious acceleration effect on queries with a certain periodicity.

[0078] It should be noted that when adjusting the storage location of query data according to the relationship between the initial data storage location and the target data storage location, the initial data storage location includes the data storage engine specified by the user corresponding to the query data. Therefore, after adjusting the storage location of the query data, it is necessary to establish the data storage engine of the query data at the target data storage location, so as to ensure that when the user inputs a query request, the query data at the target data storage location can be obtained according to the query engine.

[0079] Based on the above embodiments, Figure 3 is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is based on the above embodiments, as Figure 3 shown, the method further includes:

[0080] S21. Determine the target data table name of each query data according to the target data storage location.

[0081] To ensure that queries with a certain periodicity have an obvious acceleration effect, after determining the target data storage location of each query data according to the constraint conditions, extract the target data table name stored in the target data storage location of each query data, which is convenient for users to confirm query requests during periodic queries.

[0082] When the data storage method includes step S21, another implementable manner of step S30 includes:

[0083] S35. Store each query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location.

[0084] After determining the target data table names of the query data according to the target data storage location, store each query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location. Since it is possible that the determined target data storage location is the same as the initial data storage location, and it is also possible that the determined target data storage location is inconsistent with the initial data storage location. When the determined target data storage location is the same as the initial data storage location, at this time, the target data table names of the query data determined according to the target data storage location are the data table names corresponding to the first query request. At this time, the process of storing each query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location is to keep the storage location of the query data unchanged, that is, the query data is stored in the data table name corresponding to the first query request. When the determined target data storage location is inconsistent with the initial data storage location, there are two cases for determining the target data table names of the query data according to the target data storage location: Case 1: If the first query time of the query data at the initial data storage location is the same as the second query time of the query data at the target data storage location, the target data table names of the query data determined according to the target data storage location are the data table names corresponding to the first query request and the target data table names determined according to the target data storage location. At this time, the process of storing each query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location is: copy the query data at the initial data storage location to the target data storage location. At this time, the query data is stored in both the data table corresponding to the data table name of the first query request and the data table corresponding to the target data table name determined according to the target data storage location; Case 2: If the first query time of the query data at the initial data storage location is greater than the second query time of the query data at the target data storage location, the target data table names of the query data determined according to the target data storage location are the target data table names determined according to the target data storage location. At this time, the process of storing each query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location is: cut and paste the query data at the initial data storage location to the target data storage location. At this time, the query data is stored in the data table corresponding to the target data table name determined according to the target data storage location.

[0085] Figure 4 It is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is based on the corresponding embodiment, as Figure 3 shown, the method further includes: Figure 4 shown, the method further includes:

[0086] S40. Obtain the quantity of the target data table names in the second query request according to the second query request.

[0087] After adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location, to avoid multiple query data in the same query request corresponding to one storage engine and thus affecting the search efficiency, obtain the quantity of the target data table names in the second query request.

[0088] It should be noted that the query data requested by the second query request is the same as the query data requested by the first query request. However, due to the adjustment of the storage location of the query data according to the relationship between the initial data storage location and the target data storage location, that is, the same query data may be stored in different data table names.

[0089] S50. In the same second query request, when the quantity of the target data table names meets the preset quantity, obtain the workload of the storage engine corresponding to the target data table names.

[0090] By calculating and counting the usage of each storage engine in the same second query request, that is, when the quantity of the target data table names meets the preset quantity, obtain the workload of the storage engine corresponding to the target data table names. By calculating the workload of each storage engine, avoid a situation where the workload of a certain storage engine is too high, resulting in resource contention between engines and a reduction in query performance.

[0091] S60. Adjust the storage location of the query data in the second query request according to the relationship between the workload and the preset load.

[0092] Specifically, in load balancing, mainly when the workloads of the document engine and the key-value engine are too high, it will cause resource contention with the computing engine and lead to a reduction in query performance. In the current storage and computing structure, the relational engine is less affected by the workload. Therefore, the working pressure of the relational engine is not considered for the time being. When in the same second query request, the storage engine corresponding to the target data table names is the document engine or the key-value engine, and the quantity of the target data table names meets the preset quantity, obtain the workload of the storage engine corresponding to the target data table names, and adjust the storage location of the query data in the second query request according to the relationship between the workload and the preset load.

[0093] Exemplarily, set the preset quantity to 6, the workload of the document engine to 30%, and the workload of the key-value engine to 40%. If the quantity of the obtained target data table names is greater than 6, and the workload of the document engine or the key-value engine in this second query request exceeds the preset load, the load balancing mechanism will actively intervene in the query and transfer the access work of some data to the relational engine to complete.

[0094] Based on the above embodiments, Figure 5 is a schematic flowchart of another data storage method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is based on the above embodiments. As Figure 6 shown, the method further includes:

[0095] S01. Obtain a first query request in a query cycle.

[0096] Through information collection, the query data corresponding to all first query requests in a query cycle can be counted. The first query requests include two categories: full table scan query requests and conditional filter scan query requests. The conditional filter scan query requests are classified from two aspects: classified by the number of attributes involved in the condition, they are divided into single-column conditions and multi-column conditions; classified by the condition itself, they are divided into fixed-value queries, range queries, and judgment queries.

[0097] As the core of the rule library, first analyze the classification of data scan operations in the first query request. The first type is full table scan. This type of operation is very common in query scenarios involving multi-table operations. In the query statement, it is reflected as a sub-clause under the where clause. Taking the simplest equal condition as an example, on the left side of the equal sign is table information, while on the right side, it may be either table information or a natural value. When there is no where in the query or there is no attribute condition match in the where condition, a full table scan needs to be executed. This type of full table scan also has obvious characteristics in information collection - there is no constraint information in the three-level information. For a full table scan, obviously, it is the most efficient to execute through the relational engine because after a full table scan is completed through other storage modes, the data still needs to be converted into a relational table, which will inevitably bring additional time overhead.

[0098] The second is conditional filtering and scanning. Specifically, for the execution of queries on graph data, the current method is to expand the syntax library. By adding the Cypher language for graph queries, the selection of the graph storage engine is completed. After the relevant queries are executed by the graph engine, the results are returned in the form of a relational table. In the Cypher language, query operations start with the "MATCH" clause, so the "MATCH" clause can be used as a marker for graph engine selection. In the document engine, a characteristic query is to query information about nested document structures. Although most relational engines can store nested structures in JSON format, if queries are executed on the internal information of such nested structures according to certain conditions, the efficiency is significantly lower than that of the document engine because the nested structure also has index information in the document engine to support item search, while the relational engine has weaker capabilities in this regard. Therefore, for queries on nested structures, the document engine can be locked in. After considering this type of query with storage engine characteristics, the present invention will parse the attribute conditions given in the query and first divide them into two categories: fixed values and indefinite values. In the preset experiments for rule design, the relationship between these attribute conditions and storage can be determined, and the data storage location is determined according to the following heuristic rules:

[0099] (1) When the attribute condition is an indefinite value - including the use of inequality signs, "BETWEEN", "IN", and "IS NULL", most are suitable for using a relational engine. In most cases, the amount of data returned by this type of scan is large, and it takes a lot of time and computing resources when converting across model types;

[0100] (2) When the attribute condition is a fixed value and involves multiple attributes, since document storage has the advantage of dynamic query, and at the same time, the result return value is relatively small due to the fixed value constraint, the performance will be better than relational storage when choosing document storage;

[0101] (3) In the case of a fixed value condition for a single attribute, the conditions suitable for the key-value engine are relatively strict. This is because in multi-table queries, the role of key-value storage is limited, and at the same time, the data types supported by key-value data are limited. For example, "DECIMAL" can only be stored in double form in the key-value engine, and when executing "DECIMAL" related conditions, all data can only be delivered to the computing engine for screening, which instead reduces performance. Therefore, currently, only the single-attribute constraint for integer selection is handed over to the key-value engine for processing.

[0102] When the data storage method includes step S01, another implementable manner of step S30 includes:

[0103] S36. Adjust the storage location of the query data for one query cycle according to the relationship between the initial data storage location and the target data storage location.

[0104] The adaptive optimization algorithm based on historical query information makes selection recommendations for data storage through a rule base, which can ensure that a wide range of storage selections is beneficial to global queries. At the same time, the adaptive optimization adopts a strategy of executing at regular intervals, and adjusts the storage locations of the query data in a query cycle according to the relationship between the initial data storage location and the target data storage location, ensuring that every time a query cycle is collected, the global query performance is improved through regular optimization.

[0105] Figure 6 is a schematic structural diagram of a data storage device provided by an embodiment of the present disclosure, as Figure 6 shown, the data storage device includes:

[0106] A data acquisition module 510, configured to obtain the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request according to the first query request, where one query request includes at least one piece of the query data;

[0107] A data storage location determination module 520, configured to determine the initial data storage location of each query data according to the data source and data table name, and determine the target data storage location of each query data according to the data access constraint conditions;

[0108] A storage location adjustment module 530, configured to adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location.

[0109] For the data storage device provided by the embodiment of the present disclosure, the data acquisition module obtains the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request according to the first query request. The data storage location determination unit determines the initial data storage location of each query data according to the data source and data table name, and determines the target data storage location of each query data according to the data access constraint conditions. The storage location adjustment module adjusts the storage location of the query data according to the relationship between the initial data storage location and the target data storage location. The present invention is based on a multi-modal data unified query architecture, constructs an adaptive storage optimization system for user queries, and has an obvious acceleration effect on queries with a certain periodicity.

[0110] Optionally, it further includes:

[0111] A first determination module, configured to keep the storage locations of each query data unchanged when it is determined that the determined target data storage location is the same as the initial data storage location;

[0112] A second determination module, configured to, when it is determined that the target data storage location is inconsistent with the initial data storage location, obtain a first query time when the query data is at the initial data storage location and a second query time when the query data is at the target data storage location;

[0113] An adjustment module, configured to adjust the storage location of the query data according to the relationship between the first query time and the second query time.

[0114] Optionally, it further includes:

[0115] A copying module, configured to, when the first query time is equal to the second query time, copy the query data at the initial data storage location to the target data storage location;

[0116] A cutting and pasting module, configured to, when the first query time is greater than the second query time, cut and paste the query data at the initial data storage location to the target data storage location.

[0117] A target data table name determination module, configured to determine the target data table name of each query data according to the target data storage location;

[0118] A storage location adjustment unit, configured to store each query data into a target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location.

[0119] Optionally, it further includes:

[0120] A quantity determination module, configured to obtain the quantity of the target data table names in the second query request according to the second query request;

[0121] A workload acquisition module, configured to, in the same second query request, when the quantity of the target data table names meets a preset quantity, acquire the workload of the storage engine corresponding to the target data table name;

[0122] A first storage location adjustment module, configured to adjust the storage location of the query data in the second query request according to the relationship between the workload and the preset workload.

[0123] Optionally, it further includes:

[0124] A data acquisition unit, configured to acquire a first query request in a query cycle;

[0125] A storage location adjustment unit, configured to adjust the storage location of the query data in a query cycle according to the relationship between the initial data storage location and the target data storage location.

[0126] Optionally, it further includes:

[0127] A data storage location determination unit, configured to determine the data source, data table name, and data access constraint conditions of the query data according to the query conditions in the first query request.

[0128] The device provided by the embodiments of the present invention can execute the methods provided by any embodiments of the present invention, and has corresponding functional modules and beneficial effects for executing the methods.

[0129] It should be noted that in the embodiments of the above device, the included units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for easy distinction from each other, and are not used to limit the protection scope of the present invention.

[0130] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure, as Figure 7 shown, the electronic device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device can be one or more, Figure 7 here, one processor 610 is taken as an example; the processor 610, the memory 620, the input device 630, and the output device 640 in the electronic device can be connected through a bus or other means, Figure 7 here, the connection through the bus is taken as an example.

[0131] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. The processor 610 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 620, that is, implements the methods provided by the embodiments of the present invention.

[0132] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely provided relative to the processor 610, and these remote memories may be connected to the computer device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.

[0133] The input device 630 can be used to receive input numerical or character information and generate key signal inputs related to user settings and function controls of the electronic device, and can include a keyboard, a mouse, etc. The output device 640 can include a display device such as a display screen.

[0134] Embodiments of the present disclosure also provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to implement the methods provided by embodiments of the present invention when executed by a computer processor.

[0135] Certainly, for a storage medium containing computer-executable instructions provided by embodiments of the present invention, the computer-executable instructions are not limited to the method operations as described above, and can also execute relevant operations in the methods provided by any embodiment of the present invention.

[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0137] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0138] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data storage method under a multimodal big data system, characterized in that Including: According to the first query request, obtain the data source, data table name, and data access constraint conditions corresponding to the first query request, where one query request includes at least one piece of the query data, and the data access constraint condition is a constraint condition for determining the data access method according to the query conditions in the first query request, and the query conditions include at least one of indefinite values, multi-attribute fixed values, and single-attribute fixed values; Determine the initial data storage location of each piece of the query data according to the data source and the data table name, and determine the target data storage location of each piece of the query data according to the data access constraint conditions; Adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location.

2. The method according to claim 1, wherein The adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes: When it is determined that the target data storage location is the same as the initial data storage location, keep the storage location of each piece of the query data unchanged; When it is determined that the target data storage location is different from the initial data storage location, obtain the first query time when the query data is at the initial data storage location and the second query time when the query data is at the target data storage location; Adjust the storage location of the query data according to the relationship between the first query time and the second query time.

3. The method according to claim 2, wherein The adjusting the storage location of the query data according to the relationship between the first query time and the second query time includes: When the first query time is equal to the second query time, copy the query data at the initial data storage location to the target data storage location; When the first query time is greater than the second query time, cut and paste the query data at the initial data storage location to the target data storage location.

4. The method according to claim 1, wherein Before adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location, it further includes: Determine the target data table name of each piece of the query data according to the target data storage location; The adjusting the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes: Store each piece of the query data into the target data table corresponding to the target data table name according to the relationship between the initial data storage location and the target data storage location.

5. The method according to claim 4, characterized in that, The method further includes: According to the second query request, obtain the number of target data table names in the second query request. The query data of the second query request is the same as the query data of the first query request, but due to the adjustment of the storage location of the query data according to the relationship between the initial data storage location and the target data storage location, that is, the same query data is stored in different data table names; In the same second query request, when the number of target data table names meets the preset number, obtain the workload of the storage engine corresponding to the target data table name. Adjust the storage location of the query data in the second query request according to the relationship between the workload and the preset load.

6. The method according to claim 1, characterized in that The method further includes: Obtain a first query request in a query cycle; The adjustment of the storage location of the query data according to the relationship between the initial data storage location and the target data storage location includes: Adjust the storage location of the query data in a query cycle according to the relationship between the initial data storage location and the target data storage location.

7. The method according to claim 1, characterized in that The obtaining of the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request according to the first query request includes: Determine the data source, data table name, and data access constraint conditions of the query data according to the query conditions in the first query request.

8. A data storage device under a multi-modal big data system, characterized in that, Includes: A data acquisition module, configured to obtain the data source, data table name, and data access constraint conditions of the query data corresponding to the first query request according to the first query request, where one query request includes at least one piece of the query data, and the data access constraint condition is a constraint condition for determining the data access method according to the query conditions in the first query request, and the query conditions include at least one of an indefinite value, a multi-attribute fixed value, and a single-attribute fixed value; A data storage location determination module, configured to determine the initial data storage location of each piece of the query data according to the data source and the data table name, and determine the target data storage location of each piece of the query data according to the data access constraint conditions; A storage location adjustment module, configured to adjust the storage location of the query data according to the relationship between the initial data storage location and the target data storage location.

9. An electronic device, characterized in that, Includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the data storage method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the data storage method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data query engine selection method and server

    CN107609130A

  • Data query platform, method, and equipment and storage medium

    CN110222072A