Storage method and device of object storage data, equipment and medium

By generating and storing and querying information in object storage and using index information, the problem of low retrieval efficiency of object storage data is solved, and more efficient data query capabilities are achieved.

CN120578628APending Publication Date: 2025-09-02AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700826.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Object storage has inefficient problems in data retrieval, especially lacking index support when dealing with complex queries and large-scale data.

Method used

By extracting the public attribute parameters and business attribute parameters of the storage file, generating index information, and saving it in the index library, and storing files in the object library, and using index information to locate storage files.

Benefits of technology

It improves the accuracy of data storage and query efficiency, and can flexibly and quickly retrieve object files that meet the conditions based on multiple query conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578628A_ABST
    Figure CN120578628A_ABST
Patent Text Reader

Abstract

The invention discloses a storage method and device for object storage data, equipment and a medium. Comprising the steps that a target storage file stored in an object storage mode is acquired, and public attribute parameters and service attribute parameters of the target storage file are determined; according to predetermined attribute management information, converting the public attribute parameters and the service attribute parameters into attribute numbers, and combining the attribute numbers into index information; and storing the index information in an index library, and storing the target storage file in an object storage library so as to position the storage file in the object storage library according to the object storage unique identifier contained in the index information. According to the method and the device, the public attribute parameters and the service attribute parameters of the storage file are extracted, the index information is determined according to the attribute parameters, and the storage file is stored and inquired and positioned according to the unique object storage identifier contained in the index information, so that the accuracy of data storage is improved; therefore, the capability and efficiency of data query according to the storage result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a storage method, device, equipment and medium for object storage data. Background Art

[0002] Amidst the digital transformation of the financial industry, banks are experiencing a dramatic increase in the volume of unstructured data. This massive data volume poses numerous challenges to bank data management. Object storage, as a new storage method, has gained widespread adoption in recent years. Unlike file storage, object storage utilizes a single, flat structure without a folder hierarchy. Unlike the nested, hierarchical structure used by file storage, all objects are stored in a flat address space. Furthermore, all default and custom metadata is stored alongside the objects themselves in a flat address space with unique identifiers, making this approach easier to index and access. Compared to traditional storage, object storage offers unlimited scalability. Because it is object-based, it can easily expand capacity and handle large amounts of data. Furthermore, object storage offers high reliability. It utilizes data redundancy check technology to ensure data is not lost due to drive failures. Furthermore, object storage offers the advantages of low cost and high availability.

[0003] Similarly, for stored object data, if you need fast search and query capabilities based on business scenarios, object storage provides native search capabilities. Object storage assigns a unique identifier (such as an object key) to each object, allowing users to precisely search for the object. Prefix queries can also be used to list objects in a bucket. However, due to object storage's lack of internal understanding of the data and the fact that object storage systems are typically used to store large amounts of data, native search capabilities are relatively weak, resulting in low query efficiency.

[0004] In summary, object storage has certain limitations in data retrieval, especially when dealing with complex queries, lack of index support, and large-scale data. Summary of the Invention

[0005] The present invention provides a method, apparatus, device and medium for storing object-based data, so as to solve the efficiency problem of object-based data retrieval.

[0006] According to one aspect of the present invention, a method for storing object storage data is provided, comprising:

[0007] Obtain a target storage file stored in an object storage manner, and determine public attribute parameters and business attribute parameters of the target storage file; wherein the public attributes are determined based on the common positioning characteristics of storage files in different business scenarios, and the business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file;

[0008] According to predetermined attribute management information, the public attribute parameters and the service attribute parameters are converted into attribute numbers and combined into index information; wherein the attribute management information includes at least a correspondence between candidate public attribute parameters and candidate public attribute numbers and a correspondence between candidate service attribute parameters and candidate service attribute numbers;

[0009] The index information is stored in an index repository, and the target storage file is stored in an object storage repository, so as to locate the storage file in the object storage repository according to the object storage unique identifier included in the index information.

[0010] According to another aspect of the present invention, there is provided a storage device for object storage data, comprising:

[0011] An attribute parameter determination module is used to obtain a target storage file stored in an object storage manner and determine public attribute parameters and business attribute parameters of the target storage file; wherein the public attributes are determined based on the common positioning characteristics of storage files in different business scenarios, and the business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file;

[0012] an index information determination module, configured to convert the public attribute parameters and service attribute parameters into attribute numbers according to predetermined attribute management information, and combine them into index information; wherein the attribute management information includes at least a correspondence between candidate public attribute parameters and candidate public attribute numbers and a correspondence between candidate service attribute parameters and candidate service attribute numbers;

[0013] The storage module is used to store the index information in an index library and store the target storage file in an object storage library, so as to locate the storage file in the object storage library according to the object storage unique identifier contained in the index information.

[0014] According to another aspect of the present invention, an electronic device is provided, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the object storage data storage method described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the object storage data storage method described in any embodiment of the present invention when executed.

[0019] The technical solution of the embodiment of the present invention extracts the common attribute parameters and business attribute parameters of the storage file, determines the index information based on multiple attribute parameters, and stores and queries the storage file according to the index information, thereby improving the accuracy of data storage and improving the ability and efficiency of data query based on the storage results.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a flowchart of a method for storing object storage data provided by an embodiment of the present invention;

[0023] Figure 2 is a flowchart of another object storage data storage method provided by an embodiment of the present invention;

[0024] Figure 3 is a flowchart of another object storage data storage method provided by an embodiment of the present invention;

[0025] Figure 4 1 is a schematic diagram of the structure of a storage system for object storage data provided according to an embodiment of the present invention;

[0026] Figure 5 It is the overall workflow of index structure configuration;

[0027] Figure 6 It is a configuration information model of common attributes and business attributes;

[0028] Figure 7 It is a flowchart for adding and updating indexes;

[0029] Figure 8 It is a flowchart for deleting an index;

[0030] Figure 9 It is a flowchart of query index;

[0031] Figure 10 This is a timing diagram of high-frequency query object cache;

[0032] Figure 11 It is a time sequence diagram of users uploading stored files and creating index information;

[0033] Figure 12 It is a time sequence diagram of users querying and downloading stored files;

[0034] Figure 13 1 is a schematic structural diagram of a storage device for object storage data provided according to an embodiment of the present invention;

[0035] Figure 14 The present invention is a schematic diagram of the structure of an electronic device for implementing the object storage data storage method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0037] It should be noted that the terms "candidate", "target", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.

[0038] Figure 1 The present invention provides a flowchart of a method for storing object-stored data. This embodiment is applicable to the storage and retrieval of object-stored file data. The method can be executed by a storage device for object-stored data. The storage device for object-stored data can be implemented in the form of hardware and / or software. The storage device for object-stored data can be configured in a server with computing capabilities. Figure 1As shown, the method includes:

[0039] S110: Acquire a target storage file stored in an object storage manner, and determine public attribute parameters and business attribute parameters of the target storage file.

[0040] Among them, public attributes and business attributes are predetermined according to the storage application scenario. Public attributes are determined based on the common positioning features of storage files in different business scenarios, that is, public attributes are to abstract the same parts of storage files in different business scenarios, such as the creation time and update time of file data, a number used to uniquely identify a piece of data, or other information that can locate data. Business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file, that is, business attributes are the information used when retrieving data defined by the owner of the storage file according to its own target business scenario. From the perspective of data type, they can be divided into time type, Boolean type, string type, numerical type, and floating point type. From the perspective of business data, business attributes can be report application information, such as the name of the reporting organization, subject, etc.

[0041] Among them, the target storage file is the object data stored in the object storage method. The public attribute parameters and business attribute parameters of the target storage file are determined according to the specific data information of the target storage file and the description content of the predetermined public attributes and business attributes. Object storage, also known as object-based storage, is a computer data storage architecture designed to handle large amounts of unstructured data. It manages data as objects, which is different from other storage architectures (such as file systems that manage data as a file hierarchy, while block storage manages data as sectors and blocks within tracks). Each object usually includes the data itself, varying amounts of metadata, and a globally unique identifier. When data needs to be accessed, the object storage system will use the unique identifier to find the required object.

[0042] Exemplarily, the predetermined public attributes include a creation time attribute and an object storage unique identification attribute, and the business attributes include a data type and report application information. The public attribute parameters of the target storage file refer to the specific values ​​of the public attributes of the target storage file, that is, the specific time information corresponding to the creation time attribute, and the specific number information of the object storage unique identification attribute; the business attribute parameters of the target storage file refer to the specific values ​​of the business attributes of the target storage file, that is, the type information corresponding to the data type, and the application value corresponding to the report application information.

[0043] S120: According to predetermined attribute management information, convert the common attribute parameters and the service attribute parameters into attribute numbers, and combine them into index information.

[0044] The attribute management information is used to configure and manage predetermined public attributes and service attributes, and the attribute management information includes at least the correspondence between candidate public attribute parameters and candidate public attribute numbers and the correspondence between candidate service attribute parameters and candidate service attribute numbers.

[0045] Specifically, the public attribute number corresponding to the public attribute parameter of the target storage file is determined based on the correspondence between each candidate public attribute parameter and each candidate public attribute number defined in the attribute management information. Furthermore, the business attribute number corresponding to the business attribute parameter of the target storage file is determined based on the correspondence between each candidate business attribute parameter and each candidate business attribute number defined in the attribute management information. The public attribute number and the business attribute number of the target storage file are concatenated to form index information. In other words, the index information for each storage file is composed of public attributes and business attributes.

[0046] S130: Save the index information in the index repository, and save the target storage file in the object storage repository, so as to locate the storage file in the object storage repository according to the object storage unique identifier included in the index information.

[0047] The index information of each target storage file is stored in the index repository, and each target storage file is stored in the object storage repository, where the index repository and the object storage repository each correspond to a storage area. The target storage file storage operation can be performed before step 110, and after the target storage file is stored in the object storage repository, an object storage unique identifier is generated as a public attribute of the target storage file. Therefore, the order of the target storage file storage steps is not limited.

[0048] Since the index information includes the public attributes and business attributes of each storage file, and the public attributes include positioning feature information such as the object storage unique identifier, after determining the index information to be queried by the user, the target positioning storage file can be determined from the object repository based on the object storage unique identifier in the index information as the object positioning feature.

[0049] The technical solution of this embodiment extracts the common attribute parameters and business attribute parameters of the storage file, determines the index information based on multiple attribute parameters, and stores and locates the storage file according to the index information, thereby improving the accuracy of data storage and improving the ability and efficiency of data query based on the storage results.

[0050] Figure 2 This is a flowchart of another object storage method provided by an embodiment of the present invention. This embodiment further refines the file retrieval based on index information in the above embodiment. Figure 2 As shown, the method includes:

[0051] S210: Determine query conditions according to the user's query requirements.

[0052] The query condition is determined based on the attribute management information.

[0053] The user's query requirements are determined based on the user's personalized retrieval requirements. For example, the user's query requirements are to query storage files with a storage date of a certain date and a report subject of a target subject. The content corresponding to the predetermined public attributes and business attributes is determined from the user's query requirements, and the query public attribute parameters are determined based on the content corresponding to the public attributes, and the query business attribute parameters are determined based on the content corresponding to the business attributes.

[0054] Furthermore, semantic recognition is performed on the user's query requirements, and query public attribute parameters and query business attribute parameters are determined according to the matching results of the semantic recognition results with the public attributes and business attributes.

[0055] In a feasible embodiment, the attribute management information further includes index creation configuration information for each candidate public attribute and each candidate business attribute, where the index creation configuration information includes whether to create an index or not; S210 includes:

[0056] Determine that the index creation configuration information in the attribute management information is the target public attribute and target business attribute corresponding to the index creation;

[0057] Determine query conditions based on user query requirements, target public attributes, and target business attributes.

[0058] The index creation configuration information is used to determine whether to create an index based on each candidate public attribute and each candidate business attribute, that is, whether the candidate public attribute and candidate business attribute are included in the index query conditions. The index creation configuration information indicates that the candidate public attribute and candidate business attribute corresponding to index creation are included in the index query conditions, while the index creation configuration information indicates that the candidate public attribute and candidate business attribute corresponding to no index creation are not included in the index query conditions. Users can modify the index creation configuration information for each candidate public attribute and each candidate business attribute based on query requirements. That is, users can set the query attribute information included in the index query conditions by themselves, and modify the index creation configuration information for public attributes or business attributes that have no query requirements to not create an index.

[0059] Determine the candidate public attribute for index creation in the index creation configuration information in the attribute management information as the target public attribute, and determine the candidate business attribute for index creation in the index creation configuration information in the attribute management information as the target business attribute, determine the content corresponding to the target public attribute and the target business attribute from the user's query requirements, determine the query public attribute parameters according to the content corresponding to the target public attribute, and determine the query business attribute parameters according to the content corresponding to the target business attribute.

[0060] This embodiment modifies the index creation configuration information of candidate public attributes and each candidate business attribute, thereby adaptively adjusting the query conditions and the index query conditions corresponding to the query conditions according to user needs, realizing adaptive query retrieval and avoiding the restrictions of predetermined candidate public attributes and candidate business attributes.

[0061] In a feasible embodiment, the attribute management information further includes query weights of candidate public attributes and candidate business attributes, and / or query display quantity configuration information.

[0062] Among them, the query weights of the candidate public attributes and the candidate business attributes are used to represent the sorting results of the corresponding index query results obtained according to the index query conditions. Exemplarily, if the query weight of the first candidate public attribute is greater than that of the second candidate public attribute, the first index query result corresponding to the index query condition corresponding to the first candidate public attribute is sorted before the second index query result corresponding to the index query condition corresponding to the second candidate public attribute. Furthermore, if the index query condition corresponds to multiple candidate public attributes and multiple candidate business attributes, the total query weight corresponding to the index query condition is determined based on the sum of the query weights corresponding to the candidate public attributes and the candidate business attributes, and the index query results are sorted and displayed based on the total query weight.

[0063] The query display quantity configuration information indicates the number of index query results to be displayed based on a query in the index library according to the index query criteria. For example, the index query results are sorted based on the query weights of the candidate public attributes and candidate business attributes, and a preset number of results in the sorted results are displayed based on the query display quantity configuration information.

[0064] Optionally, the query weights of the candidate public attributes and candidate business attributes may be configured according to the user's personalized display requirements, and the query display quantity configuration information may be configured according to the user's personalized display requirements and the device attributes of the query device.

[0065] This embodiment improves the sorting accuracy of index query results by configuring the query weights of candidate public attributes and candidate business attributes, which is more in line with the user's query needs. It combines the query display quantity configuration information to implement the screening of index query results, thereby improving the user's readability of index query results.

[0066] S220: Convert the query common attribute parameters and query business attribute parameters in the query condition into index query conditions according to the attribute management information.

[0067] Specifically, based on the correspondence between each candidate public attribute parameter and each candidate public attribute number defined in the attribute management information, the query public attribute number corresponding to the query public attribute parameter in the query condition is determined; and based on the correspondence between each candidate business attribute parameter and each candidate business attribute number defined in the attribute management information, the query business attribute number corresponding to the query business attribute parameter in the query condition is determined. The query public attribute number and the query business attribute number corresponding to the query condition are concatenated to form an index query condition. That is, the index query condition of each query condition is composed of public attributes and business attributes.

[0068] Furthermore, based on the above example, the index query condition does not include other public attribute parameters and other business attribute parameters corresponding to not creating an index according to the index creation configuration information.

[0069] S230: Query the index library according to the index query condition to obtain a corresponding index query result, and determine the target query object from the object storage library according to the index query result.

[0070] Since the index query conditions include positioning features corresponding to common attributes, matching is performed in the index library according to the index query conditions, and the successfully matched target index information is determined as the corresponding index query result. The target query object is determined from the object storage library based on the file data positioning features in the index query result.

[0071] Specifically, when a storage file is stored in the object repository, an object storage unique identifier corresponding to the storage file is generated and stored together with the storage file in the object repository. An index query condition is matched with the index information corresponding to each storage file in the index repository, and at least one successfully matched index information is determined as the index query result. At least one target object storage unique identifier is determined based on the common attributes included in each target index information in the index query result. Based on the target object storage unique identifier, the corresponding storage file is determined from the object repository as the target query object.

[0072] The technical solution of this embodiment increases data retrieval capabilities and improves retrieval efficiency by flexibly and quickly retrieving object files that meet the conditions based on multiple query conditions.

[0073] Figure 3 This is a flowchart of another object storage method provided by an embodiment of the present invention. This embodiment further refines the file retrieval based on index information in the above embodiment. Figure 3 As shown, the method includes:

[0074] S310: After a preset period is reached, count the user query conditions in the current period.

[0075] The preset period is a buffer update period set by the user, such as one day, one week, etc.

[0076] After the buffer update period has expired, the currently ended period is considered the current period. Statistics are then tallied for each user query condition recorded during each file data query within the current period. Specifically, the number of queries for each user query condition within the current period is tallied, and the first frequency corresponding to each user query condition is determined based on the number of queries. For example, if the preset period is one day, then after the day ends, the number of occurrences of each user query condition within the current day is tallied to calculate the first frequency.

[0077] S320: Determine a high-frequency query condition according to the user query condition.

[0078] High-frequency query conditions are determined based on the statistical results of the user query conditions. For example, based on the comparison result of the first frequency corresponding to each user query condition and the preset frequency threshold, the user query condition whose first frequency is greater than the preset frequency threshold is used as the high-frequency query condition; or, based on the statistical results of the first frequency of the user query conditions in the current period, the user query conditions are sorted in descending order, and the first preset number of user query conditions in the sorting results are determined to be high-frequency query conditions. The preset number and the preset frequency threshold can be adjusted according to the user scenario requirements, and there is no restriction on the specific values ​​here.

[0079] In a feasible embodiment, S320 includes:

[0080] Sorting the user query conditions in descending order according to the first frequency statistics results of the user query conditions in the current period, and determining the first preset number of user query conditions in the sorting results as candidate query conditions;

[0081] Determine the historical frequency statistics of candidate query conditions within the historical period;

[0082] Determining a predicted frequency of a candidate query condition based on the first frequency statistical result and the historical frequency statistical result;

[0083] The high-frequency query conditions are determined from the candidate query conditions according to the predicted frequencies.

[0084] The historical period refers to a period that has reference value for the current period in terms of date. For example, if the preset period is Tianshi, the historical period is the year-on-year period of the current period. That is, if the current period is Monday, the corresponding historical period is the previous Monday.

[0085] Specifically, the user query conditions are first sorted based on the first frequency statistics obtained by counting the user query conditions in the current cycle. The first frequencies of the user query conditions ranked higher are greater than the first frequencies of the user query conditions ranked lower. A preset number of user query conditions ranked first in the sorting results are selected as candidate query conditions. The preset number is greater than the number of cached high-frequency query conditions. The specific number can be adjusted according to time conditions and is not limited here. Candidate query conditions are conditions that users frequently query in the current cycle.

[0086] Based on the current period, the corresponding historical period is determined. The number of queries for each candidate query condition within the historical period is counted to determine the historical frequency of each candidate query condition within the historical period, which is used as the historical frequency statistics. The predicted frequency of each candidate query condition is determined based on the sum of the first frequency and the historical frequency of the candidate query condition. Candidate query conditions with predicted frequencies greater than a preset frequency threshold are considered high-frequency query conditions. Alternatively, the candidate query conditions are sorted in descending order based on the predicted frequencies, and the first second preset number of candidate query conditions in the sorted result are determined as high-frequency query conditions.

[0087] This embodiment determines the predicted frequency of the query condition by combining the query condition statistics of the current period and the historical period, thereby improving the accuracy of the predicted frequency determination; and firstly filters the candidate query conditions from the user query conditions according to the first frequency, thereby reducing the statistical workload of the query conditions of the historical period and improving the efficiency of determining the high-frequency query conditions.

[0088] In a feasible embodiment, the historical period includes a period corresponding to last week and a period corresponding to last year determined according to the current period; the historical frequency statistical result includes a second frequency statistical result corresponding to the period corresponding to last week and a third frequency statistical result corresponding to the period corresponding to last year;

[0089] Determining the predicted frequency of the candidate query condition according to the first frequency statistical result and the historical frequency statistical result includes:

[0090] Determining a first weight of the first frequency statistical result, a second weight of the second frequency statistical result, and a third weight of the third frequency statistical result;

[0091] The predicted frequency is determined according to the sum of the product of the first frequency statistical result and the first weight, the product of the second frequency statistical result and the second weight, and the product of the third frequency statistical result and the third weight.

[0092] Among them, the corresponding cycle of last week is determined according to the Gregorian calendar date of the current cycle, and the corresponding cycle of last year is determined according to the date of the holiday in the current cycle. For example, the corresponding cycle of last week is the month-on-month cycle of the current cycle, and the corresponding cycle of last year is the cycle with the same holiday as the current cycle. If the current cycle is not a holiday, the corresponding cycle of last year is empty, that is, the statistical result of the third frequency is empty.

[0093] Specifically, the historical frequency statistics of the candidate query conditions in the corresponding period last week are determined as the second frequency statistics, and the historical frequency statistics of the candidate query conditions in the corresponding period last year are determined as the third frequency statistics. The predicted frequency of the candidate query condition is determined based on the weighted values ​​of the first frequency, second frequency and third frequency of each candidate query condition.

[0094] For example, F yesterday Indicates the first frequency of the candidate query condition, i.e., year-on-year; F lastweek Indicates the second frequency of the candidate query condition, which counts the value of the same day last week, i.e., the year-on-year comparison; F last_year_holiday The third frequency of the candidate query condition is the statistical value of the same holiday last year. Define weights w1, w2, and w3 to correspond to the influence of these three frequencies, and the sum of these weights is equal to 1 (i.e. w1+w2+w3=1). The predicted frequency of the candidate query condition can be expressed as: F=w1*F yesterday +w2*F last_week +w3*F last_year_holiday The weights of different frequencies can be adjusted based on specific circumstances to reflect the degree of impact of different business scenarios on the forecast frequency. For example, if the year-on-year change has the greatest impact on the forecast, weight w1 will be larger; if the month-on-month change has the greatest impact on the forecast, weight w2 will be larger; and if the same holiday last year had the greatest impact on the forecast, weight w3 will be larger.

[0095] Based on the frequency of query conditions in each period, the weighted frequency of candidate query conditions is calculated. Finally, the data that matches the high-frequency query conditions after analysis is loaded into the cache to accelerate object access.

[0096] S330: Determine a corresponding high-frequency query object according to the high-frequency query condition, and save the high-frequency query object in a cache in the next cycle.

[0097] Based on the attribute management information, the high-frequency query public attribute parameters and high-frequency query business attribute parameters in the high-frequency query conditions are converted into high-frequency index information; based on the high-frequency index information, a query is performed in the index library to obtain the corresponding high-frequency index query results. Based on the high-frequency index query results, the corresponding storage file is determined from the object storage library as the high-frequency query object, and the high-frequency query object is stored in the cache. After receiving the user query conditions in the next cycle, it is determined whether the user query conditions match the high-frequency query conditions. If so, the corresponding query object is directly obtained from the cache; if not, the corresponding index query conditions are determined based on the user query conditions, and the corresponding query object is obtained from the object storage library through the index library.

[0098] The technical solution of this embodiment determines high-frequency query objects by statistics of user historical query conditions, and stores the high-frequency query objects in a cache, thereby improving the data access speed of users in the next cycle.

[0099] Figure 4 A schematic diagram of a storage system for object storage data provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the system includes a configuration management unit, an index management unit, a cache management unit, and an object storage management unit. This invention designs an indexing method for object storage, abstracting complex and diverse unstructured data to extract common components while retaining customizable components based on different business needs. Using indexes, data retrieval and aggregation can be performed, and hotspot file analysis and prediction can be performed. In conjunction with the cache management unit, cache preheating can be performed to accelerate object access.

[0100] First, the object data index structure is modeled through the configuration management unit, so as to configure the common attributes, business attributes, index library, etc. The overall workflow of index structure configuration is as follows: Figure 5 The index configuration and validation process includes:

[0101] (1) The user uses the public attribute management in the configuration management unit to complete the public attribute configuration of the object data index; (2) After the public attribute configuration is completed, the configuration registration success is returned; (3) The user uses the business attribute management in the configuration management unit to complete the business attribute configuration of the object data index; Among them, the configuration information model of public attributes and business attributes is as follows Figure 6As shown in the figure, the model includes various configuration information of public attributes and business attributes; (4) After the business attribute configuration is completed, the configuration registration success is returned; (5) The user uses the index library management in the configuration management unit to complete the index library configuration of the storage object data index, that is, to maintain the index library information, such as determining the index library location information; (6) After the index library configuration is completed, the index library registration success is returned; (7) The user uses the system function to call the search engine to create the index library with the configured public attributes, business attributes and index library information. The index library creation is to initialize the index library and generate the mapping structure of the index library based on the public attributes and business attributes; Among them, the search engine is a retrieval technology that uses a specific strategy to retrieve specified information from the Internet and feedback it to the user based on user needs and certain algorithms. The search engine relies on a variety of technologies, such as web crawler technology, retrieval ranking technology, web page processing technology, big data processing technology, natural language processing technology, etc., to provide users with search services based on keywords or phrases, so that users can quickly find the information they need. With the continuous development of technology, modern search engines have become more intelligent and personalized, such as natural language processing, semantic search, personalized recommendations, etc. These functions enable search engines to better meet user needs and provide more accurate and useful information. (8) After the creation is completed, the creation success information is returned.

[0102] Specifically, the configuration management unit is primarily responsible for index information configuration and configuration information push. This unit includes four components: common attribute management, business attribute management, index library management, and configuration push. Common attributes abstract common aspects of object data, such as the data's creation time, update time, a uniquely identifying number, or other information that can be used to locate data. Business attributes are defined by the owner of the object data and are used for data retrieval. They can be categorized as time, Boolean, string, numeric, or floating-point. Index creation configuration information is determined by whether or not an index is created. If the index creation configuration indicates no index creation, queries based on that business attribute are not possible and are only returned as results. The index library physically divides index data according to the owner of the object data, isolating the data. Since configuration information for other units is stored in the configuration management unit, this unit also provides a configuration push function, pushing configuration information to other units in real time. This configuration information for other units includes at least query weights for candidate common attributes and candidate business attributes, and / or query display quantity configuration information. Each index entry consists of common attributes and business attributes. The design of public attributes requires the storage of bucket names and object numbers based on the characteristics of object storage. Users can use the bucket name + object number to locate unique objects. The design of business attributes, based on the needs of the object data user, includes search criteria that may be used later.

[0103] The index management unit is responsible for indexing and storing the characteristic information of object data based on the configured public attribute information and business attribute information, and providing external retrieval services. It provides functions such as adding indexes to object data, deleting indexes, querying indexes, and predicting object access frequency.

[0104] like Figure 7 The following is a flowchart of adding and updating indexes; Figure 8 The flowchart for deleting an index is shown below; Figure 9 The figure shows a flowchart for querying an index. Parameter verification verifies the matching degree between public attribute parameters and business attribute parameters and pre-determined candidate public attributes and candidate business attributes. Analyzing parameters to determine the index storage location determines the storage location of the corresponding index information in the index repository based on the public attribute parameters and business attribute parameters. For example, multiple index repositories are created in advance based on different business scenarios. The index information stored in each index repository corresponds to a different business scenario. Therefore, the corresponding target business scenario is determined based on the public attribute parameters and business attribute parameters, and then the corresponding index repository is determined based on the target business scenario. The index entity is the information that contains all the index content.

[0105] The index management unit also has a high-frequency query object prediction function. Each time the data is queried, the query conditions are recorded. By aggregating the query conditions of the current cycle, the high-frequency query conditions are summarized. Then, based on the high-frequency query conditions, a list of objects that meet the conditions is found and loaded into the cache management unit.

[0106] The cache management unit is used to manage the object cache. After the index management unit calculates the prediction results of the high-frequency query objects, after each cycle, the cache management unit calls the index management unit to obtain the high-frequency query condition prediction results, and calls the index management unit based on the results to query the list of hit objects. Then, based on the object list, the object storage management unit is accessed to obtain the object entity, and for high-frequency query objects, the cache management unit is loaded to improve the data access speed of the access system. Figure 10 Shown is a timing diagram of a high-frequency query object cache.

[0107] The object storage management unit mainly provides public and standard data storage and management services, aiming to meet the needs of generating, using, changing and deleting data in the business process within the bank, and provides functions such as data upload, update, deletion and download. Figure 11 The following is a sequence diagram showing the user uploading a stored file and creating index information; Figure 12 The following diagram shows the sequence of a user querying and downloading a stored file.

[0108] This embodiment uses an index management unit to index object storage data, and flexibly and quickly retrieves object data that meets the conditions based on multiple query conditions, thereby increasing data retrieval capabilities and improving retrieval efficiency. It also uses object storage indexes to predict high-frequency hot files and dynamically adjust the cache object list, which is suitable for cache acceleration capabilities in object storage scenarios, thereby accelerating object access. In addition, in high-frequency file predictions, predictions are made based on the latest data, thereby improving prediction accuracy.

[0109] Figure 13 A schematic diagram of a storage device for object storage data provided by an embodiment of the present invention. Figure 13 As shown, the device includes:

[0110] Attribute parameter determination module 1301 is used to obtain a target storage file stored in an object storage manner and determine public attribute parameters and business attribute parameters of the target storage file; wherein the public attributes are determined based on the common positioning characteristics of storage files in different business scenarios, and the business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file;

[0111] Index information determination module 1302 is configured to convert the public attribute parameters and service attribute parameters into attribute numbers based on predetermined attribute management information, and combine them into index information; wherein the attribute management information includes at least a correspondence between candidate public attribute parameters and candidate public attribute numbers and a correspondence between candidate service attribute parameters and candidate service attribute numbers;

[0112] The storage module 1303 is configured to store the index information in an index repository and store the target storage file in an object storage repository, so as to locate the storage file in the object storage repository according to the object storage unique identifier included in the index information.

[0113] The technical solution of this embodiment extracts the common attribute parameters and business attribute parameters of the stored files, determines index information based on multiple attribute parameters, stores and locates the stored files based on the index information, and flexibly and quickly retrieves object files that meet the conditions based on multiple query conditions, thereby increasing the data retrieval capability and improving the retrieval efficiency.

[0114] Optionally, the device further includes a query module, which is configured to, after storing the index information in the index library and the target storage file in the object storage library, include:

[0115] A query condition determination unit, configured to determine a query condition according to a user's query requirement; wherein the query condition is determined according to the attribute management information;

[0116] A query index code determination unit, configured to convert the query common attribute parameters and the query service attribute parameters in the query condition into an index query condition according to the attribute management information;

[0117] The query object determination unit is configured to perform a query in the index repository according to the index query condition, obtain a corresponding index query result, and determine a target query object from the object repository according to the index query result.

[0118] Optionally, the device further includes a high-frequency object caching module, including:

[0119] A query condition statistics unit is used to count the user query conditions in the current period after a preset period is reached;

[0120] A high-frequency query condition determination unit, configured to determine a high-frequency query condition based on the user query condition;

[0121] The high-frequency query object cache unit is configured to determine a corresponding high-frequency query object according to the high-frequency query condition and store the high-frequency query object in a cache in a next cycle.

[0122] Optionally, a high-frequency query condition determination unit includes:

[0123] a candidate query condition determination subunit, configured to sort the user query conditions in descending order according to a first frequency statistical result of the user query conditions in a current period, and determine the first preset number of user query conditions in the sorted result as candidate query conditions;

[0124] A historical frequency statistics subunit, configured to determine a historical frequency statistics result of the candidate query condition within a historical period;

[0125] a predicted frequency determination subunit, configured to determine the predicted frequency of the candidate query condition according to the first frequency statistical result and the historical frequency statistical result;

[0126] The high-frequency query condition determination subunit is configured to determine a high-frequency query condition from the candidate query conditions according to the predicted frequency.

[0127] Optionally, the historical period includes a corresponding period last week and a corresponding period last year determined according to the current period; and the historical frequency statistical result includes a second frequency statistical result corresponding to the corresponding period last week and a third frequency statistical result corresponding to the corresponding period last year;

[0128] The prediction frequency determination subunit is specifically used to:

[0129] determining a first weight of the first frequency statistical result, a second weight of the second frequency statistical result, and a third weight of the third frequency statistical result;

[0130] The predicted frequency is determined according to the sum of a product of the first frequency statistical result and the first weight, a product of the second frequency statistical result and the second weight, and a product of the third frequency statistical result and the third weight.

[0131] Optionally, the attribute management information further includes index creation configuration information of each candidate public attribute and each candidate business attribute, and the index creation configuration information includes creating an index and not creating an index;

[0132] The query condition determination unit is specifically used to:

[0133] Determine that the index creation configuration information in the attribute management information is a target public attribute and a target business attribute corresponding to the index creation;

[0134] The query conditions are determined according to the user's query requirements, the target public attributes and the target business attributes.

[0135] Optionally, the attribute management information further includes query weights of candidate public attributes and candidate business attributes, and / or query display quantity configuration information.

[0136] The storage device for object storage data provided by the embodiment of the present invention can execute the storage method for object storage data provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0137] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.

[0138] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0139] Figure 14 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0140] like Figure 14As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0141] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0142] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the storage of method object storage data.

[0143] In some embodiments, the storage of method object storage data may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the storage of method object storage data described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the storage of method object storage data in any other suitable manner (e.g., via firmware).

[0144] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific reference products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes switch components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, switch components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0149] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0150] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0151] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0152] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for storing object storage data, characterized in that: The method includes: Obtain a target storage file stored in an object storage manner, and determine public attribute parameters and business attribute parameters of the target storage file; wherein the public attributes are determined based on the common positioning characteristics of storage files in different business scenarios, and the business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file; According to predetermined attribute management information, the public attribute parameters and the service attribute parameters are converted into attribute numbers and combined into index information; wherein the attribute management information includes at least a correspondence between candidate public attribute parameters and candidate public attribute numbers and a correspondence between candidate service attribute parameters and candidate service attribute numbers; The index information is stored in an index repository, and the target storage file is stored in an object storage repository, so as to locate the storage file in the object storage repository according to the object storage unique identifier included in the index information.

2. The method according to claim 1, characterized in that After storing the index information in the index repository and storing the target storage file in the object storage repository, the method further includes: Determining query conditions according to the user's query requirements; wherein the query conditions are determined according to the attribute management information; Converting the query common attribute parameters and the query business attribute parameters in the query condition into index query conditions according to the attribute management information; A query is performed in the index library according to the index query condition to obtain a corresponding index query result, and a target query object is determined from the object storage library according to the index query result.

3. The method according to claim 1, characterized in that The method further comprises: After the preset period is reached, the user query conditions in the current period are counted; Determining a high-frequency query condition based on the user query condition; A corresponding high-frequency query object is determined according to the high-frequency query condition, and the high-frequency query object is stored in a cache in the next cycle.

4. The method according to claim 3, characterized in that Determining a high-frequency query condition according to the user query condition includes: Sorting the user query conditions in descending order according to a first frequency statistical result of the user query conditions in the current period, and determining the first preset number of user query conditions in the sorting result as candidate query conditions; Determine historical frequency statistics of the candidate query condition within a historical period; Determining a predicted frequency of the candidate query condition according to the first frequency statistical result and the historical frequency statistical result; A high-frequency query condition is determined from the candidate query conditions according to the predicted frequency.

5. The method according to claim 4, characterized in that The historical period includes the corresponding period of last week and the corresponding period of last year determined according to the current period; the historical frequency statistical results include the second frequency statistical results corresponding to the corresponding period of last week and the third frequency statistical results corresponding to the corresponding period of last year; Determining the predicted frequency of the candidate query condition according to the first frequency statistical result and the historical frequency statistical result includes: determining a first weight of the first frequency statistical result, a second weight of the second frequency statistical result, and a third weight of the third frequency statistical result; The predicted frequency is determined according to the sum of a product of the first frequency statistical result and the first weight, a product of the second frequency statistical result and the second weight, and a product of the third frequency statistical result and the third weight.

6. The method according to claim 2, characterized in that The attribute management information also includes index creation configuration information of each candidate public attribute and each candidate business attribute, and the index creation configuration information includes whether to create an index or not; Determine query conditions based on the user's query requirements, including: Determine that the index creation configuration information in the attribute management information is a target public attribute and a target business attribute corresponding to the index creation; The query conditions are determined according to the user's query requirements, the target public attributes and the target business attributes.

7. The method according to claim 2, characterized in that The attribute management information also includes query weights of candidate public attributes and candidate business attributes, and / or query display quantity configuration information.

8. A storage device for object storage data, characterized in that: The device includes: An attribute parameter determination module is used to obtain a target storage file stored in an object storage manner and determine public attribute parameters and business attribute parameters of the target storage file; wherein the public attributes are determined based on the common positioning characteristics of storage files in different business scenarios, and the business attributes are determined based on the query requirements of the target business scenario corresponding to the target storage file; an index information determination module, configured to convert the public attribute parameters and service attribute parameters into attribute numbers according to predetermined attribute management information, and combine them into index information; wherein the attribute management information includes at least a correspondence between candidate public attribute parameters and candidate public attribute numbers and a correspondence between candidate service attribute parameters and candidate service attribute numbers; The storage module is used to store the index information in an index library and store the target storage file in an object storage library, so as to locate the storage file in the object storage library according to the object storage unique identifier contained in the index information.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the object storage data storage method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the object storage data storage method according to any one of claims 1 to 7 when executed.