Heterogeneous data storage and retrieval method, device, equipment and storage medium
By classifying attribute features and processing the storage partition of course data, the problem of too many data rows in heterogeneous data storage is solved, and more efficient storage and retrieval is achieved.
Patent Information
- Application Number
- CN202110414397.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-04-16
AI Technical Summary
The prior art has the problem of excessive number of data rows in heterogeneous data storage, especially when using the Entity-Attribute-Value (EAV) model, resulting in inefficiency in storage.
By classifying the attribute characteristics of the data to be stored, the data is divided into two categories: system attributes and non-system attributes, and stored in different storage partitions respectively, avoiding branch storage of data.
It effectively reduces the number of rows of data, improves storage efficiency, and supports rapid retrieval and aggregation operations.
Smart Images

Figure CN113282579B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to big data technology, and more particularly to a heterogeneous data storage and retrieval method, device, equipment and storage medium. Background Art
[0002] In some scenarios, such as the storage of course data, it is necessary to use a database (such as MySql) for heterogeneous storage and support transactions. Course data includes at least live courses, on-demand courses, graphic courses, paid courses, etc. They have the same state machine (such as unsubmitted, not on the shelf, removed from the shelf, pending review, rejected, approved, etc.), and can also have their own unique attributes. For example, on-demand courses may have video-related attributes, and paid courses may have price-related attributes. As the business develops, attributes will often change, and even a new type may be added (such as offline courses). For such sudden changes, there is a solution based on database to solve heterogeneous data storage.
[0003] Specifically, the attributes of the model are stored in the form of records in the data table through the Entity-Attribute-Value (EAV) database model. Even if attributes need to be added or removed, it is only necessary to delete the corresponding records in the data table without modifying the database structure.
[0004] However, in the EAV model, because each attribute of an entity type occupies a separate row, the data grows very quickly. If an entity type has 10 attributes, the traditional model occupies 1 million rows of data, while the EAV model takes up about 10 million rows. That is, a large amount of row data is generated when the EAV model is used for heterogeneous data storage. Summary of the invention
[0005] In order to solve the above technical problems, the present application hopes to provide a heterogeneous data storage and retrieval method, device, equipment and storage medium.
[0006] The technical solution of this application is implemented as follows:
[0007] In a first aspect, a data storage method is provided, the method comprising:
[0008] Obtain data to be stored including M attributes; where M is a positive integer;
[0009] Classify the M attributes according to their attribute characteristics to obtain first-category data whose attribute characteristics are system attributes and second-category data whose attribute characteristics are non-system attributes;
[0010] The first category of data is stored in a first storage partition, and the second category of data is stored in a second storage partition.
[0011] In the above scheme, each attribute in the data to be stored includes a first attribute identifier and an attribute value; the method also includes: obtaining metadata corresponding to the data to be stored; wherein the metadata includes custom attribute information, and the custom attribute information includes: custom attribute characteristics and their corresponding attribute sets, and the attribute set includes the second attribute identifier of at least one attribute; using the first attribute identifier of the target attribute and the mapping relationship between the first attribute identifier and the second attribute identifier, the corresponding custom attribute characteristics in the metadata are used as the attribute characteristics of the target attribute.
[0012] In the above scheme, storing the first category data in the first storage partition and storing the second category data in the second storage partition includes: using the mapping relationship to determine the second attribute identifier corresponding to the first attribute identifier of each attribute; using the second attribute identifier of each attribute in the first category data as the key and the attribute value as the value, and storing them in the first storage partition; using the second attribute identifier of each attribute in the second category data as the key and the attribute value as the value, and storing them in the second storage partition.
[0013] In the above scheme, the custom attribute characteristics include system attribute characteristics and multiple non-system attribute characteristics, and the second storage partition includes multiple second sub-storage partitions; storing the second-category data in the second storage partition includes: performing non-system attribute classification on the second-category data to obtain at least two third-category data; and storing the third-category data in the corresponding second sub-storage partition according to the non-system attribute type of the third-category data.
[0014] In the above scheme, the custom attribute information also includes a partition identifier corresponding to the custom attribute characteristics; the method also includes: determining a first partition identifier in the metadata based on the system attribute characteristics of the first category of data; determining a second partition identifier in the metadata based on the non-system attribute characteristics of the third category of data; storing the first category of data in the first storage partition and storing the second category of data in the second storage partition includes: storing the first category of data in the corresponding first storage partition based on the first partition identifier; storing the third category of data in the corresponding second sub-storage partition based on the second partition identifier.
[0015] In a second aspect, a data retrieval method is provided, which is applied to a full-text search engine, and is characterized in that the method includes: obtaining retrieval information; determining N first attribute identifiers based on the retrieval information; wherein N is a positive integer; searching for storage partitions based on the N first attribute identifiers to obtain corresponding N segments of retrieval data; wherein the storage partitions partition attributes according to attribute characteristics; aggregating the N segments of retrieval data according to a preset aggregation strategy to obtain retrieval results corresponding to the retrieval information.
[0016] In the above scheme, searching the storage partition according to the N first attribute identifiers to obtain the corresponding N segments of retrieval data includes: determining the N second attribute identifiers of the N first attribute identifiers based on a pre-set mapping relationship between each first attribute identifier and the corresponding second attribute identifier; searching the storage partition based on the N second attribute identifiers to obtain the N segments of retrieval data.
[0017] In the above scheme, aggregating the N segments of search data according to a preset aggregation strategy to obtain a search result corresponding to the search information includes: aggregating the N segments of search data according to the arrangement order of the N first attribute identifiers to obtain the search result.
[0018] In a third aspect, a data storage device is provided, the device comprising:
[0019] A first acquisition unit is used to acquire data to be stored including M attributes, wherein M is a positive integer;
[0020] A classification unit, used for classifying the M attribute pairs according to the attribute features of the M attributes, to obtain first-category data whose attribute features are system attributes and second-category data whose attribute features are non-system attributes;
[0021] A storage unit is used to store the first category data in a first storage partition, and store the second category data in a second storage partition.
[0022] In a fourth aspect, a data retrieval device is provided, which is applied to a full-text search engine, and the device comprises:
[0023] A second acquisition unit acquires search information;
[0024] A determination unit, configured to determine N first attribute identifiers according to the search information; wherein N is a positive integer;
[0025] A search unit searches for storage partitions according to the N first attribute identifiers to obtain corresponding N segments of search data; wherein the storage partitions are partitioned and stored according to attribute characteristics;
[0026] The aggregation unit aggregates the N segments of search data according to a preset aggregation strategy to obtain a search result corresponding to the search information.
[0027] In a fifth aspect, a data storage device is provided, comprising: a processor and a memory configured to store a computer program that can be run on the processor, wherein the processor is configured to execute the steps of the aforementioned method when running the computer program.
[0028] In a sixth aspect, a full-text search engine is provided, comprising: a processor and a memory configured to store a computer program that can be run on the processor, wherein the processor is configured to execute the steps of the aforementioned method when running the computer program.
[0029] In a seventh aspect, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program implements the steps of the aforementioned method when executed by a processor.
[0030] By adopting the above technical solution, the data to be stored is classified according to the attribute characteristics of each attribute in the data to be stored, and the data to be stored is divided into the first category of data whose attribute characteristics are system attributes and the second category of data whose attribute characteristics are non-system attributes, and then stored in the corresponding storage partitions respectively. By performing attribute partitioning storage on the data to be stored, data with the same attribute characteristics are stored in the same partition, and there is no need to store them in rows, which avoids the generation of a large amount of row data to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a schematic diagram of the overall framework of data storage and retrieval in the embodiment of the present application;
[0032] Figure 2 This is a schematic diagram of a first process of the data storage method in an embodiment of the present application;
[0033] Figure 3 A schematic diagram of a UML class diagram in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of a data classification method in an embodiment of the present application;
[0035] Figure 5 This is a schematic diagram of the preprocessing process of the attribute value in the embodiment of the present application;
[0036] Figure 6 This is a flow chart of a storage method after data classification in an embodiment of the present application;
[0037] Figure 7 This is a schematic diagram of the first process of the data retrieval method in the embodiment of the present application;
[0038] Figure 8 This is a second flow chart of the data retrieval method in the embodiment of the present application;
[0039] Fig. 9 This is a schematic diagram of the second sub-process of the data retrieval method in an embodiment of the present application;
[0040] Fig.10 This is a schematic diagram of the structure of the data storage device in the embodiment of the present application;
[0041] Fig.11 This is a first structural diagram of the full-text search engine in the embodiment of the present application;
[0042] Fig.12 This is a schematic diagram of the structure of the data storage device in the embodiment of the present application;
[0043] Fig.13 This is a second structural diagram of the full-text search engine in the embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0045] Before introducing the data storage and retrieval method, this application provides a data storage and retrieval architecture diagram. Figure 1 Schematic diagram of the overall framework of data storage and retrieval in the embodiment of the present application.
[0046] like Figure 1 As shown, the overall framework of data storage and retrieval mainly includes API layer, application layer and storage layer.
[0047] The API layer includes a RESTFUL API interface or an RPC API interface, and provides external services through the RESTFUL API interface or the RPC API interface.
[0048] Each filter in the application layer is used to implement different functions of filtering attributes, such as filtering sensitive information of text or pictures in attributes, etc. Different filters implement different filtering functions. The application layer also encapsulates the complexity of underlying storage and retrieval, and partitions attributes with different attribute characteristics, that is, classifies the data to be stored, and stores the classified data in the corresponding database of the storage layer.
[0049] Here, the data to be stored may include three different attribute features, namely, square, triangle and circle. Different partition ranges are set for different attribute features. For example, the partition range corresponding to the square is 0-255, and 0-255 can be further divided into 0-127 and 128-255; the partition range corresponding to the triangle is 256-511; and the partition range corresponding to the circle is 512-767. Different partition ranges correspond to different read / write processors, and different partition ranges correspond to different storage partitions (i.e., databases). The databases of the storage layer include OSS databases, Mysql databases, and other databases (such as redis databases and Hbase databases).
[0050] The storage layer also includes a full-text search engine (ElasticSearch, ES), which is used to obtain corresponding data from various databases and aggregate them to obtain the final search results.
[0051] The following is a detailed explanation of data storage and data retrieval.
[0052] Embodiment 1
[0053] The present application provides a data storage method. Figure 2 Schematic diagram of the first process of the data storage method in the embodiment of the present application. Figure 2 As shown, the data storage method may specifically include:
[0054] Step 201: Acquire data to be stored including M attributes; wherein M is a positive integer;
[0055] It should be noted that the data to be stored is relative to the stored data, and the data to be stored refers to data that has a different structure from the stored data.
[0056] Here, we use course data as an example. If the stored data includes on-demand courses, and the data to be stored includes paid courses, the on-demand courses and paid courses have the same attributes (such as course name, etc.), and also have their own unique attributes (such as on-demand courses have video-related attributes, and paid courses have price-related attributes). It can be seen that the data structure of the data to be stored changes at any time, so the data storage method of this application is mainly for heterogeneous data storage methods.
[0057] Step 202: classifying the M attributes according to their attribute features to obtain first-category data whose attribute features are system attributes and second-category data whose attribute features are non-system attributes;
[0058] It should be noted that the attribute characteristics of an attribute include system attributes and non-system attributes, wherein the system attributes are the common attributes of each attribute in the data to be stored, and the non-system attributes are the non-common attributes of each attribute in the data to be stored.
[0059] It should be noted that step 202 is mainly to obtain the attribute characteristics of each attribute, and then classify the M attributes according to the attribute characteristics of each attribute. Attribute characteristics generally include system attributes and non-system attributes, so here the N attributes whose attribute characteristics are system attributes are counted as the first category of data, and the P attributes whose attribute characteristics are non-system attributes are counted as the second category of data. Wherein, N+P=M, N and P are positive integers.
[0060] In some embodiments, each attribute in the data to be stored includes a first attribute identifier and an attribute value; the method also includes: obtaining metadata corresponding to the data to be stored; wherein the metadata includes custom attribute information, and the custom attribute information includes: custom attribute characteristics and their corresponding attribute sets, and the attribute set includes a second attribute identifier of at least one attribute; utilizing the first attribute identifier of the target attribute and the mapping relationship between the first attribute identifier and the second attribute identifier, the corresponding custom attribute characteristics in the metadata are used as the attribute characteristics of the target attribute.
[0061] It should be noted that this embodiment is a method for obtaining the attribute characteristics of each attribute. The method mainly customizes the attribute characteristics of each attribute in advance in the metadata corresponding to the data to be stored, and then determines the attribute characteristics of the corresponding attribute in the data to be stored according to the attribute characteristics set in the metadata.
[0062] It should be noted that each attribute in the data to be stored includes a first attribute identifier and an attribute value, and the metadata is custom attribute information for the first attribute identifier of each attribute in the data to be stored, and the custom attribute information includes at least a custom attribute feature and a second attribute identifier. Here, when determining the attribute feature of the target attribute, the metadata can be queried using the data to be stored, or the data to be stored can be queried using the metadata. Specifically, iterate the first attribute identifier of each attribute in the data to be stored, determine the second attribute identifier corresponding to the first attribute identifier based on a mapping relationship, and use the custom attribute feature corresponding to the second attribute identifier as the attribute feature of the corresponding attribute; or iterate each custom attribute information in the metadata, determine the first attribute identifier corresponding to the second attribute identifier in each custom attribute information based on a mapping relationship, and use the custom attribute feature corresponding to the second attribute identifier as the attribute feature of the corresponding attribute.
[0063] Step 203: Store the first category of data in a first storage partition, and store the second category of data in a second storage partition.
[0064] It should be noted that step 203 mainly stores the first category data and the second category data classified in the previous step into different storage partitions, so as to facilitate targeted retrieval during subsequent data retrieval. That is, if only the first category data needs to be obtained, access the first storage partition; if only the second category data needs to be obtained, access the second storage partition.
[0065] In some embodiments, this step specifically includes: using the mapping relationship to determine the second attribute identifier corresponding to the first attribute identifier of each attribute; using the second attribute identifier of each attribute in the first category of data as the key and the attribute value as the value, and storing it in the first storage partition; using the second attribute identifier of each attribute in the second category of data as the key and the attribute value as the value, and storing it in the second storage partition.
[0066] It should be noted that when storing the first type of data whose attribute feature is a system attribute, the first attribute identifier of each attribute is replaced with the second attribute identifier and used as the key, and the attribute value corresponding to the first attribute identifier is used as the value and stored in the first storage partition. Here, the second attribute identifier can be the attribute logical name.
[0067] It should be noted that when storing the second type of data whose attribute characteristics are non-system attributes, the first attribute identifier of each attribute is replaced with the second attribute identifier and used as the key, and the attribute value corresponding to the first attribute identifier is used as the value and stored in the second storage partition. Here, the second attribute identifier can be the attribute ID. The attribute identifier of the attribute is replaced with the attribute ID, that is, the attribute identifier reference, which is easier to modify later and saves a certain amount of space compared to the cumbersome attribute identifier (such as the English name of the attribute).
[0068] In some embodiments, the custom attribute characteristics include system attribute characteristics and multiple non-system attribute characteristics, and the second storage partition includes multiple second sub-storage partitions; storing the second category data in the second storage partition includes: performing non-system attribute classification on the second category data to obtain at least two third category data; according to the non-system attribute type of the third category data, storing the third category data in the corresponding second sub-storage partition.
[0069] It should be noted that, since there are many types of non-system attribute features, the second type of data needs to be classified again.
[0070] Here, the non-system attribute features at least include basic attributes or small text attributes, large text attributes and frequently updated attributes. According to the non-system attribute feature type, the second category data is divided into three third category data, namely, basic attributes or small text attributes, large text attributes and frequently updated attributes, and then stored in corresponding different second sub-storage partitions respectively.
[0071] The frequently updated attribute mentioned above may be page views.
[0072] In some embodiments, the custom attribute information also includes a partition identifier corresponding to the custom attribute feature; the method also includes: determining a first partition identifier in the metadata based on the system attribute feature of the first category of data; determining a second partition identifier in the metadata based on the non-system attribute feature of the third category of data; storing the first category of data in the first storage partition and storing the second category of data in the second storage partition includes: storing the first category of data in the corresponding first storage partition based on the first partition identifier; storing the third category of data in the corresponding second sub-storage partition based on the second partition identifier.
[0073] It should be noted that, based on the previous embodiment, in order to clearly distinguish the attribute characteristics of each attribute, this embodiment pre-sets different partition identifiers for system attribute characteristics and non-system attribute characteristics; and pre-sets different partition identifiers for different non-system attribute characteristics. Subsequently, it is possible to directly identify whether the attribute characteristics of at least two attributes are the same through the partition identifier.
[0074] Specifically, the partition identifier of the first type of data is determined as the first partition identifier in the metadata, a pre-set partition range is determined based on the first partition identifier, and the first type of data is stored in the first storage partition corresponding to the partition range. The partition identifier of the second type of data is determined as the second partition identifier in the metadata, a pre-set partition range is determined based on the second partition identifier, and the second type of data is stored in the second sub-storage partition corresponding to the partition range. Here, different partition ranges correspond to different second sub-storage partitions.
[0075] Exemplarily, if the partition identifier of the system attribute feature is set to partition 0, and the partition identifier of the non-system attribute feature is set to non-partition 0, the partition range set for the basic attribute or small text attribute of the non-system attribute feature is 1-127, the partition range set for the large text attribute of the non-system attribute feature is 128-255, the partition range set for the frequently updated attribute of the non-system attribute feature is 256-511, and the partition range set for the unique constraint attribute of the non-system attribute feature is 512-767. Here, the partition identifier of the first type of data is determined as partition 0 in the metadata, the first storage partition (such as Mysql database) corresponding to partition 0 is determined, and then the first type of data is stored in the Mysql database. In the metadata, the partition identifier of the second type of data is determined to be partition 1, the partition range corresponding to partition 1 is determined to be 1-127, and the second type of data is stored in the second sub-storage partition corresponding to the partition range 1-127 (such as a Mysql database); in the metadata, the partition identifier of the second type of data is determined to be partition 128, the partition range corresponding to partition 128 is determined to be 128-255, and the second type of data is stored in the second sub-storage partition corresponding to the partition range 128-255 (such as an OSS database); in the metadata, the partition identifier of the second type of data is determined to be partition 256, the partition range corresponding to partition 128 is determined to be 256-511, and the second type of data is stored in the second sub-storage partition corresponding to the partition range 256-511 (such as a Mysql database); in the metadata, the partition identifier of the second type of data is determined to be partition 512, the partition range corresponding to partition 512 is determined to be 512-767, and the second type of data is stored in the second sub-storage partition corresponding to the partition range 512-767 (such as a Mysql database).
[0076] Here, the execution subject of step 201 to step 203 may be a processor of a data storage device.
[0077] By adopting the above technical solution, the data to be stored is classified according to the attribute characteristics of each attribute in the data to be stored, and the data to be stored is divided into the first category of data whose attribute characteristics are system attributes and the second category of data whose attribute characteristics are non-system attributes, and then stored in the corresponding storage partitions respectively. By performing attribute partitioning storage on the data to be stored, data with the same attribute characteristics are stored in the same partition, and there is no need to store them in rows, which avoids the generation of a large amount of row data to a certain extent.
[0078] Based on the above embodiment, the present application provides an example of a UML class diagram. Figure 3 is a schematic diagram of a UML class diagram in an embodiment of the present application, such as Figure 3 As shown, specifically,
[0079] Metadata refers to three tables: entity_type (entity type), entity_attribute (entity attribute) and entity_attribute_enum (entity attribute enumeration).
[0080] Here, the attributes in entity_type (entity type) include: id (long), root_id (long), parent_id (long), site id (long), namespace (String), entity name (String), entity plural alias (String), entity singular alias (String), configuration information in JSON format (String), version number (int), status (int), creator (String), modifier (String), creation time (Date), modification time (Date).
[0081] Among them, root_id (long) is used to implement entity type inheritance. The subtype can inherit the properties of the parent type and can also override the properties of the parent type.
[0082] parent_id (long) is used to implement entity type inheritance. The subtype can inherit the properties of the parent type and can also override the properties of the parent type.
[0083] Site ID (long) is used to support logical isolation of data between different sites, such as merchant help center, merchant learning center, etc.
[0084] Configuration information (String) is in JSON array format and is used to configure filters and their execution order. Each object usually has the following options: filter (filter name), order (filter execution order), params (some personalized parameters of the filter, optional part).
[0085] Version number (int). The entity version number is equal to the largest version number in the entity attribute set. It is used to achieve compatibility of different versions of data.
[0086] The attributes in entity_attribute (entity attribute) include: id (long), parent_id (long), entity type id (long), attribute name (String), attribute alias (String), system attribute (String), data type (String), regular check (String), data partition (int), document separator (ES attribute, String), document word segmenter (ES attribute, String), document data type (ES attribute, String), whether the document is searchable (ES attribute, byte), whether the document is stored (ES attribute, byte), document nested object (ES attribute, byte), unique constraint grouping (int), unique constraint index order (int), whether it is required (byte), default value (String), remarks (String), version number (int), status (int), creator (String), modifier (String), creation time (Date), modification time (Date).
[0087] Among them, parent_id(long) is used to implement composite type attributes, and the attribute can be an object (a common object or a nested object, and different types are selected according to different query methods).
[0088] Entity type id (long) points to the primary key of the entity type table, indicating which entity type the attribute belongs to.
[0089] Attribute name (String), that is, the Chinese name of the attribute, used as a comment.
[0090] Attribute alias (String), that is, the English name of the attribute. The system will generate a camel case name based on the attribute alias (the first letter following the underscore is uppercase and the underscore is removed) and use it to interact with the client. Naming convention: a combination of lowercase letters + numbers + underscores. Numbers cannot appear in the first position. Each word is separated by an underscore.
[0091] System attributes (String) refer specifically to attributes in the entity table (also called common attributes, global attributes). Since the naming rules of each business system are different, the names need to be unified when storing (attribute aliases are mapped to system attributes), and the reverse operation is performed when retrieving (system attributes are restored to attribute aliases). For example, the primary key of the course table of the learning center is course_id (attribute alias), which can be mapped to the business primary key of the system attribute. The creator may be update_user, modifier, etc., which can be mapped to the system attribute modifier. System attributes include (parent_id, business primary key, status, creator, modifier, creation time, modification time).
[0092] Data type (String), that is, Java data type. Currently supported: byte, short, int, long, float, double, decimal, enum, date, string.
[0093] Regular expression verification (String) is to use regular expressions to verify whether the value of an attribute is legal, such as range verification, length verification, etc.
[0094] Data partition (int) is partitioned according to attribute characteristics. For example: 000-255: JSON partition. 000: Partition 0 is a system partition that does not store any data. The partition of system attributes is fixed to 0. 001-127: Basic data partition, basic type attributes or small text attributes. 128-255: OSS partition, very long text, large objects, files, pictures. 256-511: Statistics, numerical (frequently updated) partition. 512-767: Unique key partition, to ensure the uniqueness of attribute values, used in conjunction with unique constraint grouping and indexing.
[0095] The document delimiter (ES attribute, String) is used to convert the original string data (such as investment promotion, platform rules, industry standards) into an array storage according to the specified delimiter to facilitate retrieval based on single or multiple values.
[0096] The document's tokenizer (ES attribute, String), that is, the tokenizer used when building a full-text index.
[0097] Document data type (ES attribute, String). Currently supported: byte, short, integer, integer_range, long, long_range, float, float_range, double, double_range, boolean, date, date_range, ip, ip_range, keyword, text.
[0098] Unique constraint index order (int), that is, this field is valid only when the data partition is within the unique key partition range. The attributes of the same unique key partition generate hash values through the index order (sequence), similar to the database Seq_in_index, starting from 1. For example: create a unique constraint for two attributes (app_key, external activity ID) in the activity entity. The data partitions of these two attributes are both 512, the index order of app_key is 1, and the index order of external activity ID is 2. The way to calculate the hash is hash(app_key:external activity ID).
[0099] Required (byte): whether the attribute value is required.
[0100] Default value (String), when the value is empty, the default value is displayed.
[0101] Version number (int), attribute version number, to achieve the coexistence of multiple versions of the attribute. For example: in version v1, the article summary attribute and all other attributes are stored in partition 1. Due to business development needs, the length of the summary becomes longer and needs to be stored in the OSS partition. Add version v2 of the summary attribute, and the partition is 128. At this time, new and old data coexist in the system. If the version number of the data is v1 (version in the entity table), the summary information is searched from partition 1. If the version number is v2, the summary information is searched from the OSS storage of partition 128.
[0102] The attributes in entity_attribute_enum (entity attribute enumeration) include: id (long), attribute ID (long), enumeration index (int), enumeration key (String), enumeration value (String), status (int), creator (String), modifier (String), creation time (Date), modification time (Date).
[0103] Among them, the attribute ID (long) points to the primary key of the entity attribute table.
[0104] Enumeration index (int), also known as enumeration code, is a number, and the storage layer usually uses this value.
[0105] Enumeration key (String), English name, usually corresponds to the ordinal name of the Java enumeration class.
[0106] Enumeration value (String), Chinese name, corresponding to the custom enumeration name, used to describe the enumeration.
[0107] After classifying the data to be stored according to the attribute characteristics of the attributes in the above metadata, four tables are obtained, including: entity, entity_unique (entity uniqueness constraint), entity_text (entity value-JSON) and entity_long (entity value-numeric value).
[0108] Here, the attributes in entity include: id (long), parent_id (long), business primary key (long), entity type id (long), language (int), version number (int), status (int), creator (String), modifier (String), creation time (Date), modification time (Date).
[0109] Among them, parent_id (long) realizes To-One association: if a comment has multiple replies, then the parent_id of each reply points to the comment; to realize multi-language, there are multiple sub-language data under the main language data, and the parent_id of each sub-language points to the main language.
[0110] The entity type id (long) is used to point to the primary key of the entity type table.
[0111] language (int) is used to support internationalization, or multiple languages for a site.
[0112] Status (int) refers to the status of business data, and negative numbers indicate invalid data.
[0113] The attributes in entity_unique (entity uniqueness constraint) include: id (long), entity type id (long), entity id (long), data partition (int), hash (long), value_255BYTE (String).
[0114] Among them, the entity id (long) is used to point to the primary key of the entity table.
[0115] The entity type id (long) is used to point to the primary key of the entity type table.
[0116] Data partition (int) is partitioned according to attribute characteristics. For example: 000~255: JSON partition. 000: Partition 0 is the system partition and does not store any data. The partition of system attributes is fixed to 0. 001~127: Basic data partition, basic type attributes or small text attributes, attributes of the same partition will be stored together in JSON format. 128~255: OSS partition (i.e. OSS database storage), stores very long text, large objects, files, pictures, etc. When storing, only one link or a key that can uniquely represent this object is stored. 256~511: Statistical and numerical (frequently updated) partitions. 512~767: unique key partition, to ensure the uniqueness of attribute values, used in conjunction with unique constraint grouping and indexing.
[0117] Hash (long), for attribute values with the same partition, concatenate them in order and perform hash calculation (hash (ie93jd34:20001)). This field is only used to narrow the query range (when hash collision occurs), so it is also necessary to compare whether the values are equal. Just add a normal index to these three [entity type id, data partition, hash] numeric fields. At the application layer, use a write lock (SELECT...FOR UPDATE) or a read lock (LOCK IN SHARE MODE) to determine whether there is duplicate data. The filtering conditions are as follows:
[0118] WHERE entity type id = ? AND data partition = ? AND hash = ? AND value = ?
[0119] Value_255BYTE(String), in JSON format, for example: {"21":"ie93jd34","28":"20001"}.
[0120] The attributes in entity_text (entity value-JSON) include: id (long), entity id (long), data partition (int), value (JSON, 8192BYTE, String).
[0121] Among them, the entity id (long) is used to point to the primary key of the entity table.
[0122] Data partition (int) is partitioned according to attribute characteristics. For example: 000~255: JSON partition. 000: Partition 0 is the system partition and does not store any data. The partition of system attributes is fixed to 0. 001~127: Basic data partition, basic type attributes or small text attributes, attributes of the same partition will be stored together in JSON format. 128~255: OSS partition (i.e. OSS database storage), stores very long text, large objects, files, pictures, etc. When storing, only one link or a key that can uniquely represent this object is stored. 256~511: Statistical and numerical (frequently updated) partitions. 512~767: unique key partition, to ensure the uniqueness of attribute values, used in conjunction with unique constraint grouping and indexing.
[0123] The attributes in entity_long (entity value-numeric value) include: id(long), entity id(long), attribute id_0(long), attribute id_1(long), value_0(long), value_1(long).
[0124] Among them, the entity id (long) is used to point to the primary key of the entity table.
[0125] Attribute id_0 (long) is used to point to the primary key of the entity attribute table. Rule: Attributes with data partitions in the range of 256 to 511 MUST be stored here, usually for frequently updated numeric fields. Under an entity type, partitions of different attributes in the range of 256 to 511 MUST NOT be repeated.
[0126] Value_0 (long), which is the value of the attribute.
[0127] Attribute id_1 (long), optional part. When most entity types have at least 2 to 4 frequently updated numeric attributes, this table can be appropriately expanded to add multiple key-value pairs, where the key corresponds to the attribute ID and the value corresponds to the attribute value.
[0128] value_1 (long), which is the value of the attribute.
[0129] Here, entity_relationship (entity association relationship) is also included, and its attributes include: id (int), entity id-source_id (int), entity id-target_id (int) and status (int).
[0130] Among them, entity id-source_id (int) and entity id-target_id (int) are used to implement To-Many associations. For example, given a user, find the list of stores that the user follows. If the user entity type (entity_type) ID is less than the store entity type ID, the condition is: source_id = {user_id}, and the queried target_id list is the store entity IDs list. Otherwise, the condition is: target_id = {user_id}, and the queried source_id list is the store entity IDs list. Given a store, find which users follow the store, and vice versa.
[0131] Status (int), status 1 means valid, status -1 means deleted.
[0132] against Figure 3 The 1..1 that appears in the expression indicates that an object of another class has a relationship with only one object of this class, 0..* indicates that an object of another class has a relationship with zero or more objects of this class, 1..* indicates that an object of another class has a relationship with one or more objects of this class, and 0..1 indicates that an object of another class has no or only a relationship with one object of this class.
[0133] Embodiment 2
[0134] Based on the above embodiments, the present application also provides a data storage method, which includes a data classification method (i.e. Figure 4 and Figure 5 ) and the storage method after data classification (i.e. Figure 6 ).
[0135] Figure 4 Schematic diagram of the process of data classification method in the embodiment of the present application, such as Figure 4 As shown, the specific steps may include:
[0136] Step 401: Start;
[0137] Step 402: data to be stored and corresponding metadata;
[0138] Here, the data to be stored includes M attributes, and the metadata is the custom attribute information for each attribute.
[0139] Step 403: Obtain a custom attribute information set from metadata;
[0140] Here, the custom attribute information set includes custom attribute information corresponding to each attribute, the custom attribute information includes custom attribute characteristics and other attribute sets, and the other attribute sets include attribute logical names.
[0141] Step 404: Iterate the custom attribute information set;
[0142] That is, the custom attribute information corresponding to each attribute in the custom attribute information collection is iterated in turn.
[0143] Step 405: Is there a next one? If yes, go to step 406; if no, go to step 413;
[0144] Step 406: Obtain custom attribute information of the next attribute;
[0145] Step 407: searching for the first attribute identifier of the corresponding attribute in the data to be stored according to the attribute logical name in the custom attribute information;
[0146] Each of the M attributes mentioned above includes a first attribute identifier and an attribute value. The custom attribute information of the metadata includes the logical name of each attribute (i.e., the second attribute identifier). There is a mapping relationship between the logical name of each attribute and the first attribute identifier. The first attribute identifier of the corresponding attribute in the data to be stored can be searched according to the logical name of the attribute, so that the custom attribute feature corresponding to the logical name of the attribute is the attribute feature of the corresponding attribute in the data to be stored.
[0147] Here, usually after executing step 407, some preprocessing will be performed on the attribute value, namely steps 501 to 509.
[0148] Figure 5 Schematic diagram of the preprocessing process of attribute values in the embodiment of the present application. Figure 5 As shown, the specific steps include the following:
[0149] Step 501: whether to set the default value; if so, execute step 502; if not, execute step 503;
[0150] The default value refers to a default value. Here, when the default value is set and is not empty, and the attribute value of the entity attribute is empty, step 502 is executed; when the attribute value of the entity attribute is not empty, step 503 is executed.
[0151] Step 502: Convert the default value into the data type specified in the attribute metadata;
[0152] Step 503: converting the attribute value of the entity attribute into the data type specified in the attribute metadata;
[0153] Step 502 and step 503 implement data type conversion through DataType.apply. DataType is a custom data type enumeration that implements the Java Function interface. The apply method is used to convert data types, and different enumeration types have their own implementations.
[0154] Step 504: data type verification;
[0155] Step 505: Check whether the value is required by regular expression;
[0156] Step 506: Check whether the check value range is required;
[0157] Step 505 and step 506 implement data type verification through DataType.test. DataType is a custom data type enumeration that implements the Java BiPredicate function interface. The test method is used to verify the data type.
[0158] Step 507: Whether the verification is passed; if so, execute step 508; if not, execute step 509;
[0159] Step 508: Is the attribute value empty? If so, execute step 412; if not, execute step 408;
[0160] Step 509: Throwing an abnormal error message;
[0161] Step 408: Whether the attribute is a system attribute; if so, execute step 409; if not, execute step 410;
[0162] Here, the system attribute is a common attribute of each attribute in the data to be stored, and the non-system attribute is a non-common attribute of each attribute in the data to be stored.
[0163] Here, if the attribute is a system attribute, it corresponds to partition 0; if the attribute is a non-system attribute, it corresponds to partition non-0.
[0164] Step 409: Obtain the attribute logical name mapped to the attribute in the metadata and use it as the KEY;
[0165] Step 410: Obtain the attribute ID of the attribute in the metadata and use it as the KEY;
[0166] Step 411: setting the attribute value in the corresponding attribute partition set;
[0167] Step 412: Complete current attribute processing;
[0168] Step 413: Process the system attribute collection to fill the values into the system attributes of the entity;
[0169] Step 414: Iterate the non-system attribute partition set, KEY = partition number, VALUE = attribute set;
[0170] Step 415: construct an entity value object instance, set the partition, and set the JSON value;
[0171] Step 416: Set an entity value object set for the current entity;
[0172] Step 417: End.
[0173] based on Figure 4 and Figure 5 In the step, if the data to be stored includes KEY1-VALUE1, KEY2-VALUE2, KEY3-VALUE3 ... KEYn-VALUEn, the data to be stored after classification is: partition 0 includes KEY1-VALUE1 and KEY2-VALUE2; partition 1 includes KEY3-VALUE3, KEY4-VALUE4 and KEY5-VALUE5; partition 128 includes KEY6-VALUE6; partition 256 includes KEY7-VALUE7. Here, the same partition can store at least one different attribute. In addition, partitioning can be continued according to needs, which will not be elaborated here.
[0174] For non-system attribute features, data of different partitions need to be stored in different storage partitions. Therefore, the present application provides a storage method for non-system attribute feature data after classification. Figure 6 The figure is a flowchart of the storage method after data classification in the embodiment of the present application.
[0175] like Figure 6 As shown, the specific steps may include:
[0176] Step 601: Get the partition range according to the partition identifier of the attribute; if the partition range is 1 to 127, execute step 602; if the partition range is 128 to 255, execute step 603; if the partition range is 256 to 511, execute step 607; if the partition range is 512 to 767, execute step 609;
[0177] Step 602: Determine the processors in the partition range of 1 to 127;
[0178] Here, the partition range of 1 to 127 stores data related to basic attributes or small text data.
[0179] The processor of the partition range 1 to 127 operates the corresponding storage partition (such as Mysql database) through the EntityText object, specifically executing step 613. If it is determined that the primary key exists, it means that the previously stored data needs to be updated, that is, executing step 615; if it is determined that the primary key does not exist, it means that the previously stored data needs to be updated, that is, executing step 614.
[0180] Step 603: Determine the processors in the partition range of 128 to 255;
[0181] Step 604: Generate OSS unique KEY;
[0182] Step 605: Write the large text into OSS and obtain the writing result MD5;
[0183] Step 606: Use KEY and MD5 as new values and replace the original content with a link to the OSS file;
[0184] Here, the 128-255 partition range stores the relevant data of the large text attribute, and a corresponding link is stored when storing specifically.
[0185] Use the processor in the partition range of 128-255 to operate the corresponding storage partition (such as OSS database) through the EntityText object, specifically execute step 613. If it is determined that the primary key exists, it means that the previously stored data needs to be updated, that is, execute step 615; if it is determined that the primary key does not exist, it means that the previously stored data needs to be updated, that is, execute step 614.
[0186] Step 607: Determine the processors in the partition range of 256 to 511;
[0187] Step 608: Convert the EntityText object to an EntityLong object;
[0188] Here, the partition range of 256 to 511 stores data related to frequently updated attributes.
[0189] The processor using the partition range of 256 to 511 converts the EntityText object into an EntityLong object, and then operates the corresponding storage partition (such as a Mysql database), specifically executing step 613. If it is determined that a primary key exists, it indicates that the previously stored data needs to be updated, that is, executing step 615; if it is determined that no primary key exists, it indicates that the previously stored data needs to be updated, that is, executing step 614.
[0190] Step 609: Determine the processors in the partition range of 512 to 767;
[0191] Here, the partition range 512 to 767 stores uniqueness constraint related data. The processors in the partition range 512 to 767 are usually higher than the processors in other partition ranges.
[0192] Step 610: Convert the EntityText object to an EntityUnique object;
[0193] The EntityText object is converted into an EntityUnique object using a processor in the partition range of 512 to 767.
[0194] Step 611: write lock mode / read lock mode;
[0195] In this step, to ensure the global uniqueness of the joint attribute, the partition range of 512 to 767 does not support update operations.
[0196] Step 612: Generate a HASH value according to the order of the joint unique KEY in the metadata;
[0197] Step 613: Does the primary key exist? If yes, go to step 614; if no, go to step 615;
[0198] Step 614: INSERT database;
[0199] Step 615: UPDATE the database.
[0200] Based on the above steps, partitioning can be continued, and different storage partitions can be set for different partition ranges, such as Hbase database. This will not be elaborated here.
[0201] By adopting the above technical solution, the data to be stored is classified according to the attribute characteristics of each attribute in the data to be stored, and the data to be stored is divided into the first category of data whose attribute characteristics are system attributes and the second category of data whose attribute characteristics are non-system attributes, and then stored in the corresponding storage partitions respectively. By performing attribute partitioning storage on the data to be stored, data with the same attribute characteristics are stored in the same partition, and there is no need to store them in rows, which avoids the generation of a large amount of row data to a certain extent.
[0202] Embodiment 3
[0203] The present application embodiment provides a data retrieval method, Figure 7 FIG. 1 is a first flow chart of the data retrieval method in the embodiment of the present application. Figure 7 As shown, the data retrieval method is applied to a full-text search engine, and the specific steps may include:
[0204] Step 701: Obtain search information;
[0205] It should be noted that the search information is obtained with the help of a full-text search engine (ElasticSearch, ES).
[0206] Here, ES can realize the dynamic addition of heterogeneous attributes. It defines how these dynamically added attributes should be mapped to appropriate data types to achieve seamless connection with the database storing the data and external indexing.
[0207] Specifically, ES builds indexes through dynamic templates technology and automatically establishes attribute mapping relationships with various databases. That is, when a new custom attribute information is added to an attribute in the database and data is inserted, ES will automatically create the missing attribute information based on the metadata information of the attribute type to which the data belongs. It is flexible and can well implement the retrieval of heterogeneous data. Therefore, this application implements data retrieval operations through ES.
[0208] Here, in order to avoid too much attribute information (more than 1000) in a single index when building the index, you can create n indexes with the same structure and use consistent hashing calculations. In this way, all data from a fixed site will fall into the same index, ensuring that the number of attribute information in each index does not exceed the preset threshold.
[0209] Step 702: Determine N first attribute identifiers according to the search information; where N is a positive integer;
[0210] It should be noted that the ES obtains the search information and determines the N first attribute identifiers included in the search information. Specifically, the ES performs the search operation according to the N first attribute identifiers.
[0211] Step 703: searching for storage partitions according to the N first attribute identifiers to obtain corresponding N segments of search data; wherein the storage partitions are partitioned and stored according to attribute characteristics;
[0212] It should be noted that, since the data is stored in different storage partitions according to the attribute characteristics of the attributes, the storage partition can be determined according to the retrieval data to be obtained, and it can be directly obtained from the determined storage partition. There is no need to obtain unnecessary data, which improves the retrieval speed to a certain extent.
[0213] Here, when using ES to implement data retrieval, the complexity of different underlying storage technologies is shielded, and a unified retrieval entry is provided to quickly obtain corresponding data from different storage partitions.
[0214] In some embodiments, the specific steps include: determining N second attribute identifiers of the N first attribute identifiers based on a pre-set mapping relationship between each first attribute identifier and the corresponding second attribute identifier; searching for a storage partition based on the N second attribute identifiers to obtain the N segments of retrieval data.
[0215] Here, when using ES to implement data retrieval, the stored data (i.e., attribute value) cannot be directly obtained based on the second attribute identifier of the attribute. The second attribute identifier of the attribute needs to be replaced, and the replaced identifier (hereinafter referred to as the third attribute identifier) is used to find the corresponding retrieval data. Therefore, the corresponding third attribute identifier is pre-customized for the second attribute identifier of each attribute.
[0216] Specifically, input the first attribute identifier of the attribute. When searching, you need to replace the second attribute identifier corresponding to the first attribute identifier with the third attribute identifier, obtain the corresponding search data through the third attribute identifier, and then replace the third attribute identifier with the second attribute identifier to obtain the search data corresponding to the second attribute identifier, thereby obtaining the search data corresponding to the first attribute identifier.
[0217] For example, if you are searching for courses with a lecturer whose name is equal to Wang, the first attribute identifier entered is Wang, and the second attribute identifier when stored can be: ${teacher_name} = Wang. When you actually use ES to search, the physical name of the attribute is used: keyword_12 = Wang (i.e., the third attribute identifier). Here, in actual application, ${teacher_name} = Wang is replaced with keyword_12 = Wang through the placeholder replacement process.
[0218] Step 704: Aggregate the N segments of search data according to a preset aggregation strategy to obtain search results corresponding to the search information.
[0219] It should be noted that the preset aggregation strategy is a strategy of aggregating according to the arrangement order of the N first attribute identifiers, that is, the N segments of search data are aggregated according to the arrangement order of the N first attribute identifiers to obtain search results corresponding to the N first attribute identifiers.
[0220] By adopting the above technical solution, the full-text search engine is used to search N first attribute identifiers from different storage partitions respectively, to obtain N segments of search data, and the N segments of search data are aggregated to obtain search results. Since the full-text search engine provides a unified search interface, the complexity of searching and aggregating data from different storage partitions is shielded, thereby improving search efficiency.
[0221] Embodiment 4
[0222] The present application also provides a data retrieval method. Figure 8 Schematic diagram of the second process of the data retrieval method in the embodiment of the present application.
[0223] like Figure 8 As shown, the specific steps include the following:
[0224] Step 801: Retrieve DSL analysis (original retrieval);
[0225] Here, a domain specific language (DSL) is used to parse the first attribute identifier of the attribute to be retrieved.
[0226] If the first attribute identifier is Wang, the corresponding second attribute identifier after parsing can be: ${teacher_name} = Wang.
[0227] Step 802: first placeholder replacement;
[0228] The placeholder replacement process is used to replace ${teacher_name}=Wang with keyword_12=Wang (the third attribute identifier).
[0229] Step 803: Retrieve DSL security package;
[0230] Step 804: construct ES search conditions;
[0231] The search condition includes at least one condition in step 805 .
[0232] Step 805: multi-language; paging related information and aggregation related information; query the first attribute identifier includes / excludes; construct sorting by default in descending order of creation time; routing information;
[0233] Step 806: Use ES to perform search and obtain corresponding search results;
[0234] Here, the search result corresponding to keyword_12=Wang Moumou (third attribute identifier) is obtained.
[0235] Step 807: Second placeholder replacement;
[0236] The placeholder replacement process is used to replace keyword_12=Wang Moumou (the third attribute identifier) with ${teacher_name}=Wang Moumou.
[0237] Step 808: Process the search results.
[0238] Here, the metadata corresponding to each attribute in the data to be stored includes multiple custom attribute information, and different custom attribute information corresponds to different partitions. When searching, only the data of the specified partition may be searched, or the data of all partitions of the attribute may be searched. Therefore, the present application further limits the parsing content in step 801, that is, limits the partition to be searched. Fig. 9 Schematic diagram of the second sub-process of the data retrieval method in the embodiment of the present application.
[0239] like Fig. 9 As shown, the specific steps include the following:
[0240] Step 901: whether to query only the data of the specified partition; if so, execute step 902; if not, execute step 903;
[0241] Step 902: construct a set of specified partition ranges;
[0242] Step 903: construct a set of all partition ranges contained in the metadata;
[0243] Step 904: Iterate the partition range set;
[0244] Step 905: Is there a next one? If not, go to step 906; if yes, go to step 907;
[0245] Step 906: Characterization search completed;
[0246] Step 907: Search for a corresponding read processor according to the partition range;
[0247] Step 908: request task splitting;
[0248] Step 909: Determine the read processors in the partition range of 1 to 127, search the corresponding database, and obtain the first data;
[0249] Step 910: Determine a read processor in the partition range of 128 to 255, search the corresponding database, and obtain the second data;
[0250] Step 911: determine the read processors in the partition range of 265 to 511, search the corresponding database, and obtain the third data;
[0251] Step 912: Determine the read processors in the partition range of 512 to 767, search the corresponding database, and obtain fourth data;
[0252] Step 913: performing an aggregation operation on the first data, the second data, the third data and the fourth data to obtain a corresponding aggregation result.
[0253] Based on the above steps, after partitioning and storing according to the attribute characteristics of the attributes, the present application can specifically retrieve the data corresponding to the required partition range during retrieval, without having to retrieve the data corresponding to other partition ranges, thereby improving the retrieval efficiency to a certain extent.
[0254] By adopting the above technical solution, the full-text search engine is used to search N first attribute identifiers from different storage partitions respectively, to obtain N segments of search data, and the N segments of search data are aggregated to obtain search results. Since the full-text search engine provides a unified search interface, the complexity of searching and aggregating data from different storage partitions is shielded, thereby improving search efficiency.
[0255] Embodiment 5
[0256] In order to implement the method of the embodiment of the present application, based on the same inventive concept, the embodiment of the present application also provides a data storage device, Fig.10 This is a schematic diagram of the structure of the data storage device in the embodiment of the present application. Fig.10 As shown, the data storage device comprises:
[0257] The first acquisition unit 1001 is used to acquire data to be stored including M attributes, where M is a positive integer;
[0258] A classification unit 1002 is used to classify the M attribute pairs according to the attribute features of the M attributes to obtain first-category data whose attribute features are system attributes and second-category data whose attribute features are non-system attributes;
[0259] The storage unit 1003 is used to store the first category data in a first storage partition, and store the second category data in a second storage partition.
[0260] In some embodiments, each attribute in the data to be stored includes a first attribute identifier and an attribute value; the method also includes: obtaining metadata corresponding to the data to be stored; wherein the metadata includes custom attribute information, and the custom attribute information includes: custom attribute characteristics and their corresponding attribute sets, and the attribute set includes a second attribute identifier of at least one attribute; utilizing the first attribute identifier of the target attribute and the mapping relationship between the first attribute identifier and the second attribute identifier, the corresponding custom attribute characteristics in the metadata are used as the attribute characteristics of the target attribute.
[0261] In some embodiments, the device includes a storage unit 1003, which is specifically used to use the mapping relationship to determine the second attribute identifier corresponding to the first attribute identifier of each attribute; use the second attribute identifier of each attribute in the first category of data as a key and the attribute value as a value, and store it in the first storage partition; use the second attribute identifier of each attribute in the second category of data as a key and the attribute value as a value, and store it in the second storage partition.
[0262] In some embodiments, the custom attribute characteristics include system attribute characteristics and multiple non-system attribute characteristics, and the second storage partition includes multiple second sub-storage partitions; storing the second category data in the second storage partition includes: performing non-system attribute classification on the second category data to obtain at least two third category data; according to the non-system attribute type of the third category data, storing the third category data in the corresponding second sub-storage partition.
[0263] In some embodiments, the custom attribute information also includes a partition identifier corresponding to the custom attribute feature; the method also includes: determining a first partition identifier in the metadata based on the system attribute feature of the first category of data; determining a second partition identifier in the metadata based on the non-system attribute feature of the third category of data; storing the first category of data in the first storage partition and storing the second category of data in the second storage partition includes: storing the first category of data in the corresponding first storage partition based on the first partition identifier; storing the third category of data in the corresponding second sub-storage partition based on the second partition identifier.
[0264] By adopting the above technical solution, the data to be stored is classified according to the attribute characteristics of each attribute in the data to be stored, and the data to be stored is divided into the first category of data whose attribute characteristics are system attributes and the second category of data whose attribute characteristics are non-system attributes, and then stored in the corresponding storage partitions respectively. By performing attribute partitioning storage on the data to be stored, data with the same attribute characteristics are stored in the same partition, and there is no need to store them in rows, which avoids the generation of a large amount of row data to a certain extent.
[0265] Embodiment 6
[0266] In order to implement the method of the embodiment of the present application, based on the same inventive concept, the embodiment of the present application also provides a full-text search engine, Fig.11 This is a first structural diagram of the full-text search engine in the embodiment of the present application. Fig.11 As shown, the full-text search engine includes:
[0267] The second acquisition unit 1101 acquires search information;
[0268] The determining unit 1102 is configured to determine N first attribute identifiers according to the search information, wherein N is a positive integer;
[0269] The search unit 1103 searches for storage partitions according to the N first attribute identifiers to obtain corresponding N segments of search data; wherein the storage partitions are partitioned and stored according to attribute characteristics;
[0270] The aggregation unit 1104 aggregates the N segments of search data according to a preset aggregation strategy to obtain a search result corresponding to the search information.
[0271] In some embodiments, the device includes: a search unit 1103, which is specifically used to determine the N second attribute identifiers of the N first attribute identifiers based on a pre-set mapping relationship between each first attribute identifier and the corresponding second attribute identifier; search for a storage partition based on the N second attribute identifiers to obtain the N segments of retrieval data.
[0272] In some embodiments, the apparatus includes: an aggregation unit 1104, which is specifically used to aggregate the N segments of search data according to the arrangement order of the N first attribute identifiers to obtain the search result.
[0273] By adopting the above technical solution, the full-text search engine is used to search N first attribute identifiers from different storage partitions respectively, to obtain N segments of search data, and the N segments of search data are aggregated to obtain search results. Since the full-text search engine provides a unified search interface, the complexity of searching and aggregating data from different storage partitions is shielded, thereby improving search efficiency.
[0274] The present application also provides another data storage device. Fig.12 This is a schematic diagram of the structure of the data storage device in the embodiment of the present application. Fig.12 As shown, the data storage device includes: a processor 1201 and a memory 1202 configured to store a computer program that can be run on the processor;
[0275] The processor 1201 is configured to execute the method steps in the aforementioned embodiment when running a computer program.
[0276] Of course, in practical applications, Fig.12 As shown, the various components in the data storage device are coupled together via a bus system 1203. It is understood that the bus system 1203 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1203 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.12 Various buses are labeled as bus system 1203.
[0277] The present application also provides another full-text search engine. Fig.13 This is a second structural diagram of the full-text search engine in the embodiment of the present application. Fig.13 As shown, the full-text search engine includes: a processor 1301 and a memory 1302 configured to store a computer program that can be run on the processor;
[0278] The processor 1301 is configured to execute the method steps in the aforementioned embodiment when running a computer program.
[0279] Of course, in practical applications, Fig.13 As shown in FIG. 1 , the various components in the full-text search engine are coupled together via a bus system 1303. It is understood that the bus system 1303 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1303 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.13 Various buses are labeled as bus system 1303.
[0280] In practical applications, the processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a controller, a microcontroller, and a microprocessor. It is understandable that for different devices, the electronic device used to implement the functions of the processor may also be other, and the embodiments of the present application do not specifically limit this.
[0281] The above-mentioned memory can be a volatile memory (volatile memory), such as a random access memory (RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk (HDD, Hard Disk Drive) or a solid-state drive (SSD, Solid-State Drive); or a combination of the above-mentioned types of memory, and provide instructions and data to the processor.
[0282] In an exemplary embodiment, the present application also provides a computer-readable storage medium for storing a computer program.
[0283] Optionally, the computer-readable storage medium may be applied to any one of the methods in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the processor in each method in the embodiments of the present application, which will not be described in detail here for the sake of brevity.
[0284] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0285] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0286] In addition, all functional units in the embodiments of the present invention may be integrated into one processing module, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art may understand that all or part of the steps of implementing the above method embodiments may be completed by hardware related to program instructions, and the above program may be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the above storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0287] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0288] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0289] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0290] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A data storage method, characterized in that: The method comprises: Obtain data to be stored including M attributes; where M is a positive integer; Classify the M attributes according to their attribute characteristics to obtain first-category data whose attribute characteristics are system attributes and second-category data whose attribute characteristics are non-system attributes; The first type of data is stored in a first storage partition, and the second type of data is stored in a second storage partition; wherein, The second storage partition includes a plurality of second sub-storage partitions; The storing the second type of data into the second storage partition comprises: Performing non-systematic attribute classification on the second category of data to obtain at least two third categories of data; According to the non-system attribute type of the third category data, the third category data is stored in the corresponding second sub-storage partition.
2. The method according to claim 1, characterized in that Each attribute in the data to be stored includes a first attribute identifier and an attribute value; the method further includes: Obtain metadata corresponding to the data to be stored; wherein the metadata includes custom attribute information, the custom attribute information includes: a custom attribute feature and a corresponding attribute set, the attribute set includes a second attribute identifier of at least one attribute; the custom attribute feature includes a system attribute feature and a plurality of non-system attribute features; The first attribute identifier of the target attribute and the mapping relationship between the first attribute identifier and the second attribute identifier are used to take the corresponding custom attribute feature in the metadata as the attribute feature of the target attribute.
3. The method according to claim 2, characterized in that The storing the first type of data into a first storage partition and storing the second type of data into a second storage partition comprises: Determine the second attribute identifier corresponding to the first attribute identifier of each attribute by using the mapping relationship; Using the second attribute identifier of each attribute in the first category of data as a key and the attribute value as a value, and storing them in the first storage partition; The second attribute identifier of each attribute in the second category of data is used as the key and the attribute value is used as the value, and the data is stored in the second storage partition.
4. The method according to claim 2, characterized in that: The custom attribute information also includes a partition identifier corresponding to the custom attribute feature; The method further comprises: Determining a first partition identifier in the metadata according to a system attribute characteristic of the first category of data; Determining a second partition identifier in the metadata according to the non-system attribute characteristics of the third category of data; The storing the first type of data into a first storage partition and storing the second type of data into a second storage partition comprises: According to the first partition identifier, storing the first type of data into the corresponding first storage partition; According to the second partition identifier, the third type of data is stored in the corresponding second sub-storage partition.
5. A data retrieval method, applied to a full-text search engine, characterized in that: The method comprises: Get search information; Determine N first attribute identifiers according to the search information; wherein N is a positive integer; Search the storage partition according to the N first attribute identifiers to obtain the corresponding N segments of retrieval data; wherein the storage partition is to partition and store the attributes according to the attribute characteristics; the attribute characteristics include system attributes and non-system attributes, the first category data with the attribute characteristics of the system attributes is stored in the first storage partition, the second category data with the attribute characteristics of the non-system attributes is stored in the second storage partition, the second storage partition includes a plurality of second sub-storage partitions, the second category data includes at least two third category data, each third category data corresponds to a different non-system attribute type, and each second sub-storage partition is used to store third category data of one non-system attribute type; The N segments of search data are aggregated according to the arrangement order of the N first attribute identifiers to obtain search results corresponding to the search information.
6. The method according to claim 5, characterized in that The searching for storage partitions according to the N first attribute identifiers to obtain corresponding N pieces of search data includes: Determine N second attribute identifiers of the N first attribute identifiers based on a preset mapping relationship between each first attribute identifier and a corresponding second attribute identifier; The storage partition is searched based on the N second attribute identifiers to obtain the N segments of retrieval data.
7. A data storage device, characterized in that: The device comprises: A first acquisition unit is used to acquire data to be stored including M attributes, wherein M is a positive integer; A classification unit, used for classifying the M attribute pairs according to the attribute features of the M attributes, to obtain first-category data whose attribute features are system attributes and second-category data whose attribute features are non-system attributes; A storage unit is used to store the first type of data in a first storage partition and store the second type of data in a second storage partition; wherein, The second storage partition includes a plurality of second sub-storage partitions; The storage unit is specifically used to perform non-system attribute classification on the second category data to obtain at least two third category data; according to the non-system attribute type of the third category data, the third category data is stored in the corresponding second sub-storage partition.
8. A data retrieval device, applied to a full-text search engine, characterized in that: The device comprises: A second acquisition unit acquires search information; A determination unit, configured to determine N first attribute identifiers according to the search information; wherein N is a positive integer; A search unit searches for a storage partition according to the N first attribute identifiers to obtain the corresponding N segments of retrieval data; wherein the storage partition is a partitioned storage of attributes according to attribute characteristics; the attribute characteristics include system attributes and non-system attributes, first-category data with attribute characteristics of system attributes are stored in the first storage partition, and second-category data with attribute characteristics of non-system attributes are stored in the second storage partition, the second storage partition includes a plurality of second sub-storage partitions, the second-category data includes at least two third-category data, each third-category data corresponds to a different non-system attribute type, and each second sub-storage partition is used to store third-category data of one non-system attribute type; The aggregation unit aggregates the N segments of search data according to the arrangement order of the N first attribute identifiers to obtain a search result corresponding to the search information.
9. A data storage device, characterized in that: The data storage device comprises: a processor and a memory configured to store a computer program that can be run on the processor, Wherein, the processor is configured to execute the steps of the method according to any one of claims 1 to 4 when running the computer program.
10. A full-text search engine, characterized in that: The full-text search engine comprises: a processor and a memory configured to store a computer program that can be run on the processor, Wherein, the processor is configured to execute the steps of the method described in claim 5 or 6 when running the computer program.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Facet-oriented academic big data storage and query method
CN110134661A